Jobgether
Senior Software Engineer - Reliability, Infrastructure, and Tooling

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Senior Software Engineer - Reliability, Infrastructure, and Tooling
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer - Reliability, Infrastructure, and Tooling based in United Kingdom.
This is a senior engineering role focused on building reliable, scalable infrastructure for demanding real-time and AI workloads.
You’ll work closely with product engineering teams to design systems that are secure, maintainable, observable, and resilient at global scale.
The role goes beyond traditional operations, combining infrastructure engineering, reliability, tooling, and hands-on product development.
You’ll tackle complex challenges involving distributed systems, Kubernetes, networking, real-time media, and secure execution environments.
Your work will help engineering teams self-serve infrastructure and reliability capabilities without creating unnecessary operational bottlenecks.
You’ll also contribute to incident management, on-call practices, architectural decisions, and long-term reliability improvements.
This is an opportunity to work remotely with experienced engineers while contributing to infrastructure supporting billions of real-time interactions.
Accountabilities
- Ramp up on a complex global architecture involving distributed databases, messaging systems, networking infrastructure, Kubernetes, and other core platform technologies, identifying areas of reliability debt and improvement.
- Design and ship reliability-focused engineering work directly within production codebases, including load balancing, load shedding, instrumentation, scalability, and efficiency improvements.
- Build and evolve internal infrastructure and developer tooling that enables product engineering teams to independently operate reliable workloads.
- Partner closely with product development teams to co-design systems and ensure reliability, security, maintainability, and operational readiness are considered throughout development.
- Develop observability capabilities that make system behavior measurable, understandable, and actionable, using appropriate signals and visualization techniques.
- Participate in a shared on-call rotation and contribute to effective incident response, investigation, remediation, and prevention of recurring reliability issues.
- Investigate complex system-level problems across distributed infrastructure, networking, application behavior, and production environments.
- Improve configuration management and infrastructure practices across diverse systems, reducing unnecessary complexity, errors, and technical debt.
- Contribute technical perspectives to architectural discussions and help establish engineering practices that support both short-term delivery and long-term scalability.
- Support systems with demanding workloads, including real-time media, secure customer code execution, advanced networking, and other highly concurrent services.
- Collaborate with engineering partners on potentially contentious reliability and operational decisions with clarity, pragmatism, and strong technical judgment.
- Automate repetitive operational processes wherever possible to improve engineering efficiency and reduce manual intervention.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Requirements
- Strong professional experience building and operating non-trivial production applications, particularly systems involving high concurrency, distributed workloads, or complex control loops.
- Significant experience with Kubernetes or an equivalent large-scale container orchestration platform.
- Strong understanding of Linux internals and networking, with the ability to investigate issues across multiple layers of a production system.
- Proven experience using observability, monitoring, logging, and related tooling to diagnose difficult production problems.
- Experience operating large-scale, globally distributed systems, including the configuration management, reliability challenges, and technical debt that emerge as systems grow.
- Experience responding to and managing complex production incidents, including identifying root causes and implementing durable corrective actions.
- Experience operating open-source infrastructure technologies such as Kafka, ClickHouse, or comparable distributed systems.
- Strong systems-thinking skills and an ability to reason about infrastructure in terms of signals, feedback, dependencies, and control mechanisms.
- Strong communication and collaboration skills, particularly when working with partner engineering teams and navigating competing priorities.
- A pragmatic approach to engineering that balances immediate delivery needs with long-term maintainability, reliability, and operational cost.
- A strong interest in observability, reliability engineering, clean configuration, automation, and reducing operational complexity.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Nice to have
- Experience with data engineering and analytics.
- Experience with global Layer 3 networking.
- Experience operating systems with long-lived workloads such as real-time media.
- Experience in Google SRE or another high-scale reliability engineering environment.
- Experience working with compliance frameworks such as PCI.
Benefits
- $135,000–$300,000 USD compensation range.
- Equity participation as part of the overall compensation package.
- Fully remote work with opportunities for collaboration across a globally distributed organization.
- Health, dental, and vision benefits.
- Flexible vacation policy.
- Opportunity to work on infrastructure supporting large-scale real-time and AI applications.
- Opportunity to contribute to open-source projects alongside experienced engineers.
- Exposure to challenging distributed-systems problems involving real-time media, secure compute, networking, Kubernetes, and globally distributed infrastructure.
- Opportunity to influence reliability architecture, engineering practices, and internal developer tooling.
- Shared on-call practices designed to keep production experience connected across infrastructure and product engineering teams.
- Equal opportunity employment and reasonable accommodation support throughout the hiring process.
How Jobgether works
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
Why Apply Through Jobgether?
We appreciate your interest and wish you the best!
Data Privacy Notice
By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location