Rodeo
Get started

Jobgether

Senior Site Reliability Engineer

United Kingdom
Posted about 19 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Senior Site Reliability Engineer

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United Kingdom.

This is a senior infrastructure role focused on building and operating highly reliable cloud platforms in a distributed engineering environment.

You will take ownership of production infrastructure across Kubernetes, Linux, networking, virtualization, and bare-metal environments.

The role combines deep technical expertise with automation, observability, incident management, and proactive reliability engineering.

You will help shape infrastructure architecture, improve availability and performance, and establish scalable operational practices.

Working closely with engineering and cross-functional teams, you will solve complex infrastructure challenges and optimize resource utilization.

The environment is fully remote, international, and highly collaborative, with significant autonomy and ownership.

This is an opportunity to make a direct impact on ambitious cloud infrastructure projects while working with modern technologies.

Accountabilities

  • Operate, maintain, and continuously improve Linux-based infrastructure, with a strong focus on Debian and Ubuntu environments.
  • Deploy, manage, and scale production Kubernetes clusters across bare-metal, virtualized, and on-premise environments, overseeing upgrades, node pools, networking, storage, and security hardening.
  • Design, implement, and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Build and maintain infrastructure automation using Ansible, Bash, Python, Git-based workflows, and GitOps practices, including automated provisioning through PXE boot, Preseed, and cloud-init.
  • Deploy and maintain observability and monitoring platforms such as Prometheus, Grafana, Loki, ELK, and Graylog, ensuring operational data generates actionable insights.
  • Lead incident response and escalation activities, troubleshoot complex infrastructure issues, and implement improvements that increase availability and reduce latency.
  • Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
  • Optimize alerting and monitoring pipelines while establishing effective on-call schedules to provide operational coverage across time zones.
  • Create and maintain Standard Operating Procedures for recurring infrastructure operations, maintenance, troubleshooting, and incident management.
  • Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance, and data-center operations.
  • Manage virtualization and orchestration layers using technologies such as OpenStack, Proxmox, and VMware.
  • Contribute to the overall architecture and evolution of infrastructure products, ensuring solutions remain scalable and reliable.
  • Plan infrastructure capacity and resources for future initiatives based on projected demand and business growth.
  • Partner with development teams to improve system quality, optimize resource utilization, and strengthen engineering practices.
  • Collaborate with cross-functional stakeholders to align infrastructure priorities with broader product and customer needs.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Requirements

  • Expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, storage, security, and scaling.
  • Strong network engineering expertise is essential, particularly across VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Strong Linux systems administration skills, particularly with Debian and Ubuntu.
  • Solid understanding of networking fundamentals and the ability to design and operate complex network architectures.
  • Proven experience developing infrastructure automation using Ansible, Bash and/or Python, Git-based workflows, and GitOps methodologies.
  • Practical experience with observability platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Experience working with virtualization technologies including OpenStack, Proxmox, and VMware.
  • Experience with bare-metal provisioning and MAAS (Metal as a Service).
  • Strong understanding of distributed systems and container orchestration.
  • A process-oriented mindset, with the ability to create SOPs and operational procedures from the ground up.
  • Experience managing production incidents, escalation processes, and on-call rotations.
  • Ability to work independently and make sound technical decisions in a fast-paced, engineering-driven environment.
  • Strong communication and collaboration skills, combined with a high level of technical ownership and alignment with team values.
  • Fluent English is mandatory.
  • Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
  • Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations is advantageous.
  • Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls, and network policies is beneficial.
  • Familiarity with IT asset management or license tracking workflows is an advantage.
  • Experience working across multiple time zones and establishing SRE or reliability frameworks within growing organizations is highly valued.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Benefits

  • 100% remote work within the EU time zone, with CET ±2 hours preferred.
  • Flexible working hours designed to support autonomy and effective collaboration.
  • High-impact position with significant ownership and the opportunity to influence infrastructure strategy.
  • Opportunity to work with a modern technology stack spanning Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability.
  • Collaborative and international engineering environment with exposure to complex infrastructure challenges.
  • Significant autonomy to shape operational processes, reliability practices, and technical solutions.
  • Opportunity to contribute to ambitious cloud infrastructure initiatives with a strong focus on reliability, automation, and continuous improvement.

How Jobgether Works

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Location

United Kingdom

Sign up to applySee more jobs like this