Rodeo
Get started

Oracle

Senior Site Reliability Engineering Manager

United Kingdom
Posted about 21 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Senior Manager, Site Reliability Engineering – OCI Core Infrastructure

Core Infrastructure Engineering within Oracle Cloud Infrastructure (OCI) is seeking an experienced Senior Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and operational excellence of critical database and storage services that form the backbone of OCI.

This role combines people leadership with strong technical leadership. You will initially lead a team of approximately five experienced SREs and will be responsible for their development, prioritization of work, technical direction, and successful delivery. You will work closely with senior engineers and development teams to make architectural and operational decisions, resolve complex production challenges, and continuously improve the reliability and scalability of OCI services.

The ideal candidate has a strong background in Site Reliability Engineering, cloud infrastructure, storage, or database systems and has progressed into an engineering management role while maintaining strong technical depth. You should be comfortable leading experienced engineers, challenging technical assumptions, reviewing designs, and helping teams make sound engineering decisions in complex distributed environments.

You will play a key role in defining how the team operates, prioritizing reliability initiatives, managing operational risk, and ensuring that engineering investments address the most important reliability and customer-impacting problems. While this is a management role rather than a primarily hands-on engineering position, you will be expected to remain technically engaged and able to dive deeply into architecture, incidents, reliability challenges, and engineering trade-offs when needed.

Responsibilities

  • Lead, coach, and develop a team of experienced Site Reliability Engineers responsible for critical OCI database and storage infrastructure.
  • Provide technical leadership and direction across reliability, scalability, availability, performance, and operational excellence.
  • Partner with senior SREs and development teams to define priorities, plan engineering work, and ensure effective execution.
  • Guide architectural and design discussions and help engineers evaluate technical trade-offs and make sound engineering decisions.
  • Maintain sufficient technical depth to understand complex distributed systems, challenge proposed solutions, and support engineers during difficult production and reliability issues.
  • Drive improvements in service health, telemetry, observability, capacity planning, incident management, and operational readiness.
  • Identify systemic reliability risks and recurring operational problems and ensure teams develop scalable solutions rather than relying on manual operational processes.
  • Partner with development and infrastructure teams to identify and resolve cross-functional operational risks.
  • Provide leadership during major production incidents, ensuring effective technical coordination, communication, root-cause analysis, and follow-up actions.
  • Establish clear priorities and balance reliability investments, operational work, technical debt, and strategic engineering initiatives.
  • Support release readiness and ensure appropriate engineering standards are applied before changes reach production.
  • Develop the team’s technical capabilities through coaching, mentoring, knowledge sharing, and effective delegation.
  • Build a strong engineering culture focused on ownership, reliability, continuous improvement, and customer impact.
  • Collaborate with engineering leaders and stakeholders across OCI on broader reliability initiatives and engineering standards.
  • Participate in the operational leadership model supporting 24×7 highly available services.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Qualifications

  • 10+ years of relevant engineering experience, with a strong background in Site Reliability Engineering, cloud infrastructure, distributed systems, storage, database, or related infrastructure technologies.
  • Proven experience managing and developing experienced engineers, ideally within SRE, infrastructure, platform, or cloud engineering teams.
  • Strong technical foundation in operating large-scale, highly distributed production systems on cloud platforms such as OCI, AWS, GCP, or Azure.
  • Demonstrated ability to lead senior engineers and provide credible technical guidance.
  • Strong understanding of reliability engineering practices, including availability, scalability, observability, incident management, capacity planning, performance management, and operational readiness.
  • Experience troubleshooting complex production issues involving distributed systems, storage, databases, networking, or cloud infrastructure.
  • Strong understanding of automation, infrastructure as code, CI/CD, configuration management, and modern cloud-native engineering practices.
  • Working knowledge of technologies and tools such as Linux, Kubernetes, Docker, Terraform, Python, Go, Git, Jenkins, Grafana, or equivalent technologies.
  • Experience operating business-critical, 24×7 high-availability production services.
  • Ability to evaluate technical designs, identify operational risks, and guide engineering teams toward scalable and resilient solutions.
  • Strong people leadership skills, including coaching, performance management, career development, prioritization, and team planning.
  • Excellent communication and stakeholder management skills, with the ability to work effectively with senior engineers, engineering managers, and cross-functional teams.
  • Strong organizational skills and the ability to operate independently in a complex, rapidly evolving engineering environment.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Preferred Profile

We are particularly interested in candidates who started their careers as SREs, infrastructure engineers, or engineers working on large-scale distributed systems and have subsequently moved into engineering management.

The successful candidate will combine the technical credibility required to lead experienced SREs with the people leadership skills needed to build and develop a high-performing team. They will be comfortable moving between people management, technical discussions, operational priorities, and broader engineering strategy.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Site Reliability Engineering
Cloud Infrastructure
Distributed Systems
People Management
Incident Management
Capacity Planning
Observability
Infrastructure as Code
Kubernetes
Docker
Terraform
Python
Go
Linux
CI/CD
Performance Management

Location

United Kingdom

Sign up to applySee more jobs like this