Omnicell
Sr. Manager, Site Reliability

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
About this opportunity
Omnicell is building a Global Cloud Operations organization from the ground up as our business shifts from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform that hospitals depend on 24/7. The Site Reliability Engineering function is the reliability engine of that organization, and this role is the first senior SRE hire — the person who will design the practice, set the standards, and then run the plays themselves until the team is large enough to delegate.
This is not a role where reliability practices already exist and you tune them. It is a role where you define what good looks like for Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud Operations and an Engineer III SRE you will coach and grow.
The environment is hybrid. Some of our products are still hardware in hospitals communicating with cloud services; others are fully SaaS. Some customers access us over private circuits, others over the public internet. We operate in a regulated environment — HIPAA, SOC 2, and in some engagements FedRAMP — which means reliability, security, and auditability are not separable concerns. The person we hire will be comfortable with that complexity and will help the organization design for it rather than around it.
This role also anchors Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner of how that gets introduced, prioritized against foundational reliability work, and validated in a regulated environment.
What you will own
Reliability practice (the coach half)
- Define and publish SLOs and SLIs for the top 5–10 Tier-1 customer-facing services, in partnership with Product and Engineering. Establish error budget policy and the enforcement mechanism when budgets burn.
- Design the incident command structure: severity rubric, declaration criteria, war-room protocol, stakeholder communication cadence, and the postmortem template. Train the first cohort of incident commanders across Engineering and Support.
- Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet.
- Partner with the VP to migrate the interim incident response RACI — currently held by matrixed individuals across IT, Engineering, Support, and Enterprise Security — into a durable SRE-owned model.
- Establish the on-call rotation model, including fair distribution, compensation approach, paging discipline, and the handoff protocol with our existing managed services partners (IBM, HCL) who provide L1/L2 coverage.
- Develop and track operational KPIs — MTTR, SLO attainment, change-failure rate, recurrence, cost per workload — and present reliability metrics and improvement roadmaps to senior leadership in the monthly Cloud Ops executive review.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Hands-on engineering (the player half)
- Instrument Tier-1 services yourself. Write the dashboards. Write the alerts. Write the runbooks. Do not wait for the team to grow before the work starts.
- Take the pager. Commander Sev-1 and Sev-2 incidents until a broader on-call rotation is staffed. Lead blameless postmortems and drive follow-up work to resolution.
- Contribute code and infrastructure-as-code (Terraform preferred; Chef/Puppet acceptable) to the platform. Oversee the design and evolution of CI/CD pipelines — our current stack includes CodeFresh, TeamCity, GitHub Actions, and Octopus Deploy, and we are consolidating over time.
- Administer and scale our Kubernetes platform, including secure and compliant cluster configurations. Working knowledge of Docker, Helm, and Service Mesh (Istio or Linkerd) expected.
- Run chaos and failover exercises (Chaos Monkey, LitmusChaos, or equivalent). Validate that what we think is resilient actually is.
AI-driven operations
- Architect Omnicell's AIOps direction: evaluate and introduce ML-based anomaly detection, predictive alerting, automated root cause analysis, and LLM-assisted runbook or triage pipelines.
- Make informed build-versus-buy calls across the AIOps landscape. Integrate AI-assisted tooling into the observability and incident response stack where it adds measurable value; resist the hype where it does not.
- Ensure AI-assisted operations meet the auditability and explainability bar required in a HIPAA and SOC 2 environment.
Coaching and team-building
- Coach one Engineer III SRE who joins shortly after you do. Pair on incidents. Review their design proposals. Help them grow toward senior. This is a formal, named relationship, not a side duty.
- Design the next 2–4 SRE hires. Write the requisitions, run the interview loops, make the calls. Your operating assumption is that the team grows under your direction over the next 12–18 months.
- Represent SRE in architecture reviews, product launch readiness reviews, and the monthly executive Cloud Ops metric review. Be the person in the room who knows what reliability costs and what it is worth.
- Partner with Enterprise Security, Compliance, and Architecture to ensure platform services meet regulatory and security requirements in a healthcare environment.
What success looks like in the first six months
Concrete outcomes this role will be evaluated against in the first half-year. These are drawn from the Cloud Ops 90-day plan and its extension into the following quarter.
- Month 1: SLOs drafted for the top 5 Tier-1 services with Product sign-off. Severity rubric published. First live tabletop Sev-1 run against the interim RACI.
- Month 2: Observability platform selection finalized. Instrumentation standard published. Engineer III SRE hired and onboarded.
- Month 3: On-call rotation live. First real Sev-1 commanded under the new structure with a blameless postmortem completed and follow-ups tracked.
- Month 4–6: Error budget policy in effect for the first 3 services. First incident review at executive level. Interview loop running for the next SRE hires. Initial AIOps evaluation and pilot scope defined.
Required knowledge and skills
- Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment — with demonstrated experience building a practice from zero or near-zero: you have set SLOs, defined incident command, and introduced error budget thinking to an organization that did not have it.
- Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.
- Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, GitHub Actions, Jenkins, TeamCity, or equivalent).
- Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.
- Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
- Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations. Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd).
- Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like.
- Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be applied in a regulated environment.
- Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice and documented follow-up discipline.
- Ability to coach and mentor, with direct evidence of growing junior and mid-level engineers. You are not a manager in this role, but you are a formal coach.
- Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.
- Excellent communication and stakeholder management skills; ability to translate complex technical concepts for non-technical audiences.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Basic requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience
- 8+ years of experience in software or platform engineering, with at least 4 of those in an SRE, DevOps, or platform reliability role.
- Proven Experience advising and influencing senior technical or operations leaders using data driven recommendations.
- At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.
Preferred knowledge and skills
- Masters Degree
- Prior experience in healthcare, clinical workflows, or another regulated vertical.
- Experience transitioning from MSP-heavy operations to internal-first, or integrating managed service providers (IBM, HCL, or similar) into an SRE operating model.
- Exposure to hybrid hardware-plus-cloud products, where device reliability and cloud reliability are jointly owned.
- Experience building or integrating AIOps platforms for automated incident triage and remediation.
- Familiarity with large language model APIs or agentic AI frameworks applied to on-call automation or runbook generation.
- Experience deploying and managing stateful distributed services in Kubernetes.
- Hands-on experience with security scanning and intrusion detection systems in regulated environments (HIPAA, SOC 2, or equivalent).
- Experience with messaging systems such as Kafka or RabbitMQ.
- Familiarity
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills