Rodeo
Get started

IAM Cloud

Senior Site Reliability Engineer (Observability)

Kirklees
£80k – £95k/yr
Posted about 13 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Introduction

IAM Cloud builds software that simplifies IT management in the cloud. We are a fully employee owned, bootstrapped company - which means we answer to our customers and our team, not investors. With around 35 employees, we serve more than 3,000 organisations worldwide.

Being employee owned shapes how we work. Decisions are long-term, the culture is collaborative, and the people who build the company share in its success. We have leaned deliberately into AI and agentic workflows across the business as a core part of how a focused team competes with companies ten times our size. We are fully remote with employees in the UK, Ireland, Spain, Germany and Argentina and have been remote for years, as a deliberate choice about how good work gets done.

About the role

IAM Cloud is an Azure-first SaaS company with a small, fully remote engineering team. We run our own platform end to end: infrastructure, pipelines, releases and production operations all sit with the same handful of engineers. That means ownership is broad, documentation matters, and reducing toil is treated as a first-class goal rather than a side project.

We recently added a Senior Cloud Platform Engineer to own our CI/CD and infrastructure-as-code foundations. This role is the reliability counterpart: the person who owns how we see, measure and respond to what production is doing. You will own monitoring, observability, alerting and on-call at IAM Cloud, as the person who sets the standards, writes the policy and makes the engineering team better at running its own services.

Today our telemetry lives in Azure Monitor, Log Analytics and Application Insights, with Azure Managed Grafana for dashboards and alerting. It works, but it has grown organically. Alert thresholds are inherited rather than designed, retention and cost are not governed, instrumentation is inconsistent across services, and our on-call tooling is reaching end of life and must be replaced on a fixed deadline in the first half of 2027. We want someone who can come in, take a clear-eyed look at all of that, and turn it into something deliberate.

Where we are heading

Over the medium term we intend to move more of our workloads onto Kubernetes (AKS), with Prometheus alongside Azure Monitor and OpenTelemetry as the instrumentation standard. That is a direction rather than a dated plan, and it is not the focus of this role in its first year. What matters now is that whoever owns observability here has a view on what good looks like on that platform, so that the standards you set today carry over cleanly when we get there.

What you will own (outcomes)

  • Define and evolve the observability strategy, target architecture and standards for the platform - what we instrument, how, where it lands, how long we keep it and what it costs.
  • Make Grafana the place engineers go to understand production: dashboards and alerting designed around services and customer journeys, not around whichever metrics happened to be available.
  • Establish service level objectives for our customer-facing services, with error budgets that engineering and product actually use to make decisions.
  • Rationalise alerting so that every page is actionable, owned and has a runbook - and so that the on-call engineer's phone is quiet when nothing is wrong.
  • Lead and champion the move to a modern incident response and on-call platform (incident.io or similar) integrated with Microsoft Teams, and design the on-call model around it: rota, escalation, and the post-incident review process that feeds back into the backlog.
  • Champion observability-first thinking across the engineering team, so developers instrument their own services well without needing you in the loop.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

First-quarter deliverables

  • A written telemetry governance policy covering Log Analytics table plans, retention tiers, sampling and cardinality controls - with a baseline of current ingestion and spend and a target.
  • An alert rationalisation pass across Azure Monitor and Grafana: every remaining alert mapped to a service, an owner, a severity and a runbook; noisy or unowned alerts removed or downgraded.
  • Selection, delivery and cutover of a new incident response and on-call platform (incident.io or similar, Teams-integrated) to replace our current tooling before its end-of-life date, including parallel running and a documented on-call model.
  • SLOs and SLIs defined for the top customer-facing services, with dashboards and burn-rate alerting.

Looking further out: as our Kubernetes and Prometheus direction firms up, you will own the observability standards for it - collector topology, log routing, dashboards and alerting - so that it lands well-instrumented from day one.

Day to day

  • Design and implement OpenTelemetry instrumentation across our.NET services, and the collector pipelines that shape, sample and route telemetry.
  • Own Azure Monitor, Log Analytics and Application Insights configuration, managed as code alongside the rest of our estate (Bicep first).
  • Own Grafana end to end: data sources, dashboards-as-code, alert rules and the conventions the team follows when adding to them.
  • Run the incident response and on-call platform once it is in place: routing, escalation policies, integrations with Azure Monitor and Grafana alerting, and the Teams workflow around an incident.
  • Contribute observability requirements and standards to our Kubernetes and Prometheus direction as it takes shape, alongside the platform engineer who owns that work.
  • Build and maintain Grafana dashboards and alert rules that answer real operational questions, and retire the ones that don't.
  • Lead incident response when it matters, run blameless post-incident reviews, and turn findings into engineering work.
  • Track and reduce reliability metrics that matter: MTTR, alert volume per on-call shift, incident recurrence, telemetry cost per service.
  • Evaluate tooling and AI-assisted operations capabilities on their merits, and give the team a clear recommendation rather than a shortlist.
  • Write things down: standards, runbooks, decision records, and the "why" behind thresholds.
  • Take part in a light-touch on-call rotation with the rest of the engineering team.

What we're looking for

What we're looking for (must-haves)

  • 6+ years in engineering, with 3+ in a reliability, observability or production-operations role where you owned outcomes rather than executed tickets.
  • You have defined SLOs, SLIs and error budgets for real services and can talk through how they changed decisions.
  • You have reduced alert noise - you can describe an alerting estate you inherited, what you cut, what you kept, and how you knew it was safe.
  • Hands-on OpenTelemetry: instrumenting services, running collectors, and making deliberate choices about sampling and cardinality.
  • Telemetry cost and retention governance: you understand that observability has a bill, and you have managed it.
  • Deep Azure Monitor / Log Analytics / Application Insights experience, including KQL, and strong Grafana skills - dashboards, alerting and managing both as code.
  • Infrastructure as code - Bicep is our standard; strong Terraform experience is fully transferable.
  • Clear written and spoken communication at C1 level or above, or native-level business English. You will be setting policy for a remote team; if it isn't written down, it doesn't exist.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Nice to have

  • Experience selecting, implementing or migrating incident-management and on-call platforms - strongly desirable given the first-quarter deliverables.
  • Prometheus-based monitoring and alerting, including exporters, recording rules and cardinality management.
  • Kubernetes in production, ideally AKS, and the observability patterns that go with it.
  • A software development background, ideally .NET, so instrumentation conversations with developers are peer to peer.
  • Familiarity with AI-assisted operations and coding tools, and a considered view on where they help and where they don't.
  • Azure certifications (AZ-104, AZ-400, AZ-305) or the CKA.
  • Experience in a small or scale-up environment where you were the observability function.

Why us?

How we work

  • Fully remote across the UK. We meet in person a few times a year for team events, but day-to-day work is remote.
  • Small, senior team. You will work directly with engineering leadership and have real input into technical decisions.
  • Async-friendly. We use writing as our default mode of communication and avoid meetings when a written update will do.
  • On-call is shared and light. We invest in making the platform boring rather than relying on heroics.

Our Offer

  • £80,000 - £95,000 salary (dependent on skills & experience)
  • Guaranteed £2,000 pay rise every year you're with us, separate from any merit or promotion increases.
  • Time off and flexibility
    • 26 days holiday plus public holidays, with the option to swap public holidays for days that are more meaningful to you.
    • Your birthday and work anniversary day off, every year.
    • Up to 4 weeks per year working from anywhere in the world.
    • Genuinely flexible, fully remote working - we trust you to manage your own time.
  • Health, family, and wellbeing – UK employees
    • Generous Becoming a Parent leave, designed to support every kind of family.
    • Medical scheme including remote GP, dental, optical, and diagnostics.
    • Access to Support Room - confidential counselling, therapy, and coaching.
    • Help@Hand employee assistance programme.
    • 5x salary life insurance through Unum.
  • Money and growth
    • Up to 10% pension match (UK employees).
    • Annual learning and
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Location

7 Northumberland St, Huddersfield HD1 1RL, UK

Sign up to applySee more jobs like this