PRACYVA
Site Reliability Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Senior Site Reliability Engineer (GCP)
Manchester, Leeds or Edinburgh | Hybrid | Full-time Contract: Inside IR35
ROLE PURPOSE
Improve the reliability, scalability, security, and operational excellence of cloud-hosted services on Google Cloud Platform. This is a hands-on engineering role spanning SRE practices, production Kubernetes, infrastructure automation, observability, CI/CD, incident response, and continuous service improvement.
About the role
The Senior Site Reliability Engineer will work with Cloud Platform, Software Engineering, Product, Security, and Service teams to design and operate resilient services. The role will use software engineering and automation to reduce operational toil, improve service health, and turn incident learning into measurable reliability improvements.
Key responsibilities
- Lead reliability, availability, scalability, and performance improvements across GCP-hosted applications, platforms, and shared services.
- Define and operate service level indicators, service level objectives, and error-budget practices that connect technical health to customer impact.
- Design actionable observability using Dynatrace, including instrumentation, dashboards, distributed tracing, service health views, and SLO-based alerting.
- Build modular, reusable, and maintainable Terraform code for secure cloud infrastructure, platform services, and environment provisioning.
- Administer production Kubernetes environments, covering cluster lifecycle, workload deployment, capacity, upgrades, networking, and platform troubleshooting.
- Automate operational activities and repetitive support work using Python, Groovy, Bash, or PowerShell to reduce toil and improve consistency.
- Build and enhance CI/CD pipelines using Jenkins, Azure DevOps, GitHub Actions, or equivalent tooling, with automated quality and deployment controls.
- Lead incident response, complex troubleshooting, root-cause analysis, and post-incident reviews; ensure corrective actions are tracked to completion.
- Embed security, resilience, monitoring, and supportability into platform and application designs from the outset.
- Coach engineers, contribute reusable standards and patterns, and promote SRE and operational excellence across engineering communities.
Core outcomes
- Reliable and measurable production services with clear ownership, SLOs, and actionable alerts.
- Less manual effort through automation, self-service, and repeatable engineering patterns.
- Faster restoration, stronger incident learning, and sustained reduction in repeat failures.
Essential skills and experience
Capability
Required experience
- Google Cloud Platform
- Strong hands-on GCP engineering and operations experience across production services. Relevant GCP certification is preferred.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
-
Site Reliability Engineering
- Demonstrable SRE experience improving availability, resilience, scalability, supportability, and operational performance.
-
Dynatrace & observability
- Instrumentation, dashboards, metrics, logs, traces, service health modeling, and SLO-based alerting using Dynatrace.
-
Terraform
- Strong Infrastructure as Code capability, including modular design, reusable components, state management, code review, and maintainability.
-
Kubernetes & containers
- Production Kubernetes administration and Docker knowledge, including deployments, upgrades, networking, capacity, and troubleshooting.
-
CI/CD engineering
- Hands-on Jenkins, Azure DevOps, GitHub Actions, or comparable tooling for build, test, infrastructure, and deployment automation.
-
Scripting & coding
- Practical automation using Python, Groovy, Bash, or PowerShell, with sound software engineering and source-control practices.
-
Incident management
- Major incident response, structured troubleshooting, root-cause analysis, problem management, and post-incident improvement.
-
Cloud security
- Working knowledge of IAM, least privilege, secrets, secure configuration, policy controls, and operational risk management.
-
Networking & APIs
- Strong understanding of Linux/Unix, DNS, TCP/IP, routing, load balancing, firewalls, VPC connectivity, HTTP, and API troubleshooting.
-
DevOps & toil reduction
- Automation-first mindset and experience replacing repetitive operational activity with reliable engineering solutions.
-
Stakeholder collaboration
- Ability to explain service risk and technical options clearly and work effectively across engineering, product, security, and operations teams.
Technical environment
- Cloud: Google Cloud Platform, cloud-native services, and managed platform capabilities.
- Observability: Dynatrace, telemetry, dashboards, SLI/SLO measures, logs, metrics, and distributed traces.
- Platform: Kubernetes, Docker, Linux/Unix, cloud networking, and APIs.
- Automation: Terraform, Python, Groovy, Bash, and PowerShell.
- Delivery: Jenkins, Azure DevOps, GitHub Actions, Git, and automated release controls.
Operational and engineering expectations
Reliability and observability
- Create service health models that combine application, infrastructure, and customer-impact signals.
- Use SLOs and error budgets to guide prioritisation between reliability improvement and feature delivery.
- Tune alerts to be actionable, reduce noise, and provide clear runbooks and escalation paths.
- Perform production-readiness reviews covering capacity, dependencies, failure modes, recovery, and monitoring.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Platform automation and delivery
- Create reusable Terraform modules and automated environment patterns that are secure, reviewable, and repeatable.
- Integrate infrastructure and application changes into CI/CD with validation, approvals, auditability, and rollback options.
- Automate cluster operations, routine maintenance, evidence collection, and common remediation activities.
- Use code reviews, testing, and version control to maintain engineering quality across infrastructure and operational tooling.
Incident and problem management
- Coordinate technical diagnosis during major incidents and communicate impact, findings, and recovery actions clearly.
- Complete evidence-based root-cause analysis and convert learning into engineering backlog improvements.
- Identify recurring failure patterns and remove underlying causes rather than relying on repeated manual intervention.
- Improve runbooks, diagnostics, and recovery automation based on operational experience.
Desirable experience
- Professional GCP certification or equivalent evidence of advanced Google Cloud capability.
- OpenTelemetry, Google Cloud Monitoring, service mesh, or distributed tracing experience.
- GitOps, policy as code, secrets management, and automated compliance controls.
- Hybrid or multi-cloud environments and large-scale enterprise cloud networking.
- Financial services or another highly regulated, security-conscious environment.
- On-call operations, resilience testing, disaster recovery, and capacity engineering.
Candidate profile
- Hands-on engineer who can move confidently between code, cloud infrastructure, Kubernetes, and production diagnostics.
- Uses evidence and service data to prioritise improvements and make reliability trade-offs transparent.
- Communicates calmly during incidents and builds constructive relationships across technical and non-technical teams.
- Promotes learning, automation, and continuous improvement while maintaining strong security and control standards.
Job title matches
- Site Reliability Engineer (SRE)
- Cloud SRE
- Google Cloud SRE
- DevOps Engineer with Observability
- Cloud Platform Engineer
- Observability Engineer
Success measures
- Improved availability, service performance, and operational resilience.
- Reduced alert noise, incident recurrence, and manual operational toil.
- Consistent, secure infrastructure delivery through reusable automation.
- Clear service ownership, measurable SLOs, and effective operational readiness.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London