UK Health Security Agency
Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Job Summary
An SRE engineer will apply engineering principles to remediate infrastructure and operational problems. The primary focus will be on automation and CI/CD; ensuring our services run reliably, are scalable, and perform optimally in production environments. The role will monitor and manage these aspects while taking responsibility for multiple cloud infrastructure services. Observability of systems will be key to prioritising the operational service improvements and performance improvements to meet and exceed SLOs (Service Level Objectives).
Job Description
Working with the HPC & SRE Team to:
- Ensure services are stable, scalable, performant and automated
- Respond to incidents, troubleshooting issues, and restoring services as quickly as possible
- Prioritise operational service improvements to meet or increase SLO, minimising downtime
- Ensure that effective monitoring/alerting is in place to proactively identify issues using tools and dashboards. Reducing times to respond to issues.
- Leverage automation to streamline tasks, reduce overhead on repeatable operations, reduce manual intervention and improve efficiency. Write code that is maintainable, clear, and concise
- Optimise system performance using strong problem-solving skills to identify bottlenecks with an engineering mindset
- Ensure systems can handle current and future workloads through automation and capacity planning
- Continuously improve services through observability, and identify ways to improve observability practices
- Follow SRE principles. Guide and educate stakeholders to adopt implemented principles
- Provide technical documentation for engineers. Providing training, where appropriate
- Working closely with engineering and technology teams to improve operational processes, reduce manual tasks, ensure seamless collaboration/knowledge sharing, reduce risks and adapt to new ways of working.
Detailed Job Description And Main Responsibilities
We are seeking a highly motivated and experienced Site Reliability Engineer (SRE) to join our HPC & SRE engineering team. As an SRE, you will play a critical role in ensuring the stability, scalability, and performance of our services. You will combine software engineering and systems engineering to build, improve and run reliable, scalable production systems.
Key Responsibilities:
Service Reliability & Performance
- Ensure services are stable, scalable, and performant through engineering best practices and system design.
- Proactively identify and address system bottlenecks using advanced problem-solving and performance tuning techniques.
- Conduct capacity planning and implement solutions to ensure systems can support current and future workloads.
Incident Response & Troubleshooting
- Respond swiftly to production incidents, ensuring minimal downtime and quick restoration of services.
- Perform root cause analysis and postmortems, implementing lessons learned to prevent recurrence.
Monitoring, Alerting & Observability
- Contribute to the design and implementation of effective monitoring and alerting systems using tools and dashboards.
- Improve observability of services, ensuring issues are identified and addressed before impacting users.
- Continuously refine monitoring practices to reduce alert fatigue and improve response times.
Automation & Tooling
- Develop automation to eliminate manual, repetitive tasks and improve operational efficiency.
- Write clear, maintainable, and well-tested code to support automation efforts and system tooling.
- Drive initiatives to reduce operational toil and improve reliability through Infrastructure as Code (IaC).
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Service Level Objectives (SLOs) & Operational Improvements
- Contribute to the definition, tracking, and continuous improvement of SLOs, SLIs, and error budgets.
- Identify and prioritize operational improvements that align with business goals and user experience.
SRE Best Practices & Advocacy
- Helping to evangelize SRE principles across the organization.
- Collaborate with stakeholders to integrate reliability practices into the development lifecycle.
Collaboration & Knowledge Sharing
- Work closely with software engineering, DevOps, and infrastructure teams to streamline deployment and operational workflows.
- Improve cross-functional collaboration and promote a culture of shared responsibility for service reliability.
Documentation & Training
- Maintain accurate technical documentation, runbooks, and post-incident reports.
- Provide training and mentorship to engineering teams on best practices and tools.
This list is not exhaustive.
Person specification
Essential Criteria:
- Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar role
- Coding skills in programming/scripting languages such as Python, PowerShell or Bash
- Understanding of Linux/Unix & Windows systems, networking, and distributed systems
- Experience with observability tools (e.g., Prometheus, Grafana, Datadog) and alerting systems
- Understanding of infrastructure automation (e.g., Terraform, Ansible, PowerShell, Helm)
- Excellent communication and collaboration skills
- Possesses problem solving skills and the ability to respond to sudden unexpected demands
Desirable Criteria:
- Experience with CI/CD pipelines, cloud platforms (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes)
- Experience with post-incident reviews
- Previous involvement in driving adoption of SRE practices across an organization
- Experience delivering training or mentoring junior engineers
Salary Information
Alongside your salary of £41,983, UK Health Security Agency contributes £12,162 towards you being a member of the Civil Service Defined Benefit Pension scheme. Find out what benefits a Civil Service Pension provides (opens in a new window).
- Learning and development tailored to your role
- An environment with flexible working options
- A culture encouraging inclusion and diversity
- A Civil Service pension with an employer contribution of 28.97%
Diversity and Inclusion
The Civil Service is committed to attract, retain and invest in talent wherever it is found. To learn more please see the Civil Service People Plan (opens in a new window) and the Civil Service Diversity and Inclusion Strategy (opens in a new window).
Selection process details
This vacancy is using Success Profiles and will assess your Behaviours, Experience and Technical skills.
Stage 1: Application & Sift
You will be required to complete an application form. You will be assessed on the listed 7 essential criteria, and this will be in the form of a:
- Application form (‘Employer/ Activity history’ section on the application)
- 1000 word supporting statement.
You will receive a joint score for your application form and statement. (The application form is the kind of information you would put into your C.V –please be advised you will not be able to upload your CV. Please complete the application form in as much detail as possible). Please do not email us your CV.
Longlisting:
In the event of a large number of applications we will longlist into 3 piles of:


Get help with your application
Your very own career expert that helps elevate your application to the next level.
- Meets all essential criteria
- Meets some essential criteria
- Meets no essential criteria
If used, the pile(s) ‘Meets all essential criteria’ will proceed to shortlisting.
Shortlisting:
In the event of a large number of applications we may conduct an initial sift, on the lead criteria of:
- Experience as a Site Reliability Engineer, DevOps Engineer, Operations Engineer or similar role
Desirable criteria may be used in the event of a large number of applications/large amount of successful candidates.
If you are successful at this stage, you will progress to interview & assessment.
Feedback will not be provided at this stage.
Stage 2: Interview
You will be invited to a face to face interview at Canary Wharf, London.
If face to face interviews are planned, in exceptional circumstances, we may be able to offer a remote interview.
Behaviours and technical skills will be tested at interview through questioning and a presentation.
There Will Be a Presentation Required, The Subject Being:
- Automating a complex operational process.
The Behaviours Tested During The Interview Stage Will Be:
- Changing and improving. (Lead behaviour)
- Working together.
- Managing a quality service.
- Delivering at pace.
Interview dates are to be confirmed.
Location
This role is being offered as hybrid working based at one of our Core HQ's or Scientific Campuses.
We offer great flexible working opportunities at UKHSA and operate using a hybrid working model where business needs allow. This provides us with greater flexibility about how and where we work, to get the best from our workforce. As a hybrid worker, you will be expected to spend a minimum of 60% of your contractual working hours (approximately 3 days a week pro rata, (averaged over a month) working at one of UKHSA's core HQs (Birmingham, Leeds, Liverpool, and London), or scientific campus sites (Colindale, Porton or Chilton).
Our core HQ offices are modern and newly refurbished with excellent city centre transport links and benefit from co-location with other government departments such as the Department for Health and Social Care (DHSC).
Security Clearance Level Requirement
Successful candidates must pass a basic disclosure and barring security check before they can be appointed.
Successful candidates must meet the security requirements before they can be appointed. The level of security needed is Security Check.
For meaningful National Security Vetting checks to be carried out individuals need to have lived in the UK for a sufficient period of time. You should normally have been resident in the United Kingdom for the last 5 years as the role requires Security Check (SC) clearance. UK residency less than the outlined periods may not necessarily bar you from gaining national security vetting and applicants should contact the Vacancy Holder/Recruiting Manager listed in the advert for further advice.
Eligibility Criteria
External: This vacancy is open to all external applicants (anyone) from outside the Civil Service as well as internal applicants.
Qualifications And Registrations
For roles where specific qualifications or registrations are required, successful applicants will be asked to provide appropriate evidence. Employment cannot commence until satisfactory documentation has been received and verified.
Reserve List
If more than the required number of suitable candidates pass the interview criteria, you may be kept on a reserve list for 12 months subject to your agreement. You may be contacted, in merit
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location