Willis Re
Site Reliability Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Job Title: Site Reliability Engineer (SRE) - Cloud & Infrastructure
Overview
Willis Re is expanding its Global business in 2026 and the Cloud and Infrastructure team will need to support this growth by designing, provisioning, and supporting the platforms to enable this.
We are seeking an experienced Site Reliability Engineer (SRE) to join our Cloud & Infrastructure team. The successful candidate will be responsible for designing, operating, automating, and continuously improving enterprise-scale Azure platforms, ensuring high availability, resiliency, security, and performance.
The role combines software engineering, cloud architecture, infrastructure automation, and operational excellence to improve service reliability and reduce operational overhead through automation and engineering best practices.
Key Responsibilities
Platform Reliability & Operations
- Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
- Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Proactively monitor platform health and performance using observability tooling.
- Perform root cause analysis and implement permanent fixes for recurring incidents.
- Participate in incident management and on-call support rotations where required.
- Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
- Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
- Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.
Azure Cloud Engineering
- Design, deploy, and manage Azure infrastructure services including:
- Virtual Networks
- Application Gateways
- API Management
- Azure Kubernetes Service (AKS)
- Azure Firewall
- Azure Storage
- Key Vault
- Azure Monitor
- Azure AI Services
- Implement cloud platform standards and best practices.
- Support multi-region Azure deployments and platform modernisation initiatives.
- Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.
Infrastructure as Code (Terraform)
- Develop and maintain Terraform modules and reusable infrastructure patterns.
- Implement Infrastructure as Code (IaC) standards and governance controls.
- Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
- Manage Terraform state securely and consistently across environments.
Azure DevOps & Automation
- Build and maintain Azure DevOps CI/CD pipelines.
- Automate infrastructure provisioning and application deployments.
- Implement testing, security scanning, policy compliance, and release gates.
- Support GitOps and platform engineering practices.
- Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
Resilience, Disaster Recovery & Failover
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
- Design and implement highly available Azure architectures.
- Develop and maintain disaster recovery and business continuity capabilities.
- Implement and test:
- Regional failover strategies
- Active/Passive architectures
- Active/Active deployments
- Traffic Manager and Front Door failover patterns
- Database resiliency and replication
- Backup and recovery solutions
- Conduct regular resilience and recovery testing exercises.
- Identify and reduce single points of failure across platforms.
- Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.
Security & Governance
- Ensure platforms are secure-by-design.
- Work closely with Security and Architecture teams to implement:
- Zero Trust principles
- RBAC controls / Managed Identities
- Network segmentation
- Secrets management
- Support compliance requirements and operational audits.
- Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors.
- Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.
Continuous Improvement
- Drive automation and reduction of manual operational tasks.
- Improve deployment reliability and platform observability.
- Contribute to architecture standards, runbooks, and operational documentation.
- Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.
About You
At Willis Re we work in a fast-paced evolving environment with a growth mindset and outcome focus.
Core Skills & Experience
- Proven experience in senior positions
- Ability to articulate complex ideas and scenarios into business language
- Ability to manage and prioritise workload
- Worked in a fast-paced complex environment and demonstrate hands-on experience
- Microsoft Azure Expert
- Azure Administration and Architecture
- Networking and Connectivity
- Virtual Machines and Platform Services
- Azure Monitor and Log Analytics
- Azure Identity and Access Management
- Azure Networking, NSGs, Firewalls, Load Balancing
- Azure Backup and Disaster Recovery
- Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
- Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Azure DevOps)
- Exposure and understanding of building, deploying, and managing API Gateways
- Strong understanding of Azure security controls, governance, and compliance frameworks.
- Familiarity with monitoring, logging, and diagnostic tooling.
- Strong FinOps expertise
- Scripting skills (PowerShell, Bash, or Python).
- Strong understanding of DevOps practices, tooling, and SDLC methods
- Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
- Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.
- Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.
- Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.
- Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.
- Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.
- Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.
- Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.
- Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Nice to have
- Experience with Observability tools such as DataDog
- Understanding of Zero Trust Architecture
- Snowflake experience
- GitHub Enterprise
About Willis Re
We combine specialist broking with analytics, modeling, and research to help insurers optimize risk transfer, strengthen balance sheets, and achieve sustainable growth. Our approach is relationship-driven, transparent, and outcome-focused.
At the heart of Willis Re is a focus on delivering the most cutting-edge analytical solutions to enable more informed, better decision-making for risk selection, portfolio optimization, and capital management.
The launch of Willis Re brings a strategic advantage of being unhindered by legacy, an ability to leverage data, statistical models, and advanced technologies with the best knowledge and expertise to deliver more efficient and effective reinsurance outcomes. This places Willis Re in a unique position to build a truly analytically driven business, focused on creating solutions for the reinsurance industry that are future-led and forward-thinking.
Willis Re will also leverage recognized technical expertise from WTW’s Insurance Consulting & Technology business, including their advanced modeling and analytical capabilities. Alongside this will be WTW’s Research Network, an award-winning business supporting and influencing science to improve the understanding and quantification of risk.
Willis Re is committed to embracing a diverse, inclusive, and flexible work environment. We provide equal opportunity to all qualified individuals regardless of race, colour, religion, age, gender, gender expression, national origin, veteran status, disability, orientation, or any other legally protected categories. If you have a need that requires accommodation, please email us at talentacquisition@willisre.com.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location