Rodeo
Get started

Q1 Technologies, Inc.

Site Reliability Engineer (Hyper V)

London
Posted about 15 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud

London UK (Onsite)
Contract role

Overview

We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.

The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.

Key Responsibilities

  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Required Technical Skills

Site Reliability Engineering

  • Strong understanding of Site Reliability Engineering principles and operational excellence.
  • Experience with infrastructure reliability, service availability, resiliency, and performance optimization.
  • Storage Space Direct and failover clustering technical expertise. (Storage Spaces Direct enables you to build highly available, software-defined storage by pooling local disks (SSDs, NVMe drives, and HDDs) across multiple Windows Server nodes in a cluster. Instead of relying on an external SAN, S2D uses the servers' local storage to create a resilient shared storage pool)
  • Experience managing production-critical infrastructure environments with high availability requirements.
  • Experience with incident management, problem management, RCA, and continuous operational improvement.
  • Knowledge of monitoring, observability, alerting, and performance management.

Microsoft Hyper-V (Core Expertise)

  • Deep hands-on expertise in Microsoft Hyper-V architecture, deployment, administration, troubleshooting, and optimization.
  • Extensive experience in operating enterprise private cloud environments on Hyper-V.
  • Strong experience supporting enterprise-scale VDI deployments on Hyper-V.
  • Hyper-V Failover Clustering and high-availability architecture.
  • Storage integration including SAN, NAS, Storage Spaces Direct (S2D), Cluster Shared Volumes (CSV), and storage optimization.
  • Networking within Hyper-V environments including virtual switches, VLANs, NIC Teaming, QoS, and network performance tuning.
  • System Center Virtual Machine Manager (SCVMM).

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Automation & Platform Engineering

  • Strong PowerShell scripting and automation experience.
  • Experience automating infrastructure deployment, operational tasks, health checks, and reporting.
  • Familiarity with Infrastructure as Code concepts and configuration management.
  • Experience developing reusable operational tooling to improve reliability and reduce manual effort.

Preferred Skills

  • Windows Server 2016/2019/2022 administration.
  • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
  • Exposure to hybrid cloud and private cloud platforms.
  • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
  • Experience supporting enterprise VDI environments.
  • Understanding of ITIL Incident, Problem, Change, and Release Management.
  • Experience working in regulated industries such as Banking or Financial Services.

Experience & Qualifications

  • 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.

Success Measures

  • Improved platform availability and reliability.
  • Reduced infrastructure incidents through automation and proactive monitoring.
  • Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Increased infrastructure automation and operational efficiency.
  • Consistent achievement of service reliability objectives and operational KPIs.
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Microsoft Hyper-V
Site Reliability Engineering
PowerShell
Storage Spaces Direct
Failover Clustering
VDI
Infrastructure as Code
SCVMM
Incident Response
Capacity Planning
Disaster Recovery
Monitoring and Observability
Windows Server
Root Cause Analysis
Private Cloud
Virtualization

Location

London, England, United Kingdom

Sign up to applySee more jobs like this