Rodeo
Get started

Carbon3.ai

Technical Lead - HPC

United Kingdom: (Hybrid - Visit to office / site locations required)
Posted 6 months ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Role Summary

This is a hands-on Technical Lead who can support the build, operation and scale Era4’s sovereign AI/HPC infrastructure. We need someone who can operate at the intersection of HPC platform engineering, GPU infrastructure, Linux systems, Kubernetes/Slurm, automation, observability and production incident response.

You will act as a technical authority, helping shape how GPU clusters are deployed, monitored, automated, supported and improved. You will work closely with SRE, platform engineering, infrastructure, vendors and customer-facing teams to ensure our platform is reliable, observable, scalable and ready for production workloads.

Key Responsibilities

HPC & GPU Platform Leadership

  • Lead infrastructure, supporting GPU, compute, storage and networking platforms.
  • Technical authority across platform engineering, SRE, infrastructure and customer-facing teams.
  • Mentor engineers and help define technical standards, best practices and operational excellence.

Platform Reliability & Incident Response

  • Lead technical investigations during major incidents and production outages.
  • Improve observability, monitoring and alerting across the platform.
  • Drive root-cause analysis and implement long-term reliability improvements.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Automation & Platform Engineering

  • Improve platform scalability, deployment processes and operational efficiency.
  • Oversee runbook creation, automation and development of inhouse Agent capabilities to support Operations
  • Contribute to the design, build, enhancement of future capabilities

Customer & Technical Engagement

  • Support customer onboarding, complex technical escalations and platform adoption.
  • Work with vendors, partners and internal teams to resolve infrastructure issues.
  • Translate complex technical challenges into clear, actionable communication.

AI/HPC Infrastructure Evolution

  • Contribute to Technical and Operational Roadmaps
  • Contribute to the deployment and optimisation of GPU clusters, Kubernetes environments and next-generation AI infrastructure.

Experience

  • You do not need to tick every technology box, but you must bring hands-on experience in production infrastructure and clear depth in HPC, GPU, AI infrastructure, research computing, Neocloud, cloud HPC or high-density compute environments.

Requirements

  • Linux engineering background
  • Experience in production of leading teams supporting HPC, GPU infrastructure, AI infrastructure, research computing, cloud HPC, Neocloud, platform engineering or SRE.
  • Experience operating or supporting production infrastructure across compute, networking, storage and observability.
  • Hands-on experience with at least one workload or orchestration layer such as Slurm, Kubernetes, Run:ai, LSF, PBS or equivalent.
  • Experience with Open Source monitoring and troubleshooting using tools such as Prometheus, Grafana, OpenTelemetry, Loki or equivalent.
  • Proven involvement in major incidents, on-call, escalation, root-cause analysis or production troubleshooting.
  • Ability to mentor engineers, influence technical direction and lead by technical credibility.
  • Comfortable working with internal teams, customers, suppliers and vendor engineering teams.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Why Join Era4

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

Diversity & Inclusion

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

HPC Platform Engineering
GPU Infrastructure
Linux Systems
Kubernetes
Slurm
Automation
Observability
Incident Response
SRE
Prometheus
Grafana
OpenTelemetry
Loki
Root-Cause Analysis
Technical Leadership
Platform Scalability

Location

United Kingdom

Sign up to applySee more jobs like this