Rodeo
Get started

Fuse Energy

AI Inference Engineer

London
Posted 2 days ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Fuse Energy

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. Reporting directly to the CTO, you'll own the layer above kernels and hardware: how models actually get served, scaled and delivered against committed performance targets. Few companies can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of our offering.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Responsibilities

  • Define Fuse's inference serving strategy and architecture from first principles
  • Design and build the serving stack: request routing, batching, scheduling and autoscaling for high-throughput, latency-sensitive inference workloads
  • Own model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)
  • Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plans
  • Act as direct technical owner of inference performance and reliability
  • Work closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layer
  • Set the standards, tooling and benchmarks this function will run on as it grows

Requirements

  • 4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experience
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
  • Strong systems thinking, able to reason about the full path from incoming request to served response across a large cluster
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
  • A track record of making high-stakes architecture calls and owning the outcome
  • Comfort operating without a playbook: this is a founding role shaping a new function, not joining an established one
  • Bonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Benefits

  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office-based employees

As we hire globally, benefits vary by location.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Inference Serving
GPU Optimization
CUDA
vLLM
TensorRT-LLM
SGLang
Triton Inference Server
Quantization
Speculative Decoding
KV-Cache Management
Autoscaling
Systems Architecture
Kubernetes
Slurm
Distributed Systems
Capacity Planning

Location

London, England, United Kingdom

Sign up to applySee more jobs like this