Rodeo
Get started

Jobgether

Senior Inference Engineer

UK
Posted about 12 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Senior Inference Engineer

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Inference Engineer based in United Kingdom.

This is an opportunity to become the first dedicated engineer responsible for building and owning an inference platform from the ground up.

You’ll work closely with the CTO to turn large language models into reliable, production-grade query-to-response systems running at scale.

The role combines hands-on infrastructure engineering with performance optimization across GPU-based environments.

You’ll work with modern serving technologies such as vLLM, SGLang, and TensorRT-LLM while solving real-world challenges around latency, cost, and throughput.

As the platform matures, you’ll take increasing ownership of its technical direction and collaborate with Product on the future inference roadmap.

You’ll join a remote-first, open-source-oriented startup where engineering ownership, speed, and measurable customer impact are highly valued.

This role is ideal for a senior engineer who wants significant autonomy and the opportunity to define how production inference infrastructure is built and scaled.

Accountabilities

  • Build and deploy production-grade LLM inference systems across one or multiple GPU machines, owning the complete pipeline from customer query to served response.
  • Design, implement, and operate model-serving infrastructure using technologies such as vLLM, SGLang, and TensorRT-LLM.
  • Optimize inference workloads for scale, balancing latency, throughput, reliability, and infrastructure costs.
  • Apply techniques such as quantization, batching, caching, and intelligent request routing to improve inference performance.
  • Develop robust production infrastructure using Python or Golang, with an emphasis on maintainable, scalable engineering rather than configuration-only work.
  • Establish the initial inference platform in close collaboration with the CTO and take ownership of its evolution as the organization scales.
  • Define and execute the technical roadmap for inference infrastructure, identifying opportunities to improve performance, reliability, and developer or customer experience.
  • Partner with Product to translate evolving customer and market requirements into practical inference-platform capabilities.
  • Evaluate emerging inference technologies and approaches and determine where they can create meaningful improvements.
  • Collaborate with engineering stakeholders to establish reliable operational practices for production GPU infrastructure.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Requirements:

  • Significant professional experience building and operating production software or infrastructure systems, with strong hands-on engineering capabilities.
  • Demonstrated experience deploying and serving large language models in production, ideally using vLLM, SGLang, TensorRT-LLM, or comparable inference frameworks.
  • Practical expertise optimizing inference workloads through quantization, batching, caching, routing, or similar techniques.
  • Strong programming skills in Python or Golang, with a track record of writing and maintaining production-quality code.
  • Strong understanding of production inference architectures, including the journey from user request through model execution to a reliable served response.
  • Excellent problem-solving skills and the ability to independently investigate complex performance, scalability, and reliability challenges.
  • Strong communication skills, with the ability to explain sophisticated technical concepts clearly to engineers, product stakeholders, and other audiences.
  • Comfortable taking significant ownership, working with ambiguity, and making pragmatic technical decisions in a fast-moving environment.
  • Familiarity with Docker and Kubernetes is a plus.
  • Hands-on experience with generative AI technologies such as PyTorch or Transformers is advantageous.
  • Knowledge of the GPU software stack, including CUDA, NCCL, drivers, and related libraries, is beneficial.
  • Understanding of model architectures and fine-tuning techniques is a plus.
  • Experience with NVIDIA Dynamo is an additional advantage.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Benefits:

  • Competitive compensation package including equity.
  • Health, dental, vision, and life insurance, with coverage for eligible dependents where available.
  • Benefits adapted to the country of employment.
  • Flexible working schedule focused on outcomes rather than fixed working hours.
  • High degree of workplace flexibility, supporting remote work and changing personal circumstances.
  • Remote-first environment with a globally distributed team.
  • Significant ownership over the architecture, implementation, and long-term roadmap of the inference platform.
  • Direct collaboration with senior technical leadership and Product teams.
  • Opportunity to work on production-grade GPU and AI infrastructure at scale.
  • Exposure to modern LLM serving, inference optimization, Kubernetes, cloud infrastructure, and open-source technologies.
  • A culture built around ownership, action, continuous improvement, open-source collaboration, and technically ambitious engineering.
  • Opportunity to help define emerging standards for AI infrastructure and production inference.

How Jobgether Works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

LLM inference
vLLM
SGLang
TensorRT-LLM
Python
Golang
GPU infrastructure
Performance optimization
Quantization
Batching
Caching
Kubernetes
Docker
PyTorch
Transformers
CUDA

Location

United Kingdom

Sign up to applySee more jobs like this