Rodeo
Get started

Sundayy

Machine Learning Engineer

United Kingdom
£140k – £200k/yr
Posted about 16 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

About The Company

Inworld is a leading product-oriented research laboratory composed of top-tier AI researchers and engineers dedicated to advancing the frontiers of artificial intelligence. Our core focus is on developing best-in-class real-time multimodal models and the only real-time orchestration platform optimized for handling thousands of queries per second. With substantial backing, having raised over $125 million from prominent investors such as Lightspeed, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta, and Stanford, we have established ourselves as a pioneer in the AI industry. Our innovative technology has powered experiences for renowned companies including NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella, and Bible Chat. Recognized globally, we have been named one of CB Insights’ 100 Most Promising AI Companies and ranked among LinkedIn’s Top 10 Startups in the USA, underscoring our leadership and potential in the AI ecosystem.

About The Role

We are seeking a highly skilled and motivated AI Systems Engineer to join our dynamic team. In this role, you will be instrumental in designing, optimizing, and deploying high-performance AI inference systems at scale. Your work will directly impact the development of our multimodal models and real-time orchestration platform, ensuring they operate with minimal latency, maximum throughput, and exceptional reliability. You will collaborate closely with research teams to take cutting-edge models from concept to production, containerize and optimize them, and ensure seamless deployment across distributed systems. The ideal candidate thrives in an environment of ambiguity, demonstrates rapid learning, and possesses a passion for building scalable, high-performance AI infrastructure. Your expertise will help push the boundaries of what’s possible in real-time AI applications, contributing to the future of intelligent systems that are both powerful and efficient.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Qualifications

  • PhD in Computer Science, Physics, Mathematics, or an equivalent practical experience in building backend or ML systems
  • Deep understanding of modern serving frameworks and techniques such as vLLM or TRT-LLM
  • Hands-on experience with model acceleration methods including quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in programming languages such as C++, CUDA, Rust, or highly optimized Python
  • Experience with profiling code and optimizing performance on NVIDIA GPUs
  • Knowledge of distributed systems and scaling solutions, including Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference
  • Experience with handling thousands of concurrent connections reliably
  • Contributions to open-source inference engines or technical deep-dives in relevant areas
  • Full-cycle ownership capability from research model to containerization, optimization, and production deployment

Responsibilities

  • Design, develop, and optimize high-performance AI inference systems for real-time applications
  • Implement and improve model acceleration techniques to enhance inference speed and efficiency
  • Build and maintain distributed systems capable of scaling to thousands of concurrent queries
  • Containerize research models and ensure their reliable deployment in production environments
  • Collaborate with research teams to translate innovative models into scalable production solutions
  • Profile and optimize code to maximize GPU performance and resource utilization
  • Contribute to open-source projects and technical documentation to advance the field
  • Troubleshoot and resolve system bottlenecks, latency issues, and reliability challenges
  • Stay updated with the latest advancements in AI hardware and software optimization techniques

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Benefits

  • Competitive base salary within the range of £140,000 – £200,000, commensurate with experience and location
  • Equity options providing ownership in a fast-growing AI company
  • Comprehensive benefits package including health, dental, and vision insurance
  • Flexible working arrangements and supportive work environment
  • Opportunities for professional growth and continuous learning in cutting-edge AI research
  • Potential relocation support for candidates interested in moving to the San Francisco Bay Area in the future

Equal Opportunity

Inworld is committed to creating an inclusive environment for all employees. We are an equal opportunity employer and do not discriminate based on race, ethnicity, gender, sexual orientation, age, disability, or any other protected characteristic. We believe diversity enhances innovation and are dedicated to fostering a workplace where everyone can thrive and contribute to our mission of advancing AI technology.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

vLLM
TRT-LLM
Quantization
Distillation
C++
CUDA
Rust
Python
NVIDIA GPU Optimization
Kubernetes
Ray
Distributed Systems
Model Acceleration
Containerization
AI Inference
Multimodal Models

Location

United Kingdom

Sign up to applySee more jobs like this