Sundayy
ML Engineer, Infrastructure

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
About The Company
PhysicsX is a pioneering deep-tech company rooted in numerical physics and inspired by the high-performance environment of Formula One racing. Our mission is to accelerate hardware innovation by developing an AI-driven simulation software stack tailored for engineering and manufacturing sectors across advanced industries. By enabling high-fidelity, multi-physics simulation through AI inference throughout the entire engineering lifecycle, PhysicsX empowers engineers to achieve unprecedented levels of optimization and automation in design, manufacturing, and operational processes. Our innovative solutions serve a diverse client base, including leading organizations in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive industries. We are committed to pushing the boundaries of technology and fostering a collaborative environment where talented individuals can thrive and make meaningful impact.
About The Role
The Principal ML Infrastructure Engineer at PhysicsX will play a critical role in extending and maintaining the infrastructure that supports our research model training, fine-tuning, and deployment pipelines. Embedded within the Research team, you will collaborate closely with research scientists, ML engineers, and data engineers to ensure the efficient and reliable training of large physics models at scale. Your responsibilities will include designing distributed training systems, optimizing data pipelines, building model serving infrastructure, and enhancing developer tooling to accelerate research and deployment cycles. This role offers an exciting opportunity to influence the core technical foundation of our AI-driven simulation platform, contributing to cutting-edge advancements in physics-based modeling and simulation. The ideal candidate will possess a strong background in ML infrastructure, systems engineering, and cloud computing, with a passion for solving complex technical challenges in a fast-paced, innovative environment.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Qualifications
- 5+ years of experience building and operating ML infrastructure at scale
- Deep expertise in distributed training techniques, including NCCL, FSDP, DDP, and pipeline parallelism
- Strong systems fundamentals: Linux, networking (NVLink, InfiniBand), storage I/O, profiling, and performance optimization
- Production experience with Kubernetes and SLURM for GPU cluster orchestration
- Proficiency in Python and ML frameworks, with a preference for PyTorch
- Experience with cloud GPU infrastructure, ideally on platforms like CoreWeave or similar HPC-focused clouds
- Ability to scope, prioritize, and deliver complex projects effectively
- Excellent problem-solving, analytical, and troubleshooting skills
- Strong collaboration and communication skills, capable of bridging research and engineering teams
Responsibilities
- Design, implement, and operate distributed training infrastructure for neural operator architectures on large-scale GPU platforms
- Optimize training pipelines for throughput, fault tolerance, and cost efficiency, including checkpointing and multi-node synchronization strategies
- Develop and maintain experiment tracking, observability, and monitoring systems for training and hyperparameter tuning
- Address data loading bottlenecks by optimizing data pipelines for large mesh datasets, ensuring efficient I/O from cloud storage
- Work with heterogeneous data sources, formats, and resolutions to streamline data ingestion and processing
- Build and manage model serving infrastructure supporting zero-shot inference, uncertainty quantification, and deployment in customer environments
- Design robust model packaging pipelines, ensuring reproducibility and reliable deployment with fine-tuning capabilities
- Enhance developer experience through reliable CI/CD, debugging tools, and fast iteration cycles
- Collaborate with the broader infrastructure team to establish shared standards and best practices across the organization


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Benefits
- Opportunity to shape an AI-native engineering company at a formative stage, working on impactful projects
- Collaborative environment with high-caliber engineers, scientists, and operators
- Flat organizational structure encouraging innovation and idea-sharing
- Sustainable work pace with hybrid working model combining office and remote work
- Equity options to share in the company’s growth and success
- 10% employer pension contribution
- Free office lunches to support energy and focus
- Enhanced parental leave: 3 months full pay paternity and 6 months full pay maternity leave
- YellowNest nursery scheme to assist with childcare costs
- 25 days of annual leave plus public holidays
- Private medical insurance fully covered for employees
- Wellhub subscription for wellness and fitness activities
- Regular eye tests and health checks
- Support for personal development and continuous learning
- Employee Assistance Programme for confidential wellbeing support
- Bike-to-Work scheme and Season ticket loans for sustainable commuting
- Octopus EV salary sacrifice scheme for electric vehicle ownership
Equal Opportunity
PhysicsX is an equal opportunity employer committed to fostering a diverse and inclusive workplace. We value and encourage applications from individuals of all backgrounds, regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation, or gender identity. We are dedicated to creating an environment where everyone can thrive and contribute to our shared success. Additionally, we actively support initiatives to increase diversity in tech, including sponsorship programs for underrepresented groups in STEM fields. All application data is handled confidentially and used solely to monitor our diversity and inclusion efforts in compliance with applicable legislation.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location