Cosine
ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)
Location: London; full in-office working as default
Start date: ASAP
Compensation: £80,000 - £110,000 Base Salary & £80,000 - £110,000 Share options.
Help build the software engineers of the future
Cosine is building autonomous AI engineers that plan, write and ship code inside real development workflows. Our agents work across complex software systems, and our Lumen models are trained to do more than produce code that looks correct. They are built to understand existing architectures, follow established patterns and produce software that engineers can actually maintain.
We develop our agent tooling entirely in-house and post-train open-source models for reliable, enterprise-grade coding performance. Our products are designed for on-premise, VPC and fully air-gapped environments, including security-critical settings where control, privacy and robustness are non-negotiable.
In 2024, Cosine achieved a 72% score on OpenAI’s SWE-Lancer benchmark, placing us among the strongest real-world software-engineering AI systems evaluated.
We’re now looking for an ML Systems Engineer to help train the next generation of Lumen models.
This is a highly hands-on role at the intersection of machine learning, software engineering, data and infrastructure. You’ll build the environments in which models learn to write software, develop the pipelines that generate and curate training data, and run the fine-tuning and reinforcement-learning workloads that shape model behaviour.
If you’re excited by the idea that the future quality of coding agents will be determined not just by model architecture, but by the quality of their data, environments and reward functions, this is an opportunity to work directly on that problem.
The role
You’ll work closely with ML researchers, infrastructure engineers and product teams to decide:
- What Lumen should learn next.
- How to generate the right training data.
- How to design environments that reflect real software-engineering work.
- How to reward models for producing useful, maintainable code.
- How to measure whether a new training run genuinely improves the experience for engineers.
The systems you build will sit directly inside our model-training loop. Models will write code, use tools, run tests and interact with real repositories. Your work will determine how those interactions are generated, evaluated and fed back into future training.
This is not a narrow research role and it is not traditional MLOps. You’ll move between custom PyTorch code, distributed data pipelines, Dockerised services, RL environments, evaluation infrastructure and production-quality software.
What you’ll do
-
Build the training systems behind Lumen
- Contribute to the end-to-end training of software-engineering models.
- Implement supervised fine-tuning pipelines using curated code and conversation datasets.
- Build reinforcement-learning loops in which models write code, run tests and use development tools.
- Develop custom PyTorch dataloaders, training objectives and evaluation workflows.
- Run and analyse fine-tuning and RL experiments across large, modern open-source models.
-
Create the data that teaches models to engineer
- Develop synthetic data-generation pipelines for future RL and fine-tuning runs.
- Design systems for producing, filtering, transforming and sampling large-scale datasets.
- Work with object storage, dataset sharding and data-quality checks.
- Investigate which examples, tasks and sampling strategies lead to better model behaviour.
- Turn model failures into concrete improvements to training data and future experiments.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
-
Build reliable RL infrastructure
- Design, build and deploy containerised services that support model training and evaluation.
- Use Docker and orchestration platforms such as Kubernetes to operate RL infrastructure.
- Build environments where agents can modify repositories, run tests, use tools and receive meaningful feedback.
- Improve the reliability, reproducibility and observability of large-scale training runs.
- Work across Python, Go and the surrounding infrastructure required to run these systems.
-
Improve how we evaluate SWE models
- Help maintain and extend evaluation suites for code models.
- Build evaluations around unit tests, benchmark suites, repository-level tasks and real engineering workflows.
- Analyse model failure modes, reward-hacking risks and regressions.
- Develop better ways to measure code quality, maintainability, correctness and architectural fit.
- Feed evaluation results back into model, data and infrastructure decisions.
-
Shape the next training direction
- Work with research, infrastructure and product teams to identify the highest-value training problems.
- Turn broad goals such as “make Lumen better at X” into clear experiments and measurable outcomes.
- Develop opinionated reward functions for software-engineering agents.
- Document decisions, communicate trade-offs and help the wider team understand what the results mean.
- Take ownership of projects from initial idea through to deployment and iteration.
What we’re looking for
You may come from software engineering, ML infrastructure, data engineering, applied machine learning or a closely related field.
You should be comfortable with:
-
Strong software engineering fundamentals
- Typically 3–5 years of experience, or equivalent evidence of strong technical ability.
- Reading, debugging and writing non-trivial production code.
- Working primarily in Python and Go.
- Caring about correctness, maintainability and code quality as much as model metrics.
- Taking ownership of ambiguous technical problems and turning them into working systems.
-
Training frameworks and ML systems
- Practical experience with at least one of PyTorch, TensorFlow or JAX.
- Implementing custom training loops, losses or dataloaders.
- Understanding the relationship between data, objectives, evaluation and model behaviour.
- A willingness to work close to the code rather than treating training infrastructure as a black box.
-
Containers and cloud infrastructure
- Experience with Docker and container-management or orchestration platforms such as Kubernetes.
- Experience with at least one major cloud platform, such as GCP, AWS or Azure.
- An understanding of how to build services that are reliable, observable and straightforward to operate.
- Comfort working across application code, infrastructure and deployment systems.
-
Data engineering instincts
- Experience working with large-scale datasets and object storage.
- Understanding of sharding, filtering, sampling and dataset versioning.
- The judgement to recognise that data quality can matter as much as model architecture.
- An interest in building data pipelines that are reproducible and useful for experimentation.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
- Clear communication and ownership
- The ability to explain technical decisions and experimental results clearly.
- Comfort documenting trade-offs and walking others through your reasoning.
- A practical approach to deciding what to build, what to measure and what to leave out.
- The curiosity to investigate failures rather than dismissing them as noise.
You do not need to have worked on every part of the stack. We care more about strong engineering fundamentals, technical judgement and the ability to learn quickly than about matching every keyword.
Nice to have
You don’t need all of these, but experience in the following areas would help you get up to speed quickly:
- Synthetic data generation for code or language models.
- Training LLMs in distributed environments.
- Reinforcement learning, preference optimisation or reward modelling.
- Data tooling such as SQL, Apache Iceberg or DuckDB.
- LLM-as-a-judge systems or automated code evaluation.
- Reward-hacking detection, robustness evaluation or safety-focused training.
- Open-source contributions to LLM tooling, ML infrastructure or RL libraries.
- Experience working with repository-level code tasks or software-engineering benchmarks.
What success looks like
In your first few months, you will:
- Become a trusted engineering partner to the research team.
- Contribute to reliable data-generation, training and evaluation workflows.
- Ship improvements to the infrastructure used in Lumen fine-tuning and RL runs.
- Help identify why models fail on real software-engineering tasks.
- Turn those failures into concrete changes to data, environments, reward functions or evaluation.
- Build a strong understanding of the systems required to train capable SWE agents.
Longer term, you’ll help define how Cosine trains software-engineering models: what they learn, how they practise, what they are rewarded for and how we decide whether they are genuinely getting better.
Why this role matters
Coding agents are moving quickly, but producing plausible code is not the same as being a good software engineer.
The models we build need to work within existing codebases. They need to understand context, respect architectural decisions, write maintainable code, use tools effectively and recover when their first attempt fails.
That behaviour will not emerge from model scale alone. It will come from better environments, better data, better evaluations and better incentives.
You’ll work directly on all four.
Your contributions will influence the Lumen models used in Cosine’s self-serve and enterprise products, including deployments in organisations with demanding security and infrastructure requirements. You’ll have close proximity to the research, infrastructure and product decisions that determine where the system goes next.
This is a role for someone who wants to work on the full stack of modern model training, while staying grounded in the standards of production software engineering.
Why join Cosine
- Direct impact: Your work will shape the next generations of Lumen software-engineering models.
- Real scale: Work with large open-source models, long context lengths and multi-node training runs.
- Full-stack ML engineering: Move between PyTorch, distributed systems, data curation,
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location