Rodeo
Get started

Ocho

Machine Learning Engineer

Belfast
£80k – £110k/yr
Posted about 13 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Machine Learning Engineer (LLM & Agent Evaluation)

Mid to senior ML Engineering role at a fast-scaling, AI-first software business

Build scalable evaluation infrastructure and multi-agent architectures at the core of the product

Belfast based, hybrid with async-friendly global team

Salary: competitive, reflecting experience, with equity

UK work authorisation required

About the Company

Our client is a fast-scaling AI software business powering enterprise automation for Fortune 500 clients including major names across financial services and healthcare technology. Their small, elite Data Science and AI team builds and deploys cutting-edge ML and agentic AI systems at scale, with a culture built around intellectual curiosity, hands-on leadership and pragmatic startup thinking. Leaders stay close to the code, debate ideas openly and move fast without corporate inertia. This is a team where exceptional engineers thrive.

The Role

A newly created individual contributor position for an ML Engineer who wants to own evaluation end to end. You will design robust evaluation frameworks, build automated scoring and regression testing pipelines, and track quality across model, prompt and agent behaviour changes over time. A core part of this role involves building the infrastructure that converts expensive frontier agent tokens into optimised internal neural inference, a genuine and proprietary competitive advantage. Working closely with engineering teams and data scientists, you will analyse edge-case failure modes, build real-time quality dashboards and ensure high-confidence deployment workflows as the system scales.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Key Responsibilities

  • Design evaluation frameworks and metrics covering accuracy, safety, latency and cost across agent and LLM systems
  • Build automated scoring pipelines, rubric-based grading and LLM-as-judge systems that scale beyond manual review
  • Design and build chained multi-agent architectures that underpin the core product capability
  • Stand up automated regression suites that catch quality drops from model, prompt or agent-logic changes before they reach production
  • Bridge high-cost frontier models into low-cost, optimised internal inference systems
  • Build dashboards and reporting that track model and agent quality over time and across releases
  • Identify failure modes and edge cases across diverse scenarios, working with the team to prioritise fixes
  • Partner closely with engineers building agent capabilities and with the senior data scientist for deeper analytical support

What You'll Need

Essential:

  • Bachelor's degree in Computer Science, Machine Learning, Statistics or a related field, or equivalent practical experience
  • 3 or more years of experience in ML engineering, NLP or applied data science with hands-on exposure to LLM or agent-based systems
  • Deep practical experience building, chaining and evaluating autonomous agent workflows and frontier LLMs
  • Practical experience building or operating evaluation frameworks, automated scoring or benchmark systems for ML and LLM outputs
  • Strong Python skills and comfort building data pipelines for evaluation datasets
  • Solid understanding of NLP and modern LLM capabilities including prompting techniques, agentic workflows and retrieval
  • Working experience with Google Cloud Platform including Vertex AI and BigQuery, or equivalent AWS or Azure experience
  • Experience with inference optimisation and bridging frontier models into lower-cost internal systems

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Desirable:

  • Experience with LLM-as-judge techniques, rubric design or human-in-the-loop evaluation programmes
  • Familiarity with agent architectures and the specific failure modes of multi-step and agentic systems
  • Experience operating evaluation systems at scale in a production environment

Why Apply?

  • Competitive salary plus equity, with package tailored to UK, NI or European candidates
  • Work on genuinely novel evaluation and agent infrastructure at the competitive core of a scaling AI product
  • Hands-on, intellectually driven team that debates ideas openly and challenges assumptions constructively
  • Async-friendly culture with approximately three syncs per week, designed to protect deep focus and work-life balance
  • Fast-moving startup environment with full operational autonomy and no corporate inertia
  • Belfast based with a globally distributed, elite AI engineering team behind you

Interested?

For a confidential conversation about this opportunity, connect with Justin Donaldson on LinkedIn or submit your CV via the link below.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Machine Learning Engineering
LLM Evaluation
Agentic AI
Python
NLP
Google Cloud Platform
Vertex AI
BigQuery
Inference Optimisation
Multi-agent Architectures
Automated Scoring Pipelines
Regression Testing
Prompt Engineering
Retrieval Augmented Generation
Data Pipelines
Rubric Design

Location

Belfast, Northern Ireland, United Kingdom

Sign up to applySee more jobs like this