Rodeo
Get started

Medly AI

AI/ML Data Scientist

London
Posted about 20 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

About Medly AI

We're the fastest growing EdTech startup in London, on a mission to change education forever.

This is a moment where AI-native companies are reshaping entire industries, from Cursor in code, to Midjourney in image generation, to Harvey in law. Medly is building that company for education.

Since launch, we've reached over 400,000 students with retention rates that beat Duolingo, and we've delivered personalised learning with real results. We're backed by impact-focused VCs and supported by partners including UCL, Innovate UK, Microsoft, and Google.

The role

We're looking for an AI/ML Data Scientist to join our London team. You'll work directly alongside our founders and engineers who've previously built at Meta, Atlassian, and more, and you'll own how we measure and improve Medly's AI models.

This is the person who tells us the truth about our models. Hundreds of thousands of students rely on our AI to tutor them, mark their work, and generate content pitched at the right level for their exam board. Right now, judging whether a change made that better or worse is the hardest problem we have. You'll build and improve the benchmarks, evals, and experiments that answer it, and you'll be the reason we can ship quickly without breaking what works.

You'll have real ownership from day one. This is not a role where you'll be handed a spec.

What you'll do

  • Own our AI benchmarking: help improve and design subject-specific evals that measure tutoring, marking, and content generation quality against what a good teacher would actually say
  • Improve the eval harnesses and offline test suites that let us ship model and prompt changes with confidence, and catch regressions before students do
  • Run experiments and A/B tests on live AI features, and make the call on what the results actually mean
  • Build and curate high-quality datasets for fine-tuning and evaluation, including designing labelling schemes and rubrics that other people can apply consistently
  • Investigate model failures end to end: find where the AI falls down on notation, working, partial credit, or a specific exam board, and turn that into a fix
  • Work with our learning developers and teachers to turn pedagogical judgement into something measurable
  • Get hands-on with our product data: spot patterns, surface insights, and shape what we build next
  • Contribute to the wider codebase where it makes sense (Python, Postgres, AWS, some Next.js)

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

What we're looking for

  • At least 2 years in a data science, ML, or applied research role, post-graduation
  • Hands-on experience evaluating or benchmarking AI/ML systems: you've built evals, designed metrics, or run structured model comparisons, and you can talk about where your own metrics misled you
  • Strong Python for data work (pandas, numpy, notebooks, and the ability to write code other people can run)
  • Confident with SQL and relational databases
  • Solid grounding in statistics and experiment design: sampling, significance, confidence intervals, and knowing when a result doesn't mean what it looks like
  • Working knowledge of LLMs and how they're trained, evaluated, prompted, and fine-tuned
  • A degree in Computer Science, Maths, Physics, Statistics, Engineering, Data Science, or a similarly quantitative field
  • Genuine curiosity about how AI systems behave in the wild, not just on a leaderboard

Bonus points

  • Experience with eval frameworks and LLM-as-judge setups, including their failure modes
  • Fine-tuning, RLHF, DPO, or preference data collection
  • Built annotation or human-evaluation pipelines, and managed the people doing the labelling
  • Worked on AI products in production where quality was subjective and hard to measure
  • Any teaching, tutoring, or exam-marking experience, or familiarity with the UK curriculum (GCSE, A-Level, AQA/Edexcel/OCR)
  • Published research, open-source contributions, or writing about evaluation
  • Cloud infrastructure experience (AWS)
  • Something you've built in your own time that you'd love to show us

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Right to work

You must have the right to work in the UK. We're not able to offer visa sponsorship for this role.

Who thrives here

  • You care about quality and take pride in work that holds up under scrutiny
  • You're comfortable being the person who says the model got worse, with the evidence to back it
  • Curious about how things work, and willing to dig into other people's code and data to find out
  • You take initiative. If something looks off, you flag it or fix it rather than waiting to be told
  • You can explain a technical result to someone who isn't technical, and be trusted on it
  • Comfortable in a startup environment. Things move very quickly and priorities shift often

What you'll get

  • Real ownership of how Medly measures AI quality, on a product used by hundreds of thousands of students
  • Direct work with founders and senior engineers with experience at Meta, Atlassian, and beyond
  • Hybrid setup: 4 days in our central London office, one day work from home
  • A team that takes your ideas seriously and ships them
  • Room to grow into ownership of Medly's wider AI and data function
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

AI Benchmarking
LLM Evaluation
Python
SQL
Statistics
Experiment Design
Fine-tuning
A/B Testing
Data Analysis
Prompt Engineering
Dataset Curation
Postgres
AWS
Next.js
Pandas
Numpy

Location

London, England, United Kingdom

Sign up to applySee more jobs like this