Rodeo
Get started

Cloud People

AI Quality Assurance Engineer

London
£55k – £65k/yr
Posted about 18 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

AI Quality Assurance Engineer

💰 £55,000 to £65,000

📍 London based, office roughly once a fortnight, with some travel to the UAE.

Company & role

This role sits with a global IT solutions provider standing up something genuinely different: an autonomous rapid prototyping pod for banking, insurance and fintech clients. A small, elite team that wins its own work, takes an ambiguous client problem, and turns it into a working AI prototype in four to five weeks, then ships it to production with the evidence to prove it's safe.

That last part is you. You will own the quality gate across every prototype and production release, making sure nothing ships without measured, documented evidence that it is defensible to client compliance and model risk teams. In short, you turn "it appears to work" into auditable quality. That means both solid full stack testing and, crucially, the harder problem of evaluating non deterministic AI output for accuracy, safety and reliability.

This is a dedicated QA seat, deliberately separate from development, in a team that believes developers should develop and quality should have its own owner.

Why This Role Stands Out

Quality here is not an afterthought bolted on at the end. Banks demand a full evidence log of what has been tested before anything goes live, and you are the person who produces it, which makes this role central to whether the pod ships at all.

You get to work at the frontier of AI quality, building evaluation frameworks, golden datasets and red team tests for LLM and agentic systems, a genuinely scarce and growing skill set. The pod is autonomous and protects your focus: no side of desk duties, no ticket churn from other teams, just your remit. Working pattern is mostly remote, office roughly once a fortnight.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

There is a growing UAE dimension to the business too. Nothing is expected, but if working in or relocating to the UAE would ever appeal, they will back you to do it.

Key Responsibilities

  • Own the quality gate across prototypes and production releases, ensuring nothing ships without documented, defensible evidence
  • Design test strategy and plans covering happy paths, edge cases, integration, performance and security
  • Build and maintain automated test suites across unit, integration and end to end, running in CI and CD to enable fast, confident releases
  • Validate AI products before client release using golden datasets, real world scenarios, regression suites and agreed thresholds for accuracy, relevance, safety, latency and cost
  • Evaluate non deterministic LLM and agent output using semantic evaluation, LLM as judge scoring, human review for high risk cases and repeat testing to detect instability and drift
  • Test RAG systems for retrieval quality, answer faithfulness, citation accuracy and fallback behaviour when knowledge is missing or ambiguous
  • Run adversarial and red team testing for prompt injection, jailbreaks, data leakage, insecure tool use and biased output before deployment
  • Own client release gates, blocking release where critical risks remain open, evaluation scores fall below threshold, or monitoring and rollback plans are incomplete
  • Manage defects end to end, from clear reproduction through triage, root cause analysis and driving improvements back into test plans and CI gates
  • Ensure production observability is in place for prompts, responses, traces, evaluation scores, latency, cost and drift

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Ideal Experience

Essential

  • 4 or more years in QA, testing or quality engineering
  • 2 or more years of test automation using Playwright, Cypress, Selenium or similar, writing maintainable test code
  • Full stack testing experience across user interface, API, database and integration scenarios, with a real understanding of how components interact
  • Agile testing experience within short sprints and continuous integration
  • A strong analytical mindset, comfortable with root cause analysis, metrics and designing effective test plans

Desirable

  • Financial services, banking, insurance or fintech experience, or other regulated environments
  • AI and machine learning testing, including evaluating agent and LLM behaviour, model evaluation and quality metrics for non deterministic systems
  • Experience with a modern AI testing stack such as Promptfoo, DeepEval, RAGAS, LangSmith, Braintrust or Langfuse
  • Performance and load testing with k6 or JMeter
  • Security testing knowledge including OWASP and penetration testing basics
  • Dataset management and evaluation framework experience
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Quality Assurance
Test Automation
Playwright
Cypress
Selenium
Full Stack Testing
AI Testing
LLM Evaluation
RAG Systems
Red Teaming
CI/CD
API Testing
Agile Testing
Root Cause Analysis
Prompt Engineering
Observability

Location

London, England, United Kingdom

Sign up to applySee more jobs like this