Queen Square Recruitment
Artificial Intelligence Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Senior AI Evaluation Engineer (Contract)
London, UK (Hybrid) | Long-Term Contract | Inside IR35
About the Role
We are seeking an experienced Senior AI Evaluation Engineer to join a major UK retail transformation programme, helping deliver next-generation Generative AI solutions that enhance customer experiences and business operations.
Working as part of a multidisciplinary AI engineering team, you will define and implement robust evaluation frameworks for enterprise-scale LLM applications, Retrieval-Augmented Generation (RAG) platforms, and agentic AI systems. You'll play a key role in ensuring AI solutions are accurate, reliable, safe, and production-ready through automated testing, benchmarking, and continuous quality improvement.
This is an excellent opportunity to work on cutting-edge AI technologies including microservices, event-driven architectures, cloud-native platforms, and large-scale conversational AI systems within a fast-paced retail environment.
Key Responsibilities
- Define and implement end-to-end evaluation strategies for Generative AI applications including LLM-powered conversational systems, RAG solutions, and multi-agent AI workflows.
- Design evaluation frameworks measuring response relevance, faithfulness, groundedness, hallucination rate, retrieval accuracy, latency, consistency, and customer experience quality.
- Create and maintain benchmark datasets representing real-world retail customer journeys, support interactions, operational queries, and business scenarios.
- Evaluate and benchmark GPT models, open-source LLMs, prompt strategies, embedding models, and retrieval configurations to optimise production performance.
- Develop automated evaluation pipelines using Python and modern evaluation frameworks, enabling continuous regression testing across prompts, models, and RAG pipelines.
- Conduct prompt engineering experiments, A/B testing, prompt versioning, and response optimisation to improve accuracy and reduce hallucinations.
- Perform human-in-the-loop evaluation alongside business stakeholders to validate AI behaviour and identify opportunities for continuous improvement.
- Design automated scoring methodologies using LLM-as-a-Judge, rule-based validation, and qualitative evaluation techniques.
- Validate RAG systems through retrieval quality assessment, knowledge grounding verification, citation validation, chunking optimisation, and semantic search evaluation.
- Deliver comprehensive safety, risk, and compliance testing including prompt injection protection, jailbreak testing, PII detection, bias assessment, toxicity evaluation, and Responsible AI governance.
- Build dashboards and reporting to monitor AI quality metrics, evaluation trends, production performance, and model drift.
- Collaborate with AI engineers, platform teams, product owners, and business stakeholders to translate evaluation findings into actionable improvements.
- Integrate evaluation frameworks into CI/CD pipelines supporting continuous deployment of production AI services.
- Produce technical documentation covering evaluation methodologies, testing standards, governance processes, and quality assurance frameworks.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Essential Skills & Experience
- Strong commercial experience evaluating Large Language Models (LLMs) in production environments.
- Excellent understanding of Retrieval-Augmented Generation (RAG), vector search, embeddings, prompt engineering, and conversational AI.
- Experience designing enterprise AI evaluation frameworks and quality assessment methodologies.
- Strong Python development skills with experience building automated evaluation and experimentation pipelines.
- Experience with LangSmith, DeepEval, RAGAS, PromptTools, or similar LLM evaluation platforms.
- Experience benchmarking GPT models and open-source LLMs including Llama, Mistral, Gemini, Claude, or similar.
- Strong knowledge of LLM output evaluation, automated scoring, and NLP quality assessment techniques.
- Experience designing benchmark datasets and regression testing frameworks.
- Knowledge of Responsible AI principles, AI safety, governance, and compliance testing.
- Experience working with cloud platforms such as AWS or Azure.
- Familiarity with microservices, event-driven architectures, Docker, Kubernetes, and CI/CD pipelines.
- Strong analytical and problem-solving skills with the ability to translate model behaviour into measurable business improvements.
- Excellent communication skills and experience working with cross-functional engineering and business teams.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Technology Stack
Python • OpenAI • Azure OpenAI • LangChain • LangGraph • LangSmith • DeepEval • RAGAS • PromptTools • FastAPI • PostgreSQL • pgvector • Pinecone • FAISS • Docker • Kubernetes • AWS • Azure • GitHub Actions • MLflow • Pandas • NumPy • scikit-learn • CI/CD
What's on Offer
- Long-term contract opportunity
- Hybrid working in London
- Inside IR35 engagement
- Opportunity to work on enterprise-scale AI transformation within one of the UK's leading retail environments
- Exposure to the latest Generative AI, LLM evaluation, RAG, and cloud-native technologies
- Collaborative engineering culture focused on innovation, quality, and continuous improvement
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location