Intergral
AI Evaluation & Benchmarking Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
AI Evaluation & Benchmarking Engineer
Location: Remote (United Kingdom)
Salary: £60,000–£75,000 DOE
Contract: Full-time, 40 hours per week
Employer: Intergral UK
Reports to: Director of Engineering
You must already have the right to work in the UK, as we're unable to sponsor visas for this role.
About the Role
We're looking for an AI Evaluation & Benchmarking Engineer to determine how we measure the quality of OpsPilot, and to build the systems that do it.
More important than any individual technology is how you approach measurement. We're looking for someone who questions whether a system is actually achieving the outcome it was designed for, works out how to measure that objectively, and builds what's needed to keep measuring it as the product changes.
You'll have real autonomy over the technical approach. We set the goals and check in regularly, but the expertise on how to get there is yours — we're hiring you because we need someone who can take an ambiguous problem and deliver a working system without being handed the steps. This is a hands-on engineering role with no reports and no QA silo. Our engineering structure is flat: you'll report to the Director of Engineering and work alongside the other engineers as a peer.
This isn't a traditional QA role. You won't be manually testing tickets, acting as a release gatekeeper or simply checking whether features technically work.
About OpsPilot
OpsPilot is an AI-led observability platform helping engineering and operations teams move from monitoring data to evidence-backed operational understanding — so they can investigate problems faster and act with greater confidence.
Intergral has more than 20 years of experience in application performance and observability, with an established customer base built around FusionReactor. OpsPilot is expanding beyond its historic Java and ColdFusion roots into the wider observability market, supporting OpenTelemetry-based metrics, logs and traces.
At the heart of OpsPilot is Coworker, an AI operations capability that continuously investigates telemetry, identifies situations that need attention and provides evidence-backed findings and recommended next steps.
What You'll Do
Evaluate the agent and the product it runs on
Coworker's findings are only as good as the system underneath them. An investigation can fail because the model reasoned badly, because retrieval surfaced the wrong evidence, because ingestion dropped a trace, or because the finding was presented in a way no engineer could act on. Evaluating the agent in isolation would tell us very little, so this role covers both.
On the agentic side, you'll:
- Build automated evaluations for our agentic AI capabilities.
- Create realistic synthetic scenarios, datasets and workloads with meaningful ground truth and evaluation criteria.
- Measure task success, diagnostic accuracy, evidence quality, reliability, consistency, latency and cost.
- Account for the non-deterministic nature of AI systems through repeated runs, variance analysis and determining whether changes are meaningful rather than noise.
- Benchmark models, prompts, tools, retrieval strategies and agent workflows against repeatable baselines.
- Use relevant industry benchmarks and standards, including established SRE practices and emerging AI-agent, AIOps and incident-response benchmarks, and build our own where existing approaches don't represent real operational work.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Across the wider product, you'll:
- Extend evaluation across important customer journeys, APIs, backend services and UI.
- Use OpenTelemetry, including its Semantic Conventions, to make benchmark environments representative and portable.
- Use product telemetry and correlate benchmark results with metrics, logs and traces to understand why failures occur.
- Build end-to-end measurements focused on customer outcomes rather than isolated components.
- When we change a model, prompt, tool or agent workflow, we want to know what became better, what became worse and why — including the impact on quality, reliability, latency and cost.
Turn what we learn into continuous improvement
- Turn failures and real-world problems into new evaluation scenarios.
- Identify recurring failure patterns and capability gaps.
- Test potential improvements against established baselines, holdouts and unseen scenarios.
- Detect regressions, benchmark overfitting and improvements that don't generalise.
- Longer term, this evaluation system becomes the harness for controlled self-improvement: identifying weaknesses, testing potential changes and objectively determining whether they should be retained. That's the direction of travel rather than the first year's work, but it's why we're building this properly.
Work with the rest of engineering
- Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product.
- Make evaluation failures clear, reproducible and actionable.
- Make straightforward fixes yourself where that's the most efficient approach.
- Build tooling that makes evaluations easy for other engineers to create, run and understand.
- Use AI-assisted engineering where it improves the speed or quality of your work.
- Benchmarking should provide continuous feedback that helps engineering improve the product. This role is not a release gatekeeper.
What We're Looking For
Above everything else: the ability to take an ambiguous technical problem, develop an approach and deliver a working system independently.
We're more interested in demonstrated ability than an exact number of years, but we'd generally expect around 3+ years of relevant technical experience. Your background might be as an AI engineer, software engineer, SRE, platform engineer, performance engineer, SDET or similar.
Alongside that, we're looking for:
- Practical experience working with LLMs, AI agents or AI evaluation.
- Strong software engineering skills, particularly Python or a similar language.
- Experience building automated evaluation, benchmarking, testing or experimentation infrastructure.
- Experience creating synthetic workloads, datasets or evaluation scenarios.
- An understanding of non-deterministic evaluation, including repeated measurement, variance and distinguishing meaningful changes from noise.
- The ability to turn complex system behaviour into measurable criteria.
- Comfort working across APIs, distributed systems and multiple layers of a software product.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Desirable
Any of the following would be useful, but none are required:
- Agentic AI evaluation, tool use and multi-step workflows.
- LLM evaluation frameworks and model-based evaluation techniques.
- Automated experimentation or self-improving systems.
- Dataset, ground-truth and holdout evaluation design.
- Statistical experimentation and performance benchmarking.
- OpenTelemetry, metrics, logs and distributed tracing.
- SRE, incident response or observability.
- Production SaaS and distributed systems.
What Success Looks Like
We have the beginnings of an evaluation harness, but the design and expertise are what we're hiring for.
We'd expect the first few months to go into the core evaluation harness and a starting corpus for Coworker's investigation quality, then extend outward across the rest of the product as that proves itself. How you sequence it is your call.
By six months, we should be able to objectively answer questions such as:
- Is Coworker getting better at investigating operational problems?
- Where does it perform well or poorly, and why?
- Are its conclusions supported by the right evidence?
- How do different models, tools and agent configurations compare?
- What quality, latency and cost trade-offs are we making?
- How do we compare against relevant external benchmarks?
- Have improvements introduced regressions elsewhere?
- Where should we focus improvement next?
Our evaluation corpus should keep growing as we encounter new problems.
Success isn't measured by the number of tests written or percentage test coverage. It's measured by our ability to understand how well OpsPilot is doing its job, where it isn't, and whether the changes we're making are actually making it better.
What We Offer
A small company rather than a large one, with the trade-offs that implies. Under ten people in engineering, a flat structure, and decisions made in a conversation rather than across three meetings. You'll have genuine influence over how this is done, and very little bureaucracy to work through to get there.
- Fully remote within the UK.
- Flexible working hours.
- 25 days holiday plus bank holidays.
- Real autonomy over your technical approach and how you deliver the role.
Our Interview Process
Straightforward: usually two or three conversations, with no technical coding tests.
If this sounds like the kind of challenge you're looking for, we'd love to hear from you.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location