Valarian
AI Engineer (Harness)

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Valarian Technologies is a dual-use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We're rethinking security beyond traditional military domains, addressing asymmetric threats that impact our technological advantage, economic strength, and democratic institutions.
We build Acra – the platform foundation for everything we do as a dual-use technology company. The platform’s name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure-agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include: increased operational resiliency for mission-critical systems and functions; enabling organisations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero-trust / least-privilege environments for M&A, attorney-client privileged communications, etc. And we’ve only scratched the surface.
At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you’ll be a part of collaborative, high-performing teams, creating cutting-edge software, platforms, and infrastructure.
The Role
As an AI Harness Engineer at Valarian, you will own the experimental design, evaluation methodologies, and benchmark infrastructure for our AI models and autonomous agentic workloads. In an emerging domain with no standard playbook, you will serve as the bridge between rigorous data science, statistical validation, and production agent scaffolding.
You will design robust evaluation datasets, architect calibrated LLM-as-a-judge pipelines, and build metrics that accurately capture multi-step agent performance under non-deterministic conditions. Your insights and error analyses will directly drive iterative improvements to our harness scaffolding, tool orchestration, and system safety guardrails.
What you’ll do:
- Architect Evaluation Runtimes & Harnesses: Design and maintain scalable Python evaluation harnesses that measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance across multi-step agentic systems.
- Design Experiments & Validate Results: Apply rigorous statistical methods—including hypothesis testing, confidence intervals, bootstrapping, and power analysis—to determine whether performance changes are true system improvements or stochastic noise.
- Build & Calibrate LLM-as-a-Judge Pipelines: Develop automated judging systems validated against human ground truth. Measure alignment using agreement metrics (Cohen’s kappa, correlation, precision/recall) and systematically detect judge failure modes such as position bias, verbosity bias, self-preference, and prompt sensitivity.
- Curate Benchmark Datasets & Rubrics: Define sampling strategies, detailed annotation guidelines, and scoring rubrics. Measure inter-annotator agreement and manage label noise to ensure benchmark integrity over time.
- Deep Error & Trajectory Analysis: Conduct hands-on failure analysis on agent runs, cluster root causes into actionable error categories, and clearly communicate findings to both engineering teams and stakeholders.
- Drive Harness Scaffolding Iteration: Translate evaluation results directly into architectural improvements across agent scaffolding, prompt design, tool execution loops, context window compaction, retries, and safety guardrails.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
What we are looking for:
- Grounding in Statistics & Experimental Design: Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic machine learning systems.
- LLM-as-a-Judge Validation: Proven experience building, auditing, and validating LLM judge setups against human annotations, including tracking agreement metrics and mitigating judge biases.
- Evaluation Dataset Engineering: Strong track record designing evaluation datasets, crafting precise scoring rubrics, managing label noise, and measuring inter-annotator agreement.
- Evaluation of Non-Deterministic Multi-Step Systems: Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs rather than relying solely on single-shot benchmarks.
- Metric Design & Integrity: Exceptional judgment in defining metrics tied to real-world outcomes, with a keen eye for detecting benchmark gaming, target drift, or shortcut learning.
- Rigor in Error Analysis: Ability to dissect complex execution logs, categorize failure modes, and synthesize clear, data-driven recommendations.
- Hands-On LLM Integration: Practical proficiency with LLM APIs, prompt engineering, structured outputs, function/tool calling, and retrieval systems in Python.
- Agent Architecture Awareness: Clear understanding of planning loops, tool orchestration, memory management, and multi-agent coordination, including their trade-offs and common failure modes.
- Harness & Scaffolding Contribution: Ability to write clean, maintainable Python code to iterate on harness components such as guardrails, retry logic, state handling, and context management.
Nice to have:
- Experience with agent and evaluation frameworks (e.g., LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, or Ragas).
- Exposure to running evaluation pipelines or agent workloads within secure, enclave, or Kubernetes environments.
Benefits
Our benefits are designed to ensure our employees feel taken care of and are proud to be a part of the Valarian team. We are committed to consistently enhancing our benefit package, taking into account the overall well-being and needs of our teammates. Here are the key benefits accessible to all employees at Valarian Technologies:


Get help with your application
Your very own career expert that helps elevate your application to the next level.
- Equity – because you have the right to own what you’re building.
- A competitive salary – because we value your unique skills.
- Employer pension contributions – because you deserve a secure future.
- Private health insurance - because your health is important to us.
- Hybrid work setup – because everyone has different needs.
- Rewarding company retreats and meetups that respect your work/life balance – because we love getting to know each other!
Life at Valarian
Our culture is built on inclusivity, compassion, and flexibility – we want everyone to be empowered to achieve their goals at Valarian.
The work we do is vital, but so are the connections that make it happen. We thrive on the shared energy, spontaneous conversations, and mutual trust built when we spend time together. We operate on a hybrid model, gathering in our London office 3 days a week to support one another and collaborate.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Valarian Technologies Limited is an equal opportunity employer and welcomes applications from individuals regardless of race, colour, religion, sex, sexual orientation, gender, identity or expression, national origin, age, disability, genetic information, marital status, veteran, amnesty, or any other legally protected characteristic.
We are committed to ensuring a fair and inclusive recruitment process and providing employment opportunities to all applicants. Decision recruitment, hiring, and employment are based solely on qualifications, skills, and experience relevant to the job requirements.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location