Nominal
Staff Research Engineer, Agent Evals & Post-training

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
ABOUT NOMINAL
Our mission is to accelerate how the world engineers new hardware. Nominal's connected test and operations platform powers the world's most advanced hardware programs and its most ambitious startups, from spacecraft, racecars, and autonomous vehicles to next-generation defense and energy programs. Our customers include Anduril, Shield AI, Hermeus, Albedo, Shinkei, and Pratt Miller Motorsports, as well as U.S. Navy and U.S. Air Force programs. Now we're expanding across the entire hardware lifecycle, building the foundation, AI-native applications, and agents that accelerate innovators' work and change what's possible to build.
We're backed by Sequoia, General Catalyst, Founders Fund, Lux Capital, and Lightspeed, and our team comes from SpaceX, Apple, Palantir, Anduril, Applied Intuition, and other leading companies.
ABOUT HARDWARE INTELLIGENCE
We're the team behind Nominal's agents, AI-native applications, and MCP, and its forward-leaning AI bets. Our mission is to unlock the bottlenecks of the hardware lifecycle with AI. Our agents reason over physical reality, from high-rate telemetry and test campaigns to designs and simulations, where real test results are the ground truth their work is checked against. We believe opinionated AI, built for the real work of hardware programs, will change how the world engineers.
We're collaborative, iterative, and high-agency, and we're human-centered and customer-focused. We build with the newest AI tools every day, and because those tools keep changing, so do we: we stay curious and keep looking for the better way. Our team spans data science and ML, distributed systems, search, and knowledge systems, and we obsess over how agents can be genuinely useful to the engineers who rely on them.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
THE ROLE
As a Staff Research Engineer supporting agent evals & post-training, you'll define how Nominal measures its agents, for customers and for ourselves, and build the path from evals to post-trained models when hardware needs them. Evals come first; post-training follows when air-gapped deployment or cost makes it the right investment.
💼 WHAT YOU'LL DO
- Build eval suites for our agents, the MCP tool layer, and our internal company agent, grounded in real hardware tasks.
- Invent new benchmarks for agents working over physical engineering data, where none exist today.
- Benchmark our agents against frontier agents using our MCP, and show where and why ours win.
- Build the eval infrastructure: datasets, model-based and human graders, and regression gates in CI.
- Turn evals into reward signals and training data, and lead post-training (fine-tuning, RL, distillation) when we need models we can run anywhere, including air-gapped environments.
- Set Nominal's strategy for measurement and model improvement, working with hardware experts on what "correct" means.
🚀 WHAT YOU'LL BRING
- 8+ years in ML engineering or research, including evals or post-training work you led in production.
- Statistical rigor: you design evals that don't fool you, and you know when a difference is real.
- Deep experience evaluating LLM or agent systems, including model-graded evals and their limits.
- Hands-on post-training experience: fine-tuning, RL from feedback, or distillation on real tasks.
- The judgment to know when to measure, when to train, and when a better prompt or tool is the answer.
- A track record of setting technical direction across a team and raising the bar for the engineers around you.
- You build with modern AI coding agents (Claude Code, Cursor, Codex) every day, and stay curious and open to better ways of working. The tools keep changing, and so do we.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
⚡️ NICE TO HAVE
- You've built evals or post-training at a frontier lab or an AI-native company.
- You've published benchmarks or eval methods that others use.
- You've trained or served open-weight models in restricted, on-prem, or air-gapped environments.
- You've worked in test, reliability, or verification engineering for physical systems.
BENEFITS/PERKS
- 🏥 100% coverage of medical, dental, and vision insurance
- 🏖️ Unlimited PTO and sick leave
- 🍽️ Free lunch, snacks, and coffee
- 🚀 Professional Development Stipend
- ✈️ Annual company retreat
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin.
ITAR Requirements
To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here [https://www.pmddtc.state.gov/?id=ddtc_kb_article_page&sys_id=24d528fddbfc930044f9ff621f961987].
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location