Rodeo
Get started

Braintrust

Full Stack with Deep AI coding experience

United Kingdom
Posted 1 day ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Sr. AI/LLM Engineer

Reunion Marketing

Altitude Intelligence
Full-time

You will own the LLM application layer of Altitude Intelligence — the chat, reports, and analysis engine our teams and automotive dealer clients rely on for performance answers. We are looking for a senior engineer who has operated production LLM systems where mistakes carry real consequences, and who sees beyond the implementation — understanding how what we build affects client outcomes and where it fits in the bigger picture.

About This Role

Altitude Intelligence is the AI layer of Reunion Marketing's platform. It turns live client performance data and our marketing methodology into answers, reports, and analyses that our teams and automotive dealer clients act on — through three product surfaces.

You own the LLM application layer behind all three. In practice, that means the defining questions are yours:

  • What does the model retrieve, and how?
  • When is the right tool prompting, tool-use, or retrieval?
  • Where does the human belong in the loop?
  • What does "ready for clients" mean — and how do we prove it?

It is a senior, hands-on role with unusual visibility: every output of this system either protects a client relationship, wins one, or grows one, and we expect engineering decisions to be made with that in mind.

The Opportunity

  • Production data at scale. Live performance data across five marketing product lines for hundreds of automotive dealerships, not a demonstration dataset.
  • Daily users who depend on the output. Strategists, client success, sales, and external clients consume what this system produces as part of their working day.
  • Meaningful stakes. An incorrect answer does not stay inside a demo; it can reach a client conversation. The engineering standard follows from that.
  • Genuine ownership. The product direction is set and shipping, and many of the significant architectural decisions are still open for this hire to make and defend.

The Perspective We Expect

This role calls for more than strong implementation. The engineer we hire will:

  • Understand where each piece of work fits in the product and what it changes for our clients and teams — retention, expansion, hours returned — and let that understanding shape the engineering decisions.
  • See the system end to end: how a retrieval choice shows up in a report, how a prompt change lands in a client conversation, how today's shortcut becomes next quarter's incident.
  • Treat an incorrect number in front of a client as the most expensive defect the system can produce, and design the grounding, evaluation, and review gates accordingly.
  • Explain technical trade-offs to strategists and executives in their terms, and be willing to defend a delayed release when shipping would put client trust at risk.
  • Judge their own success by the accuracy, adoption, and time savings the system delivers.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

What You'll Own

  • The answer pipeline — from question or scheduled trigger, through data retrieval and reasoning, to a grounded, client-safe artifact: structured outputs, versioned prompts managed as reviewed artifacts, per-product context assembly, and multi-tenant scoping that is never bypassed.
  • Retrieval architecture — a hybrid retrieval problem spanning relational product data, client history, and our Knowledge Vault, the curated knowledge base of methodology and strategy material that grounds the system's answers. Responsibilities include the ingestion pipeline, hybrid search across embedding and lexical indexes with reranking, relevance evaluation, and selecting the appropriate retrieval method for each use case.
  • Generation workflows — the multi-step flows behind report and analysis generation, built to be durable, observable, retryable, and cancellable rather than fragile request handlers with hidden state.
  • Evaluation as an engineering discipline — we run large structured evaluation banks against live data and gate releases on them. You will own and grow that system: offline suites, regression gating in CI, factuality and groundedness checks, cost and latency tracking, and converting reviewer feedback into measurable quality improvement.
  • Production reliability — hallucinations, retrieval misses, tool-use failures, output drift, cost spikes, model deprecations, and the rollbacks that follow. We expect instrumentation before conjecture, and decisions in writing.
  • Architectural judgment — prompting versus retrieval versus tool-use versus fine-tuning, human review versus automated release, iteration speed versus client-safe controls. Well-supported positions are expected; unsupported ones are not.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

What We're Looking For

  • 6+ years of software engineering experience, including 2+ years shipping production LLM systems that were customer-facing, with real consequences when they failed. Internal demonstrations and proofs of concept do not meet this bar.
  • Strength in a modern application language used for production LLM systems, with a record of well-structured production systems rather than glue scripts.
  • Deep, hands-on familiarity with the LLM stack: prompt design and structured outputs; retrieval architecture (vector, lexical, hybrid, reranking, multi-tenant scoping); tool-use and agent design, including when not to use them; evaluation offline, online, and as a regression gate; and a defensible position on fine-tuning, including when to avoid it.
  • An evaluation suite you designed that caught a genuine regression in production, and the ability to walk through what it caught and what shipped, or did not, because of it.
  • Production debugging instincts for LLM failures — specific failure modes you have fixed, and a habit of instrumenting before guessing.
  • A record of owning the quality bar on a multi-team AI product, including incident response, rolling back a prompt change, and defending a delayed release to stakeholders who wanted to ship.
  • Strong written and architectural communication — a one-page proposal that a non-AI engineer, a strategist, and a CTO can all read and respond to.

Nice to Have

  • Production experience with Anthropic Claude specifically, including tool-use, prompt caching, and streaming.
  • Multi-tenant client data and the access-control patterns that come with it.
  • Domain experience in marketing analytics, SEO/SEM, paid media, local search, or answer-engine visibility.
  • Experience building a system of record that an organization runs on, rather than a feature inside one.
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

LLM Application Development
Retrieval Augmented Generation
Prompt Engineering
Vector Databases
Hybrid Search
Agent Design
Evaluation Frameworks
Production Debugging
Multi-tenant Architecture
Software Engineering
Anthropic Claude
Marketing Analytics

Location

United Kingdom

Sign up to applySee more jobs like this