Rodeo
Get started

FactTrace

Principal Engineer: Retrieval Geometry & Vector Quantization

Cambridge
Posted about 11 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

About FactTrace

FactTrace is building the world's most precise content provenance infrastructure. From our base in Cambridge UK, we trace how language travels, mutates, and resurfaces across different systems — mapping the lineage of every claim from its origin through every paraphrase, reframing, and reinterpretation that follows.

We are an early-stage, research-intensive company operating at the intersection of representation learning, information retrieval, and applied cryptography. Working towards a reference corpus exceeding the 100 million documents and continues to scale. The systems we are designing today will underpin a forthcoming Zero-Knowledge cryptographic layer — and they demand mathematical structure of the highest order.

This is not a role for engineers who want to wrap foundation models in thin orchestration. We are designing proprietary representational spaces from first principles, and we are hiring the engineer who will lead that effort.

The Mission

You will architect the next generation of retrieval — systems that move beyond the implicit compromises of the current AI stack and set a new standard for what retrieval can be.

Standard dense retrieval groups text by topical proximity; the geometry conflates what a passage is about with what it actually asserts. That gap is the entire problem space. We are constructing intent-aware vector spaces in which the model can natively distinguish between a faithful paraphrase, a direct refutation, and a subtle distortion — preserving factual intent as a first-class geometric quantity.

The retrieval pipeline you build must answer a single question with sub-second latency against a 100M+ document corpus: where did this claim originate, and how has it evolved? Answering it well means retrieving with precision that conventional architectures cannot reach, and doing so without the computational cost that has, until now, been the price of that precision.

What You'll Lead

  • Single-stage, precision-first retrieval architecture. You will design retrieval pipelines that achieve cross-encoder-grade semantic precision in a single forward pass, eliminating the two-stage retrieve-and-rerank paradigm as a computational bottleneck. This is the central research and engineering challenge of the role: producing the discriminative geometry of a cross-encoder at the latency and throughput of a bi-encoder, at corpus scale. Your work here defines our retrieval ceiling.
  • Vector quantization and binarization at scale. You will pioneer compression regimes for our embeddings that preserve full semantic resolution at massive scale — moving beyond naïve dimensionality reduction into learned quantization, product and residual quantization, and binarized representations that retain the geometric structure our downstream systems depend upon. Compression without distortion is the design constraint, not an aspiration: the representational fidelity you preserve is the input to our cryptographic layer, where any geometric degradation would propagate irrecoverably.
  • Intent-aware representation learning. You will own the training objectives, contrastive structures, and hard-negative regimes that produce our proprietary vector spaces — drawing on metric learning and Natural Language Inference to encode entailment, contradiction, and neutral stance as native geometric relationships rather than downstream classification tasks.
  • The mathematical foundation. Every retrieval index, quantization scheme, and embedding you ship is also cryptographic substrate. You will work in close partnership with our cryptography track to ensure the geometry you produce is pristine, well-conditioned, and amenable to the Zero-Knowledge proofs that will operate over it. This is representation learning with a mathematical mandate that few teams in the world are positioned to pursue.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

What We're Looking For

We expect deep, demonstrable expertise across representation learning, metric learning, and Natural Language Inference, plus the engineering judgment to ship these systems at corpus scale. Specifically:

  • Representation & metric learning fluency. Substantial experience designing contrastive, triplet, and supervised metric learning objectives; familiarity with hard-negative mining, debiased contrastive losses, and representation collapse — and how to design around them.
  • NLI as a training signal. Strong grasp of entailment, contradiction, and neutral relation modelling, and a track record of using NLI signals to shape representational geometry rather than treating them as a classification head.
  • Retrieval systems at scale. Experience with ANN indexing (HNSW, IVF, ScaNN, or comparable), and a clear thesis on single-stage retrieval — why and how it can surpass two-stage retrieve-and-rerank pipelines.
  • Vector quantization & binarization. Working knowledge of PQ, OPQ, residual quantization, and learned binary codes; understanding of the precision–compression frontier and how to reason about it rigorously rather than empirically.
  • Production engineering. Comfortable owning systems from research prototype to high-throughput production, with strong Python, PyTorch (or JAX), and a deep understanding of the GPU memory and latency profiles of retrieval workloads.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

A research profile is welcome but not required; we weight demonstrated, shipped systems above publication counts. What is required is the conviction that the current retrieval stack is not the ceiling — and the ability to build the one above it.

Why This Role

  • A research-grade problem at production scale. Few teams are simultaneously pushing the frontier of intent-aware representation and operating 100M+ document retrieval under hard latency budgets. You will.
  • Direct cryptographic impact. The geometry you produce will be reasoned over by Zero-Knowledge proofs — a rare and demanding constraint that elevates every design decision you make.
  • Cambridge, at the centre of it. We are rooted in one of the world's deepest concentrations of ML, cryptography, and information-retrieval talent.

To Apply

Through linkedin and also send to HR@facttrace.ai a short note on a single retrieval or representation-learning system you've built that you're proud of, and (optionally) links to relevant work. Tell us what you would build first.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Representation Learning
Metric Learning
Natural Language Inference
Vector Quantization
Approximate Nearest Neighbor Indexing
PyTorch
JAX
Python
Binarization
Contrastive Learning
Information Retrieval
Zero-Knowledge Proofs

Location

Cambridge, England, United Kingdom

Sign up to applySee more jobs like this