Singapore AI Safety Hub
Research Engineer / Research Scientist, Remote Compute Accounting

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
The team
Some of the most consequential decisions of the AI era will depend on answering one question: can we verify what is happening inside AI datacenters? Answering it could unlock international cooperation, help countries protect their sovereignty, and enable trustworthy adoption of AI in high-stakes industries. Building tools that can earn trust across borders is an urgent technical and political challenge.
SASH’s Verification team builds and tests tools to verify agreements about AI. We prototype new verification mechanisms, lead international research collaborations, and work with policymakers around the world to show what these tools can do and inform how they’re used.
You’d join a small team combining technical expertise with international AI policy experience. Our technical team is led by Pascal Berrang, an Associate Professor at University of Birmingham, and brings experience from the Singapore Government. Our policy team brings experience from Oxford and the Centre for the Governance of AI, while our partners include experts from the Future of Life Institute and the University of Oxford.
The problem
The hardest question in verification is not whether a datacenter is running the model it says it is. It is whether it is also running something else. A cluster can serve declared inference all day and train an undeclared model on the capacity left over.
We are interested in software-only approaches to this problem. If a verifier can keep a cluster's accelerators demonstrably occupied (i.e., their high-bandwidth memory holding the state they claim to hold, their compute answering challenges that cannot be faked or precomputed in advance), then there is no spare capacity for an undeclared training run. Physics supplies the lever: information cannot travel faster than light, and some computations cannot be performed faster than the hardware allows, so the time a cluster takes to answer is itself evidence about what hardware it has and what that hardware is doing.
The nearest concrete version is a memory-state challenge protocol. The prover commits frequently to workload checkpoints, the verifier fires N unique nonces simultaneously, and each device must return a function of its claimed memory state within a deadline (a function designed to demand serious memory bandwidth and matrix multiplication). Round-trip times are recorded and a random sample is proven genuine. A related family, proof of useful work, pushes further: the declared workload's own computation becomes the evidence, so the accounting costs almost nothing to produce.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Compression is one of the open questions: weights or a KV cache engineered for fast lossless decompression would let a prover reconstruct state it was not actually holding, freeing the hardware for whatever else it wanted to do. Showing that workload state cannot be decompressed inside the challenge window, or modifying workloads so that it cannot, would unlock the approach. Beyond that, a prover may hold secret memory capacity in error-correcting-code cells or overprovisioned regions; the challenge has to be designed so honest response time stays feasible; and the deadline has to be short enough to block cheating while surviving network latency and kernel launch times.
Your work
You would own the question of how much of a cluster can be accounted for from outside it, and how little room can be left over.
Concretely, that means designing challenge protocols and then measuring whether real accelerators can satisfy them: microbenchmarking, writing challenge kernels, characterising what high-bandwidth memory can and cannot do under adversarial conditions, and deciding whether the numbers support a claim someone could stake an agreement on. The work sits between two disciplines: measurement is GPU performance engineering, the rest is constructing a hardness assumption and then trying to break it. We expect most candidates to be strong on one side and curious about the other. You'd work alongside cryptographers and hardware-security colleagues, and you'd be expected to publish, including when the answer is that something does not work.
Representative projects
- Implement a candidate challenge function and characterise honest-response latency across real cluster topologies.
- Establish how much spare capacity a cluster could retain while still passing every challenge, then design that margin down.
- Run the adversarial compression study: how much of a frontier model's weights or KV cache can be losslessly decompressed inside a realistic challenge window, and how much hardware does that free up?
- Design and evaluate post-training that measurably raises the cost of compressing weights, without measurably hurting the model.
- Quantify how much memory a prover could plausibly hide, and what protocol changes make hiding it useless.
- Extend the approach from memory occupancy to compute, through proof-of-useful-work schemes where the declared workload's own computation carries the evidence.
- Take the protocol end to end against a live inference cluster and write up what broke.
- Design a red-team/blue-team competition for the protocol: cheap to run, because they need no new hardware.
- Red-team our own timing assumptions.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
About you
We are hiring this role at a range of levels. We care more about ability and trajectory than years of experience.
You may be a good fit if you:
- Have GPU performance or systems measurement experience (kernels, microbenchmarks, memory hierarchies, distributed timing) or a cryptographic background and a real appetite for getting your hands on hardware.
- Are rigorous about measurement, including about the ways your own measurement can be fooled.
- Can hold an adversary in your head while you optimise something.
- Are comfortable working on a protocol that may turn out not to work, and saying so in public if it does not.
- Thrive in ambiguous, early-stage environments where defining the problem is part of the job.
Strong candidates may also have:
- CUDA or equivalent low-level accelerator programming.
- Compression, information theory, or model quantisation experience.
- Background in timing-based protocols, proofs of space or time, or distance bounding.
- Experience characterising hardware you don’t have documentation for.
- Familiarity with LLM inference internals: batching, KV cache management, scheduling.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Location