Rodeo
Get started

Astek

Senior LLM Researcher - Alignment

London
Posted about 24 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Senior LLM Researcher - Alignment

For one of our clients, we are looking for a Senior LLM Researcher - Alignment to help build reliable, safe, and high-performing language models.

In this role, you will design and implement alignment, sanitization, evaluation, and guardrailing pipelines that enable the delivery of trustworthy AI at scale. You will also play a key role in advancing Arabic-focused language model capabilities, including dialectal coverage and culturally relevant safety considerations.

Key Responsibilities

  • Own alignment strategies, including:
    • Supervised fine-tuning (SFT)
    • Preference optimization, such as DPO and RLHF
    • Reward modeling
    • Rejection sampling
    • Constrained decoding
  • Design and build end-to-end alignment pipelines covering data curation, training, preference optimization, and evaluation.
  • Develop data pipelines for filtering, augmentation, preference-data collection, and quality control.
  • Build model sanitization workflows, including:
    • PII removal
    • Data redaction
    • Prompt and response scrubbing
    • Harmful-pattern filtering
  • Implement runtime guardrails using policy models, classifiers, rule engines, retrieval-based safety checks, and rate controls.
  • Establish red-teaming and adversarial testing programs, including incident taxonomies and response playbooks.
  • Define and maintain evaluation frameworks covering safety, robustness, task performance, drift, latency, and telemetry.
  • Expand Arabic language coverage across dialects, domains, and culturally relevant safety scenarios.
  • Partner with product, platform, and engineering teams to move research into production.
  • Mentor engineers and researchers while developing internal documentation, playbooks, and best practices.
  • Lead major AI initiatives that improve model helpfulness, harmlessness, faithfulness, reliability, and cost efficiency.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Requirements

  • Proven research or industry experience training and aligning large language models.
  • Strong hands-on experience with SFT and preference optimization, including RLHF, DPO, or reward modeling.
  • Strong machine learning theory and engineering capabilities.
  • Demonstrated experience in LLM safety and security, including guardrails, classifiers, adversarial prompting, or PII handling.
  • Experience designing evaluation frameworks for model quality, safety, robustness, and reliability.
  • Track record of taking research from experimentation through production, supported by clear metrics and documentation.
  • Strong programming skills in Python and experience with modern deep learning frameworks, particularly PyTorch.
  • Ability to work cross-functionally and communicate complex research clearly.
  • Experience mentoring engineers, researchers, or technical teams.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Nice to Have

  • Experience with red-teaming, AI policy design, or compliance in multilingual environments.
  • Familiarity with Arabic NLP, including dialectal variation.
  • Knowledge of retrieval-augmented generation, constrained decoding, or tool-use agents.
  • Experience with inference optimization, quantization, low-latency serving, or observability.
  • Familiarity with distributed training environments and large-scale model deployment.
  • Experience with systems such as Hugging Face, Slurm, LlamaFactory, NeMo, RLHF/DPO frameworks, vector databases, and policy or guardrail services.
  • Experience working with H200 or similar GPU infrastructure.
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Location

London, England, United Kingdom

Sign up to applySee more jobs like this