Rodeo
Get started

Astek

Senior LLM Researcher — Alignment (Arabic)

London
Posted about 16 hours ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Senior LLM Researcher

We are looking for a Senior LLM Researcher to help build reliable, safe and high-performance language models, with a particular focus on Arabic and multilingual AI.

You will design and implement LLM alignment and sanitization pipelines, develop robust safety guardrails, establish rigorous evaluation frameworks, and help translate AI research into production-ready capabilities.

This is an opportunity to work at the intersection of LLM research, AI safety, alignment and production engineering, contributing to high-impact AI initiatives and shaping the development of trustworthy language models.

What You'll Do

LLM Alignment & Post-Training

  • Design and implement end-to-end LLM alignment pipelines covering data curation, supervised fine-tuning (SFT), preference optimization and rigorous evaluation.
  • Apply alignment techniques including DPO, RLHF, reward modelling, rejection sampling and constrained decoding.
  • Develop and improve preference datasets, including data collection, filtering, augmentation and quality control.
  • Establish reproducible research processes, baselines and measurable performance metrics.

Model Safety, Sanitization & Guardrails

  • Build model sanitization pipelines covering PII removal, data redaction, prompt/response scrubbing and harmful-pattern filtering.
  • Design and integrate runtime safety mechanisms including policy models, classifiers, rule engines and retrieval-based safety checks.
  • Develop and maintain red-teaming and adversarial testing processes.
  • Establish incident taxonomies and response playbooks to identify and address model safety risks.

Evaluation & Reliability

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

  • Build automated evaluation frameworks covering task performance, safety, robustness and model behaviour.
  • Develop continuous testing and reporting integrated into CI and model development workflows.
  • Monitor model quality, safety, drift and telemetry.
  • Support A/B testing and canary releases and measure improvements in helpfulness, harmlessness and faithfulness.

Arabic & Multilingual AI

  • Improve LLM alignment for Arabic language use cases, dialectal coverage and culturally relevant safety considerations.
  • Contribute to the development of preference data and evaluation methodologies appropriate for Arabic and multilingual models.
  • Help identify and address language-specific safety and alignment challenges.

Research, Leadership & Collaboration

  • Stay current with research and best practices in LLM alignment, safety and post-training.
  • Translate research findings into production-ready solutions.
  • Partner with product and platform teams to deploy aligned models and measure real-world impact.
  • Lead major AI initiatives and contribute to the AI roadmap.
  • Mentor engineers and researchers on alignment development, evaluation and safety engineering.
  • Create internal documentation, playbooks, experiments and best-practice guidelines.

What We're Looking For

Required

  • Proven research or industry experience working on LLM training, post-training or alignment.
  • Hands-on experience with Supervised Fine-Tuning (SFT) and at least one preference-optimization approach such as DPO or RLHF.
  • Strong understanding of ML theory and practical ML engineering.
  • Demonstrated experience in LLM safety, security or responsible AI, including guardrails, classifiers, adversarial prompting, red-teaming or PII handling.
  • Experience designing or implementing LLM evaluation frameworks and measuring model performance, safety or robustness.
  • Track record of translating research or experimentation into production-ready AI systems.
  • Strong Python and PyTorch experience.
  • Experience with modern LLM tooling and frameworks, particularly Hugging Face.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Nice to Have

  • Experience with Arabic NLP, Arabic language models or multilingual LLMs.
  • Experience with red-teaming, AI safety or policy design in multilingual environments.
  • Knowledge of RAG, constrained decoding and tool-use/agentic systems.
  • Experience with inference optimisation, quantization, low-latency model serving or observability.
  • Experience with distributed model training environments such as Slurm and frameworks such as LlamaFactory or NeMo.
  • Experience working with preference-data and alignment tooling.

Technical Environment

Python | Rust | PyTorch | Hugging Face | SFT | RLHF | DPO | Reward Models | Slurm | LlamaFactory | NeMo | LLM Evaluation | Red Teaming | Guardrails | RAG | Vector Databases | Optimised Inference | Observability

Why Join?

  • Work on challenging LLM alignment and AI safety problems.
  • Help shape the development of Arabic-oriented and multilingual AI.
  • Work across research, engineering and product teams.
  • Lead high-impact AI initiatives and influence technical direction.
  • Contribute to building trustworthy, reliable and production-ready AI systems.

Please note: This role is based in London and requires the successful candidate to work from the London office.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Location

London, England, United Kingdom

Sign up to applySee more jobs like this