Rodeo
Get started

Applied Computing

Reinforcement Learning Researcher

London
Posted about 15 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Applied Computing

Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable abundance for a growing planet.

The hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data. We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational data and optimising in real time for any metric. Decisions get faster, operations get safer, and carbon intensity falls.

We’ve raised over $32 million, including one of the largest seed rounds for an AI company in the UK. We’re just getting started.

What You’ll Own

  • Orbital’s learning-based optimisation and control stack
  • RL + control hybrid systems for industrial processes
  • Safe and constrained policy learning frameworks
  • Simulation environments and digital twin integrations
  • Research → production translation for RL systems
  • Benchmarking standards for decision-making systems

Must-Have Qualifications

  • PhD in Computer Science, Robotics, Control, Applied Mathematics, or related field
  • First-author publications in:
    • Reinforcement Learning
    • Control systems
    • Sequential decision-making
  • 3+ years of hands-on RL research experience
  • Strong foundation in:
    • Reinforcement Learning (online + offline)
    • Optimisation and control theory (MPC, dynamic programming, etc.)
    • Deep learning (PyTorch)
  • Experience with:
    • Real-world deployment of ML systems
    • Simulation environments or digital twins
    • Working with noisy, real-world data

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

How We Work

  • Research is judged by production impact, not paper count
  • We optimise for real systems, not benchmarks alone
  • We value safe, reliable decision-making over theoretical elegance
  • Physics, control, and learning are treated as one system

What This Role Is Not

  • Not toy RL environments (Atari, MuJoCo-only thinking)
  • Not unconstrained policy learning without safety guarantees
  • Not offline research disconnected from deployment
  • Not a support role; this position owns core optimisation IP

Core Responsibilities

  1. Design & Implement RL-Based Decision Systems

    • Process optimisation (yield, efficiency, cost reduction)
    • Control policy learning (setpoint optimisation, constraint handling)
    • Sequential decision-making under uncertainty
    • Work across:
      • Model-free RL (policy gradients, actor-critic, offline RL)
      • Model-based RL (world models, planning-based methods)
      • Hybrid approaches combining RL with optimisation / MPC
  2. Build Physics-Constrained RL Systems

    • Embed domain knowledge into policy learning:
      • Hard constraints (safety, operating limits, regulatory bounds)
      • Soft constraints (efficiency, degradation, economic trade-offs)
      • Physics-informed reward shaping and transition models
    • Ensure policies:
      • Respect physical feasibility
      • Generalise across operating regimes
      • Remain stable under real-world disturbances

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job
  1. Offline RL, Simulation & Digital Twin Integration

    • Develop RL systems that work in data-scarce and risk-sensitive environments:
      • Offline RL from historical plant data
      • Simulation-based training via digital twins
      • Sim-to-real transfer strategies
    • Handle:
      • Distribution shift
      • Partial observability
      • Sparse / delayed rewards
  2. Safety, Robustness & Interpretability

    • Design safe RL systems for production environments:
      • Constrained RL / safe exploration
      • Policy validation before deployment
      • Fail-safe mechanisms and fallback strategies
    • Ensure outputs are:
      • Interpretable to engineers and operators
      • Auditable and explainable
      • Reliable under sensor faults and regime changes
  3. Production-Grade Deployment

    • Deploy RL systems into real-world infrastructure:
      • Containerised deployment (Docker, AWS / Azure)
      • Integration with control systems (APC, DCS, advisory layers)
      • Real-time inference and monitoring
    • Build pipelines for:
      • Continuous policy evaluation
      • Safe rollout and rollback
      • Online / batch policy updates
  4. Benchmarking & Validation

    • Define evaluation standards for RL systems:
      • Offline policy evaluation
      • Counterfactual analysis
      • Comparison vs MPC, heuristics, and operator baselines
    • Ensure:
      • Measurable economic impact
      • Reproducible results
      • Defensible performance claims
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Reinforcement Learning
Control Theory
Deep Learning
PyTorch
MPC
Dynamic Programming
Simulation Environments
Digital Twins
Offline RL
Policy Gradients
Actor-Critic
Physics-Informed Modeling
Docker
AWS
Azure
Sequential Decision-Making

Location

London, England, United Kingdom

Sign up to applySee more jobs like this