Reinforcement Learning Researcher

Applied Computing
London

Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable abundance for a growing planet.


The hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data. We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational data and optimising in real time for any metric. Decisions get faster, operations get safer, and carbon intensity falls.


We’ve raised over $32 million, including one of the largest seed rounds for an AI company in the UK. We’re just getting started

What You’ll Own

  • Orbital’s learning-based optimisation and control stack

  • RL + control hybrid systems for industrial processes

  • Safe and constrained policy learning frameworks

  • Simulation environments and digital twin integrations

  • Research → production translation for RL systems

  • Benchmarking standards for decision-making systems

Must-Have Qualifications

  • PhD in Computer Science, Robotics, Control, Applied Mathematics, or related field

  • First-author publications in:

    • Reinforcement Learning

    • Control systems

    • Sequential decision-making

  • 3+ years of hands-on RL research experience

Strong foundation in:

  • Reinforcement Learning (online + offline)

  • Optimisation and control theory (MPC, dynamic programming, etc.)

  • Deep learning (PyTorch)

Experience with:

  • Real-world deployment of ML systems

  • Simulation environments or digital twins

  • Working with noisy, real-world data

How We Work

  • Research is judged by production impact, not paper count

  • We optimise for real systems, not benchmarks alone

  • We value safe, reliable decision-making over theoretical elegance

  • Physics, control, and learning are treated asone system

What This Role Is Not

  • Not toy RL environments (Atari,MuJoCo-only thinking)

  • Not unconstrainedpolicy learning without safety guarantees

  • Not offline research disconnected from deployment

  • Not a support role; this position owns core optimisation IP

Core Responsibilities

1. Design & Implement RL-Based Decision Systems

  • Process optimisation (yield, efficiency, cost reduction)

  • Control policy learning (setpoint optimisation, constraint handling)

  • Sequential decision-making under uncertainty

Work across:

  • Model-free RL (policy gradients, actor-critic, offline RL)

  • Model-based RL (world models, planning-based methods)

  • Hybrid approaches combining RL with optimisation / MPC

2. Build Physics-Constrained RL Systems

Embed domain knowledge into policy learning:

  • Hard constraints (safety, operating limits, regulatory bounds)

  • Soft constraints (efficiency, degradation, economic trade-offs)

  • Physics-informed reward shaping and transition models

Ensure policies:

  • Respect physical feasibility

  • Generalise across operating regimes

  • Remain stable under real-world disturbances

3. Offline RL, Simulation & Digital Twin Integration

Develop RL systems that work in data-scarce and risk-sensitive environments:

  • Offline RL from historical plant data

  • Simulation-based training via digital twins

  • Sim-to-real transfer strategies

Handle:

  • Distribution shift

  • Partial observability

  • Sparse / delayed rewards

4. Safety, Robustness & Interpretability

Design safe RL systems for production environments:

  • Constrained RL / safe exploration

  • Policy validation before deployment

  • Fail-safe mechanisms and fallback strategies

Ensure outputs are:

  • Interpretable to engineers and operators

  • Auditable and explainable

  • Reliable under sensor faults and regime changes

5. Production-Grade Deployment

Deploy RL systems into real-world infrastructure:

  • Containerised deployment (Docker, AWS / Azure)

  • Integration with control systems (APC, DCS, advisory layers)

  • Real-time inference and monitoring

Build pipelines for:

  • Continuous policy evaluation

  • Safe rollout and rollback

  • Online / batch policy updates

6. Benchmarking & Validation

Define evaluation standards for RL systems:

  • Offline policy evaluation

  • Counterfactual analysis

  • Comparison vs MPC, heuristics, and operator baselines

Ensure:

  • Measurable economic impact

  • Reproducible results

  • Defensible performance claims

Posted 2026-08-07

Recommended Jobs

Personal Tax Assistant - London (hybrid) c. £38, 000

Buckley Consulting
London

Personal Tax Assistant London (hybrid) c£38,000 The well respected and very busy private client tax team of this Top 20 firm can offer you a really broad range of high quality work, a friendly a…

View Details
Posted 2025-09-05

Speech and Language Therapy

London

PSL Recruitment Services is urgently recruiting a Locum Band 6 or Band 7 Adult Speech & Language Therapist (SLT) for a full-time position within an acute London based NHS Trust. The role involves wo…

View Details
Posted 2026-01-06

Customer Service Agent

CHERRY PICK PEOPLE
Central London

The City & Hybrid Working £30,000 plus lots of benefits Are you looking for a new and exciting opportunity, where you can further develop your career within the property industry? Do you have …

View Details
Posted 2025-07-30

2026 Research Analyst - Competition

The Brattle Group
London

The Brattle Group is looking for a Research Analyst (RA) to work out of our London office on a permanent basis.   ABOUT THE BRATTLE GROUP AND ECONOMIC CONSULTING The Brattle Group is a leadi…

View Details
Posted 2026-06-21

School Business Manager - Secondary School - Trafford,...

Marchant Recruitment
London

School Business Manager – Progressive Secondary School – Trafford, Greater Manchester Start Date: As soon as possible Contract: Full-time, Permanent Salary: Competitive salary dependent o…

View Details
Posted 2026-03-10

Personal Trainer (Newly Qualified / Early Career)

Love Recruitment
London

Personal Trainer (Newly Qualified / Early Career) Common Purpose Training, Mayfair, London £25,000 Base + £30,000+ OTE | Full-Time | Career Development Pathway Looking to start your PT career …

View Details
Posted 2026-03-28

A-level Economics Teacher - January 2026

Marchant Recruitment
Hounslow, Greater London

Economics Teacher – Hounslow, West London &##128176; Inspire Future Economists and Business Leaders at a Successful Academy We are seeking an enthusiastic, academically ambitious, and committe…

View Details
Posted 2025-11-19

Year 4 Teacher - Southwark - Independent School

Marchant Recruitment
London

Join a forward-thinking Independent School in Southwark as a Year 4 Teacher from January 2026. This Independent School seeks a high-energy Year 4 Teacher who will drive curriculum depth, nurture inde…

View Details
Posted 2025-11-06