Senior Agentic RL Engineer
About Recurvia
Recurvia is a frontier reinforcement learning lab building AI systems that improve other AI systems.
Our mission is to accelerate machine learning development using machine learning itself. Our work spans fundamental reinforcement learning research, large-scale experimentation, and product development.
Our current products are the first applications of this research programme:
• Paramorph — a foundation model, trained with multi-agent reinforcement learning, that controls and optimises the training of other AI models.
• Metrana — an AI-powered observability platform for large-scale reinforcement learning and foundation model training workloads.
We are backed by Amadeus Capital Partners, the European Innovation Council and Amazon Web Services. Following the award of a €2.5 million EIC Accelerator grant, we are expanding the team to accelerate development and commercialisation of our technology platform.
About the Role
We are looking for a Senior Agentic RL Engineer to lead the agentic core of Metrana and contribute to a new stealth project in inference-time optimisation.
You will own the Metrana agent end to end: designing the agent harness, building evaluation frameworks, running RL-based post-training, and developing autonomous research capabilities that let the system investigate and act on training workloads with minimal supervision. Alongside this, you will help shape an early-stage project applying test-time and inference-time scaling techniques to optimise model serving.
We are looking for someone who has made the transition from classic reinforcement learning research to post-training and agentic RL — someone grounded in RL theory who now spends their time on GRPO runs, tool-use agents, evals and harness design rather than gridworlds and game environments. The role sits closer to the engineering end of the research–engineering spectrum than our Senior RL Research Engineer position, but still demands solid reinforcement learning knowledge and research judgement.
You will work closely with research scientists and engineers across the company, shaping both the scientific direction of our agentic work and the systems built from it. This is an opportunity to join a deep reinforcement learning lab at an early stage and define how production agents are built, evaluated and improved.
Key Responsibilities
• Design and build the Metrana agent harness, including tool integration via MCP, context and memory management, and orchestration of multi-agent workflows.
• Develop rigorous evaluation frameworks and agent evaluation methodologies covering capability, reliability and long-horizon task performance.
• Run RL-based post-training using GRPO, RLVR and related techniques, including reward design and synthetic data generation for agentic tasks.
• Build autonomous research agents that automate experimentation, analysis and hypothesis generation across our research programme.
• Contribute to a stealth project on inference-time optimisation, applying test-time scaling and related techniques to model serving.• Develop tool-use, code and computer-use agent capabilities with robust long-horizon planning.
• Collaborate with engineers across the company to ship agentic systems into production environments.
Required Skills & Experience
• MSc, PhD or equivalent industry experience in computer science, machine learning, AI or a related discipline.
• Solid grounding in reinforcement learning theory, with demonstrable experience applying it to post-training of large language models.
• Hands-on experience with GRPO, RLVR, RLHF or related post-training methods.
• Experience designing and building agent harnesses, including tool use, MCP or similar protocols, and evaluation loops.
• Experience building evals — designing tasks, metrics and infrastructure to measure agent capability and reliability.
• Excellent Python skills and a track record of building production-quality software.
• Strong experience with PyTorch and the modern machine learning ecosystem.
• Ability to drive research and engineering projects independently, from concept through evaluation.
Desirable Skills & Experience
• Experience building autonomous research agents or research automation pipelines.
• Experience with multi-agent systems, code agents, computer-use agents or other long-horizon tool-use systems.
• Experience with test-time and inference-time scaling techniques, or inference optimisation more broadly.
• Experience with synthetic data generation for agent training and evaluation.
• Interest in self-improving agent architectures and recursive approaches to ML development.
• Publications at leading machine learning conferences or journals.
• Contributions to open-source agent frameworks, RL libraries or related projects.
• Experience in a research-driven environment such as a frontier AI lab, advanced AI startup or leading academic group.
What We Offer
Competitive Compensation
• Competitive salary and meaningful equity participation.
Frontier AI Research
• Work on novel agentic and reinforcement learning technologies with the potential to influence how future AI systems are trained, evaluated and deployed.
• Substantial compute resources and the freedom to pursue ambitious experimental research.
Research Impact
• Support for publication at leading venues such as NeurIPS, ICML and ICLR, where appropriate.
• Opportunities to contribute to patents, open-source projects and novel technological advances.
Join at a Pivotal Stage
• Backed by Amadeus Capital Partners, the European Innovation Council and AWS.
• Join early enough to shape both the technology and the company itself.