About the Role
RL has become the de-facto method of post-training LLMs. We are looking for a talented researcher to weave together pre-LLM and post-LLM approaches to learning from experience.
Responsibilities
- Develop new RL algorithms for post-training language models.
- Adapt ideas from pre-LLM reinforcement learning to modern LLM and agentic settings.
- Establish empirical baselines and evaluation protocols for measuring sample efficiency, robustness, generalization, and reward exploitation in LLM RL.
- Analyze failure modes of RL-trained models.
- Collaborate with researchers on environments, evals, interpretability, reward modeling, and infrastructure.
- Own and develop a research agenda within Vmax.
Requirements
- PhD or equivalent experience in machine learning, reinforcement learning, or a closely related field.
- Track record of research excellence.
- Deep understanding of modern machine learning, especially reinforcement learning, representation learning, and large language models.
- Strong familiarity with LLM post-training methods.
- Experience designing and running rigorous ML experiments.
- Experience with large-scale ML infrastructure, distributed training, experiment tracking, data pipelines, and debugging unstable training runs.
- Expertise with Python and at least one major ML framework such as PyTorch or JAX.
- Ability to work independently on open-ended research problems.
Skills
- Reinforcement learning
- Model-based RL
- Temporal abstraction
- Value-based learning
- Large language models (LLMs)
- Sample efficiency
- Robustness
- Generalization
- Reward exploitation
- Reward hacking
- Mode collapse
- Over-optimization
- Exploration failures
- Distribution shift
- Python
- PyTorch
- JAX
- LLM pre-training
- Reward modeling
- Verifiers
- Process supervision
- Outcome supervision
- Automated evaluation systems
- Software engineering
- Communication skills
Location
- San Francisco
Work Type
- Hybrid
Experience Level
- PhD or equivalent experience
Education Level
- PhD
Salary/Compensations
- $300,000 - $500,000 USD
About the Company
- Vmax is an applied research lab developing AI capable of open-ended learning.
- We are building systems to exceed humans in all capacities by optimising beyond the local maxima of learning from human expertise.
