Member of Technical Staff - RL Algorithms at Vmax AI Corp | CA, US | Rezi

Member of Technical Staff - RL Algorithms at Vmax AI Corp

Member of Technical Staff - RL Algorithms

Vmax AI Corp · CA, US

2 weeks ago

Member of Technical Staff - RL Algorithms

Vmax AI Corp · CA, US

20 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

RL has become the de-facto method of post-training LLMs. We are looking for a talented researcher to weave together pre-LLM and post-LLM approaches to learning from experience.

Responsibilities

  • Develop new RL algorithms for post-training language models.
  • Adapt ideas from pre-LLM reinforcement learning to modern LLM and agentic settings.
  • Establish empirical baselines and evaluation protocols for measuring sample efficiency, robustness, generalization, and reward exploitation in LLM RL.
  • Analyze failure modes of RL-trained models.
  • Collaborate with researchers on environments, evals, interpretability, reward modeling, and infrastructure.
  • Own and develop a research agenda within Vmax.

Requirements

  • PhD or equivalent experience in machine learning, reinforcement learning, or a closely related field.
  • Track record of research excellence.
  • Deep understanding of modern machine learning, especially reinforcement learning, representation learning, and large language models.
  • Strong familiarity with LLM post-training methods.
  • Experience designing and running rigorous ML experiments.
  • Experience with large-scale ML infrastructure, distributed training, experiment tracking, data pipelines, and debugging unstable training runs.
  • Expertise with Python and at least one major ML framework such as PyTorch or JAX.
  • Ability to work independently on open-ended research problems.

Skills

  • Reinforcement learning
  • Model-based RL
  • Temporal abstraction
  • Value-based learning
  • Large language models (LLMs)
  • Sample efficiency
  • Robustness
  • Generalization
  • Reward exploitation
  • Reward hacking
  • Mode collapse
  • Over-optimization
  • Exploration failures
  • Distribution shift
  • Python
  • PyTorch
  • JAX
  • LLM pre-training
  • Reward modeling
  • Verifiers
  • Process supervision
  • Outcome supervision
  • Automated evaluation systems
  • Software engineering
  • Communication skills

Location

  • San Francisco

Work Type

  • Hybrid

Experience Level

  • PhD or equivalent experience

Education Level

  • PhD

Salary/Compensations

  • $300,000 - $500,000 USD

About the Company

  • Vmax is an applied research lab developing AI capable of open-ended learning.
  • We are building systems to exceed humans in all capacities by optimising beyond the local maxima of learning from human expertise.