Research Engineer - Distributed Training at Prime Intellect | San Francisco | Rezi

Research Engineer - Distributed Training at Prime Intellect

Research Engineer - Distributed Training

Prime Intellect · San Francisco

1 weeks ago

Research Engineer - Distributed Training

Prime Intellect · San Francisco

14 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Prime Intellect is building the open superintelligence stack, providing infrastructure for ambitious AI teams. Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment for post-training at frontier scale. We are developing open frontier AI, including open-source models for long-horizon tasks and the full-stack platform our research team uses. The next generation of AI companies needs the ability to transform their workflows into owned superintelligence.

Responsibilities

  • Build and optimize distributed training infrastructure for pre-training and large-scale RL workloads using prime-rl.
  • Improve end-to-end training efficiency across compute, memory, networking, and scheduling.
  • Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
  • Work on distributed training systems for data, tensor, and pipeline parallel workloads.
  • Help shape the architecture of the RL training stack, including async rollout and post-training systems.
  • Contribute to open-source libraries and internal infrastructure for frontier-scale model training.
  • Collaborate with researchers and infrastructure engineers to translate bottlenecks into systems improvements.
  • Stay current with training systems, inference systems, compiler/runtime tooling, and hardware-aware optimization techniques.

Requirements

  • Strong systems engineering experience in AI/ML infrastructure, particularly large-scale model training or inference.
  • Deep familiarity with PyTorch and distributed training frameworks (e.g., PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray).
  • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.
  • Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong understanding of GPU architecture, profiling, and performance debugging.
  • Ability to identify stack bottlenecks and drive improvements from first principles.
  • Comfort working in a fast-moving environment with ambiguous problems and high ownership.

Skills

  • CUDA / Triton kernels
  • Compiler or runtime optimization for ML systems
  • RL training infrastructure
  • Rollout systems
  • Asynchronous training pipelines
  • Multi-node GPU clusters
  • High-performance networking
  • Open-source ML systems or infrastructure projects
  • Technical writing and publishing

Location

  • San Francisco
  • Remote

Work Type

  • Remote
  • In-person

Experience Level

  • Systems engineering experience in AI/ML infrastructure
  • Experience with large-scale model training or inference
  • Experience optimizing training performance
  • Hands-on experience with large-scale training techniques
  • Experience writing or optimizing CUDA / Triton kernels
  • Experience with compiler or runtime optimization for ML systems
  • Experience working on RL training infrastructure
  • Experience with multi-node GPU clusters

Salary/Compensations

  • $150-350k

Benefits

  • Cash Compensation Range of $150-350k
  • Equity incentives
  • Flexible work arrangements
  • Option to work remotely or in-person
  • Visa sponsorship
  • Relocation assistance
  • Quarterly team off-sites
  • Hackathons
  • Conferences
  • Learning opportunities

About the Company

  • Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team.
  • Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models.
  • We are building open frontier AI: open-source models trained end to end for long-horizon tasks like autonomous research, and the full-stack platform our own research team uses to build them.
  • The next generation of AI companies, enterprises, and research teams do not just need more GPUs. They need the ability to turn their own workflows, tools, data, and feedback loops into superintelligence they own.
  • We train open frontier models and ship the same stack to our customers. Its spans the full stack of training, deploying and continuously improving models — compute, large-scale RL, environments, sandboxes, evals, and deployment.
  • Prime Intellect has raised $150M in total funding from Founders Fund, Radical Ventures, NVIDIA, and exceptional AI, infrastructure, and enterprise operators — including Andrej Karpathy, Dwarkesh Patel, and leaders and founders from Ramp, Perplexity, Harvey, Mercor, Zapier, Datadog, Semianalysis, Cognition, OpenAI, Thinking Machines, Together AI, SemiAnalysis, LangChain, Browserbase, Cloudflare, Sierra, Databricks, Airbnb, OpenRouter, Standard Intelligence, Fleet, Core Auto, and more.
  • We are looking for people who want to build at the intersection of frontier research, real infrastructure, and go-to-market for a category that does not fully exist yet.