ML Infra Engineer, Modeling at Physical Intelligence | CA, US | Rezi

ML Infra Engineer, Modeling at Physical Intelligence

ML Infra Engineer, Modeling

Physical Intelligence · CA, US

3 weeks ago

ML Infra Engineer, Modeling

Physical Intelligence · CA, US

23 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this ML Infra Engineer, Modeling role.

Rezi rewrites your resume against Physical Intelligence's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the ML Infra Engineer, Modeling posting at Physical Intelligence — free, in seconds.

About the Role

This is a hands-on, high-leverage role at the intersection of ML, software engineering, and scalable infrastructure. You will help scale and optimize our training systems and core model code, owning critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines. You will work closely with researchers and model engineers to translate ideas into experiments and production training runs.

Responsibilities

  • Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging.
  • Scale JAX-based training across TPU and GPU clusters with minimal friction.
  • Profile and improve memory usage, device utilization, throughput, and distributed synchronization.
  • Build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • Translate research needs into infra capabilities and guide best practices for training at scale.
  • Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics.

Requirements

  • Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
  • Hands-on large-scale training experience in JAX (preferred), PyTorch.
  • Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines.
  • Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS).
  • Ability to debug and optimize performance bottlenecks across the training stack.
  • Strong cross-functional communication and ownership mindset.
  • Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels).
  • Experience operating close to hardware (GPU/TPU performance tuning).
  • Background in robotics, multimodal models, or large-scale foundation models.
  • Experience designing abstractions that balance researcher flexibility with system reliability.

Skills

  • JAX
  • PyTorch
  • Distributed training
  • Cloud platforms (SLURM, Kubernetes, GCP TPU/GKE, AWS)
  • ML systems
  • GPU/TPU performance tuning

About the Company

  • Physical Intelligence is bringing general-purpose AI into the physical world.
  • We are a group of engineers, scientists, roboticists, and company builders developing foundation models and learning algorithms to power the robots of today and the physically-actuated devices of the future.
  • The ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast.
  • The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs.

Equal Opportunity

  • Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.