Software Engineer - ML Infrastructure at Epsilon Labs, Inc. | San Francisco | Rezi

Software Engineer - ML Infrastructure at Epsilon Labs, Inc.

Software Engineer - ML Infrastructure

Epsilon Labs, Inc. · San Francisco

Yesterday

Software Engineer - ML Infrastructure

Epsilon Labs, Inc. · San Francisco

a day ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Software Engineer to build the ML infrastructure and data systems necessary for training and deploying state-of-the-art models for medical imaging. This role involves owning data pipelines, distributed training infrastructure, and inference/evaluation systems, requiring a blend of ML systems, data engineering, distributed training, and production deployment expertise.

Responsibilities

  • Build and optimize distributed training infrastructure for foundation models on large-scale medical imaging, including parallelism and checkpointing for volumetric CT/MR training.
  • Build the reinforcement learning training stack, enabling large-scale online, multi-reward RL.
  • Build high-throughput data loading and preprocessing to keep GPUs saturated on large volumetric and multimodal datasets.
  • Design and implement robust data pipelines for collecting, processing, and storing large-scale multimodal medical imaging data.
  • Build centralized data storage solutions with standardized formats for efficient retrieval and training.
  • Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery.
  • Contribute to production serving and deployment pipelines, including model rollout, canary deployments, and monitoring.

Requirements

  • 5+ years building ML infrastructure, data pipelines, or ML systems in production.
  • Strong Python skills and expertise in PyTorch or JAX.
  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and optimizing large GPU jobs.
  • Hands-on experience with data pipeline technologies (e.g., Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk) and schema design.
  • Experience with distributed systems, cloud infrastructure (AWS/GCP), and containerization (Docker/Kubernetes).
  • Track record of building scalable data systems and shipping production ML infrastructure.
  • Ability to move quickly and handle competing priorities in a fast-paced environment.

Skills

  • Python
  • PyTorch
  • JAX
  • Distributed Training
  • FSDP
  • DeepSpeed
  • Megatron-style parallelism
  • Data Pipeline Technologies
  • Spark
  • Airflow
  • BigQuery
  • Snowflake
  • Databricks
  • Chalk
  • Schema Design
  • Distributed Systems
  • Cloud Infrastructure
  • AWS
  • GCP
  • Containerization
  • Docker
  • Kubernetes
  • Reinforcement Learning
  • High-performance inference
  • vLLM
  • SGLang
  • TensorRT
  • Triton
  • A/B testing
  • Vision-language models (VLMs)
  • Multimodal architectures
  • Medical Imaging formats
  • DICOM
  • Healthcare data standards
  • MLOps
  • Model deployment pipelines
  • Privacy-preserving data systems
  • HIPAA compliance

Experience Level

  • 5+ years

About the Company

  • Tackling critical challenges in medical imaging and diagnostics using AI and clinical practice.
  • Building technology that directly impacts patient outcomes.
  • Possess one of the industry's most comprehensive and diverse medical imaging datasets.
  • Have proven product-market fit with a substantial customer pipeline.