About the Role
We are seeking a Software Engineer to build the ML infrastructure and data systems necessary for training and deploying state-of-the-art models for medical imaging. This role involves owning data pipelines, distributed training infrastructure, and inference/evaluation systems, requiring a blend of ML systems, data engineering, distributed training, and production deployment expertise.
Responsibilities
- Build and optimize distributed training infrastructure for foundation models on large-scale medical imaging, including parallelism and checkpointing for volumetric CT/MR training.
- Build the reinforcement learning training stack, enabling large-scale online, multi-reward RL.
- Build high-throughput data loading and preprocessing to keep GPUs saturated on large volumetric and multimodal datasets.
- Design and implement robust data pipelines for collecting, processing, and storing large-scale multimodal medical imaging data.
- Build centralized data storage solutions with standardized formats for efficient retrieval and training.
- Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery.
- Contribute to production serving and deployment pipelines, including model rollout, canary deployments, and monitoring.
Requirements
- 5+ years building ML infrastructure, data pipelines, or ML systems in production.
- Strong Python skills and expertise in PyTorch or JAX.
- Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and optimizing large GPU jobs.
- Hands-on experience with data pipeline technologies (e.g., Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk) and schema design.
- Experience with distributed systems, cloud infrastructure (AWS/GCP), and containerization (Docker/Kubernetes).
- Track record of building scalable data systems and shipping production ML infrastructure.
- Ability to move quickly and handle competing priorities in a fast-paced environment.
Skills
- Python
- PyTorch
- JAX
- Distributed Training
- FSDP
- DeepSpeed
- Megatron-style parallelism
- Data Pipeline Technologies
- Spark
- Airflow
- BigQuery
- Snowflake
- Databricks
- Chalk
- Schema Design
- Distributed Systems
- Cloud Infrastructure
- AWS
- GCP
- Containerization
- Docker
- Kubernetes
- Reinforcement Learning
- High-performance inference
- vLLM
- SGLang
- TensorRT
- Triton
- A/B testing
- Vision-language models (VLMs)
- Multimodal architectures
- Medical Imaging formats
- DICOM
- Healthcare data standards
- MLOps
- Model deployment pipelines
- Privacy-preserving data systems
- HIPAA compliance
Experience Level
- 5+ years
About the Company
- Tackling critical challenges in medical imaging and diagnostics using AI and clinical practice.
- Building technology that directly impacts patient outcomes.
- Possess one of the industry's most comprehensive and diverse medical imaging datasets.
- Have proven product-market fit with a substantial customer pipeline.
