Research Scientist - Vision Foundation Models at Epsilon Labs, Inc. | San Francisco, California | Rezi

Research Scientist - Vision Foundation Models at Epsilon Labs, Inc.

Research Scientist - Vision Foundation Models

Epsilon Labs, Inc. · San Francisco, California

Today

Research Scientist - Vision Foundation Models

Epsilon Labs, Inc. · San Francisco, California

9 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Research Scientist with deep expertise in vision foundation models to join our ML Research team. You will be at the forefront of developing and deploying state-of-the-art vision models for medical imaging applications. This role focuses on pretraining and scaling vision encoders for radiology diagnosis across X-ray, CT, and MRI, with a growing emphasis on 3D volumetric modeling. You will work with one of the largest and most diverse medical imaging datasets in the industry, pushing the boundaries of what's possible in AI-assisted diagnosis while maintaining the rigor required for clinical deployment.

Responsibilities

  • Design, train, and scale vision foundation models for radiology applications across X-ray, CT, and MRI modalities, implementing self-supervised, contrastive, masked image modeling, and joint-embedding predictive (JEPA) frameworks.
  • Extend 2D pretraining recipes to volumetric CT and MR data, addressing long sequence lengths, anisotropic spacing, and multi-sequence studies.
  • Evaluate model performance rigorously across academic benchmarks, internal offline datasets, and live production data.
  • Contribute hands-on to all stages of model development including dataset curation, architecture design, distributed training, and production deployment.
  • Stay current with cutting-edge research in computer vision and medical imaging AI.
  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for training robust medical imaging models at scale.

Requirements

  • 6+ years of academia/industry experience in computer vision/machine learning
  • Deep expertise in training vision encoder models at scale (e.g. ViT, ConvNeXt).
  • Strong foundation in self-supervised pretraining, including contrastive, masked image modeling, self-distillation, and JEPA-style objectives.
  • Experience training on volumetric or spatiotemporal data (video, 3D medical imaging)
  • Track record of implementing complex models from research papers and adapting them to new domains
  • Proficiency in PyTorch or JAX, with experience training models on multi-GPU/distributed systems
  • Hands-on experience with medical imaging applications, particularly radiology (X-ray, CT, MRI)
  • Strong software engineering skills and ability to write production-quality code
  • Publications at top-tier conferences (CVPR, ICCV/ECCV, NeurIPS, ICLR, MICCAI)
  • Experience with 3D medical image processing and retrieval tasks
  • Familiarity with CT and MR acquisition (windowing, multi-sequence protocols, voxel spacing)
  • Experience with long-context training techniques (sequence parallelism, efficient attention)
  • Knowledge of vision-language models and multimodal learning
  • Experience with model interpretability and explainability methods
  • Understanding of clinical evaluation metrics, clinical workflows, and healthcare data (DICOM, HL7, etc.)

Skills

  • Vision foundation models
  • Computer vision
  • Machine learning
  • Vision encoder models
  • ViT
  • ConvNeXt
  • Self-supervised pretraining
  • Contrastive learning
  • Masked image modeling
  • Self-distillation
  • JEPA objectives
  • Volumetric data training
  • Spatiotemporal data training
  • PyTorch
  • JAX
  • Distributed systems training
  • Medical imaging applications
  • Radiology
  • X-ray
  • CT
  • MRI
  • Software engineering
  • Production-quality code
  • 3D medical image processing
  • 3D medical image retrieval
  • CT acquisition
  • MR acquisition
  • Long-context training
  • Sequence parallelism
  • Efficient attention
  • Vision-language models
  • Multimodal learning
  • Model interpretability
  • Model explainability
  • Clinical evaluation metrics
  • Clinical workflows
  • Healthcare data (DICOM, HL7)

Experience Level

  • 6+ years of academia/industry experience

About the Company

  • We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics.
  • Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes.
  • We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.