About the Role
We are seeking a Research Engineer to join our ML Research team, responsible for data quality and evaluation across our modeling efforts. This role involves building systems to ensure clean training data for VLM and classifier models and designing methods to measure the quality of radiology report generation, impacting patient outcomes.
Responsibilities
- Build data filtering and curation pipelines to maintain clean VLM and classifier training sets, detecting various quality issues at scale.
- Develop model-based data quality signals to identify ambiguous or high-value cases for human review.
- Collaborate with radiologists and annotators to define quality criteria and translate clinical judgment into scalable filters.
- Design evaluation methodology for report generation, measuring clinical accuracy, hallucination rates, and reporting style.
- Build and maintain clinical benchmark sets, stratified by modality, pathology, and difficulty, ensuring no train/eval contamination.
- Develop and validate model-based evaluators against radiologist judgment and track correlation with production outcomes.
- Implement continuous evaluation and regression testing for rapid model change assessment.
- Work across the research stack to identify bottlenecks and develop tooling to accelerate research scientists' progress.
Requirements
- 2+ years of industry or research experience in ML, data engineering, or a related area.
- Strong Python and solid software engineering fundamentals.
- Experience building tooling and data pipelines from scratch.
- Strength in data quality (dataset curation, filtering, deduplication, label-noise detection, data-centric ML) or evaluation (designing metrics, eval harnesses for generative models, LLM-as-judge, NLG/factuality evaluation).
- Demonstrated agency in identifying problems and driving them to completion.
- Comfort working in an ambiguous, fast-moving research environment.
- Ability to collaborate closely with research scientists.
Skills
- Python
- Software Engineering
- Data Engineering
- ML
- Dataset Curation
- Filtering
- Deduplication
- Label-Noise Detection
- Data-Centric ML
- Metrics Design
- Evaluation Harnesses
- Generative Models
- LLM-as-judge
- NLG
- Factuality Evaluation
- Medical Imaging
- Clinical Data
- DICOM
- Radiology Reports
- Clinical NLP
- Vision-Language Models
- Multimodal Training
- Human-in-the-loop Annotation
- Inter-annotator Agreement
- Clinical Accuracy Metrics
- Entity Extraction
- Relation Extraction
- RadGraph-style Scoring
- Data Pipeline Tooling
- Experiment Tooling
- Spark
- Airflow
- Databricks
Location
- Remote
Work Type
- Full-time
Experience Level
- 2+ years
About the Company
- Tackling critical challenges in medical imaging and diagnostics using AI and clinical practice.
- Building technology that directly impacts patient outcomes.
- Possesses one of the industry's most comprehensive and diverse medical imaging datasets.
- Has proven product-market fit with a substantial customer pipeline.
