Data Engineer at Diffractive Labs | GB | Rezi

Data Engineer at Diffractive Labs

Data Engineer

Diffractive Labs · GB

1 months ago

Data Engineer

Diffractive Labs · GB

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Diffractive Labs is building the infrastructure to accelerate the discovery of novel magnetic materials through closed-loop AI and physical experimentation. As our Data Engineer, you will own the data architecture that powers this research. This role requires a builder who understands both large-scale machine learning pipelines and the messy reality of physical lab data. You will serve as the critical link between experimental results generated at the bench and the models evaluating them, ensuring our research team always has the exact, high-quality datasets required to push the frontier of materials science. We are looking for someone with a rigorous, experimental mindset who thrives in an interdisciplinary environment and operates with a high degree of technical ownership.

Responsibilities

  • Drive the overarching data architecture across our training stack, mapping out data requirements with ML researchers and evaluating new external sources to fill knowledge gaps.
  • Design and deploy the ingestion pipelines that capture physical experimental data directly from our wet lab instruments and feed it seamlessly into our model training workflows.
  • Construct robust, reproducible systems for processing, standardizing, and versioning diverse scientific corpora, creating a highly reliable foundation for the research team.
  • Develop custom evaluation datasets and reinforcement learning environments specifically calibrated for the properties and behaviors of magnetic materials.
  • Build internal tooling that allows machine learning researchers and physical scientists to effectively query, inspect, and audit the data feeding into pretraining, midtraining, and RL runs.
  • Continuously integrate emerging techniques in synthetic data generation, data selection, and data-efficient training into our production systems.

Requirements

  • 3+ years of engineering experience focused on large-scale data pipelines, ideally within an applied ML, scientific, or LLM training environment.
  • High proficiency in Python and modern workflow orchestration frameworks (e.g., Dagster, Airflow, Prefect, or similar).
  • Demonstrated experience with dataset lineage, versioning, and reproducibility tooling (such as DVC, Delta Lake, or custom equivalents).
  • A track record of collaborating directly with machine learning researchers, translating complex modeling needs into scalable pipeline architecture and back again.
  • Strong DevOps fundamentals, including hands-on experience with containerization (Docker, Kubernetes) and CI/CD deployment.

Skills

  • Python
  • Dagster
  • Airflow
  • Prefect
  • DVC
  • Delta Lake
  • Docker
  • Kubernetes
  • CI/CD

Location

  • London-based

Work Type

  • Flexible approach to how and where you work

Experience Level

  • 3+ years of engineering experience

Salary/Compensations

  • Competitive salary

Benefits

  • Generous equity
  • Benefits

About the Company

  • Diffractive is building the AI Material Scientist that autonomously learns from real-world experimentation to push the boundaries of scientific discovery.
  • We're early, moving fast, and working on problems that genuinely matter.
  • You'll join a small, high-calibre team where your work has real impact from day one.

Equal Opportunity

  • Diffractive is an equal opportunities employer.
  • We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities.
  • If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.