About the Role
We are seeking a DevOps Engineer to build and own the infrastructure for an AI-driven materials discovery platform. You will work with ML researchers and software engineers to accelerate scientific breakthroughs by making model training, experimentation, and deployment fast, reliable, and reproducible. This is a foundational hire where you will set the patterns for future development.
Responsibilities
- Design, provision, and manage cloud infrastructure (AWS/GCP) using infrastructure-as-code (Terraform, Pulumi, or equivalent).
- Own GPU compute environments for model training and inference, including cluster configuration, job scheduling, and cost optimisation.
- Build and maintain CI/CD pipelines for rapid model iteration, automated testing, and safe deployments.
- Support ML workflow orchestration, including experiment tracking, training run management, and data pipeline reliability.
- Ensure reproducibility across research and production environments through containerisation and rigorous environment management.
- Define monitoring, alerting, and incident response processes.
- Implement security best practices: secrets management, IAM, network segmentation, vulnerability scanning.
- Build internal tooling and documentation for researcher self-service infrastructure.
Requirements
- 4+ years in a DevOps, Platform Engineering, or SRE role.
- Strong proficiency with at least one major cloud provider and its core services (compute, storage, networking, IAM).
- Hands-on experience with infrastructure-as-code and container orchestration (Kubernetes or equivalent).
- Solid CI/CD pipeline experience (GitHub Actions, GitLab CI, or similar).
- Proficient in Python and Bash; comfortable reading and writing code across a polyglot stack.
- Deep Linux systems knowledge and strong networking fundamentals.
- A bias for building things properly the first time, even under early-stage constraints.
Skills
- Infrastructure-as-code
- Container orchestration (Kubernetes)
- CI/CD pipelines
- Python
- Bash
- Linux systems
- Networking fundamentals
- GPU cluster management
- ML training workloads
- MLOps tooling
- Experiment tracking (MLflow, Weights & Biases)
- Workflow orchestration (Airflow, Prefect, Argo)
- Data versioning (DVC)
- Scientific computing
- HPC environments
Location
- London-based
Work Type
- Flexible approach to how and where you work
Experience Level
- 4+ years
Salary/Compensations
- Competitive salary
Benefits
- Generous equity
- Benefits
About the Company
- Diffractive is building the AI Material Scientist that autonomously learns from real-world experimentation to push the boundaries of scientific discovery.
- We're early, moving fast, and working on problems that genuinely matter.
- You'll join a small, high-calibre team where your work has real impact from day one.
Equal Opportunity
- Diffractive is an equal opportunities employer.
- We are committed to creating an inclusive environment for all employees and welcome applications from people of all backgrounds, experiences, and identities.
- If you require any adjustments or accommodations at any point during the interview process please let us know - we will be happy to help.
