About the Role
The Principal ML Infrastructure Engineer will extend and operate the infrastructure that powers our research model training, fine-tuning, and serving pipelines. You will be embedded within our Research function, partnering directly with ML engineers and research scientists to ensure they can train Large Physics Models efficiently and reliably at scale.
Responsibilities
- Design and operate distributed training infrastructure for neural operator architectures on our large NVIDIA DGX B200 platform.
- Optimize training pipelines for throughput, fault tolerance, and cost efficiency.
- Build and maintain experiment tracking and observability systems.
- Solve data loading bottlenecks for large-scale mesh datasets.
- Optimize data pipelines for efficient I/O from cloud storage.
- Work with heterogeneous data sources of varying formats and resolutions.
- Build serving infrastructure for pre-trained LPMs, supporting both zero-shot inference and uncertainty quantification.
- Design and implement model packaging pipelines for customer deployment.
- Ensure reproducibility: any model checkpoint should be deployable with consistent behaviour.
- Improve developer experience for the Research team with fast iteration cycles, reliable CI/CD, clear debugging tools.
- Collaborate with the broader Infrastructure team on shared patterns and standards.
Requirements
- Ability to scope and effectively deliver projects, prioritising activity as needed.
- Problem-solving skills and the ability to analyse issues, identify causes, and recommend solutions quickly.
- Excellent collaboration and communication skills, especially in a research setting.
- 5+ years of experience building and operating ML infrastructure at scale.
- Deep expertise in distributed training: you've debugged NCCL hangs, optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism.
- Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization.
- Production experience with Kubernetes and SLURM for job orchestration on GPU clusters.
- Proficiency in Python and ML frameworks (PyTorch strongly preferred).
- Experience with cloud GPU infrastructure; ideally CoreWeave or similar GPU/HPC-focused clouds.
Skills
- Distributed training
- NCCL
- FSDP
- DDP
- Pipeline parallelism
- Linux
- Networking
- NVLink
- InfiniBand
- Storage I/O
- Profiling
- Performance optimization
- Kubernetes
- SLURM
- Python
- PyTorch
- Cloud GPU infrastructure
- Geometric deep learning
- Neural operators
- HPC
- CFD
- FEA
- Model serving infrastructure
- Weights & Biases
- MLflow
- Prometheus
- Grafana
- Model packaging
- Model registries
- Versioning
Location
- Shoreditch office
Work Type
- Hybrid
Experience Level
- 5+ years of experience building and operating ML infrastructure at scale
Benefits
- Equity options
- 10% employer pension contribution
- Free office lunches
- Enhanced parental leave (3 months full pay paternity and 6 months full pay maternity leave)
- YellowNest nursery scheme
- 25 days of Annual Leave (+ Public Holidays)
- Private medical insurance (100% employee cover)
- Wellhub Subscription
- Eye tests
- Personal development support
- Employee Assistance Programme (EAP)
- Bike2Work scheme
- Season ticket loan
- Octopus EV salary sacrifice
About the Company
- PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software.
- We are building an AI-driven simulation software stack for engineering and manufacturing across advanced industries.
- By enabling high-fidelity, multi-physics simulation through AI inference across the entire engineering lifecycle, PhysicsX unlocks new levels of optimization and automation in design, manufacturing, and operations — empowering engineers to push the boundaries of possibility.
- Our customers include leading innovators in Aerospace & Defense, Materials, Energy, Semiconductors, and Automotive.
Equal Opportunity
- We value diversity and are committed to equal employment opportunity regardless of sex, race, religion, ethnicity, nationality, disability, age, sexual orientation or gender identity.
- We strongly encourage individuals from groups traditionally underrepresented in tech to apply.
- To help make a change, we sponsor bright women from disadvantaged backgrounds through their university degrees in science and mathematics.
- We collect diversity and inclusion data solely for the purpose of monitoring the effectiveness of our equal opportunities policies and ensuring compliance with UK employment and equality legislation.
- This information is confidential, used only in aggregate form, and will not influence the outcome of your application.
