About the Role
Lightning AI is seeking AI Platform Support Engineers to join their EMEA Customer Experience team. This role involves supporting ML engineers with large-scale training and inference workloads on cloud infrastructure, Kubernetes, and GPU platforms. You will act as a technical partner, diagnosing failures, improving reliability, and guiding customers through complex distributed systems challenges.
Responsibilities
- Partner directly with customer engineering teams running training and inference workloads in production
- Help customers diagnose and resolve complex distributed systems and ML infrastructure issues
- Act as a technical advisor during high impact incidents and platform degradation events
- Translate infrastructure level issues into actionable guidance for ML engineers
- Build credibility with customers through strong technical reasoning and clear communication
- Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems
- Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues
- Analyze logs, metrics, traces, and system behavior to isolate root causes
- Debug containerized workloads running across Kubernetes and bare metal GPU environments
- Support customers scaling workloads across multi node GPU systems
- Diagnose performance bottlenecks involving compute, memory, networking, or storage
- Identify recurring patterns across customer issues and drive long term reliability improvements
- Contribute to post incident reviews and operational improvements
- Build internal tooling, automation, documentation, and runbooks
- Partner closely with infrastructure, networking, and platform engineering teams
- Help improve observability, operational visibility, and troubleshooting workflows
- Improve the customer experience through better processes and technical guidance
Requirements
- Strong software engineering and systems troubleshooting background
- Experience with Kubernetes and containerized environments
- Linux systems knowledge, including networking, storage, process management, and performance tuning
- Experience with cloud infrastructure and distributed systems
- Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry
- Hands on experience operating machine learning workloads in production or research environments
- Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL
- Familiarity with GPU infrastructure and orchestration
- Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure
- Understanding of the operational challenges involved in running ML systems at scale
- Strong communication skills and ability to work directly with highly technical customers and engineering teams
- Comfortable operating in fast moving, highly ambiguous environments
- Enjoys solving complex technical problems collaboratively
- Experience with large scale model training or distributed inference systems
- Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms
- Experience with InfiniBand, RDMA, or high-performance networking
- Experience operating bare metal infrastructure
- Familiarity with storage systems commonly used in ML environments
- Experience working at an AI infrastructure, cloud, MLOps, or developer tooling company
- Contributions to platform engineering, developer infrastructure, or operational tooling projects
- Experience writing automation, tooling, or scripts in Python or similar languages
Skills
- PyTorch
- Kubernetes
- GPU orchestration
- Distributed systems
- Cloud infrastructure
- Linux
- Networking
- Storage
- Process management
- Performance tuning
- Observability tools
- Debugging tools
- Prometheus
- Grafana
- OpenTelemetry
- CUDA
- NCCL
- Python
Location
- London, UK
Work Type
- Hybrid
- Full-time
Experience Level
- Mid-level
- Senior
Salary/Compensations
- £75,000—£95,000 GBP
Benefits
- Comprehensive Health Coverage: Medical, dental, and vision coverage
- Meaningful Equity: RSUs
- Retirement Savings: Pension contributions (U.K.)
- Flexible Time Off: Unlimited PTO, company holidays, and floating holidays
- Company-Wide Winter Break: Two weeks
- Paid Parental & Family Leave
- Professional Development: Annual learning and development allowance
- Wellness Benefits: Wellness and work-from-home stipends
- Sabbatical Program: Four weeks of paid sabbatical leave after four years of service
- Flexible schedules
- In-Office Meals: Complimentary meals
About the Company
- Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems.
- Through a merger with Voltage Park, Lightning AI combines developer-first software with cost-efficient, large-scale compute.
- The company serves solo researchers, startups, and large enterprises globally.
- Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London.
- The company is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
Equal Opportunity
- Lightning AI is committed to fostering an inclusive and diverse workplace.
- We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic.
- We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.
