About the Role
This role embeds engineers directly with teams running large-scale training and inference on AI clusters. Responsibilities include onboarding, job tuning, debugging, and owning the underlying infrastructure and automation to ensure reliability. The position requires deep engagement within customer environments to diagnose and resolve complex issues, improve performance, and drive product improvements.
Responsibilities
- Serve as the primary technical point of contact for teams running large-scale training and inference workloads.
- Own onboarding end to end, including environment setup, orchestration choice, storage layout, and initial successful scaling.
- Work within customer environments to diagnose failures such as NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, and container/driver mismatches.
- Reproduce, isolate, fix, and document identified issues.
- Profile and improve distributed training performance on live workloads, focusing on MFU, idle GPU time, and time-to-first-successful-run.
- Own reliability outcomes for assigned customer accounts.
- Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) and diagnose fabric-level issues.
- Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health.
- Automate recurring deployment problems, including cluster provisioning, GPU health checks, preflight validation, self-healing, and firmware/driver lifecycle management.
- Lead incident response for complex, multi-layer failures.
- Own customer-facing communication during incidents and lead postmortem analysis and systemic fixes.
- Influence the product roadmap by identifying rough edges and building necessary solutions.
- Document customer technical goals, including training focus, scaling roadmap, and constraints.
- Ensure customer engineers are the first point of contact for issues.
- Improve customer reliability and throughput numbers through specific changes.
- Convert recurring field problems into automation, checks, or documentation.
- Advocate internally for roadmap changes based on customer needs.
Requirements
- Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent).
- Understanding of GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes.
- Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training.
- Ability to diagnose slow all-reduce operations, identify degraded links, and reason about congestion control.
- Working knowledge of how large training and inference jobs run at the systems level.
- Expert-level Linux experience, including kernel tuning, driver management, cgroup/namespace internals, container runtimes, and performance profiling.
- Strong experience running Kubernetes in production with GPU workloads, including device plugins, topology-aware scheduling, multi-cluster, and custom operators.
- Experience with Slurm or other HPC schedulers is valued.
- Strong engineering skills in Python, Go, or Bash, with experience building production-grade tools and services.
- Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).
- Hands-on experience building monitoring and alerting for GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics).
- Ability to engage with customer infrastructure teams on architecture and articulate tradeoffs to leadership.
- Proven track record leading incident response for complex distributed systems.
- Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS).
- Experience embedded with external engineering teams (solutions architecture, professional services, deployed SRE, technical account ownership).
- Experience operating production inference, including autoscaling, batching, KV cache behavior, cold starts, and multi-tenant GPU sharing.
- Contributions to relevant OSS projects, benchmarks, postmortems, or published deep-dives.
- Experience working across heterogeneous providers and regions.
Skills
- GPU cluster operation
- Distributed training performance tuning
- InfiniBand/RoCE/NVLink fabric management
- Linux kernel tuning
- NVIDIA driver and CUDA toolkit management
- Container runtime management
- Performance profiling
- Kubernetes (GPU workloads)
- Slurm
- Python
- Go
- Bash
- Infrastructure-as-Code (Terraform, Helm, Ansible)
- GPU telemetry monitoring (DCGM, nvidia-smi)
- Incident response
- High-performance parallel file systems
- Production inference operation
- OSS contributions
Location
- North America Remote
- SF-Hybrid
Work Type
- Full-Time
- Remote
- Hybrid
Experience Level
- Senior
Benefits
- Meaningful equity
- Comprehensive benefits for dependents
- Healthcare coverage
- Dental coverage
- Vision coverage
- 401(k)
- Unlimited PTO
About the Company
- Andromeda Cluster was founded by Nat Friedman and Daniel Gross to provide early-stage startups with access to scaled AI infrastructure.
- The company builds systems, networks, and orchestration layers to make AI infrastructure more accessible.
- Andromeda works with AI labs, data centers, and cloud providers to deliver compute globally.
- Their platform routes training and inference jobs across global supply, offering flexibility and efficiency in the AI market.
- The long-term vision is to build the liquidity layer for global AI compute.
- They are expanding to find talent in AI infrastructure, research, and engineering.
Equal Opportunity
- Andromeda Cluster is an equal opportunity employer.
- We celebrate diversity and are committed to creating an inclusive environment for all employees.
- We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
