About the Role
This role is responsible for the technical aspects of managing provider relationships, ensuring prospective compute providers meet Andromeda's quality standards, and assisting them in achieving compliance. You will define and build the qualification framework, including acceptance tests, benchmarks, and technical standards, from scratch. This role is the technical counterpart to the Provider Technical Program Manager, focusing on technical judgment and upstream validation to prevent issues before they impact customers.
Responsibilities
- Vet prospective compute providers by assessing cluster architecture, GPU hardware, network fabric, storage, and orchestration against Andromeda's quality metrics.
- Define the qualification bar, including building the acceptance test suite, benchmark methodology, and quality thresholds.
- Run validation tests such as burn-in testing, fabric validation, NCCL and application-level benchmarks, and storage performance testing.
- Guide providers through technical onboarding, remediating gaps and tuning configurations.
- Partner with compute procurement to identify and qualify new providers, performing technical due diligence.
- Build and maintain technical relationships with existing providers, acting as their technical point of contact.
- Feed learnings back into provider-facing standards and internal documentation.
Requirements
- Deep HPC experience in designing, building, or operating GPU clusters at scale.
- Strong fabric knowledge, ideally including InfiniBand and RoCE.
- Experience with distributed orchestration and the HPC software stack (Slurm, Kubernetes, OpenMPI or equivalent).
- Proficiency in Linux systems engineering.
- Data-center literacy, including power, cooling, cabling, and physical-layer realities.
- Benchmarking judgment for large-scale training workloads.
- Ability to write precise, testable standards for external engineering teams.
- Credibility with external engineering teams, including the ability to deliver a failing grade while maintaining relationships.
- Comfort with ambiguity and building new frameworks.
Skills
- HPC
- GPU Clusters
- InfiniBand
- RoCE
- Distributed Orchestration
- Slurm
- Kubernetes
- OpenMPI
- Linux Systems Engineering
- Data Center Operations
- Benchmarking
- Technical Standards Development
- Provider Relationship Management
- Technical Due Diligence
Location
- North America Remote
- SF-Hybrid
Work Type
- Full-Time
Experience Level
- First HPC Architect
- Build this function from the ground up
Benefits
- Competitive compensation
- Meaningful equity
- Comprehensive benefits for you and your dependents
- Healthcare, dental, and vision coverage
- 401(k)
- Unlimited PTO
About the Company
- Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to scaled AI infrastructure.
- The company provides a managed cluster service and is building the systems, network, and orchestration layer for AI infrastructure.
- Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute globally.
- The platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in the AI market.
- The long-term vision is to build the liquidity layer for global AI compute.
Equal Opportunity
- Andromeda Cluster is an equal opportunity employer.
- We celebrate diversity and are committed to creating an inclusive environment for all employees.
- We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
