About the Role
Inferact is seeking a hands-on cluster administration engineer to manage and operate high-performance GPU compute infrastructure, ensuring its health, availability, observability, and usability for engineering teams. This role involves taking ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response, while collaborating with leadership to standardize compute operations across providers.
Responsibilities
- Own and operate high-performance GPU compute infrastructure.
- Ensure infrastructure is healthy, available, observable, and usable around the clock.
- Take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response.
- Work closely with engineering leadership and infrastructure owners to standardize compute provisioning, operation, debugging, and scaling across providers.
- Automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.
- Improve cluster utilization, reduce idle or unavailable GPU capacity, and debug scheduling or resource contention issues.
- Manage secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.
- Build monitoring, alerting, runbooks, health checks, or remediation workflows.
- Standardize provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
- Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
- Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
- Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
- Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.
- Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
- Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.
- Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.
- Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
- Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.
- Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.
- Carried real operational responsibility for infrastructure used by many engineers or researchers.
Skills
- Linux systems administration
- Networking
- Processes
- Storage
- Package management
- Shell scripting
- Logs
- Access control
- System debugging
- GPU server operation
- Driver management
- GPU health monitoring
- Node failure diagnostics
- Memory error diagnostics
- Scheduler issue diagnostics
- Hardware diagnostics
- SLURM
- Kubernetes
- Bash
- Python
- Ansible
- Terraform
- Helm
- InfiniBand
- RoCE
- NVLink / NVSwitch
- RDMA
- NCCL
- NFS
- Lustre
- Ceph
- Distributed filesystems
- High-throughput storage systems
- Secure access management
- Identity management
- Permissions management
- SSH
- VPNs
- Bastion hosts
- Secrets management
- Infrastructure security hygiene
- Monitoring
- Alerting
- Runbooks
- Health checks
- Remediation workflows
Location
- San Francisco, California
- Remote in the US
Work Type
- Full-time
Experience Level
- Hands-on experience administering large compute clusters
- Experience operating GPU servers
- Experience with cluster scheduling and resource allocation
- Experience operating GPU compute across providers
- Familiarity with high-performance GPU networking
- Experience with storage for HPC or ML workloads
- Managed GPU or HPC infrastructure
- Operated Kubernetes clusters for ML or GPU workloads at meaningful scale
- Carried real operational responsibility for infrastructure
Education Level
- Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
Salary/Compensations
- $200,000 - $400,000 USD + equity
Benefits
- Generous health, dental, and vision benefits
- 401(k) company match
About the Company
- Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
- Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
Equal Opportunity
- We sponsor visas on a case-by-case basis.
