About the Role
We are building the fastest GPU compiler in the world, which requires running an enormous amount of untrusted, freshly generated kernels on real silicon quickly and safely. This role will build the serverless GPU container service that makes this possible across multiple vendors and accelerator types at an unprecedented scale and fidelity.
Responsibilities
- Extend our sandboxing stack to new vendors and accelerators
- Build and maintain GPU virtualization below the runtime, including gVisor work at the driver and ioctl level
- Make sandboxes first-class citizens on spot capacity, including preemption-aware scheduling, checkpointing, and rescheduling
- Support multi-GPU and multi-node sandboxes, including interconnect and RDMA paths
- Own live migration end-to-end, including socket-preserving migration
- Guarantee measurement and profiling fidelity and isolation
- Work directly with compiler, post-training, and kernel teams to ensure throughput is not capped by sandboxes
Requirements
- Strong low-level systems engineering background: Linux kernel internals, containers, namespaces, cgroups, syscall interception or hypervisors
- Experience in GPU systems engineering: drivers, runtimes, or scheduling on accelerator fleets
- Comfortable with distributed systems failure modes: preemption, partial failure, checkpoint/restore, and dealing with states that cannot be lost
- Proficient in Go, C/C++, or Rust
- Strong bias toward building solutions when vendor support is lacking
Skills
- Go
- C/C++
- Rust
- Linux kernel internals
- Containers
- Namespaces
- cgroups
- Syscall interception
- Hypervisors
- GPU systems engineering
- GPU drivers
- GPU runtimes
- Accelerator fleet scheduling
- Distributed systems
- Preemption handling
- Checkpoint/restore
- gVisor
- Firecracker
- Kata
- QEMU/KVM
- CRIU
- Live migration
- Connection-preserving failover
- NCCL/RCCL
- RDMA
- InfiniBand
- Vendor interconnects
- Spot capacity management
- Bare-metal provisioning
- Fleet management
- Security
- Isolation boundaries
- Untrusted code execution
Location
- San Francisco
Work Type
- Full-time
- Onsite
Experience Level
- Member of Technical Staff
Salary/Compensations
- $275,000-$315,000
Benefits
- Meaningful equity
About the Company
- SF Tensor is building the future of high-performance compute for AI by rethinking and rebuilding the stack from hardware to cloud. They aim to make compute faster, cheaper, and more available, enabling portability across different clouds and chips.
- Backed by Susa Ventures, Y Combinator, and other notable investors and individuals.
- The team has experience in pre-training foundation models on large GPU clusters, designing NVL72 clusters, and designing a TOP500 supercomputer.
