About the Role
Build the systems that make AI inference fast, reliable, and cost-efficient at global scale. You’ll design the control plane that schedules and autoscales our enormous queue of tokens over a diverse fleet of machines, spread all over the world.
Responsibilities
- Design and implement high-performance schedulers (admission control, queuing, priority, fairness, preemption, bin packing).
- Build global routing and traffic management (latency-aware dispatch, predictive autoscaling, failover strategies, cache-aware routing).
- LLM-specific routing optimizations, e.g. KV caching that lets us trade memory for compute, across the cascading layers of GPU RAM, CPU RAM, and NVMe flash.
- Build deep observability: trace every millisecond of our systems, and catch failures early enough that we can make things right before customers even notice.
Requirements
- Strong distributed systems fundamentals (concurrency, networking, databases, queues, performance engineering).
- Eagerness to work with agents coupled with a healthy dose of skepticism.
- Focus attention on the highest-level design decisions first, and debate and justify each decision.
- Use finite review time on the critical pieces of the system that must be built with good mental models to do good work.
Skills
- ML inference stacks (vLLM/SGLang)
- GPUs/accelerators
- High-RPS systems (e.g. trading, messaging)
Location
- San Francisco
Work Type
- Onsite
Benefits
- All meals are on us
- Studio Display at each desk
- Six different ways to make coffee or tea in the office
About the Company
- Sail is the foundation of useful, agentic AI.
- We are here to take a big swing at the most ambitious engineering challenge of our careers.
- Everyone working at Sail will become an expert; nothing less will do in our immensely competitive market.
