About the Role
In this role, you'll own token processing down to the lowest layers of the stack. You'll develop new request scheduling strategies, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.
Responsibilities
- Modify and extend state-of-the-art inference engines like vLLM and SGLang, and work on our own internal engine.
- Understand every microsecond of GPU time spent during a forward pass.
- Explain every kernel launch on an nsys profile.
- Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
- Write and debug GPU kernels to excel in specific regimes, such as cascade attention.
Requirements
- Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
- Interest in MLSys research.
- Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
- Great interpersonal and technical communication.
- Do not use LLMs to write prose.
- Desk-reject slopful cover letters and resumes.
Skills
- LLM mechanics
- KV cache
- Mixture-of-experts
- Prefill vs. decode phases
- MLSys research
- Speculative decoding
- Sparse attention
- Triton
- CUTLASS
- ThunderKittens
- GPU programming
- Interpersonal communication
- Technical communication
Location
- San Francisco
Work Type
- Onsite
Benefits
- All meals are on us
- Studio Display (or two) at their desk
- Investing in anything that saves us time or energy
- Six different ways to make coffee or tea in the office
About the Company
- Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs).
- Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.
