Member of Technical Staff - Inference at Sail Research | CA, US | Rezi

Member of Technical Staff - Inference at Sail Research

Member of Technical Staff - Inference

Sail Research · CA, US

Yesterday

Member of Technical Staff - Inference

Sail Research · CA, US

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

In this role, you'll own token processing down to the lowest layers of the stack. You'll develop new request scheduling strategies, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.

Responsibilities

  • Modify and extend state-of-the-art inference engines like vLLM and SGLang, and work on our own internal engine.
  • Understand every microsecond of GPU time spent during a forward pass.
  • Explain every kernel launch on an nsys profile.
  • Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
  • Write and debug GPU kernels to excel in specific regimes, such as cascade attention.

Requirements

  • Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
  • Interest in MLSys research.
  • Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
  • Great interpersonal and technical communication.
  • Do not use LLMs to write prose.
  • Desk-reject slopful cover letters and resumes.

Skills

  • LLM mechanics
  • KV cache
  • Mixture-of-experts
  • Prefill vs. decode phases
  • MLSys research
  • Speculative decoding
  • Sparse attention
  • Triton
  • CUTLASS
  • ThunderKittens
  • GPU programming
  • Interpersonal communication
  • Technical communication

Location

  • San Francisco

Work Type

  • Onsite

Benefits

  • All meals are on us
  • Studio Display (or two) at their desk
  • Investing in anything that saves us time or energy
  • Six different ways to make coffee or tea in the office

About the Company

  • Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs).
  • Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.