About the Role
We are seeking an Engineering Manager to lead our GPU Kernel Engineering team. This player-coach role requires hands-on experience in kernel development and the ability to lead a team of elite GPU engineers. You will own the technical direction for a team working at the intersection of GPU architecture, ML systems, and production inference, optimizing models and reducing latency and cost for AI companies.
Responsibilities
- Lead, grow, and mentor a team of GPU kernel engineers, owning hiring, performance, and career development.
- Set technical direction for the kernel roadmap, balancing short-term inference wins with long-term architectural investments.
- Partner closely with the Chief Scientist, VP Engineering, and peer engineering leads to align kernel work with Baseten's broader inference stack strategy.
- Drive cross-functional collaboration between the kernel team and Model Performance, Capacity, and Infrastructure teams.
- Establish and maintain a high technical bar for kernel quality, performance, and correctness.
- Review kernel designs and implementations, providing feedback on GPU architecture decisions, memory hierarchy tradeoffs, and optimization strategies.
- Guide the team's approach to profiling and bottleneck identification using tools like Nsight Systems, Nsight Compute, and Torch Profiler.
- Stay current on the NVIDIA GPU ecosystem and translate architectural advancements into team priorities.
- Build processes for a highly technical, distributed team to ship with velocity and rigor.
- Represent the kernel team's work to senior leadership and external audiences.
- Contribute to Baseten's open-source GPU library presence and technical brand.
Requirements
- Proven experience leading a team of GPU or ML systems engineers, with a track record of hiring and developing strong technical talent.
- Deep personal background in GPU kernel engineering, with experience writing and shipping production CUDA kernels.
- Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology.
- Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem.
- Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership.
Skills
- CUDA kernel development
- GPU architecture
- ML systems
- Production inference
- NVIDIA GPU ecosystem
- Nsight Systems
- Nsight Compute
- Torch Profiler
- Triton
- CUTLASS
- CuTe DSL
- LLM inference kernels
- Attention variants
- GEMMs
- Quantization (FP8/FP4)
- MoE routing
Location
- Remote
Work Type
- Full-time
Experience Level
- Manager
- Senior
Benefits
- Competitive compensation
- Meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy
- Company wide Winter Break (offices closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
About the Company
- Baseten powers mission-critical inference for dynamic AI companies.
- We unite applied AI research, flexible infrastructure, and seamless developer tooling.
- We enable companies operating at the frontier of AI to bring cutting-edge models into production.
- We recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital.
- Join us and help build the platform engineers turn to to ship AI products.
Equal Opportunity
- Baseten is committed to fostering a diverse and inclusive workplace.
- We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
- We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law.
