Member of Technical Staff, GPU Kernels at San Francisco Tensor Company | CA, US | Rezi

Member of Technical Staff, GPU Kernels at San Francisco Tensor Company

Member of Technical Staff, GPU Kernels

San Francisco Tensor Company · CA, US

Today

Member of Technical Staff, GPU Kernels

San Francisco Tensor Company · CA, US

4 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are building the fastest GPU compiler in the world, capable of searching a wider space of optimizations by proving correctness at the end. This role focuses on establishing and pushing the hardware's capabilities before the search phase, optimizing kernels for various hardware and workloads.

Responsibilities

  • Write and hand-optimize kernels for real workloads to establish performance ceilings.
  • Profile at the microarchitectural level, analyzing SM and CU utilization, warp stalls, memory bank conflicts, register pressure, and instruction throughput.
  • Debug issues down to clock behavior, thermal throttling, and driver paths.
  • Translate hard-won knowledge into machine-searchable structures for compiler exploration.
  • Work at the ISA-level below PTX, reasoning about SASS and cubins to emit schedules not expressible in PTX.
  • Build performance models, microbenchmarks, and tooling to predict kernel behavior.
  • Collaborate with the formal correctness team to ensure aggressive kernels ship with formal proofs.
  • Dedicate approximately two-thirds of time to kernel development and one-third to translating learnings into the compiler search space.

Requirements

  • Track record of hand-writing kernels that match or beat vendor libraries.
  • Comfortable reading PTX, SASS, GCN/CDNA ISA, or equivalent machine-level assembly.
  • Fluent with low-level profiling tools such as Nsight Compute, Nsight Systems, rocprof, or omniperf.
  • Solid systems programming skills in C++ and CUDA or ROCm/HIP.
  • Working understanding of how high-level ML operations map onto hardware, including framework layer impacts.

Skills

  • GPU Kernel Engineering
  • GPU Compiler Optimization
  • Microarchitectural Profiling
  • Low-level Debugging
  • ISA-level Programming
  • Performance Modeling
  • C++
  • CUDA
  • ROCm/HIP
  • MLIR
  • LLVM
  • Codegen
  • Instruction Selection
  • Scheduling
  • Superoptimization
  • Program Synthesis
  • Formal Verification
  • Search-based Compilation
  • Distributed AI Training
  • High-speed Interconnects
  • HPC
  • Driver Development
  • Computer Architecture
  • Hardware Design

Location

  • San Francisco

Work Type

  • Full-time
  • Onsite

Experience Level

  • Member of Technical Staff

Salary/Compensations

  • $285,000-$315,000

Benefits

  • Meaningful equity
  • Relocation assistance

About the Company

  • SF Tensor is building the future of high-performance compute for AI by optimizing the entire stack from hardware to cloud.
  • We are developing a Kernel Optimizer and a Model Foundry to make compute faster, cheaper, and more available.
  • Backed by Susa Ventures, Y Combinator, and other notable investors and industry leaders.
  • Our team has experience in pre-training foundation models, designing GPU clusters, and building supercomputers.