Member of Technical Staff, GPU Compiler at San Francisco Tensor Company | CA, US | Rezi

Member of Technical Staff, GPU Compiler at San Francisco Tensor Company

Member of Technical Staff, GPU Compiler

San Francisco Tensor Company · CA, US

Today

Member of Technical Staff, GPU Compiler

San Francisco Tensor Company · CA, US

4 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We build the fastest GPU compiler in the world. This role involves building the machine that performs the search for the fastest possible code transformations, working on the entire pipeline from ingestion to backend code generation and executable binaries. You will have access to unique tooling due to our ownership of the stack down to the ISA, enabling the creation of advanced kernels.

Responsibilities

  • Build and extend MLIR dialects and passes to optimize training and inference workloads.
  • Work in our LLVM backend on instruction selection, scheduling, register allocation, and direct cubin emission, and establish equivalent capabilities on AMD, TPU, and Trainium targets.
  • Expand search-based compiler infrastructure, including agent- and RL-driven program search and formal correctness proofs.
  • Implement classic compiler optimizations tuned for large-scale training.
  • Create hybrid codegen paths for cases where direct MLIR lowering is not practical.
  • Own testing, benchmarking, and performance regression systems, including bit-identical hardware models.
  • Collaborate closely with the research team and customer workloads to identify optimization opportunities.

Requirements

  • Deep experience in compiler infrastructure (LLVM, MLIR, or similar).
  • Strong background in GPU architecture and low-level optimization (CUDA, ROCm, or similar).
  • Hands-on experience with at least one of: PTX/SASS, GCN/RDNA assembly, or other GPU ISAs.
  • Familiarity with ML compiler stacks (XLA, TVM, Triton, torch.compiler, or similar).
  • Solid systems programming skills in C++ and/or Rust.
  • Proven track record of building production-grade compiler infrastructure.

Skills

  • Compiler infrastructure
  • GPU architecture
  • Low-level optimization
  • CUDA
  • ROCm
  • PTX/SASS
  • GCN/RDNA assembly
  • GPU ISAs
  • ML compiler stacks
  • XLA
  • TVM
  • Triton
  • torch.compiler
  • C++
  • Rust
  • Autotuning
  • Search-based optimization
  • Formal verification
  • Proof assistants
  • SMT solvers
  • LLVM backend development
  • Distributed systems
  • Multi-device compilation
  • Open-source compiler contributions
  • Large-scale training infrastructure
  • StableHLO

Location

  • San Francisco

Work Type

  • Full-time
  • Onsite

Experience Level

  • Member of Technical Staff

Salary/Compensations

  • $285,000-$315,000

Benefits

  • Meaningful equity

About the Company

  • SF Tensor is building the future of high-performance compute for AI by rethinking and rebuilding the compute stack from hardware to cloud. They aim to make compute faster, cheaper, and more available, enabling portability across different clouds and chips.
  • Backed by Susa Ventures, Y Combinator, and other notable investors and executives.
  • The team has experience in pre-training foundation models on large GPU clusters, designing GPU clusters, and building supercomputers.