Kernel Engineer (Custom Silicon), Hardware at River AI Inc. | Austin, TX, US | Rezi

Kernel Engineer (Custom Silicon), Hardware at River AI Inc.

Kernel Engineer (Custom Silicon), Hardware

River AI Inc. · Austin, TX, US

1 weeks ago

Kernel Engineer (Custom Silicon), Hardware

River AI Inc. · Austin, TX, US

9 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Kernel Engineer (Custom Silicon), Hardware role.

Rezi rewrites your resume against River AI Inc.'s job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Kernel Engineer (Custom Silicon), Hardware posting at River AI Inc. — free, in seconds.

About the Role

We are seeking exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. You will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture, bridging the gap between high-level compilation and raw hardware capability.

Responsibilities

  • Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code.
  • Author and optimize core deep learning primitives (GEMM/MatMul, Attention mechanisms, Convolutions, and element-wise layers) directly targeted at our custom hardware.
  • Hand-craft and automate instruction scheduling, register allocation, and software pipelining to maximize ALU utilization and hide execution latency on our silicon.
  • Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
  • Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
  • Benchmark generated assembly against hardware simulators and silicon, utilizing hardware performance counters to eliminate performance gaps and ensure mathematical correctness.

Requirements

  • Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming.
  • Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels.
  • Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and complex memory hierarchies (caches, SRAM, HBM/DRAM).
  • Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks.
  • Strong mathematical foundation in linear algebra operations and deep learning primitives.
  • A highly collaborative mindset to push boundaries and co-design effectively with hardware and compiler teams.
  • Deep familiarity with implementing microarchitectural optimizations for Tensor Cores, matrix multiply-accumulate units, or custom vector extensions.
  • Experience utilizing advanced C++ template metaprogramming or code-generation techniques to automate the creation of heavily parameterized kernel variants.
  • Advanced experience with low-level hardware profiling tools, execution tracing, and utilizing performance counters to identify cache misses, pipeline stalls, and ALU bubbles.

Skills

  • C++
  • Code Generation
  • Meta-programming
  • Assembly Code
  • Deep Learning Primitives
  • GEMM
  • MatMul
  • Attention Mechanisms
  • Convolutions
  • Element-wise Layers
  • Instruction Scheduling
  • Register Allocation
  • Software Pipelining
  • Tiling
  • Double-buffering
  • Data Movement
  • Hardware Simulation
  • Performance Profiling
  • Hardware Performance Counters
  • CUDA
  • Triton
  • CUTLASS
  • Computer Architecture
  • Vector Units
  • Execution Pipelines
  • Register Files
  • Memory Hierarchies
  • Linear Algebra
  • Template Metaprogramming

Location

  • Austin, Texas
  • Palo Alto, California

Work Type

  • Onsite

Experience Level

  • 5+ years of practical industry experience

Education Level

  • Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field

Salary/Compensations

  • $200,000 - $420,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Relocation support

About the Company

  • At River, our mission is to create personal AI owned and shaped by each individual.
  • We are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.
  • We are scientists, engineers, and builders from the industry's top tech companies and AI labs.
  • We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

Equal Opportunity

  • We sponsor visas. We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.