Software Engineer, AI Kernels & Performance Optimization — MTIA Software at Meta | Bellevue, WA, US | Rezi

Software Engineer, AI Kernels & Performance Optimization — MTIA Software at Meta

Software Engineer, AI Kernels & Performance Optimization — MTIA Software

Meta · Bellevue, WA, US

2 weeks ago

Software Engineer, AI Kernels & Performance Optimization — MTIA Software

Meta · Bellevue, WA, US

17 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Meta designs and deploys its own AI systems. MTIA is Meta's family of in-house AI accelerator ASICs, running recommendation and ranking workloads in production across Meta's data centers today and expanding into generative AI inference and training. The MTIA Software team builds the entire stack a chip vendor would normally supply, co-designing it with silicon teams. The AI kernel and optimization software development team drives the layer where architecture meets arithmetic, focusing on performance and programmability at scale. We are hiring an experienced kernel and performance engineer to take on this work at a senior level, owning the performance of workloads that serve billions of people.

Responsibilities

  • Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators, taking ownership from architectural analysis through production deployment.
  • Profile and root-cause performance across the full stack and drive the fixes to the right layer.
  • Build and extend kernel authoring frameworks, templates, and libraries.
  • Deliver and maintain broad kernel coverage for PyTorch operators across recommendation, ranking, and generative AI workloads.
  • Partner with silicon architecture and design teams on hardware/software co-design.
  • Work with compiler, runtime, framework, and product-facing teams to land end-to-end wins on production models.
  • Investigate numerics and precision trade-offs, and design software mitigations that recover performance or accuracy lost to hardware limitations.
  • Set technical direction for a kernel domain, write design documents, and mentor engineers on performance methodology and accelerator programming.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience.
  • 6+ years of professional experience in high-performance computing, accelerator kernel development, compiler backends, or systems performance engineering.
  • Proficiency in C++ and Python, including low-level systems programming, templates and generic programming, and comfort reading and writing performance-critical code.
  • Demonstrated experience writing and optimizing kernels for a parallel architecture.
  • Working knowledge of computer architecture as it applies to performance.
  • A rigorous, measurement-driven approach to performance: the ability to build a roofline or analytical model, profile against it, and explain the residual gap.

Skills

  • C++
  • Python
  • Low-level systems programming
  • Templates and generic programming
  • Performance-critical code
  • Kernel development
  • Compiler backends
  • Systems performance engineering
  • Parallel architecture kernel optimization
  • Computer architecture
  • Memory hierarchies and bandwidth
  • Latency hiding
  • Occupancy and scheduling
  • Vectorization
  • Synchronization
  • Roofline modeling
  • Profiling
  • Mentoring engineers
  • Technical direction
  • Distributed execution
  • Collective communication
  • Low-precision numerics
  • Quantization
  • MLIR
  • LLVM
  • TVM
  • XLA
  • Halide
  • Polyhedral scheduling
  • Transformer kernel design
  • Attention kernel design
  • FlashAttention-class algorithms
  • KV-cache management
  • Paged and chunked attention
  • Linear and state-space attention variants
  • MoE routing and expert dispatch
  • Open-source contribution
  • Pre-silicon software development
  • Architectural simulators
  • FPGA emulation
  • Performance modeling
  • Hardware/software co-design
  • PyTorch internals
  • Torch.compile / Inductor
  • Custom operator integration
  • Inference serving stacks
  • vLLM
  • SGLang
  • High-performance kernel libraries
  • Frameworks
  • CUTLASS
  • cuBLAS
  • cuDNN
  • CUTE
  • Triton
  • Helion
  • ThunderKittens
  • oneDNN
  • Composable Kernel
  • AI skill development
  • Prompt/context engineering
  • Agent orchestration
  • AI tools integration

Experience Level

  • Senior level
  • 6+ years of professional experience
  • 8+ years of experience in accelerator software, HPC, or ML systems performance (or equivalent with an advanced degree)

Education Level

  • Bachelor's degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience.

Salary/Compensations

  • $183,997/year to $257,000/year

Benefits

  • bonus
  • equity
  • benefits

About the Company

  • Meta designs and deploys its own AI systems.
  • MTIA — the Meta Training and Inference Accelerator — is Meta's family of in-house AI accelerator ASICs, running recommendation and ranking workloads in production across Meta's data centers today and expanding into generative AI inference and training as successive silicon generations land.
  • The MTIA Software team is part of the AI & Compute Foundation (ACF) organization within Meta Infrastructure.
  • Because the hardware is ours, the software is ours too: we build the entire stack a chip vendor would normally supply — compiler and LLVM toolchain, runtime, kernel authoring frameworks and libraries, developer tooling, and deep PyTorch integration — and we co-design it with the silicon teams generation over generation.
  • Within that stack, the AI kernel and optimization software development team drives the layer where architecture meets arithmetic.
  • Our mission is performance and programmability at scale: hit roofline enablements on the workloads that matter, and make kernel authoring accessible enough that the whole organization can close coverage gaps without funneling every problem through a handful of experts.
  • We do this by shipping high-performance kernel libraries with broad PyTorch operator coverage, by building the C++ and Python kernel authoring frameworks and DSL surfaces that others build on, and by writing production kernels against new architectures long before first silicon — turning hardware proposals into measured roofline evidence while the design can still change.