About the Role
Meta is seeking a Software Engineer for the MTIA Software Tooling team, responsible for developing and maintaining the tooling ecosystem for Meta's in-house AI accelerator ASICs. This role involves designing and building developer tools for debugging, profiling, and monitoring AI workloads on MTIA hardware, working at the intersection of compilers, runtime, hardware, and ML frameworks.
Responsibilities
- Own the technical vision and roadmap for key areas of MTIA's developer tooling ecosystem, focusing on debugging, workload error analysis, and product / fleet reliability
- Design and develop debugging tools — including graph-mode debugging, kernel-level diagnostics, and multi-rank fault isolation
- Contribute to overall MTIA SW tooling infrastructure — enabling profiling, performance debugging, memory sanitization, and reliability analysis for MTIA training and inference workloads
- Collaborate closely with MTIA compiler, runtime, kernel, and hardware teams to integrate tooling hooks throughout the MTIA software stack
- Drive AI-native tooling approaches by leveraging automation and LLM-guided diagnostics to improve developer productivity and reduce time-to-root-cause
- Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and prioritize tooling investments
- Advise on tooling best practices, debugging methodologies, and systems-level analysis for accelerator software; communicate architectural decisions clearly through design documents and cross-team reviews
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 6+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field
- Experience building debugging, profiling, or diagnostic tools for complex software/hardware systems
- Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation
- Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware)
- Experience leading the technical design and delivery of tooling or infrastructure projects from inception through production deployment
- Experience using data-driven methods and experimentation to evaluate and validate tooling effectiveness and systems performance improvements
- Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs) including performance profiling, memory analysis, and runtime debugging using their toolchains (cuda-gdb, nsight-compute, nsight-systems, cuda-memcheck)
- Demonstrated cross-stack debugging ability, including Linux kernel and driver-level debugging, with capacity to trace issues across application, OS, and hardware boundaries
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience with distributed systems debugging or profiling (multi-device, multi-node/multi-rank)
- Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF, ftrace, coredump analysis, hardware performance counters) and familiarity with binary formats and debugging metadata (ELF/DWARF)
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- 7+ years of experience in systems software, developer tooling, or accelerator software development (or equivalent with advanced degree)
- Track record of building developer tools adopted by large engineering populations, ideally with contributions to open-source projects (gdb, LLVM sanitizers, Valgrind, Triton, etc.)
Skills
- C++
- Python
- Low-level systems programming
- Scripting for tool automation
- Systems software engineering
- Performance engineering
- Developer tooling
- Debugging tools
- Profiling tools
- Diagnostic tools
- Compiler
- Runtime
- Hardware
- ML frameworks
- ML framework internals
- AI compiler stacks
- Accelerator ecosystems
- Performance profiling
- Memory analysis
- Runtime debugging
- Cross-stack debugging
- Linux kernel debugging
- Driver-level debugging
- Distributed systems debugging
- Distributed systems profiling
- Linux debugging
- Linux profiling infrastructure
- Binary formats
- Debugging metadata
- Responsible AI practices
- Ethical AI practices
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
- 6+ years
- 7+ years
Education Level
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
Salary/Compensations
- $183,997/year to $257,000/year
Benefits
- bonus
- equity
- benefits
About the Company
- Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem for Meta's in-house AI accelerator ASICs. The Tooling team provides debugging, profiling, memory analysis, and monitoring capabilities for the whole MTIA Ecosystem, advancing ML accelerator tooling by leveraging Meta's full-stack ownership from silicon specs to fleet observability.
