About the Role
Lead the integration of diverse AI models including VLA, Vision, and Multimodal architectures by utilizing a proprietary kernel programming language to ensure accuracy and performance while maintaining a developer-ready stack.
Responsibilities
- Design and implement efficient kernels on FuriosaAI’s kernel programming stack targeting Tensor Contract Processor architectures.
- Diagnose and optimize kernel performance using profiling tools and roofline analysis for RNGD-accelerated AI models.
- Develop and apply automated kernel generation and optimization for AI workloads.
- Build diagnostic tools and testbeds for robust kernel validation.
- Drive end-to-end programming enablement on RNGDs by creating reproducible guides and reference implementations.
Requirements
- BS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.
- Experience in low-level systems programming targeting XPU architectures.
- Experience collaborating across engineering, research, and product teams.
- MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.
- Experience in optimizing high-performance kernels on AI accelerators for AI products.
- Understanding of XPU architecture and software-hardware co-optimization strategies.
- Experience in open-source or research projects on AI model architectures such as Diffusion, Mamba, and VLA.
- Experience in designing efficient deep learning architectures and developing algorithms for AI applications.
Skills
- Kernel programming
- vISA
- TCL
- Tensor Contract Processor architecture
- Performance profiling
- Roofline analysis
- Automated kernel generation
- XPU architecture
- Software-hardware co-optimization
- Deep learning architecture design
Education Level
- BS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field
- MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field
