About the Role
As a Systems Research Engineer Intern specialized in GPU Programming, you will develop and optimize GPU-accelerated kernels and algorithms for ML/AI applications. You will co-design GPU kernels and model architecture with the modeling and algorithm team, and contribute to the co-design of efficient GPU architectures and programming models with hardware and software teams. Your research skills will ensure our AI infrastructure stays at the forefront of innovation.
Responsibilities
- Optimize and fine-tune GPU code for performance and scalability
- Collaborate with cross-functional teams to integrate GPU-accelerated solutions
- Stay up-to-date with the latest advancements in GPU programming techniques and technologies
Requirements
- Strong background in GPU programming and parallel computing (e.g., CUDA, Triton)
- Knowledge of ML/AI applications and models
- Knowledge of performance profiling and optimization tools for GPU programming
- Excellent problem-solving and analytical skills
Skills
- GPU Programming
- Parallel Computing
- CUDA
- Triton
- ML/AI applications
- ML/AI models
- Performance profiling
- Optimization tools
Location
- Remote
Work Type
- Internship
- Full-time
Experience Level
- Internship
Salary/Compensations
- $58 to $63 per hour
Benefits
- Housing stipends
- Competitive benefits
About the Company
- Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems.
- The company aims to lower the cost of modern AI systems by co-designing software, algorithms, and models.
- Together AI has contributed to open-source research, models, and datasets, including FlashAttention, Mamba, FlexGen, Petals, Mixture of Agents, and RedPajama.
Equal Opportunity
- Together AI is an Equal Opportunity Employer and offers equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
