About the Role
The Model Shaping team at Together AI focuses on products and research for tailoring open foundation models to downstream applications. This role involves building services for ML developers to select and improve models using domain-specific data, and developing efficient model training and evaluation methods.
Responsibilities
- Design, implement, and optimize core components of Together's large-scale training infrastructure.
- Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
- Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
- Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
- Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
- Rapidly enable support for newly released open-source foundation models on the Together platform.
- Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.
Requirements
- Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
- Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
- Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
- Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
- Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
- Passion for staying current with advances in AI research and applying them to real-world systems.
- Excitement about translating cutting-edge research into production systems that deliver customer impact.
Skills
- Python
- PyTorch
- ML systems fundamentals
- GPU architecture
- Mixed-precision training
- Distributed training paradigms
- Data parallelism
- Tensor parallelism
- Pipeline parallelism
- Expert parallelism
- CUDA
- Triton
- NCCL
- NVSHMEM
- FSDP
- DeepSpeed
- Megatron-LM
Work Type
- full-time
Salary/Compensations
- $200,000 - $290,000
Benefits
- Competitive compensation
- Startup equity
- Health insurance
About the Company
- Together AI is a research-driven artificial intelligence company.
- We believe open and transparent AI systems will drive innovation and create the best outcomes for society.
- Our mission is to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
- We have contributed to leading open-source research, models, and datasets to advance the frontier of AI.
- Our team has been behind technological advancement such as FlashAttention, ATLAS, RedPajama, and Mamba.
- Join a passionate group of researchers in our journey in building the next generation AI infrastructure.
Equal Opportunity
- Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
