About the Role
Lead a Site Reliability Engineering team responsible for Together AI's production infrastructure, focusing on bare-metal, inference, and virtual clusters platforms. This player-coach role involves 50-60% management and 40-50% hands-on technical work, including coding, architectural discussions, and incident response. A key focus is transitioning the team from reactive, manual operations to systemic, automation-first work. The infrastructure includes bare-metal GPU compute, public-cloud Kubernetes, and Kubernetes-with-virtualization; depth in at least one area is required.
Responsibilities
- Lead and develop a team of approximately 10 SRE engineers across multiple function areas, partnering with technical leads.
- Drive the team's shift from manual operations to systemic, automated, scalable infrastructure, reducing toil.
- Stay hands-on by coding, reviewing architecture, leading incidents, and participating in technical decisions.
- Build coaching and feedback rhythms to develop engineers, focusing on incident leadership, on-call habits, and systemic problem-solving.
- Strengthen on-call practices and incident response, including blameless postmortems with engineering follow-through.
- Partner with other SRE Engineering Managers to shape org-wide practices, hiring, and operational maturity.
- Own hiring, leveling, and career development for engineers in your region.
- Plan capacity, prioritize work across function areas, and represent SRE in broader engineering conversations.
- requirements
Requirements
- Prior experience managing SRE, infrastructure, or platform engineering teams, ideally including leading through reliability or culture turnarounds.
- Deep technical credibility in at least one of: bare-metal infrastructure with Ansible-based config management, Kubernetes on public cloud, or Kubernetes with virtualization.
- Strong Kubernetes and Terraform fundamentals with hands-on production experience.
- Genuine player-coach orientation, contributing as an engineer in addition to management.
- Experience leading teams through serious production incidents and on-call rotations (PagerDuty or equivalent).
- Track record of coaching engineers and shifting team culture through engineering systems.
- Comfort operating in a matrix structure where technical direction is shared with tech leads.
- Adaptability to shifts in management/individual contributor balance based on team needs.
- Based in (or willing to relocate to) San Francisco, with regular in-office presence.
Skills
- SRE management
- Infrastructure management
- Platform engineering management
- Bare-metal infrastructure
- Ansible-based config management
- Kubernetes on public cloud
- Kubernetes with virtualization
- Kubernetes
- Terraform
- Incident leadership
- On-call management
- Coaching
- Systemic problem-solving
- Hiring
- Leveling
- Career development
- Capacity planning
- Prioritization
- Architectural review
- Coding
Location
- San Francisco
Work Type
- Full-time
- Onsite
Experience Level
- Experienced Manager
Salary/Compensations
- US base salary range: $250,000 - $325,000
- Equity
- Individual compensation determined by experience, skills, and job-related knowledge
Benefits
- Competitive compensation
- Startup equity
- Health insurance
- Other competitive benefits
About the Company
- Together AI is a research-driven artificial intelligence company.
- Believes open and transparent AI systems drive innovation and create best outcomes for society.
- Mission: significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
- Contributed to leading open-source research, models, and datasets to advance AI.
- Team behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama.
- Building the next generation AI infrastructure.
Equal Opportunity
- Together AI is an Equal Opportunity Employer and offers equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
