Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Staff Software Engineer - AI Compute, Together Cloud role.
Rezi rewrites your resume against Together AI's job description. Free.

Tailor your resume to this Staff Software Engineer - AI Compute, Together Cloud role.
Rezi rewrites your resume against Together AI's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Staff Software Engineer - AI Compute, Together Cloud posting at Together AI — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Staff Software Engineer - AI Compute, Together Cloud posting at Together AI — free, in seconds.
About the Role
Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle. The Together Cloud team builds the flagship IaaS product providing high-performance, AI-ready GPU clusters and the virtual infrastructure layer powering Together's products. As a Staff Software Engineer focusing on AI Compute, you will set technical direction and build major components of the next generation AI cloud platform, a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware. This is an architect-and-build role, defining architecture for automated data center bootstrapping, high-performance virtualization, and fault-tolerant control planes, and owning the hardest parts in code and design. Your designs will span the IaaS layer up to the global management plane, setting standards and growing engineers to accelerate shipping.
Responsibilities
- Own the GPU and network virtualization stack, including hypervisor, kernel, and SDN work for performance, portability, and isolation.
- Own the in-DC IaaS layer, architecting and roadmapping services, Kubernetes operators, and libraries for provisioning and managing compute, storage, and networks.
- Lead the build-out of the IaaS layer for a new data center, from hardware bring-up to customer-facing API.
- Design the GPU scheduling and global management plane, the distributed control plane for on-demand and reserved clusters.
- Architect monitoring and automated remediation for fault tolerance to ensure continuous operation through hardware failures.
- Set technical direction across teams by leading design reviews, resolving architectural disagreements, and defining standards.
- Grow the team by mentoring engineers and deepening expertise in virtualization, DC networking, and GPU infrastructure.
- Set the engineering bar by creating testing frameworks, tools, and developer documentation.
- Drive ambiguous work from definition to production.
Requirements
- 7+ years of professional software development experience.
- Expert-level proficiency in at least one backend language (Golang desired).
- Experience writing high-performance, well-tested, production-quality code.
- Track record of owning the architecture of large distributed systems from inception to production at scale.
- Deep experience building and operating globally distributed, high-performance microservice architectures across cloud providers (AWS, Azure, GCP).
- Expert systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and global scale.
- Demonstrated technical leadership beyond personal contributions, including mentoring senior engineers and leading design reviews.
- Excellent communication and diplomacy skills for writing design docs and collaborating with stakeholders.
- Experience building and operating reliable, customer-facing production systems at scale.
- Experience owning infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD).
Skills
- Golang
- Distributed systems architecture
- Microservice architectures
- Cloud computing (AWS, Azure, GCP)
- Compute, networking, and storage systems
- Concurrency
- Memory management
- Performant I/O
- Infrastructure automation (Terraform, Ansible)
- Observability (Prometheus, Grafana)
- CI/CD (GitHub Actions, ArgoCD)
- Kubernetes internals
- VMs/hypervisors (QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV)
- DC networking (VLAN, VXLAN, VPN, VPC, OVS/OVN)
- Cluster API
- High-performance compute, networking, and/or storage
- GPU virtualization
- InfiniBand virtualization
- IaaS or PaaS systems
- DPUs/SmartNICs
- GPU programming
- NCCL
- CUDA
Location
- Remote
Work Type
- Full-time
Experience Level
- Staff
Salary/Compensations
- $260,000 - $300,000
Benefits
- Startup equity
- Health insurance
- Flexibility in terms of remote work
About the Company
- Together AI is a research-driven artificial intelligence company focused on open and transparent AI systems.
- The company's mission is to lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
- Together AI has contributed to leading open-source research, models, and datasets.
- The team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama.
- The company is building the next generation AI infrastructure.
Equal Opportunity
- Together AI is an Equal Opportunity Employer and offers equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.