Staff Platform Engineer, Service Infrastructure at Together AI | CA, US | Rezi

Staff Platform Engineer, Service Infrastructure at Together AI

Staff Platform Engineer, Service Infrastructure

Together AI · CA, US

3 weeks ago

Staff Platform Engineer, Service Infrastructure

Together AI · CA, US

22 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Together AI is hiring a Staff Platform Engineer to drive its service infrastructure strategy within the Product Foundations engineering organization. This role focuses on evolving core infrastructure strategy by understanding team needs, creating reusable patterns for infrastructure problems, and coordinating across platform owners to ensure reliability and consistency.

Responsibilities

  • Own the technical direction for service infrastructure within Product Foundations, including Kubernetes, AWS, Terraform, CDNs, ALBs, DNS, IAM, service networking, and related operational patterns.
  • Up-level existing Product Foundations services by improving reliability, operability, deployment safety, infrastructure consistency, and production readiness.
  • Partner deeply with API Platform and UI Platform on networking, DNS, CDN, load balancing, delivery, and gateway patterns for critical customer-facing interfaces.
  • Work closely with Infrastructure, Networking, and Security teams to bring company-wide platform standards into Product Foundations and contribute PF requirements back into shared frameworks.
  • Help drive cross-company infrastructure initiatives that Product Foundations depend on or help maintain, including Terraform CI/CD, Kubernetes networking, zero-trust service communication, policy-as-code, and cross-DC/provider networking.
  • Build and evolve reusable service infrastructure primitives, including Helm charts, Terraform modules, GitHub Actions/GitOps workflows, service scaffolding, runbooks, and documentation.
  • Establish durable technical standards through design docs, architecture reviews, mentorship, and hands-on implementation that help Together scale services across teams, regions, and cloud environments.

Requirements

  • 7+ years of professional experience in platform engineering, service infrastructure, SRE, distributed systems, cloud infrastructure, or related roles.
  • Deep production experience with Kubernetes, including EKS, Helm, ArgoCD/Argo Rollouts, ingress, autoscaling, secrets, service identity, networking, and progressive delivery.
  • Strong Terraform experience, including module design, infrastructure CI/CD, policy enforcement, production applies, and safe self-service workflows.
  • Experience operating networking and edge infrastructure such as CDNs, ALBs/NLBs, DNS, TLS, ingress/egress controls, and traffic management.
  • Proficiency in one or more programming languages used for infrastructure tooling and automation, such as Go, Python, TypeScript, or similar.
  • AWS experience, ideally including EKS, IAM, VPC networking, load balancing, Route 53, CloudFront, ECR, and related service infrastructure.
  • Direct experience with observability systems, including metrics, logs, traces, dashboards, alerting, SLOs, and incident response.
  • Proven ability to lead cross-functional technical initiatives across product engineering, infrastructure, networking, and security teams.
  • Strong written communication skills, with experience producing clear design docs, migration plans, operational guidance, and technical standards.
  • Staff-level judgment: ability to define ambiguous problems, make pragmatic tradeoffs, influence without authority, and improve systems and teams.

Skills

  • Kubernetes
  • EKS
  • Helm
  • ArgoCD
  • Argo Rollouts
  • Ingress
  • Autoscaling
  • Secrets management
  • Service identity
  • Service networking
  • Progressive delivery
  • Terraform
  • Module design
  • Infrastructure CI/CD
  • Policy enforcement
  • Self-service workflows
  • CDN
  • ALB
  • NLB
  • DNS
  • TLS
  • Ingress controls
  • Egress controls
  • Traffic management
  • Go
  • Python
  • TypeScript
  • AWS
  • IAM
  • VPC networking
  • Load balancing
  • Route 53
  • CloudFront
  • ECR
  • Observability systems
  • Metrics
  • Logs
  • Traces
  • Dashboards
  • Alerting
  • SLOs
  • Incident response
  • Cross-functional leadership
  • Technical writing
  • Design documentation
  • Migration planning
  • Operational guidance
  • Technical standards

Experience Level

  • Staff

Salary/Compensations

  • $240,000 - $280,000

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Other competitive benefits

About the Company

  • Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems through co-design of software, hardware, algorithms, and models.
  • The company has contributed to leading open-source research, models, and datasets, with team members involved in advancements like FlashAttention, Hyena, FlexGen, and RedPajama.
  • They aim to build the next generation of AI infrastructure.

Equal Opportunity

  • Together AI is an Equal Opportunity Employer and offers equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
  • Privacy policy available at https://www.together.ai/privacy.