Member of Technical Staff - Platform at ai& | JP | Rezi

Member of Technical Staff - Platform at ai&

Member of Technical Staff - Platform

ai& · JP

1 weeks ago

Member of Technical Staff - Platform

ai& · JP

8 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

The Platform team turns raw heterogeneous compute into a serving platform. ai& owns its data centers and runs AMD, NVIDIA, and Tenstorrent silicon side by side. Your job is everything between the bare metal and the inference engines: cluster orchestration, node lifecycle, scaling, networking, observability, and reliability. The platform must let a small team operate hundreds of nodes across multiple sites without heroics, and it must scale by an order of magnitude over the next two years as new sites come online. You will work directly with the inference team, which owns the engines and serving gateway, and the data center team, which owns power, cooling, and physical deployment. You own the layer that makes their work composable.

Responsibilities

  • Run Kubernetes across GPU clusters in ai&-owned data centers.
  • Own node lifecycle from bring-up and burn-in through drain and repair, across multiple accelerator vendors.
  • Build the capacity and scheduling machinery that places inference workloads across heterogeneous silicon and multiple sites.
  • Enable bringing a new site from empty racks to serving traffic on a predictable timeline.
  • Define and hold SLOs for the platform.
  • Build the observability stack (metrics, logs, tracing, alerting) and failure isolation.
  • Operate high-bandwidth fabrics for multi-node inference.
  • Solve model weight distribution for shipping models.
  • Own CI/CD and GitOps for the fleet.
  • Implement infrastructure as code, reproducible node images, and safe rollouts.

Requirements

  • Operated large multi-cluster Kubernetes environments, ideally with GPU scheduling, device plugins, and topology-aware placement.
  • Strong Linux fundamentals; ability to reason about NUMA, PCIe, NICs, and storage.
  • Ability to debug systems from symptoms to root cause.
  • Understanding of L2/L3 networking, and ideally RDMA fabrics (InfiniBand or RoCE) in production.
  • Experience with Terraform or equivalent infrastructure as code tools.
  • Experience with GitOps workflows.
  • Experience carrying a pager for critical systems.
  • Experience with Go or Python.
  • Experience with Prometheus-family observability tools.
  • Comfort automating repetitive tasks.
  • Mission-driven approach to engineering.
  • Clear communication, hands-on execution, and valuing collective success.

Skills

  • Production Kubernetes at scale
  • Systems depth
  • Networking fundamentals
  • Infrastructure as code
  • Ownership under load
  • Go or Python
  • Prometheus-family observability
  • Team spirit

Location

  • Tokyo
  • SF
  • Austin
  • Toronto
  • Worldwide

Work Type

  • Full-time

About the Company

  • ai& is a new global AI technology company dedicated to meeting the world's growing demand for AI.
  • Vision: to serve as a premier AI lab specializing in localization, and to act as a global infrastructure and compute provider.
  • Building a unified, optimized global platform that integrates next-generation data centers and infrastructure, heterogeneous compute serving, and advanced model services.
  • Believes that the most effective way to build and scale AI is to own the stack from top to bottom.
  • Empowers small teams with the autonomy needed to tackle significant challenges.
  • Approach: deconstruct large problems into manageable components and solve complex issues collaboratively.
  • Seeks highly motivated, mission-driven individuals who demonstrate strong personal agency.
  • Values curiosity as the foundation of talent.
  • Looking for people eager to develop alongside evolving technology and expanding business.