Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Senior AI Infrastructure Engineer, Kubernetes role.
Rezi rewrites your resume against Firmus Technologies's job description. Free.

Tailor your resume to this Senior AI Infrastructure Engineer, Kubernetes role.
Rezi rewrites your resume against Firmus Technologies's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Senior AI Infrastructure Engineer, Kubernetes posting at Firmus Technologies — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Senior AI Infrastructure Engineer, Kubernetes posting at Firmus Technologies — free, in seconds.
About the Role
The Senior Kubernetes Engineer, AI Infrastructure will own the technical design and delivery of the backend infrastructure powering the Firmus Kubernetes platform. This hands-on role is responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities in GPU-accelerated bare-metal environments. You will solve complex platform engineering problems, set Kubernetes engineering standards, and provide technical sign-off for platform designs, collaborating across teams to create a secure, resilient, multi-tenant platform.
Responsibilities
- Define and own the Kubernetes platform reference architecture across management and workload clusters.
- Build and maintain backend services, APIs, controllers, operators, and automation for Kubernetes cluster provisioning, configuration, upgrade, scaling, and retirement.
- Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies.
- Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh.
- Define persistent storage and data service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster recovery mechanisms.
- Integrate and productionize NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement.
- Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management.
- Build platform security into the architecture through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management.
- Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance.
- Lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.
- Set engineering standards, design patterns, review practices, and operational readiness criteria.
- Mentor senior engineers and resolve cross-team technical decisions while remaining directly involved in implementation.
Requirements
- 7+ years of progressive infrastructure, systems, or platform engineering experience.
- Substantial ownership of production Kubernetes platforms.
- At least 3 years operating at senior staff, principal, or equivalent level.
- Deep knowledge of Kubernetes internals.
- Demonstrated experience designing, building, and operating highly available, large scale and multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
- Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills.
- Expert Linux systems knowledge.
- Strong Kubernetes networking expertise.
- Strong infrastructure automation and GitOps experience.
- Practical experience with Kubernetes security and governance.
- Experience implementing production observability.
- Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements.
- Experience with distributed storage and data services.
- CKA-level expertise is expected.
- Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
- Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.
Skills
- Kubernetes
- Go
- Rust
- Python
- Bash
- Linux
- CNI
- CSI
- GitOps
- Terraform
- Ansible
- Argo CD
- Flux
- GitHub Actions
- GitLab CI
- Jenkins
- RBAC
- OPA Gatekeeper
- Kyverno
- Prometheus
- Grafana
- OpenTelemetry
- Loki
- Elasticsearch
- NVIDIA GPU Operator
- RDMA
- Ceph
Location
- San Francisco Bay Area
Work Type
- Full-time
Experience Level
- Senior
- Principal
Education Level
- Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent practical engineering experience.
About the Company
- Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
- Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
- We design, build, and operate a new class of digital infrastructure – the AI Factory.
- Our model-to-grid technology approach pushes the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction.
- Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale.
- It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings.
- We are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Equal Opportunity
- At Firmus, we are committed to building a diverse and inclusive workplace.
- We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
- Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering.