Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Senior Kubernetes Platform Engineer role.
Rezi rewrites your resume against Firmus Technologies's job description. Free.

Tailor your resume to this Senior Kubernetes Platform Engineer role.
Rezi rewrites your resume against Firmus Technologies's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Senior Kubernetes Platform Engineer posting at Firmus Technologies — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Senior Kubernetes Platform Engineer posting at Firmus Technologies — free, in seconds.
About the Role
AI FactoryOS Operations runs the proprietary operating system for the AI Factory, governing GPU telemetry, cooling, power, and grid interaction. This role ensures the 24/7 reliability of AI FactoryOS, Firmus AI Cloud, and their associated platforms, meeting defined service levels. The function focuses on engineering, building automated remediation and operational tooling to transform manual responses into software-defined capabilities, and running shared services essential for the estate's operation. It collaborates with engineering teams to implement permanent fixes and design improvements based on real-world performance and failure analysis.
Responsibilities
- Responsible for the reliable operation, automation, and continuous improvement of the multi-tenant Kubernetes platform across all Firmus sites.
- Build operational tooling, guarded remediation, and orchestration to automate fleet-scale operations.
- Contribute operator and controller requirements, and code where agreed, to AI Infrastructure's platform backlog with supporting production evidence.
- Execute Kubernetes cluster lifecycle across the fleet, including provisioning, patching, upgrades, and decommissioning.
- Run deployment and upgrade mechanisms through agreed staged or canary paths, managing estate version compliance and retirement coordination.
- Operate and recover Kubernetes control planes supporting live tenant workloads, including etcd state, certificate rotation, failed upgrades, and corrupted resources.
- Operate the virtual cluster platform and multi-tenant isolation patterns, and the automated tenant onboarding and release pipeline.
- Operate the GPU integration layer for Kubernetes, including device plugins, GPU scheduling, and driver coordination.
- Diagnose and resolve scheduling failures, CNI and CSI faults, admission rejections, and resource contention.
- Drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
- Run Kubernetes baselines in production, managing admission policy, workload identity, and network policy content.
- Provide senior technical expertise for Kubernetes faults across the estate, diagnosing complex issues and driving permanent fixes.
- Mentor engineers in frontline diagnosis, documenting operational procedures, runbooks, and performance results.
- Lead technical recovery during major Kubernetes incidents, driving changes to eliminate repeat causes.
- Participate in the after-hours escalation roster for the Kubernetes estate.
Requirements
- Strong skills in platform and infrastructure engineering.
- Substantial ownership of production Kubernetes platforms in a 24/7 environment.
- Deep experience operating Kubernetes at fleet scale, including cluster lifecycle, upgrades, and multi-cluster management.
- Strong experience with Kubernetes control plane internals, including etcd, the API server, controllers, schedulers, and certificate and credential rotation.
- Strong experience writing Kubernetes controllers, operators, or admission logic in a production setting.
- Experience with multi-tenant or virtual cluster patterns, including tenant isolation at the Kubernetes layer.
- Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling, and driver coordination.
- Strong skills in infrastructure automation, infrastructure-as-code, and GitOps practices.
- Change delivered through peer review, automated testing, and progressive rollout.
- Strong experience with scripting or programming for operational automation and tooling (e.g., Go, Python, Bash).
- Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, post-incident review, and the production of executable runbooks.
- Solid understanding of Kubernetes security fundamentals, including admission control, workload identity, and network policy.
- Clear technical judgment and communication skills, with the ability to produce documentation, design notes, and escalations.
- Experience operating Kubernetes for GPU or HPC workloads at scale.
- Experience with automated tenant or customer onboarding pipelines in a multi-tenant platform.
- Experience with vendor Kubernetes distributions or reference architectures for accelerated computing.
- Familiarity with DPU or SmartNIC-based networking as it relates to Kubernetes CNI design.
- Experience contributing to open-source Kubernetes ecosystem projects.
- A Bachelor's degree in computer science, engineering, or a related discipline, or an equivalent combination of relevant experience and training.
Skills
- Platform engineering
- Infrastructure engineering
- Kubernetes
- Cluster lifecycle management
- Multi-cluster management
- Kubernetes control plane internals
- etcd
- API server
- Controllers
- Schedulers
- Certificate rotation
- Credential rotation
- Kubernetes controllers
- Kubernetes operators
- Admission logic
- Multi-tenant patterns
- Virtual cluster patterns
- Tenant isolation
- GPU-enabled Kubernetes
- Device plugins
- GPU scheduling
- Driver coordination
- NVIDIA GPU Operator
- Infrastructure automation
- Infrastructure-as-code
- GitOps
- OpenTofu
- Terraform
- Ansible
- Argo CD
- Scripting
- Programming
- Go
- Python
- Bash
- Incident response
- On-call participation
- Post-incident review
- Runbook production
- Kubernetes security
- Admission control
- Workload identity
- Network policy
- Technical judgment
- Communication skills
- Documentation
- Design notes
- HPC workloads
- Automated tenant onboarding pipelines
- Vendor Kubernetes distributions
- Accelerated computing
- DPU networking
- SmartNIC networking
- Kubernetes CNI design
- Open-source Kubernetes ecosystem projects
Location
- Australia
- Singapore
Work Type
- 24/7 environment
- On-call
Experience Level
- 8+ years of experience overall
- Senior
Education Level
- Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training
About the Company
- Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation.