About the Role
Voltage Park is seeking a skilled Infrastructure Operations Engineer to join their 24/7 team, ensuring the stability, scalability, and performance of compute, storage, and platform infrastructure. This role is crucial for delivering high-performance environments for AI/ML workloads at scale.
Responsibilities
- Design, build, and roll out new platforms and patterns to minimize incidents and enable customer-facing and internal features.
- Deploy updates and improvements to support both Voltage Park’s internal and end customer use cases.
- Collaborate with colleagues in Infrastructure Engineering, Network Operations, Customer Success, and Software and Platform Development Teams.
- Participate in an evenly distributed on-call rotation in a primary/secondary pattern.
Requirements
- 8+ years working with Linux as a server/hosting platform (Ubuntu experience is a plus).
- 5+ years of experience with AWS.
- 2+ years of experience with Kubernetes and strong container fundamentals.
- 2+ years of experience with Terraform and Ansible.
- 2+ years of experience with network-attached storage management (NFS, Ceph, or other protocols; VAST storage systems experience is a plus).
- Experience with monitoring systems (Prometheus, ELK stack).
- Familiarity with the GitOps workflow.
- Software development experience using Python, Go, Bash, or other languages for automation and system integration.
- Deep networking fundamentals (datacenter-level networks, 400Gb Ethernet, and Infiniband experience is a plus).
- Experience building and delivering complex systems.
- Ability to navigate tradeoffs between design, risk, cost, and outcomes.
- Comfortable navigating ambiguity.
- Strong written and oral communication skills.
- Experience with bare metal hardware troubleshooting and provisioning (Dell hardware experience is a plus).
- Experience with GPU servers, both bare metal and virtualized.
- Deep experience with network switches, routers, and firewalls (SONiC switches, Palo Alto firewalls, and Juniper Networks experience is a plus).
- Experience with VAST storage systems.
Skills
- Linux
- AWS
- Kubernetes
- Container fundamentals
- Terraform
- Ansible
- Network-attached storage management
- NFS
- Ceph
- Monitoring systems
- Prometheus
- ELK stack
- GitOps workflow
- Python
- Go
- Bash
- Networking fundamentals
- Datacenter networks
- 400Gb Ethernet
- Infiniband
- Bare metal hardware troubleshooting
- Bare metal provisioning
- GPU servers
- Network switches
- Network routers
- Network firewalls
- SONiC switches
- Palo Alto firewalls
- Juniper Networks
- VAST storage systems
Location
- Hybrid/On-site (Seattle, NYC, or San Francisco)
- Fully Remote (U.S.)
Work Type
- Hybrid
- On-site
- Remote
- 24/7
Experience Level
- 8+ years
- 5+ years
- 2+ years
About the Company
- Voltage Park offers scalable compute power, on-demand and reserved bare metal AI infrastructure using NVIDIA GPUs, with world-class service, performance, and value.
- Their mission is to make accessible AI computing for all, with flexible, affordable GPU solutions powering everyone from builders to enterprises.
- Through a merger with Lightning AI, they combine developer-first software with cost-efficient, large-scale compute, providing tools for experimentation, training, and production inference with security, observability, and control.
- The company culture values collaboration with friendly, motivated, execution-focused colleagues, a high degree of autonomy, wearing multiple hats, and the importance of good documentation.
Equal Opportunity
- Voltage Park is an equal opportunity employer and makes employment decisions on the basis of merit.
- All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic under federal, state, or local law.
- Accommodation is available during the job application process upon notification to the recruiter.
