Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Principal Cloud Platform Engineer role.
Rezi rewrites your resume against SambaNova's job description. Free.

Tailor your resume to this Principal Cloud Platform Engineer role.
Rezi rewrites your resume against SambaNova's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Principal Cloud Platform Engineer posting at SambaNova — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Principal Cloud Platform Engineer posting at SambaNova — free, in seconds.
About the Role
As a Principal Cloud Platform Engineer, you will specialize in our AI Inferencing Service, ensuring its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability.
Responsibilities
- Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning
- Standing-up and automating AI infrastructure in new regions
- Participating in a shared primary/secondary on-call rotation, and leading incident response
- Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization
- Finding and eliminating performance bottlenecks
- Designing auto-scaling policies that handle variable inference loads
- Managing cloud and on-prem infrastructure as code in Terraform and Ansible
- Building CI/CD pipelines that safely deploy new model versions and service updates
- Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend
- Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work
Requirements
- B.S. in Computer Science, Computer Engineering, or related field
- 3+ years of experience in a Site Reliability Engineering, DevOps
- Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)
- Strong programming and scripting skills in languages like Python, Go, Rust, or Java
- Proven experience with containerization and orchestration technologies (Docker and Kubernetes)
- Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)
- Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)
- Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)
- Strong Linux/Unix system administration fundamentals
Skills
- Python
- Go
- Rust
- Java
- Docker
- Kubernetes
- Prometheus
- Grafana
- ELK Stack
- Datadog
- Terraform
- CloudFormation
- Jenkins
- GitHub Actions
- ArgoCD
- Linux/Unix system administration
- Hybrid cloud and on-premise infrastructure management
- ML/AI inferencing services
- GPU-accelerated computing
- NVIDIA GPUs
- vLLM
- SGLang
- Ray
- MLOps
- SQL databases
- NoSQL databases
- Redis
- Memcached
Location
- United States
- Asia
- Europe
- Latin America
Work Type
- On-premises
- Cloud
Experience Level
- Principal
Education Level
- B.S. in Computer Science, Computer Engineering, or related field
Salary/Compensations
- $144,000—$189,000 USD
Benefits
- Equity
- 95% premium coverage for employee medical insurance
- 77% premium coverage for dependents
- Health Savings Account (HSA) with employer contribution
- Dental insurance
- Vision insurance
- Short-term Disability insurance
- Long-term Disability insurance
- Basic Life insurance
- Voluntary Life insurance
- AD&D insurance
- Flexible Spending Account (FSA) options (Health Care, Limited Purpose, Dependent Care)
- Headspace subscription
- Gympass+ membership
- One Medical membership
- Counseling services with an Employee Assistance Program
About the Company
- SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations.
- Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy.
- Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets.
Equal Opportunity
- SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.