About the Role
As a Senior Site Reliability Engineer focused on Managed Gateways, you'll architect and maintain resilient, scalable infrastructure for Kong's mission-critical managed services. You will also serve as the technical representative for this product to enterprise customers, ensuring reliability and performance for their connected applications.
Responsibilities
- Lead, mentor, and inspire a team of Site Reliability Engineers for Kong's Managed Gateway offerings.
- Architect and implement robust, scalable, and fault-tolerant cloud-native systems using Kubernetes, Golang, and major cloud providers.
- Own the end-to-end operational lifecycle, including proactive monitoring, alerting, incident response, and blameless post-mortems.
- Implement automation, self-service tooling, and streamlined workflows for deploying and managing API gateways.
- Define, track, and report on key SLOs and SLIs for Managed Gateways.
- Champion technical debt prevention and advocate for architectural best practices.
- Collaborate with Product, engineering, and Customer Success to influence roadmap decisions and ensure operational readiness.
- Partner with enterprise customers to drive end-to-end onboarding and implementation of Cloud Gateways.
- Productize recurring implementation patterns into repeatable playbooks and platform capabilities.
- Utilize cross-cloud expertise (AWS, GCP, Azure) to manage unique customer topologies.
- Serve as the technical owner of the customer relationship through implementation.
- Act as the escalation point for technically complex accounts for Customer Success.
- Provide feedback on real-world implementation patterns and customer constraints to the Product team.
Requirements
- Extensive experience as a Site Reliability Engineer, focusing on highly available and distributed systems.
- Deep expertise with Kubernetes and cloud-native architectures across multiple public cloud providers (AWS, GCP, Azure).
- Strong proficiency in Golang or similar modern programming languages for automation and tool development.
- Proven track record in building and maintaining CI/CD pipelines and infrastructure as code (Terraform, Ansible).
- In-depth knowledge of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack, Datadog).
- Experience with managed services, API gateways, or similar network infrastructure is highly desirable.
- Ability to take ownership of systems and treat reliability as a first-class feature.
- Ability to operate with a sense of urgency and drive quick, effective resolutions.
- Ability to thrive in a collaborative environment, actively sharing knowledge.
- Ability to adapt to shifting plans and work effectively in a fast-paced, distributed team environment.
Skills
- Kubernetes
- Golang
- AWS
- GCP
- Azure
- Terraform
- Ansible
- Prometheus
- Grafana
- ELK stack
- Datadog
- API Gateways
- Service Mesh technologies (Istio, Linkerd)
- Database administration (PostgreSQL, Cassandra)
- Open-source SRE tools
Location
- North America
Work Type
- Full-time
Experience Level
- Senior
Education Level
- Relevant cloud certifications (e.g., AWS Certified DevOps Engineer, CKA)
About the Company
- Kong Inc. is a leading developer of API and AI connectivity technologies, building the infrastructure for the agentic era.
- Kong's unified API and AI platform, Kong Konnect, enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI models.
- Trusted by the Fortune 500 and startups alike.
