About the Role
We are seeking a Senior Cloud Reliability Engineer to take ownership of highly available, cloud-native production environments. This role focuses on operating critical production systems, managing incidents, and driving improvements in existing cloud infrastructure, combining deep AWS and Kubernetes expertise with SRE practices.
Responsibilities
- Own the reliability and operational health of production AWS and Kubernetes environments.
- Participate in on-call rotations and own high-severity production incidents from detection through resolution.
- Lead root cause analysis and post-incident reviews, implementing permanent corrective actions.
- Identify architectural, reliability, security, performance, and operational weaknesses in cloud environments.
- Propose and implement improvements based on AWS and Kubernetes best practices.
- Design, maintain, and continuously improve AWS infrastructure and production Kubernetes/EKS environments.
- Automate infrastructure provisioning and operational workflows using Terraform and configuration-management tools.
- Improve deployment and GitOps processes using tools such as Argo CD.
- Build and improve monitoring, logging, tracing, dashboards, and actionable alerting using Prometheus, Grafana, ELK and related observability technologies.
- Improve scalability and workload management using Kubernetes autoscaling technologies such as Karpenter or KEDA.
- Support distributed and event-driven environments, including technologies such as Kafka.
- Develop automation and operational tooling using Python, Bash, Go, or similar languages.
- Strengthen cloud security, resilience, disaster recovery, and production-readiness practices.
- Work directly with enterprise customers for technical troubleshooting, escalations, incident discussions, and explaining infrastructure issues.
- Collaborate with engineering teams, bringing independent ideas and challenging existing technical approaches.
Requirements
- Strong professional experience in Site Reliability Engineering, Cloud Reliability, Platform Engineering, or a comparable production-focused role.
- Senior-level hands-on AWS expertise, including design, troubleshooting, and improvement of AWS architectures.
- Senior-level Kubernetes experience, ideally operating Amazon EKS in production.
- Strong hands-on experience with Terraform and Infrastructure as Code.
- Proven experience participating in an on-call rotation and owning production incidents.
- Demonstrable experience with high-severity incident response, root cause analysis, postmortems, MTTR reduction, and permanent remediation.
- Proven experience communicating directly with external or enterprise customers in a technical capacity.
- Strong observability experience with Prometheus, Grafana, ELK, or equivalent production observability stacks.
- Strong Linux and cloud networking fundamentals.
- Experience automating operational processes using Python, Bash, Go, or similar scripting/programming languages.
- Experience with CI/CD and GitOps environments.
- Ability to independently identify infrastructure weaknesses and translate them into practical technical improvements.
- Strong communication skills and professional-level English.
- Comfortable explaining complex technical issues to both engineering teams and customers.
- Hands-on production experience with Argo CD and GitOps-based deployment workflows.
- Hands-on experience with Kafka in distributed production environments.
- Production experience with Kubernetes autoscaling using Karpenter and/or KEDA.
- Strong hands-on experience with Ansible for infrastructure/configuration automation.
- Practical experience with AWS security tooling and cloud security best practices.
- A relevant AWS certification.
Skills
- AWS
- Kubernetes
- EKS
- Terraform
- Infrastructure as Code
- Incident Response
- Root Cause Analysis
- Postmortems
- MTTR Reduction
- Prometheus
- Grafana
- ELK
- Observability
- Linux
- Cloud Networking
- Python
- Bash
- Go
- CI/CD
- GitOps
- Argo CD
- Kafka
- Karpenter
- KEDA
- Ansible
- AWS Security Tooling
- Cloud Security Best Practices
Location
- EU
- US
Work Type
- Fully remote
- B2B Contract
- Full-time
Experience Level
- Senior
Education Level
- AWS certification
Benefits
- Fully remote position
- Opportunity to work on complex, production-critical cloud environments
- High level of technical ownership and autonomy
- Real influence over cloud architecture, reliability practices, automation, and platform improvements
- International engineering environment
- Direct collaboration with experienced technical teams and enterprise customers
- Long-term cooperation and opportunities to introduce new technologies and engineering practices
Equal Opportunity
- We are dedicated to creating and sustaining an inclusive, respectful workplace for all - regardless of gender, ethnicity, or background.
- We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast-paced, supportive team.
