About the Role
As a Site Reliability Engineer (SRE) in our Application Hosting Team, you will be the technical backbone of our product platform for Managed Nextcloud, Nextcloud Workspace, IONOS GPT, and other web services running on our Kubernetes platform. You will collaborate with experienced colleagues to design new services and products that remain performant and fault-tolerant even under high load.
Responsibilities
- Further develop the infrastructure/platform of our products.
- Integrate new products/web services into our Kubernetes and cloud infrastructure.
- Ensure the stable and secure operation of our product platform.
- Perform in-depth analysis and optimization of our primarily containerized and Kubernetes-based application infrastructure.
- Provision and manage infrastructure declaratively and reproducibly using tools like Terraform, GitLab CI/CD, and ArgoCD.
- Analyze and resolve complex problems in a distributed system landscape.
- Continuously improve our platform.
- Develop and maintain our monitoring, logging, and alerting solutions (e.g., with Prometheus, Grafana, ELK-Stack) to proactively identify bottlenecks and error sources.
Requirements
- Several years of experience as a Site Reliability Engineer or in a related role (Linux System Administrator, Platform Engineer, DevOps Engineer, Full Stack Developer) in a Linux and Kubernetes environment.
- Very good knowledge and several years of experience with the Linux operating system, container technologies, and specifically Kubernetes.
- Experience with Infrastructure as Code (preferably Terraform), CI/CD pipelines (e.g., GitLab CI/CD or GitHub Actions), and using Helm Charts.
- Proficiency in at least one programming or scripting language (e.g., Go, Python, Bash) for automation and monitoring tasks.
- Experience with operating and troubleshooting highly available and distributed production environments, including monitoring, alerting, and log analysis of distributed applications (e.g., Prometheus, Grafana, FluentD, ELK, VictoriaMetrics, icinga).
- Proactive, solution-oriented, and independent working style.
- Ability to systematically analyze and sustainably resolve complex technical problems.
- Good German and English language skills.
Skills
- Linux
- Kubernetes
- Infrastructure as Code
- Terraform
- CI/CD pipelines
- GitLab CI/CD
- GitHub Actions
- Helm Charts
- Go
- Python
- Bash
- Monitoring
- Logging
- Alerting
- Prometheus
- Grafana
- ELK-Stack
- FluentD
- VictoriaMetrics
- icinga
- ArgoCD
Location
- Germany
Work Type
- Hybrid
Experience Level
- Several years of experience
Benefits
- Hybrid work model
- Flexible working hours with trust-based working time
- Subsidized canteen and various free drinks at some locations
- Modern office spaces with very good transport links
- Various employee discounts for activities and products
- Employee events such as summer and winter parties, as well as workshops
- Numerous further training and development opportunities
- Various health offers, such as sports and health courses
About the Company
- IONOS is the leading European digitalization partner for small and medium-sized enterprises (SMEs).
- IONOS has over six million customers and operates with a globally available platform in 18 markets across Europe and North America.
- With its Web Presence & Productivity offerings, the company acts as a 'One-Stop-Shop' for all digitalization needs - from domains and web hosting to classic website builders and do-it-yourself solutions, from e-commerce to online marketing tools.
- Additionally, IONOS offers cloud solutions for companies looking to move to the cloud as they develop their business.
Equal Opportunity
- We value diversity and welcome all applications – regardless of, for example, gender, nationality, ethnic and social origin, religion, disability, age, or sexual orientation and identity, physical characteristics, marital status, or any other criterion irrelevant to the application under applicable law.
