About the Role
As a Site Reliability Engineer, you will be responsible for ensuring the stability, performance, and scalability of mission-critical systems. You will play a key role in the design, implementation, and maintenance of a highly reliable and efficient infrastructure, leveraging expertise in modern monitoring, observability, and cloud-native technologies.
Responsibilities
- Design, build, and operate a high-performance monitoring stack based on Prometheus, Grafana, Loki, and Alertmanager for distributed Fog and Edge infrastructures.
- Develop specialized dashboards for monitoring hardware health, resource utilization, and network connectivity.
- Design and implement alerting strategies for autonomous operating models, disconnected, and air-gapped environments.
- Implement local data retention concepts, log rotation mechanisms, and efficient data management strategies.
- Continuously optimize the performance of the monitoring stack considering limited CPU, memory, and storage resources.
- Integrate Kubernetes monitoring solutions, analyze operational and performance metrics, and identify optimization opportunities.
- Automate the deployment and management of monitoring components using Infrastructure as Code (IaC) principles and YAML-based configurations.
- Contribute to the continuous evolution of a robust, scalable, and highly available observability platform for modern Edge and Fog computing environments.
Requirements
- Strong experience with Prometheus, including PromQL, Federation, Remote Write, and Local Retention Management.
- Advanced knowledge of Grafana, including dashboard development, alerting, and provisioning as code.
- Experience with Loki, including LogQL, retention management, and compaction.
- Strong knowledge of Alertmanager, including standalone routing, alert inhibition, and escalation/notification strategies.
- Experience in monitoring and operating Kubernetes environments.
- Solid understanding of container platforms and cloud-native architectures.
- Experience working in resource-constrained environments.
- Experience with optimization of CPU, memory, and storage utilization.
- Understanding of offline, air-gapped, and disconnected operating scenarios.
- Advanced proficiency in YAML.
- Strong understanding of IT security principles and security awareness.
- Strong analytical and structured approach to problem-solving.
- Experience in Site Reliability Engineering (SRE) or infrastructure operations.
- Excellent German language skills, both written and spoken (C1 level).
Skills
- Monitoring & Observability
- Prometheus
- PromQL
- Federation
- Remote Write
- Local Retention Management
- Dashboarding & Visualization
- Grafana
- Alerting
- Provisioning as Code
- Log Management
- Loki
- LogQL
- Retention Management
- Compaction
- Alerting & Incident Management
- Alertmanager
- Standalone routing
- Alert inhibition
- Escalation and notification strategies
- Kubernetes & Cloud-Native Technologies
- Kubernetes monitoring
- Container platforms
- Cloud-native architectures
- Edge & Fog Computing
- Resource-constrained environments
- CPU, memory, and storage optimization
- Offline, air-gapped, and disconnected operating scenarios
- YAML
- IT security principles
- Site Reliability Engineering (SRE)
- Infrastructure as Code (IaC)
- Terraform
- Ansible
- GitOps
- Linux system administration
- OpenTelemetry
- Highly available infrastructures
Location
- Germany
Work Type
- Mobile work options
- Hybrid work model
- Part-time
Experience Level
- Site Reliability Engineering (SRE)
- Infrastructure operations
Salary/Compensations
- Competitive salary package
Benefits
- Flexible working hours
- Mobile work options
- Hybrid work model
- Extensive training programs, courses, and career development initiatives
- Modern work environment
- Open and collaborative work atmosphere
- Health promotion offers and initiatives
- Company pension scheme
- Employee discounts
- Public transport ticket options
About the Company
- At T-Systems, we offer business customers the right system solutions for their digital business.
- With our portfolio we ensure that digital transformation reduces complexity, saves costs and makes day-to-day work easier.
- We focus on the areas connectivity, digital, cloud & infrastructure as well as security - Let's power higher performance!
Equal Opportunity
- People with disabilities will take priority in case of equal qualifications.
