Site Reliability Engineer (SRE) (m/w/d) at Deutsche Telekom | Berlin, Baden-Württemberg, DE | Rezi

Site Reliability Engineer (SRE) (m/w/d) at Deutsche Telekom

Site Reliability Engineer (SRE) (m/w/d)

Deutsche Telekom · Berlin, Baden-Württemberg, DE

2 weeks ago

Site Reliability Engineer (SRE) (m/w/d)

Deutsche Telekom · Berlin, Baden-Württemberg, DE

20 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Site Reliability Engineer, you will be responsible for ensuring the stability, performance, and scalability of mission-critical systems. You will play a key role in the design, implementation, and maintenance of a highly reliable and efficient infrastructure, leveraging expertise in modern monitoring, observability, and cloud-native technologies.

Responsibilities

  • Design, build, and operate a high-performance monitoring stack based on Prometheus, Grafana, Loki, and Alertmanager for distributed Fog and Edge infrastructures.
  • Develop specialized dashboards for monitoring hardware health, resource utilization, and network connectivity.
  • Design and implement alerting strategies for autonomous operating models, disconnected, and air-gapped environments.
  • Implement local data retention concepts, log rotation mechanisms, and efficient data management strategies.
  • Continuously optimize the performance of the monitoring stack considering limited CPU, memory, and storage resources.
  • Integrate Kubernetes monitoring solutions, analyze operational and performance metrics, and identify optimization opportunities.
  • Automate the deployment and management of monitoring components using Infrastructure as Code (IaC) principles and YAML-based configurations.
  • Contribute to the continuous evolution of a robust, scalable, and highly available observability platform for modern Edge and Fog computing environments.

Requirements

  • Strong experience with Prometheus, including PromQL, Federation, Remote Write, and Local Retention Management.
  • Advanced knowledge of Grafana, including dashboard development, alerting, and provisioning as code.
  • Experience with Loki, including LogQL, retention management, and compaction.
  • Strong knowledge of Alertmanager, including standalone routing, alert inhibition, and escalation/notification strategies.
  • Experience in monitoring and operating Kubernetes environments.
  • Solid understanding of container platforms and cloud-native architectures.
  • Experience working in resource-constrained environments.
  • Experience with optimization of CPU, memory, and storage utilization.
  • Understanding of offline, air-gapped, and disconnected operating scenarios.
  • Advanced proficiency in YAML.
  • Strong understanding of IT security principles and security awareness.
  • Strong analytical and structured approach to problem-solving.
  • Experience in Site Reliability Engineering (SRE) or infrastructure operations.
  • Excellent German language skills, both written and spoken (C1 level).

Skills

  • Monitoring & Observability
  • Prometheus
  • PromQL
  • Federation
  • Remote Write
  • Local Retention Management
  • Dashboarding & Visualization
  • Grafana
  • Alerting
  • Provisioning as Code
  • Log Management
  • Loki
  • LogQL
  • Retention Management
  • Compaction
  • Alerting & Incident Management
  • Alertmanager
  • Standalone routing
  • Alert inhibition
  • Escalation and notification strategies
  • Kubernetes & Cloud-Native Technologies
  • Kubernetes monitoring
  • Container platforms
  • Cloud-native architectures
  • Edge & Fog Computing
  • Resource-constrained environments
  • CPU, memory, and storage optimization
  • Offline, air-gapped, and disconnected operating scenarios
  • YAML
  • IT security principles
  • Site Reliability Engineering (SRE)
  • Infrastructure as Code (IaC)
  • Terraform
  • Ansible
  • GitOps
  • Linux system administration
  • OpenTelemetry
  • Highly available infrastructures

Location

  • Germany

Work Type

  • Mobile work options
  • Hybrid work model
  • Part-time

Experience Level

  • Site Reliability Engineering (SRE)
  • Infrastructure operations

Salary/Compensations

  • Competitive salary package

Benefits

  • Flexible working hours
  • Mobile work options
  • Hybrid work model
  • Extensive training programs, courses, and career development initiatives
  • Modern work environment
  • Open and collaborative work atmosphere
  • Health promotion offers and initiatives
  • Company pension scheme
  • Employee discounts
  • Public transport ticket options

About the Company

  • At T-Systems, we offer business customers the right system solutions for their digital business.
  • With our portfolio we ensure that digital transformation reduces complexity, saves costs and makes day-to-day work easier.
  • We focus on the areas connectivity, digital, cloud & infrastructure as well as security - Let's power higher performance!

Equal Opportunity

  • People with disabilities will take priority in case of equal qualifications.