About the Role
As a Site Reliability Engineer specializing in Observability, you will ensure the transparency, stability, and reliability of distributed cloud, edge, and Kubernetes platforms. You will design, implement, and evolve modern observability solutions, providing the foundation for reliable operation of complex cloud-native infrastructures.
Responsibilities
- Design, implement, and operate high-performance pipelines for metrics, logs, and traces using OpenTelemetry, distributed tracing, and modern telemetry architectures.
- Integrate the OpenTelemetry Collector, OTLP, SDKs, and auto-instrumentation into existing platforms for end-to-end observability.
- Implement and operate Jaeger as the central distributed tracing platform.
- Develop correlation mechanisms across logs, metrics, and traces to rapidly identify root causes and provide end-to-end visibility.
- Apply modern sampling strategies, context propagation techniques, and efficient approaches to trace storage and high-performance analysis of large-scale telemetry data.
- Integrate security, infrastructure, and platform events into centralized SIEM solutions.
- Develop standardized telemetry and instrumentation concepts for Kubernetes and cloud-native platforms.
- Automate integration and operational processes using Go, Python, or Shell scripting.
- Contribute to the continuous evolution of a scalable, highly available, and future-proof observability platform.
Requirements
- Extensive experience with OpenTelemetry, including the OpenTelemetry Collector, OTLP, SDKs, and auto-instrumentation.
- Strong expertise in distributed tracing, including span correlation, context propagation, and sampling strategies.
- Experience deploying and operating Jaeger, including trace storage, indexing, and query optimization.
- Solid knowledge of SIEM integration, security event management, and the use of CEF and Syslog-based protocols.
- Proficiency in Go, Python, or Shell scripting for telemetry integration, automation, and operational tooling.
- Experience with Kubernetes observability, cloud-native platforms, and modern telemetry architectures.
- Strong understanding of IT security principles and security best practices.
- Strong analytical and structured approach to problem-solving.
- Experience in Site Reliability Engineering (SRE) or infrastructure operations.
- Excellent German language skills, both written and spoken (CEFR C1 level).
Skills
- OpenTelemetry
- OpenTelemetry Collector
- OTLP
- SDKs
- auto-instrumentation
- Distributed tracing
- Span correlation
- Context propagation
- Sampling strategies
- Jaeger
- Trace storage
- Indexing
- Query optimization
- SIEM integration
- Security event management
- CEF
- Syslog
- Go
- Python
- Shell scripting
- Kubernetes observability
- Cloud-native platforms
- Telemetry architectures
- IT security principles
- Security best practices
- Site Reliability Engineering (SRE)
- Infrastructure operations
Location
- Germany
Work Type
- Mobile work options
- Hybrid work model
- Part-time
Salary/Compensations
- Competitive salary package
Benefits
- Flexible working hours
- Mobile work options
- Hybrid work model
- Extensive training programs, courses, and career development initiatives
- Modern work environment equipped with the latest technologies and tools
- Open and collaborative work atmosphere
- Health promotion offers and initiatives
- Company pension scheme
- Employee discounts
- Public transport ticket options
About the Company
- At T-Systems, we offer business customers the right system solutions for their digital business.
- With our portfolio we ensure that digital transformation reduces complexity, saves costs and makes day-to-day work easier.
- We focus on the areas connectivity, digital, cloud & infrastructure as well as security - Let's power higher performance!
Equal Opportunity
- People with disabilities will take priority in case of equal qualifications.
