Site Reliability Engineer (f/m/d) – Observability & Internal Tools at Bertelsmann-Jobs | BE, DE | Rezi

Site Reliability Engineer (f/m/d) – Observability & Internal Tools at Bertelsmann-Jobs

Site Reliability Engineer (f/m/d) – Observability & Internal Tools

Bertelsmann-Jobs · BE, DE

2 days ago

Site Reliability Engineer (f/m/d) – Observability & Internal Tools

Bertelsmann-Jobs · BE, DE

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Develop observability and automation of 'tools we use' into a platform capability that supports and empowers every other engineer in the company, making them more productive.

Responsibilities

  • Take responsibility for the observability stack: Drive Prometheus, Grafana, and Forgejo forward, evolving them into a top-tier monitoring ecosystem.
  • Design meaningful and actionable alerts and define SLOs to shift from reactive firefighting to proactive stability.
  • Evaluate and implement modern open-source alternatives to proprietary software, influencing technology adoption and integration.
  • Integrate security engineering directly into the development and delivery process, identifying vulnerabilities proactively.
  • Dive deep into Linux systems and distributed tools, balancing bold experimentation with a stable production environment.

Requirements

  • Demonstrable experience implementing metrics, logs, and traces, with the ability to derive actionable insights from data.
  • Proficiency in Linux and understanding of distributed systems communication and failure modes.
  • A mindset for automation, using code to manage infrastructure and eliminate manual tasks.
  • An ownership mentality, embracing the 'You build it, you run it' philosophy.
  • Experience with GCP or AWS in a production environment at scale (nice-to-have).
  • Active engagement in the open-source community or a portfolio of self-hosted projects (nice-to-have).
  • Experience with blameless post-mortems and structured root-cause analysis (nice-to-have).

Skills

  • Observability
  • Linux
  • Systems Engineering
  • Automation
  • Ownership
  • Metrics
  • Logs
  • Traces
  • GCP
  • AWS
  • Open Source
  • Post-Mortems
  • Root-Cause Analysis

Location

  • Remote
  • Berlin
  • Cologne

Work Type

  • Remote
  • On-site

Experience Level

  • Builder mentality
  • Systems thinking

Benefits

  • 30 days of vacation
  • Time off on Christmas Eve and New Year's Eve
  • Smart Fridays (4-day week option)
  • Mobility options (Deutschlandticket & JobRad)
  • Sports and health offerings
  • Mental health support
  • Corporate Benefits
  • RTL+ access

About the Company

  • Develop platforms rather than simply subscribing to services.
  • Believe that deep internal know-how and a consistent open-source approach enable flexibility and performance beyond Enterprise SaaS solutions.
  • Develop in-house, steer systems, and take end-to-end responsibility for the entire tech stack.
  • Use AI to accelerate thinking, not replace it.
  • Design systems, steer output, and take responsibility for shipped products.
  • Operate with speed where sensible and care where it matters.

Equal Opportunity

  • smartclip is committed to creating a diverse and inclusive environment. All qualified applicants will receive consideration for employment without regard to race, ethnicity, nationality, age, gender, gender identity, religion, sexual orientation, disability, or any other diverse characteristics.