Site Reliability Engineer (f/m/d) – Observability & Internal Tools at Bertelsmann-Jobs | Berlin | Rezi

Site Reliability Engineer (f/m/d) – Observability & Internal Tools at Bertelsmann-Jobs

Site Reliability Engineer (f/m/d) – Observability & Internal Tools

Bertelsmann-Jobs · Berlin

Yesterday

Site Reliability Engineer (f/m/d) – Observability & Internal Tools

Bertelsmann-Jobs · Berlin

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

You will be the guardian and architect of smartclip’s internal infrastructure, turning observability and automation into a platform capability that empowers every engineer in the company.

Responsibilities

  • Own the Observability Stack: Take the lead on Prometheus, Grafana, and Forgejo, evolving them into a world-class monitoring ecosystem.
  • Engineer for Reliability: Design actionable alerting and define SLOs to move from reactive firefighting to proactive stability.
  • Champion Open Source: Evaluate and implement cutting-edge open-source alternatives to proprietary software.
  • Secure the Pipeline: Integrate security engineering directly into the delivery process to find vulnerabilities proactively.
  • Master the Metal: Navigate Linux systems and distributed tooling, balancing experimentation with production stability.

Requirements

  • Looking for a builder bored by simple configuration who thrives on systems thinking.
  • Proven track record of implementing metrics, logs, and traces.
  • Comfortable in the terminal and understanding of distributed systems communication and failure.
  • Automate infrastructure and eliminate toil using code.
  • Embrace the "you build it, you run it" philosophy and take pride in system stability.
  • Hands-on experience with GCP or AWS at a production scale.
  • Active contributions to the open-source community or a portfolio of self-hosted projects.
  • Experience in conducting blameless post-mortems and driving root-cause analysis.
  • A portfolio, side project, or demo repo showing production-ready, thought-through, and completed work.

Skills

  • Observability Expertise
  • Linux & Systems Engineering
  • Automation Mindset
  • Ownership Culture
  • GCP or AWS
  • Open-source contributions
  • Blameless post-mortems
  • Root-cause analysis

Location

  • Remote

Work Type

  • Remote
  • On-site (for specific events)

Experience Level

  • Mid-level

Benefits

  • Ownership over tickets and real responsibility
  • No unnecessary bureaucracy or micromanagement
  • Test what works, fail fast, learn faster approach
  • High standards, low ego, direct feedback, honest collaboration
  • Investment in growth: Hackathons, conferences, community
  • 30 days of vacation + Dec 24 & 31 off
  • Smart Fridays (4 days week possible)
  • Mobility (Germany ticket & JobRad)
  • Sports & health offerings
  • Mental health support
  • Corporate benefits
  • RTL+ access

About the Company

  • We engineer platforms using deep in-house expertise and a strong commitment to open source.
  • We design, steer, and own our stack end-to-end.
  • We use AI to accelerate thinking, not replace it.
  • We design the system, steer the output, and take responsibility for what we ship.
  • We are fast where it makes sense and careful where it matters.

Equal Opportunity

  • smartclip is committed to creating a diverse and inclusive environment. All qualified applicants will receive consideration for employment without regard to race, ethnicity, nationality, age, gender, gender identity, religion, sexual orientation, disability, or any other diverse characteristics.