About the Role
Trainline is seeking a mid-level Site Reliability Engineer to join the Reliability & Operations Engineering team. This role involves ensuring the platform's observability, reliability, scalability, and resilience by partnering with product engineering teams, responding to incidents, and continuously improving system reliability.
Responsibilities
- Develop an understanding of system architecture, dependencies, and failure modes across the Trainline platform.
- Participate in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration.
- Contribute to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience.
- Take part in the SRE on-call rotation.
- Design, build, and maintain observability using metrics, logs, events, and traces to support effective detection and diagnosis.
- Improve monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD).
- Ensure relevant operational data is surfaced quickly and clearly during live incidents.
- Make informed tooling and technology choices using SRE principles, balancing team and business needs.
- Support AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling.
- Collaborate with product engineering teams to ensure services are operationally ready and deployed safely.
- Advise on reliability and resilience practices.
- Write and maintain reliable, well-structured code and scripts to support reliability and observability goals.
- Prioritize work effectively and collaborate using agile processes to deliver against team and business goals.
Requirements
- Experience of SRE concepts such as SLI, SLO and error budgets.
- Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar.
- Experience working with cloud providers (preferably AWS).
- Experience troubleshooting Linux operating systems.
- Experience of scripting in at least one language (preferably Python).
- Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts.
- Understanding of application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling).
- Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving averages and other functions.
- Experience with build, deployment & configuration management tooling such as GitHub Actions and Terraform.
Skills
- AWS
- New Relic
- ELK stack
- Grafana
- Incident.io
- Docker
- ECS
- Terraform
- Github Actions
- Python
- SRE concepts (SLI, SLO, error budgets)
- Observability tooling
- Cloud providers (AWS)
- Linux operating systems
- Scripting
- Load balancing
- Reverse proxy concepts
- Application architecture concepts
- Time series data management
- Build, deployment & configuration management tooling
Location
- London
- Paris
- Barcelona
- Milan
- Edinburgh
- Madrid
Work Type
- Hybrid
- Remote (Work from Abroad Policy)
Experience Level
- Mid-level
Benefits
- Private healthcare & dental insurance
- Generous work from abroad policy
- 2-for-1 share purchase plans
- EV Scheme
- Extra festive time off
- Excellent family-friendly benefits
- Personal learning budgets
- Regular learning days
About the Company
- Trainline is Europe's number 1 downloaded rail app, enabling millions of travellers to find and book the best value tickets across carriers, fares, and journey options.
- The company collaborates with 270+ rail and coach companies in over 40 countries.
- Trainline is a FTSE 250 company driven by over 1,000 Trainliners from 50+ nationalities.
- The company's mission is to create a world where travel is simple, seamless, eco-friendly, and affordable.
- Trainline is focused on growth in the UK and Europe.
Equal Opportunity
- We know that having a diverse team makes us better and helps us succeed. And we mean all forms of diversity - gender, ethnicity, sexuality, disability, nationality and diversity of thought. That's why we're committed to creating inclusive places to work, where everyone belongs and differences are valued and celebrated.
