Site Reliability Engineer / Production Support at Monument Bank Limited | London, England, GB | Rezi

Site Reliability Engineer / Production Support at Monument Bank Limited

Site Reliability Engineer / Production Support

Monument Bank Limited · London, England, GB

5 days ago

Site Reliability Engineer / Production Support

Monument Bank Limited · London, England, GB

6 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Monument’s SRE role is the single point of ownership for production environment incidents. You will directly oversee the offshore Production Support team, run on-call and incident response, and ensure fast detection, triage, and restoration of services. This role involves actively debugging incidents, driving permanent fixes, and using AI tools for automation. It offers a rare opportunity to own production reliability at a pre-IPO challenger bank.

Responsibilities

  • Directly oversee the offshore Production Support team and act as the single point person for incidents.
  • Run on-call and incident response, ensuring fast detection, triage, and restoration.
  • Maintain observability standards (logs, metrics, traces) and alert quality.
  • Understand key system flows, services, partners, and teams to debug incidents.
  • Lead reliability engineering efforts, including resilience patterns, performance tuning, and capacity planning.
  • Facilitate post-incident reviews and track action items to completion.
  • Utilize AI tools for automated alert correlation, root cause analysis, and runbook generation.
  • Identify routine tasks for automation and execute automation plans.

Requirements

  • Strong SRE or production support experience with accountability for incident response in a production environment.
  • Deep understanding of observability tools, alerting, logging, and distributed systems debugging.
  • Experience managing and working with offshore support teams.
  • Hands-on experience with reliability engineering: resilience patterns, performance tuning, capacity planning.
  • Active use of AI tools for incident triage, automation, and runbook generation.
  • Ability to understand complex system flows across multiple services and third-party integrations.
  • Experience in financial services or similarly regulated environments is a strong advantage.

Skills

  • SRE
  • Production Support
  • Incident Response
  • Observability Tools
  • Alerting
  • Logging
  • Distributed Systems Debugging
  • Reliability Engineering
  • Resilience Patterns
  • Performance Tuning
  • Capacity Planning
  • AI Tools
  • Automation
  • Runbook Generation

Location

  • London (Oxford Circus)

Work Type

  • Hybrid: 2 days per week

About the Company

  • Monument is building a financial brand for the mass affluent, professionals, entrepreneurs, and ambitious savers, aiming to make wealth management simpler, smarter, and more human.
  • The company holds over £7 billion in client savings, serves more than 100,000 clients, and was named the UK's fastest growing fintech in 2025.
  • Monument's values shape decision-making, interpersonal interactions, and client service, emphasizing ambitious goals, learning from failures, collaboration, diverse perspectives, and continuous improvement.