About the Role
As a Site Reliability Engineer (SRE), you will help ensure our platforms are fast, reliable, and scalable. You’ll design, build, and monitor systems, preventing issues before they happen. You will be embedded in the XP engineering group, which builds the foundations for other engineers, including shared UI frameworks, BFF layer, and internal developer platform.
Responsibilities
- Manage and automate resources in Azure, Datadog, NGINX & Cloudflare.
- Develop, deploy and monitor Kubernetes and Serverless resources.
- Build, manage, and evolve IAC using Terraform, Helm and Go CRDs.
- Improve systems, processes, and technologies; consulting stakeholders to enhance platform performance.
- Design and maintain comprehensive monitoring and alerting strategies using Datadog.
- Get involved in new application architecture & design processes.
- Design solutions to reduce toil, automate repetitive tasks and streamline workflows.
- Create SLIs and SLOs; increasing application visibility.
- Align with the Product team on SLAs and core service objectives.
- Collaborate with core foundational teams to build and maintain reusable, automated solutions.
- Optimize CI/CD pipelines and developer workflows using Azure DevOps, Github, Octopus Deploy, MirrorD, and Flux.
- Lead incident response; taking an active part in communications, investigations, remediation and post-mortems.
Requirements
- Experience managing public cloud environments.
- Proficient in contributing to IaC technologies involving expertise in writing, managing, and optimising infrastructure with tools such as Terraform.
- Experience using CI/CD tools, building pipelines, templates and troubleshooting.
- Proficient with containerisation technologies such as Kubernetes and Docker.
- Experience building and deploying frontend applications with tools such as CodeMagic.
- Knowledge or Experience with Google Firebase.
- Experience working with a cloud monitoring solution.
- Proficiency in at least one scripting language such as Python, PowerShell, Go.
- Great communication skills with the ability to collaborate effectively.
- Proven experience with a software development background.
Skills
- Azure
- Datadog
- NGINX
- Cloudflare
- Kubernetes
- Serverless
- Terraform
- Helm
- Go CRDs
- CI/CD
- Azure DevOps
- Github
- Octopus Deploy
- MirrorD
- Flux
- Python
- PowerShell
- Docker
- CodeMagic
- Google Firebase
Location
- London, Old Street
Work Type
- Hybrid
- 2 Days in Office
Experience Level
- Software development background
Benefits
- Private Healthcare including dental and opticians services through Vitality
- Worldwide travel insurance through Vitality
- Anniversary Rewards (£250, £500, £750, 4-week fully paid sabbatical)
- Salary Sacrifice Pension Scheme up to 7% match
- 28 days holiday (plus bank holidays)
- Annual Learning and Wellbeing Budget
- Enhanced Parental Leave
- Cycle to Work Scheme
- Season Ticket Loan
- 6 free therapy sessions per year
- Dog Friendly Offices
- Free drinks and snacks in our offices
About the Company
- Capital on Tap started because small businesses were underserved by big banks.
- Today we're a financial platform offering a best-in-class business credit card, SME-focused spend management platform, a savings product, and a growing suite of tools.
- We have 1,000+ employees, £20bn in annual card spend, 200,000+ customers, and 17,000+ Trustpilot reviews averaging 4.7 stars.
- We are profitable and just getting started.
Equal Opportunity
- We welcome, consider and encourage applications from anyone who shares our commitment to inclusivity.
- Join us in creating a space where authenticity thrives, and everyone can do their best work.
