About the Role
As a Site Reliability Engineer (SRE), you will help ensure our platforms are fast, reliable, and scalable. You’ll design, build, and monitor systems, preventing issues before they happen. You will be embedded in the Data teams, working alongside business teams and assisting with data ingestion, exports, model deployments, data warehouse management, and database engineering.
Responsibilities
- Manage and automate resources in Azure, Datadog, NGINX & Cloudflare.
- Develop, deploy and monitor Kubernetes and Serverless resources.
- Build, manage, and evolve IAC using Terraform, Helm and Go CRDs.
- Improve systems, processes, and technologies; consulting with stakeholders to enhance platform performance.
- Design and maintain comprehensive monitoring and alerting strategies using Datadog.
- Get involved in new application architecture & design processes.
- Design solutions to reduce toil, automate repetitive tasks and streamline workflows.
- Create SLIs and SLOs; increasing application visibility.
- Align with the Product team on SLAs and core service objectives.
- Collaborate with core foundational teams to build and maintain reusable, automated solutions.
- Optimize CI/CD pipelines and developer workflows using Azure DevOps, Github, Octopus Deploy, MirrorD, and Flux.
- Lead incident response; taking an active part in communications, investigations, remediation and post-mortems.
- Help design and build systems that improve the reliability, resiliency and maintainability of our data systems and products.
- Provide assistance and support to other teams focused on database related applications methodologies and system resources.
Requirements
- Experience managing public cloud environments.
- Proficient in contributing to IaC technologies involving expertise in writing, managing, and optimising infrastructure with tools such as Terraform.
- Experience using CI/CD tools, building pipelines, templates and troubleshooting.
- Experience working with a cloud monitoring solution.
- Experience with Data Warehouses or similar.
- Experience with Snowflake or other DBMS platforms.
- Expert in database fundamentals and SQL.
Skills
- Kubernetes
- Docker
- Python
- PowerShell
- Go
- Terraform
- Helm
- Go CRDs
- Azure DevOps
- Github
- Octopus Deploy
- MirrorD
- Flux
- Datadog
- NGINX
- Cloudflare
- SQL
Location
- London, Old Street
Work Type
- Hybrid
- 2 Days in Office
Benefits
- Private Healthcare including dental and opticians services through Vitality
- Worldwide travel insurance through Vitality
- Anniversary Rewards (£250, £500, £750, 4-week fully paid sabbatical)
- Salary Sacrifice Pension Scheme up to 7% match
- 28 days holiday (plus bank holidays)
- Annual Learning and Wellbeing Budget
- Enhanced Parental Leave
- Cycle to Work Scheme
- Season Ticket Loan
- 6 free therapy sessions per year
- Dog Friendly Offices
- Free drinks and snacks in our offices
About the Company
- Capital on Tap started because small businesses were underserved. Big banks were slow, their products weren't fit for purpose, and small business owners often couldn't access what they needed. Today we're a financial platform - not just a credit card company. We offer a best-in-class business credit card, SME-focused spend management platform, a savings product that hit £1 billion in funds within its first year, and a growing suite of tools and financial products that make running a small business easier.
- 1,000+ employees, £20bn in annual card spend, 200,000+ customers, 17,000+ Trustpilot reviews averaging 4.7 stars, and we're profitable. We’ve done a pretty good job so far, but we’re just getting started!
Equal Opportunity
- We welcome, consider and encourage applications from anyone who shares our commitment to inclusivity. Join us in creating a space where authenticity thrives, and everyone can do their best work.
