About the Role
As a Senior Site Reliability Engineer, you will own reliability and infrastructure strategy across the organization. Decisions made on system design, tooling, and process will directly affect every engineering team's ability to operate and ship at scale. This role requires a strong background in cloud infrastructure and platform engineering, with experience across infrastructure-as-code, CI/CD and release automation, observability, security, and disaster recovery. You should possess strong technical judgment, the ability to set standards for other engineers, and a proactive approach to eliminating operational risk.
Responsibilities
- Drive infrastructure and reliability strategy that connects directly to business outcomes.
- Analyze, test, and evolve systems to improve reliability and performance at an architectural/infrastructure level.
- Architect multi-region, multi-AZ infrastructure with clear failover and disaster recovery strategies.
- Define observability strategy and standards, and develop tooling and dashboards.
- Define and own SLOs and error budgets for services, and use them to prioritize reliability work.
- Take a leading role in major incidents and lead troubleshooting on complex production issues.
- Reduce operational toil by building automation and self-service platforms.
- Develop and maintain design, troubleshooting, and runbook standards.
- Mentor other engineers and raise the bar on production ownership, testing, and code review.
- Apply AI-assisted development to infrastructure problems.
- Drive cloud cost optimisation at the organizational level.
Requirements
- 5+ years of experience in SRE, Platform Engineering, DevOps or other related roles.
- Deep understanding of SRE and platform engineering principles.
- Extensive experience using, configuring, and setting standards for modern observability tools.
- Expert-level experience with cloud native and container technology such as Docker.
- Hands-on experience designing and managing Kubernetes clusters at scale.
- Deep experience defining infrastructure-as-code standards and module libraries using tools such as Terraform.
- Comfortable scripting and developing internal tooling with Bash and at least one programming language (e.g. Python, Go).
- Fluent with AI-driven development environments like Cursor, Claude Code, or Gemini.
- Experience working with Linux.
- Strong understanding of networking, distributed systems, and system architecture at scale.
- Proven experience deploying, scaling, and monitoring web applications and databases across multi-region or high-availability environments.
- Expert-level knowledge of AWS and/or GCP platforms.
- Experience driving cloud cost optimisation and platform decisions at an organisational level.
- Takes ambiguous problems and drives them to shipped outcomes.
- Balances speed with quality.
- Manages risk proactively.
- Creates clarity from ambiguity; documents decisions.
- Influences through evidence and collaboration.
- Communicates technical concepts clearly to engineers, product, and business stakeholders.
- Treats AI tools as essential infrastructure.
- Understands LLM strengths and limitations.
- Thinks in leverage: automates the repetitive, focuses human attention on judgement calls.
- Helps others adopt AI workflows.
Skills
- Cloud infrastructure
- Platform engineering
- Infrastructure-as-code
- CI/CD
- Release automation
- Observability
- Security
- Disaster recovery
- AWS
- GCP
- Docker
- Kubernetes
- Terraform
- Bash
- Python
- Go
- AI-assisted development
- Networking
- Distributed systems
- System architecture
- Web applications
- Databases
- Cloud cost optimisation
- FinOps
Location
- Global
Work Type
- Hybrid
Experience Level
- Senior
- 5+ years of experience
Education Level
- Bachelor's degree in Computer Science/Engineering or equivalent experience
Benefits
- Flexible Work Environment
- Global company, with the opportunity to work from any of our offices for 4 weeks a year
- Employee Stock Options
- CG Gives programs
- Social Initiatives
About the Company
- Cover Genius is a Series E Insurtech that protects the global customers of the world’s largest digital companies.
- Partners integrate with XCover, our award-winning insurance distribution platform, to embed protection for millions of customers worldwide each year.
- Our team and products have been recognized with dozens of awards.
- Our diverse team across 20+ countries and many language groups commits itself to diverse cultural programs.
- Our People are Bold, Authentic, Purposeful and Inspired.
- Our People are not Perfect, Traditional, Complacent or Cautious.
Equal Opportunity
- Cover Genius promotes diversity and inclusivity. We don't tolerate discrimination, demeaning treatment of anyone, or harassment due to race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or any other legally protected status.
- By submitting your application, you acknowledge that we may collect, store and process your personal data for recruitment purposes.
- To ensure a fair evaluation, we may use AI to assist in sorting applications, but all final decisions are made by our hiring team and no candidate dispositions are automated.
- We will keep your information on file for three years from the date of your application.
