Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Site Reliability Engineer role.
Rezi rewrites your resume against Cognition's job description. Free.

Tailor your resume to this Site Reliability Engineer role.
Rezi rewrites your resume against Cognition's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Site Reliability Engineer posting at Cognition — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Site Reliability Engineer posting at Cognition — free, in seconds.
About the Role
This role focuses on ensuring the production reliability of user-facing products and the platform engineering that enables rapid and confident shipping. You will own SLOs, incident response, and on-call duties, alongside CI/CD pipelines, deployment infrastructure, and developer tooling. The best SREs at Cognition understand that reliability is engineered in from the start.
Responsibilities
- Define and own SLOs, SLIs, and error budgets for Devin and Windsurf.
- Build monitoring, alerting, and observability systems for service health.
- Lead incident response with speed and clarity.
- Run blameless postmortems to drive improvements.
- Build runbooks and tooling for effective on-call.
- Own deployment pipelines, release infrastructure, and internal developer tooling.
- Reduce toil systematically through automation.
- Manage cloud infrastructure through code.
- Build reproducible, auditable, version-controlled environments.
- Model growth and forecast resource needs for capacity planning.
- Ensure infrastructure stays ahead of demand.
- Profile and improve system performance.
- Treat security as a reliability requirement.
- Ensure misconfigurations, vulnerabilities, and access failures are caught and remediated urgently.
- Partner with product and engineering teams to build reliability in from the start.
- Identify single points of failure during architecture reviews.
Requirements
- Deep experience running production systems at scale, including SLOs, error budgets, on-call rotations, and incident command.
- Strong software engineering fundamentals; involves writing real code.
- Proficiency with cloud infrastructure (AWS, GCP, or Azure).
- Proficiency with container orchestration (Kubernetes).
- Proficiency with infrastructure as code (Terraform or equivalent).
- Experience building and owning CI/CD pipelines and deployment infrastructure for fast-moving product teams.
- Strong observability instincts: ability to instrument systems, build useful dashboards, and design effective alerts.
- Track record of reducing toil systematically through automation.
- Comfort owning incidents end to end: detection, triage, mitigation, resolution, and postmortem.
- Product empathy to understand reliability from a user's perspective.
- Experience with developer-facing products or platforms is a strong plus.
Skills
- Production Reliability
- SLOs
- SLIs
- Error Budgets
- Monitoring
- Alerting
- Observability
- Incident Response
- On-Call
- Platform Engineering
- CI/CD
- Deployment Infrastructure
- Developer Tooling
- Infrastructure as Code
- Cloud Infrastructure (AWS, GCP, Azure)
- Kubernetes
- Terraform
- Capacity Planning
- Performance Tuning
- Security
- Automation
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
Salary/Compensations
- $260,000 - $300,000 + significant early-stage equity
Benefits
- Medical, Dental, Vision: Fully paid for you and your dependents
- 401(k): Company match included
- Perks: Private chef, cozy slippers, endless snacks, and more
About the Company
- An applied AI lab building end-to-end software agents.
- Makers of Devin, the first AI software engineer.
- Team includes world-class competitive programmers, former founders, and leaders from companies at the cutting edge of AI.
- Focuses on solving major world problems and building AI that reasons on real-world tasks.
- Small, highly selective team shipping products used by hundreds of thousands of developers daily.
- High ownership and high trust environment.
- Rewards engineers who are proactive, systematic, and treat reliability as a craft.
Equal Opportunity
- Cognition is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.
- We are committed to providing reasonable accommodations for candidates with disabilities throughout the hiring process.