About the Role
We are seeking a Senior Site Reliability Engineer to build and maintain highly available, scalable, and resilient systems. You will drive reliability improvements, mentor engineers, shape incident response culture, and partner with cross-functional teams to embed reliability into the development lifecycle.
Responsibilities
- Design, build, and maintain production infrastructure across cloud platforms ensuring 99.99%+ availability.
- Define and champion SLOs, SLIs, and error budgets to drive data-informed reliability decisions.
- Lead incident response as Incident Commander and conduct blameless post-mortems.
- Develop and maintain infrastructure-as-code and CI/CD pipelines for automated deployments.
- Build and improve observability platforms using tools like Prometheus, Grafana, NewRelic, Splunk, or ELK.
- Automate toil reduction through custom tooling, self-healing systems, and capacity planning.
- Architect and operate container orchestration systems at scale.
- Collaborate with security teams to embed best practices into infrastructure and pipelines.
- Mentor junior and mid-level SREs through code reviews and pair-programming.
- Contribute to on-call rotations and maintain runbooks and alerting procedures.
Requirements
- 7+ years of experience in SRE, DevOps, or platform engineering.
- Strong proficiency in at least one programming language such as Python, Go, or Java.
- Deep hands-on experience with at least one major cloud provider including networking and IAM.
- Expert-level knowledge of Kubernetes and microservices architectures.
- Demonstrated experience defining SLOs/SLIs and managing error budgets.
- Solid understanding of distributed systems concepts and fault-tolerant design.
- Proficiency with infrastructure-as-code tools like Terraform or CloudFormation.
- Experience with CI/CD platforms and GitOps workflows.
- Strong Linux systems administration and networking fundamentals.
- Proven track record of leading incident response and implementing systemic fixes.
Skills
- Python
- Go
- Java
- AWS
- GCP
- Azure
- Kubernetes
- Terraform
- CloudFormation
- Prometheus
- Grafana
- NewRelic
- Splunk
- ELK stack
- CI/CD
- GitOps
- Linux
- TCP/IP
- DNS
- Load balancing
Location
- Austin, TX
Work Type
- Hybrid
- Full-time
Experience Level
- 7+ years
Salary/Compensations
- $111,600.00 - $186,000.00
Benefits
- Competitive base salary with annual performance bonuses
- Comprehensive health, dental, and vision insurance
- Flexible hybrid/remote work model
- Annual learning and development budget
- Generous PTO policy
- Paid parental leave
- Wellness programs
- 401(k) with employer match
- Paid holidays
- Bereavement leave
- Time off to vote
- Jury duty leave
- Volunteer time off
- Military leave
About the Company
- Cox Automotive transforms the way the world buys, owns, and sells cars through groundbreaking technology.
- The company manages iconic consumer brands like Autotrader and Kelley Blue Book.
- The company operates industry-leading dealer-facing companies like vAuto and Manheim.
Equal Opportunity
- Cox Automotive is an equal opportunity employer committed to creating an inclusive environment for all employees.
- All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic.
