About the Role
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.
Responsibilities
- Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
- Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
- Lead incident response, conduct root cause analysis, and implement preventive measures
- Develop and maintain disaster recovery and business continuity plans
- Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
- Automate deployment pipelines, monitoring, and operational workflows
- Optimize cloud resource utilization and cost management
- Build and maintain internal tools and services to improve operational efficiency
- Collaborate with development teams to implement reliability best practices
- Conduct code reviews and provide technical guidance on system design
- Develop monitoring solutions, alerting systems, and observability frameworks
- Integrate security practices into CI/CD pipelines (SAST/DAST)
- Implement and maintain security controls across infrastructure and applications
- Ensure compliance with industry standards and regulatory requirements
- Conduct security assessments and vulnerability management
- Mentor junior SRE team members and promote SRE culture across the organization
- Partner with software engineering teams to improve system reliability
- Drive technical initiatives and contribute to architectural decisions
- Document processes, runbooks, and technical specifications
Requirements
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
Skills
- Java
- Python
- Node.js
- Microservices architecture
- Distributed systems
- Data structures
- Algorithms
- Design patterns
- Clean, maintainable, and testable code
- AWS services (Lambda, ECS, EC2, Fargate, S3, EBS, EFS, RDS, DynamoDB, Aurora, VPC, Route53, CloudFront, API Gateway, CloudWatch, X-Ray)
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
- GitLab (CI/CD pipelines, runners, GitOps)
- Terraform
- Docker
- Kubernetes/ECS
- Configuration management tools
- SAST tools
- DAST methodologies
- Security best practices
- OWASP Top 10
- Compliance frameworks
- Secrets management
- Identity and Access Management (IAM)
- Monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with AWS X-Ray
Location
- Richmond, VA
- San Francisco, CA
Work Type
- Full-time
- Onsite
Experience Level
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
Education Level
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
Salary/Compensations
- Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco)
Benefits
- The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment.
About the Company
- When you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems.
- We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.
Equal Opportunity
- The SF Fed is an Equal Opportunity Employer.
- The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
