About the Role
As an SRE on our team, you'll own the reliability, performance, and scale of the distributed storage and data platform systems that power Apple's services. You'll debug replication and consensus failures, tune systems at petabyte scale, and write production code that operates the platform. We firmly believe in ownership, with software engineers accountable for the code they write.
Responsibilities
- Debug replication and consensus failures
- Tune systems at petabyte scale
- Write production code that operates the platform
- Partner with development teams on system-wide architecture rather than individual components
- Take part in on-call rotations and incident response to keep critical systems healthy
- Define SLIs/SLOs
- Build observability
- Use error budgets to drive reliability decisions
- Perform data migration, disaster recovery, or capacity planning at scale
Requirements
- Experience in managing and scaling large-scale distributed systems in a private or hybrid cloud environment
- Comfortable designing, writing, and releasing production code in languages such as Go or Python
- Able to debug and reason about how distributed systems fail and perform at scale
- Willingness to take part in on-call rotations and incident response to keep critical systems healthy
- Contributions to distributed-systems internals, open-source data infrastructure, or storage/database engines
- Experience defining SLIs/SLOs, building observability, and using error budgets to drive reliability decisions
- A good grasp of Unix internals and networking fundamentals
- Experience with data migration, disaster recovery, or capacity planning at scale
Skills
- Go
- Python
- Unix internals
- Networking fundamentals
Location
- Private cloud
- Hybrid cloud
Work Type
- Full-time
Experience Level
- Large-scale distributed systems management
- Production code development
- System-wide architecture design
- Distributed systems debugging and performance analysis
- On-call rotations and incident response
- SLI/SLO definition
- Observability building
- Error budget utilization
- Unix internals understanding
- Networking fundamentals understanding
- Data migration experience
- Disaster recovery experience
- Capacity planning experience
About the Company
- At Apple, we believe that innovation flourishes in an environment where ideas are challenged, collaboration is encouraged, and technology is pushed to its limits.
- This environment is only possible when diverse minds come together, bringing unique perspectives and experiences.
- Our people and their ideas inspire innovation in everything we do.
- The Apple Services Engineering (ASE) organisation builds and provides systems and infrastructure that fuel Apple’s services — iCloud, iTunes, Siri, and Maps.
- Our team builds and operates the data platform infrastructure behind them, keeping petabyte-scale workloads fast, resilient, and reliable.
- The platform runs on large-scale distributed systems, including object stores, databases, and data pipelines, on Linux across private and hybrid cloud.
- You'll work on storage engines, distributed consensus, and data-flow internals
