About the Role
Apple's cloud AI inference platform is rapidly growing, requiring engineering managers to lead teams building components of a complex inference stack. This role is crucial for delivering generative AI inference with verifiable privacy guarantees, operating within Apple's cloud infrastructure.
Responsibilities
- Own a set of components within the inference stack.
- Hire, grow, and lead a team of engineers.
- Own delivery against a roadmap.
- Lead design reviews and make architectural decisions.
- Run on-call and incident management practices.
- Partner across time zones with ML research, hardware, platform, security, privacy, SRE, and product teams.
- Work with teams in London, Cupertino, and Seattle.
- Absorb shifting priorities on behalf of the team.
- Steward APIs and compatibility across versions and hardware generations.
- Reduce cycle time without lowering quality.
Requirements
- Experience managing software engineers, including hiring, coaching, feedback, and performance management.
- Strong software engineering background in systems, backend, distributed systems, or platform work.
- Ability to engage deeply in design trade-offs.
- Demonstrated ownership of delivery on an infrastructure or platform team (roadmap, sequencing, cross-team dependencies, shipped results).
- Agile mindset and track record of operating effectively in ambiguity.
- Ability to absorb rapidly shifting priorities without losing execution discipline or team trust.
- Excellent written communication.
- Effective working habits across geographies and time zones.
- UK/US collaboration experience.
- Genuine security and privacy mindset for systems handling sensitive user content.
- Technical credibility to engage in design reviews and read code.
- Ability to hold your own in a design review.
- Ability to tell a good argument from a confident one.
Skills
- Systems engineering
- Backend engineering
- Distributed systems
- Platform engineering
- Agile methodologies
- LLM inference
- Model serving at scale
- Batching and scheduling
- KV-cache reuse
- Paged attention
- Prefix caching
- Disaggregated serving
- Speculative decoding
- Quantisation
- Model parallelism
- GPU performance
- Custom-accelerator performance
- ML runtime internals
- Framework internals
- Production operations for latency-sensitive services
- SLOs and error budgets
- Observability
- Capacity planning
- Canary and rollback discipline
- Developer experience
- Build infrastructure
- Test infrastructure
- Swift
- C++
- Rust
- Go
- Python tooling
- Privacy-preserving systems
- Security-sensitive systems
- Attested systems
- Reasoning rigorously about logging and measurement
Location
- London
Work Type
- Onsite
Experience Level
- Engineering Manager
About the Company
- Apple's cloud AI inference platform is growing quickly.
- The organisation builds a complex inference stack.
- The generative-AI landscape is rapidly evolving.
- Private Cloud Compute is the system that lets Apple Intelligence reach beyond the device without compromising user privacy.
- It is the server software behind Apple Intelligence.
- The stack includes on-device client frameworks, cloud services for attestation, routing, and orchestration, an inference engine, and model runtimes executing across heterogeneous hardware.
- The platform addresses challenges in context and cache management, model asset management, throughput and latency, observability, and developer/test infrastructure.
- The company values a clear technical direction amidst change and defining roadmaps.
- Collaboration occurs across time zones with various teams.
- The engineering environment is challenging due to privacy constraints that limit usual shortcuts.
