About the Role
Mount Thor is seeking a software engineer to develop systems for managing a large-scale compute fleet. The Fleet team is responsible for the software layer that transforms raw compute capacity into a functional fleet, defining how machines enter, operate within, recover, and leave the fleet. This role involves connecting physical capacity to workload demand and establishing systems and policies for fleet reliability, security, and efficiency. The team aims to maximize healthy, schedulable capacity and minimize stranded capacity. Additionally, the team is a proving ground for agentic engineering, building systems that software agents can safely inspect and operate, with humans handling policy, risk management, and novel failures. The team collaborates across hardware, operating systems, networking, scheduling, security, and data center operations.
Responsibilities
- Own the technical strategy and roadmap for the fleet control plane and full machine lifecycle.
- Lead complex work across teams and systems.
- Build the distributed control plane and node-level software for inventory, configuration, health, and lifecycle state management.
- Ensure all actions are safe, observable, auditable, and recoverable.
- Automate capacity ingestion across Apple hardware generations, including provisioning, validation, configuration, updates, reimaging, diagnostics, repair, and return to service.
- Connect fleet health and capacity to workload scheduling to improve availability, placement, utilization, recovery time, and speed of capacity reaching production.
- Integrate agentic development and operation into the team's core practices.
- Utilize coding agents for investigation, implementation, testing, and operations.
- Develop interfaces for software agents to inspect state, take safe action, verify results, and escalate exceptions.
- Establish strong operational practices, including defining health signals and service objectives.
- Lead incidents and improve on-call health.
- Implement system improvements based on failure analysis.
Requirements
- Experience building and operating large production systems that other teams depend on.
- Strong software engineering skills in Go, Python, Rust, or a similar language.
- Experience with distributed systems, control planes, state machines, controllers, or durable workflows.
- Strong knowledge of Linux, macOS, or Unix systems, including boot flows, processes, networking, storage, containers, and system performance.
- Experience with bare-metal compute, machine provisioning, Kubernetes, workload schedulers, or large server fleets.
- Experience connecting node-level software to distributed control planes or automated operators.
- Deep experience using coding agents to build production software, including providing context, tools, tests, and constraints for reliable results.
- A track record of leading complex, multi-team infrastructure work from strategy through production.
- Strong operational judgment, with designs that account for partial failure, safe retries, auditability, and recovery.
Skills
- Go
- Python
- Rust
- Distributed systems
- Control planes
- State machines
- Controllers
- Durable workflows
- Linux
- macOS
- Unix systems
- Boot flows
- Processes
- Networking
- Storage
- Containers
- System performance
- Bare-metal compute
- Machine provisioning
- Kubernetes
- Workload schedulers
- Large server fleets
- Node-level software
- Automated operators
- Coding agents
- Agentic development
- Operational practices
- Incident leadership
- On-call health management
- Fleet-management systems
- Apple Silicon
- macOS infrastructure
- Host agents
- Health daemons
- Provisioning pipelines
- Automated repair systems
- Scheduling
- Capacity management
- Permissions
- Validation
- Rollback
- Human escalation
- Availability improvements
- Utilization improvements
- Provisioning speed improvements
- Recovery time improvements
Location
- Datacenter scale
Work Type
- Full-time
Experience Level
- Senior
About the Company
- Mount Thor makes Apple hardware (macOS and Apple Silicon) available and performant at datacenter scale for AI workloads.
- The company transforms consumer hardware into an infrastructure platform for elastic compute consumption via developer-friendly interfaces.
- Customers utilize Mount Thor for developing computer-use model capabilities, deploying long-running agents, and accelerating agentic engineering.
