About the Role
Mount Thor is seeking an engineer to design and enhance the data center systems that power our compute fleet. This on-site role involves direct responsibility for production infrastructure.
Responsibilities
- Own the technical direction for data center architecture and operations engineering.
- Define standards for rack design, power, cooling, cabling, network integration, serviceability, safety, and security.
- Design and automate the commissioning of new capacity.
- Build repeatable systems for installation checks, inventory validation, connectivity testing, hardware qualification, and production readiness.
- Build observability for the physical fleet.
- Collect and validate hardware, power, thermal, network, and environmental telemetry.
- Make telemetry data useful for diagnosis, capacity planning, and automated health decisions.
- Lead complex failure investigations across hardware and software boundaries.
- Work directly with production systems to isolate issues involving hardware, firmware, macOS, networking, storage, power, or cooling.
- Turn failures into engineering improvements.
- Build better diagnostics, test tools, hardware designs, maintenance procedures, and automated workflows.
- Design the systems behind maintenance and hardware lifecycle management.
- Improve preventive maintenance, repair, spare planning, vendor escalation, hardware refresh, and secure decommissioning.
- Make agentic development part of daily engineering.
- Use agents to analyze telemetry, develop tools, investigate failures, and improve documentation.
- Build structured data and workflows that agents can use safely.
- Lead production incidents and planned infrastructure changes.
- Participate in on-call coverage and coordinate with Fleet, networking, security, data center partners, and hardware vendors.
Requirements
- Designed or operated production data center, HPC, cloud, or large-scale compute infrastructure.
- Strong knowledge of data center systems, including racks, power, cooling, structured cabling, networks, and environmental monitoring.
- Deep systems knowledge across macOS, Linux, or Unix.
- Ability to troubleshoot boot flows, firmware, storage, networking, system performance, and hardware-software interactions.
- Experience designing commissioning, qualification, diagnostic, or maintenance systems for physical infrastructure.
- Strong programming skills in Python, Go, Rust, Bash, or a similar language.
- Experience building telemetry pipelines, dashboards, alerts, or diagnostic tools using metrics, logs, and time-series data.
- A record of leading root-cause analysis and turning individual failures into broader reliability improvements.
- Strong operational judgment.
- Experience using coding agents or other AI tools for engineering, investigation, and data analysis.
- Comfort working in active data center environments.
- Willingness to join an on-call rotation.
- Willingness to travel to other sites when needed.
Skills
- macOS
- Apple Silicon
- Python
- Go
- Rust
- Bash
- Data center architecture
- Data center operations
- Rack design
- Power systems
- Cooling systems
- Cabling
- Network integration
- Serviceability
- Safety standards
- Security standards
- Commissioning automation
- Installation checks
- Inventory validation
- Connectivity testing
- Hardware qualification
- Production readiness
- Physical fleet observability
- Telemetry collection and validation
- Hardware telemetry
- Power telemetry
- Thermal telemetry
- Network telemetry
- Environmental telemetry
- Failure investigation
- Hardware lifecycle management
- Preventive maintenance
- Repair planning
- Spare parts planning
- Vendor escalation
- Hardware refresh
- Secure decommissioning
- Agentic development
- AI tools for engineering
- AI tools for investigation
- AI tools for data analysis
- Incident management
- Infrastructure change management
- On-call rotation
- macOS recovery
- macOS firmware
- macOS secure boot
- macOS hardware diagnostics
- Automated device restoration
- Custom rack design
- Industrial protocols
- Vendor APIs
- Failure analysis
- Predictive maintenance
- Component-life tracking
- Data-driven spare planning
- AI-assisted operational systems
Location
- On-site
Work Type
- On-site
- Full-time
Experience Level
- Senior
About the Company
- Mount Thor makes Apple hardware (macOS and Apple Silicon) available and performant at datacenter scale for AI workloads.
- We take consumer hardware and build the infrastructure platform around it to enable consumption as elastic compute exposed through developer-friendly interfaces.
- Our customers use us to develop computer-use model capabilities, deploy long-running agents, and accelerate agentic engineering.
