Member of Technical Staff, Fleet Operations at Mount Thor | CA, US | Rezi

Member of Technical Staff, Fleet Operations at Mount Thor

Member of Technical Staff, Fleet Operations

Mount Thor · CA, US

2 weeks ago

Member of Technical Staff, Fleet Operations

Mount Thor · CA, US

21 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Mount Thor is seeking an engineer to design and enhance the data center systems that power our compute fleet. This on-site role involves direct responsibility for production infrastructure.

Responsibilities

  • Own the technical direction for data center architecture and operations engineering.
  • Define standards for rack design, power, cooling, cabling, network integration, serviceability, safety, and security.
  • Design and automate the commissioning of new capacity.
  • Build repeatable systems for installation checks, inventory validation, connectivity testing, hardware qualification, and production readiness.
  • Build observability for the physical fleet.
  • Collect and validate hardware, power, thermal, network, and environmental telemetry.
  • Make telemetry data useful for diagnosis, capacity planning, and automated health decisions.
  • Lead complex failure investigations across hardware and software boundaries.
  • Work directly with production systems to isolate issues involving hardware, firmware, macOS, networking, storage, power, or cooling.
  • Turn failures into engineering improvements.
  • Build better diagnostics, test tools, hardware designs, maintenance procedures, and automated workflows.
  • Design the systems behind maintenance and hardware lifecycle management.
  • Improve preventive maintenance, repair, spare planning, vendor escalation, hardware refresh, and secure decommissioning.
  • Make agentic development part of daily engineering.
  • Use agents to analyze telemetry, develop tools, investigate failures, and improve documentation.
  • Build structured data and workflows that agents can use safely.
  • Lead production incidents and planned infrastructure changes.
  • Participate in on-call coverage and coordinate with Fleet, networking, security, data center partners, and hardware vendors.

Requirements

  • Designed or operated production data center, HPC, cloud, or large-scale compute infrastructure.
  • Strong knowledge of data center systems, including racks, power, cooling, structured cabling, networks, and environmental monitoring.
  • Deep systems knowledge across macOS, Linux, or Unix.
  • Ability to troubleshoot boot flows, firmware, storage, networking, system performance, and hardware-software interactions.
  • Experience designing commissioning, qualification, diagnostic, or maintenance systems for physical infrastructure.
  • Strong programming skills in Python, Go, Rust, Bash, or a similar language.
  • Experience building telemetry pipelines, dashboards, alerts, or diagnostic tools using metrics, logs, and time-series data.
  • A record of leading root-cause analysis and turning individual failures into broader reliability improvements.
  • Strong operational judgment.
  • Experience using coding agents or other AI tools for engineering, investigation, and data analysis.
  • Comfort working in active data center environments.
  • Willingness to join an on-call rotation.
  • Willingness to travel to other sites when needed.

Skills

  • macOS
  • Apple Silicon
  • Python
  • Go
  • Rust
  • Bash
  • Data center architecture
  • Data center operations
  • Rack design
  • Power systems
  • Cooling systems
  • Cabling
  • Network integration
  • Serviceability
  • Safety standards
  • Security standards
  • Commissioning automation
  • Installation checks
  • Inventory validation
  • Connectivity testing
  • Hardware qualification
  • Production readiness
  • Physical fleet observability
  • Telemetry collection and validation
  • Hardware telemetry
  • Power telemetry
  • Thermal telemetry
  • Network telemetry
  • Environmental telemetry
  • Failure investigation
  • Hardware lifecycle management
  • Preventive maintenance
  • Repair planning
  • Spare parts planning
  • Vendor escalation
  • Hardware refresh
  • Secure decommissioning
  • Agentic development
  • AI tools for engineering
  • AI tools for investigation
  • AI tools for data analysis
  • Incident management
  • Infrastructure change management
  • On-call rotation
  • macOS recovery
  • macOS firmware
  • macOS secure boot
  • macOS hardware diagnostics
  • Automated device restoration
  • Custom rack design
  • Industrial protocols
  • Vendor APIs
  • Failure analysis
  • Predictive maintenance
  • Component-life tracking
  • Data-driven spare planning
  • AI-assisted operational systems

Location

  • On-site

Work Type

  • On-site
  • Full-time

Experience Level

  • Senior

About the Company

  • Mount Thor makes Apple hardware (macOS and Apple Silicon) available and performant at datacenter scale for AI workloads.
  • We take consumer hardware and build the infrastructure platform around it to enable consumption as elastic compute exposed through developer-friendly interfaces.
  • Our customers use us to develop computer-use model capabilities, deploy long-running agents, and accelerate agentic engineering.