Member of Technical Staff, Inference at Mount Thor | CA, US | Rezi

Member of Technical Staff, Inference at Mount Thor

Member of Technical Staff, Inference

Mount Thor · CA, US

2 weeks ago

Member of Technical Staff, Inference

Mount Thor · CA, US

20 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Member of Technical Staff, Inference role.

Rezi rewrites your resume against Mount Thor's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Member of Technical Staff, Inference posting at Mount Thor — free, in seconds.

About the Role

Mount Thor is building a production inference platform on Apple Silicon, offering elastic compute through developer-friendly interfaces. You will own the stack from GPU kernels and model execution to distributed scheduling, networking, and customer-facing serving systems, establishing the technical foundation for inference at the company.

Responsibilities

  • Build and optimize the inference engine, implementing model execution, scheduling, KV-cache management, and optimizations like quantization and speculative decoding.
  • Write performance-critical kernels and runtime code, optimizing operations across Metal, MLX, and compiler internals.
  • Design for Apple Silicon, considering unified memory, bandwidth, compute capabilities, and OS behavior.
  • Build distributed inference and its communication layer, implementing model sharding and optimizing collective communication.
  • Own inference scheduling and routing, building request queues, admission control, and autoscaling.
  • Ship a production inference service, building model loading, serving APIs, and integrating with fleet systems.
  • Make performance reproducible by building benchmarks and profiling tools.
  • Set the engineering direction, turning customer workloads into technical priorities and contributing to open-source projects.

Requirements

  • A record of building and optimizing production inference engines, GPU compute software, or distributed ML systems, with substantial depth in at least one and hands-on work across multiple layers.
  • Strong systems programming skills in C++ or Rust, proficiency in Python, and experience working inside performance-critical libraries and runtimes.
  • A practical understanding of transformer inference, including attention, prefill and decode, batching, KV caches, quantization, and their compute and memory costs.
  • Experience with GPU programming and performance analysis, including memory hierarchies, parallel execution, synchronization, and numerical precision.
  • Strong distributed-systems fundamentals, including scheduling, concurrency, networking, failure handling, and resource management.
  • Experience taking software from architecture through production deployment, debugging, and ongoing operation.
  • The ability to turn an ambiguous performance problem into a measured bottleneck, an implementation, and a verified improvement.
  • The judgment and ownership to establish a new technical area, prioritize the work, and explain tradeoffs clearly to engineers and customers.

Skills

  • C++
  • Rust
  • Python
  • GPU programming
  • Performance analysis
  • Distributed systems
  • Transformer inference
  • Metal
  • MLX

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Founding

About the Company

  • Mount Thor makes Apple hardware—macOS and Apple Silicon—available at datacenter scale for AI workloads.
  • We build the infrastructure around that hardware to deliver elastic compute through developer-friendly interfaces.
  • We are building an inference platform that gives developers a straightforward way to deploy and serve models on Apple Silicon, backed by deep engineering across hardware, runtimes, and distributed systems.