Software Engineer, Systems Generalist at Thinking Machines Lab | CA, US | Rezi

Software Engineer, Systems Generalist at Thinking Machines Lab

Software Engineer, Systems Generalist

Thinking Machines Lab · CA, US

3 weeks ago

Software Engineer, Systems Generalist

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking generalist infrastructure and systems engineers to build and scale the core infrastructure powering our foundation models and supporting research and product development teams. You will solve complex distributed systems problems and build robust, scalable platforms, working directly with researchers to accelerate experiments and improve infrastructure efficiency.

Responsibilities

  • Architecting and scaling core infrastructure.
  • Solving complex distributed systems problems.
  • Building robust, scalable platforms.
  • Working across the full technical stack.
  • Supporting teams that train, research, and serve AI models.
  • Building underlying infrastructure for clusters to reliably and safely train frontier models.
  • Building systems and running large Kubernetes clusters with GPU workloads.
  • Building infrastructure to support Tinker.
  • Designing and optimizing data pipelines using tools like Spark and other modern data infrastructure technologies.
  • Building scalable, reliable data infrastructure while embedding governance best practices.
  • Building tooling, systems, and frameworks to ensure well-configured, optimized developer environments.
  • Taking initiative to work across different stacks and teams to ensure successful product shipment.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large-scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Ability to thrive in a highly collaborative environment involving cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative.
  • Strong debugging across application, OS, and network layers.
  • Proficiency in Python or Rust (or similar), containers, and modern CI.
  • Experience with Kubernetes, controllers/operators, or performance profiling.
  • Familiarity with GPU/ML workflows or large-scale data/eval pipelines.

Skills

  • Python
  • Rust
  • Kubernetes
  • Slurm
  • Spark
  • CI/CD
  • Debugging
  • Performance Profiling
  • GPU/ML workflows
  • Data Pipelines

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Health benefits
  • Dental benefits
  • Vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.