DevOps / Platform Engineer at Ineffable Intelligence | United Kingdom | Rezi

DevOps / Platform Engineer at Ineffable Intelligence

DevOps / Platform Engineer

Ineffable Intelligence · United Kingdom

1 weeks ago

DevOps / Platform Engineer

Ineffable Intelligence · United Kingdom

7 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a foundational hire on the platform team, you'll be joining at a point where individual decisions shape how the platform is designed, not just maintained. You'll own the infrastructure that the mission depends on; keeping large-scale GPU compute reliable, developer environments fast and the whole platform humming. This is an opportunity for someone who wants to be a part of defining a new paradigm in AI and wants their work to directly speed up the experimentation and progress of our ambitious research.

Responsibilities

  • Manage Kubernetes clusters and deploy internal tools.
  • Hands-on with GPU scheduling tools like KAI scheduler and Kueue to keep thousands of GPUs busy and productive.
  • Address hardware failures and cluster-scale chaos, building systems to prevent recurrence.
  • Manage cloud infrastructure on Google Cloud and other major providers.
  • Implement log management and monitoring with tools like QuickWit, Grafana or Datadog.
  • Develop developer environments using tools like Tailscale, Workbrew and dev containers.
  • Write clean, well-crafted code in Python and Rust for tools and infrastructure.

Requirements

  • Experience with Kubernetes and containers.
  • Experience with GPU scheduling at scale.
  • Experience making large-scale infrastructure resilient.
  • Experience with cloud infrastructure on Google Cloud and other major providers.
  • Experience with observability tools like QuickWit, Grafana or Datadog.
  • Experience with developer environment tooling like Tailscale, Workbrew and dev containers.
  • Proficiency in Python and Rust.

Skills

  • Kubernetes
  • Containers
  • GPU scheduling
  • KAI scheduler
  • Kueue
  • Cloud infrastructure
  • Google Cloud
  • Observability
  • Log management
  • Monitoring
  • QuickWit
  • Grafana
  • Datadog
  • Tailscale
  • Workbrew
  • Dev containers
  • Python
  • Rust

Location

  • London

Work Type

  • Onsite

Experience Level

  • Foundational hire

About the Company

  • Our mission is to make first contact with superintelligence.
  • We are creating a superlearner that discovers all knowledge from its own experience, from elementary motor skills through to profound intellectual breakthroughs.
  • This superlearning capability - the ability to endlessly discover knowledge and skills, without relying on human data - will be driven by the world’s most powerful reinforcement learning algorithms.
  • The superlearner is expected to rediscover and then transcend the greatest inventions in human history, such as language, science, mathematics and technology.
  • If successful, this will represent a scientific breakthrough of comparable magnitude to Darwin: where his law explained all life, our law will explain and build all intelligence.