Senior Software Engineer, ML Infrastructure at Gridmatic Inc | San Francisco, California, US | Rezi

Senior Software Engineer, ML Infrastructure at Gridmatic Inc

Senior Software Engineer, ML Infrastructure

Gridmatic Inc · San Francisco, California, US

Today

Senior Software Engineer, ML Infrastructure

Gridmatic Inc · San Francisco, California, US

5 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure that powers Gridmatic. Our platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight schedules, battery dispatch commands must execute reliably in real time, and ML models need to train and deploy continuously as new data arrives.

Responsibilities

  • Design and build the compute platform that runs Gridmatic's production services, batch jobs, and ML training workloads
  • Own production workflow orchestration end to end, from the cluster and node pools it runs on to the abstractions our teams build on top of it
  • Make our ML iteration cycle faster and cheaper by profiling workflows, cutting latency, and improving how we use compute
  • Build the observability and developer tooling that lets engineers understand and troubleshoot their own workloads, from post-run cost reporting to monitoring and alerting
  • Improve the reliability and resiliency of production workloads, including how we handle capacity constraints and multi-region routing
  • Manage GPU and accelerator capacity: node pools, drivers, spot vs. on-demand tradeoffs, and scheduling and quota so training and batch jobs get the compute they need without overspending
  • Own autoscaling and quota management so the platform scales up under load and down to zero when idle
  • Work closely with the ML team to find pain points and quickly ship solutions
  • Make architectural decisions that shape how we build software as we grow

Requirements

  • Significant experience building and operating production infrastructure on a public cloud platform (GCP, AWS, or Azure)
  • Strong distributed systems and infrastructure skills: standing up services, scaling and debugging Kubernetes (GKE), writing Terraform, and comfort with cloud networking, IAM, and secrets management
  • Hands-on experience with workflow orchestration tools like Flyte, Temporal, or Airflow
  • Proficiency in Python, and either already know Go or have experience with a similar systems language (C++, Java, Rust) and are excited to work in Python and Go day-to-day
  • Experience working closely with ML engineers or data scientists as your customers
  • Clear communication, whether writing a design doc, reviewing code, or explaining a complex system to someone new to it
  • Experience running GPU or accelerator workloads on Kubernetes (node pools, drivers, scheduling, quota) is strongly preferred
  • Familiarity with observability tooling (Grafana + Google Cloud Monitoring)
  • Experience building internal developer platforms or tooling that other engineers rely on
  • Prior work in domains where latency and reliability have direct business consequences

Skills

  • Python
  • Go
  • Kubernetes
  • GKE
  • Terraform
  • Cloud Networking
  • IAM
  • Secrets Management
  • Flyte
  • Temporal
  • Airflow
  • Grafana
  • Google Cloud Monitoring

Location

  • GCP

Work Type

  • Full-time

Experience Level

  • Senior

Salary/Compensations

  • $209,000—$235,000 USD

Benefits

  • Continuing Education Opportunities
  • Flexible PTO
  • Medical, Dental and Vision plans with competitive employer contributions
  • Pre-Tax commuter benefits
  • $1500/year non profit donation matching program through Millie
  • Home Office Stipend
  • 401K contribution match up to 4%
  • Company-paid parental leave
  • Company Paid Life Insurance
  • Stock Option Loan Program

About the Company

  • Gridmatic's platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight schedules, battery dispatch commands must execute reliably in real time, and ML models need to train and deploy continuously as new data arrives.