Principal Member of Technical Staff, Platform Infrastructure at Edison Scientific | CA, United States | Rezi

Principal Member of Technical Staff, Platform Infrastructure at Edison Scientific

Principal Member of Technical Staff, Platform Infrastructure

Edison Scientific · CA, United States

3 weeks ago

Principal Member of Technical Staff, Platform Infrastructure

Edison Scientific · CA, United States

22 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Principal MTS, you will design, scale, and operate the core platform infrastructure powering autonomous scientific discovery. Your focus will be on orchestrating AI agents at scale, building and managing clusters for thousands of persistent, stateful workloads, developing custom resource definitions (CRDs) and operators, and ensuring the reliability and efficiency of the compute layer. This role influences platform architecture, establishes infrastructure best practices, and partners with engineering and research teams to deliver a production-grade environment.

Responsibilities

  • Architect, implement, and operate Kubernetes clusters supporting thousands of concurrent, persistent resources with high availability and efficient resource utilization.
  • Design and develop custom resource definitions (CRDs) and Kubernetes operators to model and manage domain-specific workloads.
  • Drive strategy for cluster scaling, node pool management, autoscaling policies, and resource quota frameworks.
  • Build and maintain infrastructure-as-code for reproducible, version-controlled environment management.
  • Design and implement robust scheduling, placement, and affinity strategies for heterogeneous workloads.
  • Establish and uphold best practices for observability, monitoring, alerting, and incident response.
  • Own storage and networking strategy within Kubernetes, including persistent volume management, CSI drivers, service mesh, network policies, and ingress architecture.
  • Troubleshoot complex, cross-system infrastructure issues and guide others through effective debugging and remediation.
  • Collaborate with backend, ML, and research teams to understand workload requirements and translate them into reliable infrastructure patterns.

Requirements

  • Typically, 10+ years of professional infrastructure or platform engineering experience with deep hands-on Kubernetes expertise in production environments.
  • Experience designing and implementing custom resource definitions (CRDs) and Kubernetes operators.
  • Track record of operating and scaling Kubernetes clusters supporting thousands of persistent or long-lived resources.
  • Deep understanding of Kubernetes internals and their behavior at scale.
  • Expertise with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated networking, storage, and IAM primitives.
  • Proficiency in at least one systems or backend language for operator development and infrastructure tooling.
  • Hands-on experience with infrastructure-as-code tools and GitOps workflows.
  • Strong working knowledge of container networking, storage, and security.
  • Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production.

Skills

  • Kubernetes
  • Custom Resource Definitions (CRDs)
  • Kubernetes Operators
  • Infrastructure-as-Code (Terraform, Pulumi, Crossplane)
  • GitOps
  • Cloud Infrastructure (AWS EKS, GCP GKE, Azure AKS)
  • Networking
  • Storage
  • IAM
  • Systems Programming
  • Backend Programming
  • Observability
  • Monitoring
  • Alerting
  • Incident Response
  • Container Networking (CNI, Service Mesh, Network Policies)
  • Container Storage (CSI, Persistent Volumes, StatefulSets)
  • Container Security (RBAC, Pod Security Standards, Secrets Management)

Location

  • On-site in San Francisco

Work Type

  • On-site
  • Full-time

Experience Level

  • Principal
  • 10+ years of professional infrastructure or platform engineering experience

Salary/Compensations

  • $200,000 - $350,000

Benefits

  • Competitive salary and equity
  • Full healthcare coverage (100% premium for employee and dependents)
  • Support for growing families, including a yearly new parent stipend and fertility coverage
  • 401(k) company matching
  • $300 health and wellness benefit
  • Daily lunch provided
  • Dinner provided when working late
  • Regular team offsites and company events

About the Company

  • Edison Scientific builds and commercializes AI agents for science.
  • Our mission is to build an AI scientist to accelerate scientific discovery.
  • We are assembling a team of top researchers and engineers across AI and biology.
  • We operate in a fast-moving, mission-driven culture where smart people do their best work and enjoy it.

Equal Opportunity

  • Edison Scientific is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.