Principal Member of Technical Staff, Core Infrastructure at Edison Scientific | CA, US | Rezi

Principal Member of Technical Staff, Core Infrastructure at Edison Scientific

Principal Member of Technical Staff, Core Infrastructure

Edison Scientific · CA, US

2 weeks ago

Principal Member of Technical Staff, Core Infrastructure

Edison Scientific · CA, US

19 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Principal Member of Technical Staff, Core Infrastructure role.

Rezi rewrites your resume against Edison Scientific's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Principal Member of Technical Staff, Core Infrastructure posting at Edison Scientific — free, in seconds.

About the Role

As a Principal Member of Technical Staff, you will design, scale, and operate the core platform infrastructure powering autonomous scientific discovery. Your focus will be on agent orchestration at scale, building and managing clusters for thousands of persistent, stateful workloads, developing custom resource definitions (CRDs) and operators, and ensuring the reliability and efficiency of the compute layer. This role influences platform architecture, establishes infrastructure best practices, and partners with engineering and research teams to deliver a production-grade environment.

Responsibilities

  • Architect, implement, and operate Kubernetes clusters supporting thousands of concurrent, persistent resources with high availability and efficient resource utilization.
  • Design and develop custom resource definitions (CRDs) and Kubernetes operators to manage domain-specific workloads like AI agent lifecycles and research pipelines.
  • Drive strategy for cluster scaling, node pool management, autoscaling policies, and resource quotas to handle workload growth.
  • Build and maintain infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible, version-controlled environment management.
  • Design and implement robust scheduling, placement, and affinity strategies for heterogeneous workloads.
  • Establish and uphold best practices for observability, monitoring, alerting, and incident response.
  • Own storage and networking strategy within Kubernetes, including persistent volume management, CSI drivers, service mesh, network policies, and ingress architecture.
  • Troubleshoot complex, cross-system infrastructure issues and guide others through debugging and remediation.
  • Collaborate with backend, ML, and research teams to understand workload requirements and translate them into infrastructure patterns.

Requirements

  • Typically, 10+ years of professional infrastructure or platform engineering experience with deep hands-on Kubernetes expertise in production environments.
  • Experience designing and implementing custom resource definitions (CRDs) and Kubernetes operators.
  • Track record of operating and scaling Kubernetes clusters supporting thousands of persistent or long-lived resources.
  • Deep understanding of Kubernetes internals and their behavior at scale.
  • Expertise with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated primitives.
  • Proficiency in at least one systems or backend language for operator development and infrastructure tooling.
  • Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or Crossplane) and GitOps workflows.
  • Strong working knowledge of container networking, storage, and security.
  • Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production.
  • Experience with data-intensive platforms, scientific computing, or ML/AI infrastructure (bonus).
  • Prior experience in startups or small teams with significant architectural ownership and ambiguity (bonus).
  • Experience scaling systems, teams, or platforms through periods of rapid growth (bonus).

Skills

  • Kubernetes
  • Custom Resource Definitions (CRDs)
  • Kubernetes Operators
  • Infrastructure-as-Code
  • Terraform
  • Pulumi
  • Kubebuilder
  • Operator SDK
  • controller-runtime
  • Cloud Infrastructure (AWS EKS, GCP GKE, Azure AKS)
  • Networking
  • Storage
  • IAM
  • Systems Programming
  • Backend Programming
  • GitOps
  • Container Networking (CNI, Service Mesh, Network Policies)
  • Container Storage (CSI, Persistent Volumes, StatefulSets)
  • Container Security (RBAC, Pod Security Standards, Secrets Management)
  • Observability
  • Monitoring
  • Alerting
  • Incident Response
  • Prometheus
  • Grafana
  • Datadog

Location

  • San Francisco

Work Type

  • On-site

Experience Level

  • 10+ years of professional infrastructure or platform engineering experience
  • Principal Member of Technical Staff

Salary/Compensations

  • $200,000 - $350,000

Benefits

  • Competitive salary and equity
  • Full healthcare coverage; 100% of premiums paid for employees and dependents
  • Support for growing families, including a yearly new parent stipend and fertility coverage
  • Mental health support through Rula
  • 12 weeks of paid parental leave
  • Pet care support with a yearly employer-funded stipend
  • Commuter benefits
  • 401(k) company matching
  • $300 health and wellness benefit quarterly
  • Daily lunch provided in the office
  • Dinner provided when working late
  • Regular team off-sites and company events

About the Company

  • Edison Scientific builds and deploys AI scientist agents to accelerate science and the development of new medicines.
  • We are an ambitious team run by scientists and engineers from leading institutions across biology, physics, chemistry, and AI.
  • We're a fast-moving, mission-driven culture where smart people do their best work and actually enjoy doing it.

Equal Opportunity

  • Edison Scientific is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.