About the Role
As a Member of Technical Staff, Infrastructure Engineer, you'll play a key role in designing, scaling, and operating the core platform infrastructure that powers autonomous scientific discovery. Our mission is to build an AI scientist, and you'll own the infrastructure foundation it runs on. AI agents performing long-running scientific research demand resilient scheduling, lifecycle management, and resource orchestration far beyond typical cloud-native workloads. This role will establish infrastructure best practices, and partner closely with backend engineers, ML engineers, and researchers to deliver a production-grade environment that lets science move faster. Engineering at the senior level is about technical ownership and leverage- understanding how complex systems interact, making sound architectural tradeoffs, and building foundations that allow teams and science to move faster.
Responsibilities
- Architect, implement, and operate Kubernetes clusters that support thousands of concurrent, persistent resources (agents, jobs, services) with high availability and efficient resource utilization.
- Drive the strategy for cluster scaling, node pool management, autoscaling policies, and resource quota frameworks to handle rapid workload growth.
- Design and implement robust scheduling, placement, and affinity strategies to optimize cost, performance, and fault tolerance for heterogeneous workloads (CPU, GPU, memory-intensive).
- Establish and uphold best practices around observability, monitoring, alerting, and incident response for infrastructure systems (Prometheus, Grafana, Datadog, or similar).
- Own storage and networking strategy within Kubernetes — including persistent volume management, CSI drivers, service mesh, network policies, and ingress architecture.
- Troubleshoot complex, cross-system infrastructure issues and guide others through effective debugging and remediation in distributed environments.
- Collaborate closely with backend, ML, and research teams to understand workload requirements and translate them into reliable infrastructure patterns.
Requirements
- Typically, 5+ years of professional infrastructure or platform engineering experience, with hands-on Kubernetes experience in production environments.
- Experience with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated networking, storage, and IAM primitives.
- Proficiency in at least one systems or backend language for operator development and infrastructure tooling.
- Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or Crossplane) and GitOps workflows.
- Strong working knowledge of container networking (CNI plugins, service mesh, network policies), storage (CSI, persistent volumes, StatefulSets), and security (RBAC, Pod Security Standards, secrets management).
- Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production.
Skills
- Kubernetes
- Cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS)
- Networking
- Storage
- IAM
- Systems programming
- Backend programming
- Infrastructure-as-code (Terraform, Pulumi, or Crossplane)
- GitOps
- Container networking (CNI plugins, service mesh, network policies)
- Storage (CSI, persistent volumes, StatefulSets)
- Security (RBAC, Pod Security Standards, secrets management)
- Observability
- Monitoring
- Alerting
- Incident response
- Prometheus
- Grafana
- Datadog
Location
- San Francisco
Work Type
- On-site
Experience Level
- 5+ years of professional infrastructure or platform engineering experience
- Senior level
Salary/Compensations
- $175,000 - $240,000
Benefits
- Competitive salary and equity
- Full healthcare coverage; we pay 100% of premiums for you and your dependents
- Support for growing families, including a yearly new parent stipend and fertility coverage through Carrot
- Mental health support through Rula, our in-network therapist and psychiatrist network with fast availability
- 12 weeks of paid parental leave for maternity, paternity, and adoption
- Pet care support with a yearly employer-funded stipend for your animal companions
- Commuter benefits so you can pay for transit and parking with pre-tax dollars
- 401(k) company matching
- $300 health and wellness benefit quarterly
- Lunch is on us every day you're in the office, and dinner is on us when you're working late
- Regular team off-sites and company events
About the Company
- Edison Scientific builds and deploys AI scientist agents to accelerate science and the development of new medicines.
- We are an ambitious team run by scientists and engineers from leading institutions across biology, physics, chemistry, and AI.
Equal Opportunity
- Edison Scientific is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.
