Data Center Operations Lead - Partner Site Operations at Anthropic | Austin, TX, US | Rezi

Data Center Operations Lead - Partner Site Operations at Anthropic

Data Center Operations Lead - Partner Site Operations

Anthropic · Austin, TX, US

1 months ago

Data Center Operations Lead - Partner Site Operations

Anthropic · Austin, TX, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Data Center Operations Lead - Partner Site Operations role.

Rezi rewrites your resume against Anthropic's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Data Center Operations Lead - Partner Site Operations posting at Anthropic — free, in seconds.

About the Role

Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. At our partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day-to-day data hall work. As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response. Rather than managing operations staff directly, you provide tactical direction, set priorities, and define the standards for the vendor's on-site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met. You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. Expect to build the playbook as much as you run it, not just at a site level, but defining and developing program improvements fleet-wide.

Responsibilities

  • Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self-reporting.
  • Set daily and weekly priorities and lead the operating cadence, including standups and business reviews.
  • Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance.
  • Analyze operational trends and standardize lessons across the program.
  • Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary.
  • Participate in the incident escalation on-call rotation.
  • Direct vendor response, own communications, and close out post-incident actions when designated Anthropic Incident Commander for a site-specific incident.
  • Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.

Requirements

  • Have 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability.
  • Managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action.
  • Carry hands-on technical depth in server, network, and rack-level infrastructure, enough to independently verify vendor claims and audit quality.
  • Built or substantially improved operational processes, not just run them.
  • Served in an incident command or lead-responder role and communicate clearly under ambiguity.
  • Can support non-standard hours, including an on-call rotation and availability during deployment surges and maintenance windows.
  • Experience with third-party colocation providers or partner-operated sites, delivering IT operations outcomes inside a facility someone else runs.
  • Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment.
  • Experience with GPU/accelerator or high-density liquid-cooled infrastructure.
  • Familiarity with multi-vendor sites where facilities and IT operations are performed by different partners.
  • Experience leading projects from initiation to completion across teams you didn't own.
  • Background in incident management frameworks, contract/SLA design, or EHS programs.

Skills

  • Server infrastructure
  • Network infrastructure
  • Rack-level infrastructure
  • Incident management frameworks
  • Contract/SLA design
  • EHS programs

Location

  • San Francisco

Work Type

  • Hybrid

Experience Level

  • 8+ years

Education Level

  • Bachelor's degree or an equivalent combination of education, training, and/or experience

Salary/Compensations

  • $320,000—$405,000 USD

Benefits

  • Competitive compensation
  • Benefits
  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
  • Lovely office space

About the Company

  • Anthropic’s mission is to create reliable, interpretable, and steerable AI systems.
  • We want AI to be safe and beneficial for our users and for society as a whole.
  • Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
  • We believe that the highest-impact AI research will be big science.
  • At Anthropic we work as a single cohesive team on just a few large-scale research efforts.
  • We value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles.
  • We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science.
  • We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time.
  • We greatly value communication skills.
  • Anthropic is a public benefit corporation headquartered in San Francisco.

Equal Opportunity

  • We encourage you to apply even if you do not believe you meet every single qualification.
  • Not all strong candidates will meet every single qualification as listed.
  • We think AI systems like the ones we're building have enormous social and ethical implications.
  • We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
  • We strive to include a range of diverse perspectives on our team.