About the Role
As the Senior Manager of Technical Support Engineering - Infrastructure, you will own, scale, and continuously improve CoreWeave's infrastructure support function. This is a globally distributed 24/7/365 organization of skilled engineers who resolve complex technical challenges with deep expertise, efficiency, and empathy. High-touch, expert-led support is a key differentiator for CoreWeave, giving this role significant visibility and influence across the company. You will operate at the department level, setting the technical and operational bar for the entire function and designing operating systems for scalability. You will lead with empathy, invest in individual growth, and build structures and processes to support hyper-growth without compromising culture.
Responsibilities
- Own the strategy, health, and performance of the entire 24/7/365 infrastructure support function, scaling coverage and capability across regions and domains.
- Own talent acquisition and retention by hiring, onboarding, and developing engineers through diligent performance management and coaching.
- Stay hands-on by digging into complex, customer-impacting issues alongside the team and serving as a senior technical escalation point.
- Build and facilitate enablement and career-development frameworks for onboarding, technical training, and progression paths.
- Implement quality assurance measures, including ticket reviews and best-practice playbooks, to improve resolution speed, accuracy, and consistency.
- Own the support operating model, including roles, responsibilities, escalation boundaries, coverage/staffing, and SLOs.
- Lead customer communication during critical incidents and resolve conflicts with clarity, composure, and empathy.
- Track and report on KPIs focused on team performance and customer satisfaction, and own strategic planning for team growth and scalability.
- Own the cross-functional interface between support and Product Engineering, Specialist Field Engineers, and domain teams.
- Champion the voice of the customer, turning recurring support patterns into product, tooling, and process improvements.
- Set the multi-quarter vision and operating plan for infrastructure support, and represent the function in company-level planning discussions.
Requirements
- 5+ years of people-leadership experience.
- 8+ years total in technical support/operations, including running a 24/7 support function at scale in a cloud operations environment.
- Strong background in Linux, containerization technologies, and Kubernetes.
- Understanding of virtualization and cloud computing concepts.
- Experience at a hyperscaler or cloud infrastructure provider is a strong plus.
- Ability to lead with empathy and get hands-on with the work.
- Experience building enablement and quality programs that scaled across a function or multiple teams.
- Experience designing the operating model for a support organization, including coverage/staffing, escalation boundaries, and SLOs.
- Calm, clear communication skills during critical incidents and ability to resolve conflicts effectively.
- Systems thinking and multi-quarter planning ability.
- Experience defining KPI/SLO frameworks, reporting to senior leadership, and owning capacity and growth planning.
- Experience leading a globally distributed team across time zones.
- Proven ability to lead through senior talent and set direction for a function.
- Executive-level communication skills.
- Ability to build operating systems, plans, and metrics that scale a function.
- Robust problem-solving skills and adaptability in a fast-paced, hyper-growth environment.
- Experience with program-management tools and methodologies.
- Experience supporting AI/ML, HPC, or GPU-accelerated workloads at scale is a bonus.
- Hands-on Kubernetes operations experience (CKA certification a plus) is a bonus.
- Familiarity with Slurm/SUNK, RDMA networking, distributed storage, and observability tooling such as Grafana is a bonus.
- Experience with infrastructure as it relates to Data Center Operations is a bonus.
- Must be a U.S. person (U.S. citizen or national, U.S. lawful permanent resident, refugee, or asylee) or eligible to access export controlled information without authorization, or eligible to obtain required export authorization.
Skills
- Linux
- Containerization technologies
- Kubernetes
- Virtualization
- Cloud computing
- People leadership
- Talent acquisition
- Talent retention
- Performance management
- Coaching
- Technical escalation
- Enablement programs
- Quality assurance
- Operating model design
- Customer communication
- Conflict resolution
- KPI tracking
- SLO definition
- Strategic planning
- Cross-functional collaboration
- Program management
- AI/ML workload support
- HPC workload support
- GPU-accelerated workload support
- Slurm/SUNK
- RDMA networking
- Distributed storage
- Observability tooling (e.g., Grafana)
- Data Center Operations
Location
- US-based
Work Type
- 24/7/365
- Full-time
Experience Level
- Senior
- 5+ years people leadership
- 8+ years technical support/operations
Salary/Compensations
- $198,000 to $264,000
Benefits
- Discretionary bonus
- Equity awards
- Medical insurance (100% paid)
- Dental insurance (100% paid)
- Vision insurance (100% paid)
- Company-paid Life Insurance
- Voluntary supplemental life insurance
- Short-term disability insurance
- Long-term disability insurance
- Flexible Spending Account
- Health Savings Account
- Tuition Reimbursement
- Employee Stock Purchase Program (ESPP)
- Mental Wellness Benefits through Spring Health
- Family-Forming support provided by Carrot
- Paid Parental Leave
- Flexible, full-service childcare support with Kinside
- 401(k) with a generous employer match
- Flexible PTO
- Catered lunch
- Casual work environment
About the Company
- CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence.
- Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability.
- Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.
- CoreWeave operates on a Direct-to-Expert model for technical support, routing customers to the right expert quickly with a single owner accountable for the issue end-to-end.
- The Technical Support Engineering - Infrastructure team is the front line for customers running AI and HPC workloads on the CoreWeave platform, supporting Kubernetes-powered infrastructure including GPU compute, high-performance networking and storage, Slurm/HPC clusters, and large-scale training workloads.
- Core values include: Be Curious at Your Core, Act Like an Owner, Empower Employees, Deliver Best-in-Class Client Experiences, Achieve More Together.
- The company supports an entrepreneurial outlook, independent thinking, and collaboration, fostering an environment for developing innovative solutions.
Equal Opportunity
- CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace.
- All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.
- CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship.
