About the Role
Gimlet Labs is seeking a Data Center Facilities Operations Lead to manage the critical facilities operating model for Gimlet data centers and high-density AI infrastructure deployments. This role ensures facility-side systems supporting Gimlet's compute capacity are ready, monitored, maintained, and operating within the required parameters, with a focus on infrastructure supporting liquid-cooled AI systems.
Responsibilities
- Build the facilities operations model for current and future Gimlet sites, including operating standards, escalation paths, maintenance routines, acceptance criteria, and facility readiness gates.
- Translate OEM and engineering requirements for liquid-cooled platforms into practical site operating envelopes for temperature, flow, pressure, water quality, alarms, and heat rejection.
- Own monitoring and response for facility-side telemetry, including supply and return water temperatures, delta-T, flow, pressure, leak detection, CDU status, cooling capacity margins, and BMS/DCIM alarms.
- Partner with colocation providers, facility vendors, OEMs, Site Managers, Data Center Technicians, Deployment Leads, and TPMs to ensure facilities are ready before new compute capacity is deployed.
- Create and maintain MOPs, SOPs, EOPs, maintenance windows, runbooks, inspection routines, and incident response procedures for critical facilities and liquid cooling operations.
- Coordinate preventive maintenance, repairs, and vendor response for CDUs, facility water loops, filters, valves, pumps, sensors, leak detection systems, chillers, dry coolers, CRAHs, and related infrastructure.
- Lead facility-side root cause analysis for thermal, leak, power, cooling, monitoring, and environmental events, then drive durable corrective actions.
- Build reporting that shows facility health, risk, readiness, capacity margin, recurring issues, open repairs, and operational trends across Gimlet sites.
Requirements
- Experience in data center facilities operations, critical facilities engineering, MEP operations, commissioning, facilities maintenance, or high-density infrastructure operations.
- Understanding of liquid cooling operations, facility water systems, CDUs, heat rejection, supply and return temperature management, flow, pressure, filtration, leak detection, and water quality controls.
- Experience operating or supporting BMS, DCIM, EPMS, CDU monitoring, alarm response, trend analysis, and facilities telemetry in production environments.
- Ability to write and run MOPs, SOPs, EOPs, maintenance plans, incident procedures, and vendor repair workflows with strong operational discipline.
- Clear communication with site teams, network teams, deployment TPMs, engineering, colocation providers, OEMs, and facilities vendors.
- Comfort working in active data center environments and supporting urgent facilities escalations.
- Experience supporting GPU clusters, GB200/GB300-class platforms, NVL rack-scale systems, HPC environments, AI infrastructure, or other high-density liquid-cooled compute deployments.
- Experience with data center commissioning, integrated systems testing, site acceptance testing, facility turnover, or deployment readiness reviews.
- Familiarity with power distribution, UPS/generator coordination, chilled water systems, dry coolers, CRAH/CRAC systems, CDUs, rear-door heat exchangers, and liquid cooling safety practices.
- Experience managing colocation provider obligations, service levels, maintenance windows, vendor escalations, and facilities contract deliverables.
- A track record of improving facility reliability through monitoring, preventive maintenance, incident analysis, documentation, training, and operational controls.
Skills
- Data center facilities operations
- Critical facilities engineering
- MEP operations
- Commissioning
- Facilities maintenance
- High-density infrastructure operations
- Liquid cooling operations
- Facility water systems
- CDUs
- Heat rejection
- Temperature management
- Flow management
- Pressure management
- Filtration
- Leak detection
- Water quality controls
- BMS
- DCIM
- EPMS
- CDU monitoring
- Alarm response
- Trend analysis
- Facilities telemetry
- MOPs
- SOPs
- EOPs
- Maintenance plans
- Incident procedures
- Vendor repair workflows
- Root cause analysis
- Reporting
- Communication
- Collaboration
Location
- Data centers
Work Type
- Full-time
Experience Level
- Lead
About the Company
- Gimlet is building the first multi-silicon neocloud designed for fast, efficient inference.
- The company's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling significant improvements in performance and efficiency.
- Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
- Gimlet works with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
- This provides the team access to systems research problems grounded in frontier models, cutting-edge production workloads, and emerging hardware architectures.
Equal Opportunity
- Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.
