Lead Engineer, Issue Management & Triage at Diligent Robotics | Austin, TX, US | Rezi

Lead Engineer, Issue Management & Triage at Diligent Robotics

Lead Engineer, Issue Management & Triage

Diligent Robotics · Austin, TX, US

1 weeks ago

Lead Engineer, Issue Management & Triage

Diligent Robotics · Austin, TX, US

12 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Lead Engineer, Issue Management & Triage role.

Rezi rewrites your resume against Diligent Robotics's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Lead Engineer, Issue Management & Triage posting at Diligent Robotics — free, in seconds.

About the Role

Diligent builds helpful robots that work safely and autonomously in real-world environments. As a Lead Engineer, Issue Management & Triage, you will lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering. This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet.

Responsibilities

  • Own Issue Management & Triage Systems
  • Design and own end-to-end systems for issue intake, triage, and escalation
  • Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization
  • Build Tools & Automation (Hands-On)
  • Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort
  • Contribute directly to codebases (Python, backend services) and partner with Engineering on system integrations (logs, telemetry, alerts)
  • Bridge Operations & Engineering
  • Act as the primary technical interface between the Remote Operations Center (ROC) and Engineering
  • Translate real-world issues into prioritized, categorized technical problems for resolution alignment
  • Performance Measurement & Classification Frameworks
  • Develop systems and taxonomies to systematically measure and classify robot performance, failure modes, and degradation across the fleet
  • Build dashboards and reporting systems to track trends, severity, and impact
  • Root Cause Analysis & Continuous Improvement
  • Establish best practices for Root Cause Analysis (RCA) and identify systemic issues
  • Drive long-term fixes and create feedback loops to influence improvements in hardware, software, and autonomy

Requirements

  • 7+ years in relevant technical or program management roles (e.g., engineering, incident management)
  • 3+ years of people management
  • Experience with complex, real-world systems (robotics, autonomous/distributed systems, or hardware-software products)
  • Proven track record building operational tools, systems, or infrastructure for workflows
  • Strong systems thinker, translating ambiguous operational problems into structured technical solutions
  • Experience defining metrics, taxonomies, and performance frameworks
  • Data-driven approach to prioritization and decision-making
  • Hands-on and willing to dive into technical problems when needed
  • Strong ownership and bias toward action
  • Comfortable operating in a fast-paced, scaling environment
  • Passion for improving real-world system performance and reliability

Skills

  • Python programming
  • Backend services development
  • Data pipelines
  • Telemetry systems
  • Monitoring infrastructure
  • Alerting infrastructure
  • Logging infrastructure
  • Internal tools development
  • Automation systems development
  • Scalable systems design
  • Classification systems
  • Prioritization systems
  • Workflow automation
  • Jira
  • Zendesk
  • SQL
  • Looker
  • Foxglove

Location

  • Austin, TX
  • Remote (U.S.)

Work Type

  • Remote
  • Onsite

Experience Level

  • Lead

About the Company

  • Diligent builds helpful robots that work safely and autonomously in real world environments.
  • We move quickly, solve messy problems, and care deeply about reliability at scale.