About the Role
We are seeking a Founding Infrastructure & Reliability Engineer to join the earliest technical team at a fast-moving AI video infrastructure startup. This role is focused on making the core product reliable and scalable as usage grows, and involves owning production systems, debugging failures, and making complex workflows predictable. The work spans job execution, queueing, failure recovery, observability, cost efficiency, and deployment.
Responsibilities
- Improve job execution reliability, queueing behavior, failure recovery, and production observability
- Design systems that scale expensive workflows without making the product feel slow or unpredictable to users
- Build deployment, monitoring, and incident-response workflows appropriate for an early but serious production system
- Improve cost efficiency across compute, storage, networking, and third-party services
- Create internal tools that make debugging and operating the system easier for the whole team
- Partner across the team to connect product requirements with robust technical systems
Requirements
- Experience operating production infrastructure or high-throughput backend systems
- Strong debugging ability across services, networking, storage, and cloud environments
- Familiarity with queues, object storage, databases, observability tooling, deployment automation, and cost/performance tradeoffs
- Comfort taking ownership of ambiguous reliability problems, even without a clear root cause
- Practical judgment on when to automate, simplify, or build the smallest robust solution
- Comfortable using AI tools to move faster, while using personal judgment to define the problem and drive the solution
- Must be willing to work on-site in San Francisco, CA
- Experience communicating infrastructure tradeoffs to non-infrastructure engineers
- History of taking ownership of outcomes, not just assigned tickets
- Comfort making infrastructure decisions independently, without a fully-defined spec
Skills
- Queueing systems
- Object storage
- Databases
- Observability tooling
- Deployment automation
- Cloud infrastructure
- AI-assisted development tools
Location
- San Francisco, CA
Work Type
- On-Site
- Remote
About the Company
- Hire Hangar works with fast-growing global companies.
- The company is an AI video infrastructure startup.
- The company's first product turns prompts and documents into whiteboard-style explainer videos in seconds.
- The company has a broader vision to make video generation fast, cheap, and programmable enough for AI to use video anywhere it uses text today.
