About the Role
As one of the first engineering hires at an early-stage AI video infrastructure startup, you will ensure the reliability and scalability of core systems as usage grows. This involves debugging production issues, enhancing observability, and optimizing system performance and cost-effectiveness.
Responsibilities
- Improve reliability of job execution, queueing behavior, and failure recovery.
- Build monitoring, alerting, and incident-response workflows for a production system.
- Design systems that scale expensive workflows without compromising speed or predictability.
- Improve cost efficiency across compute, storage, and third-party services.
- Build internal tools to simplify debugging and system operations.
- Partner with the broader team to align product needs with robust technical solutions.
Requirements
- Experience operating production infrastructure or high-throughput backend systems.
- Strong debugging skills across services, networking, storage, and cloud environments.
- Familiarity with queues, object storage, databases, and observability tooling.
- Comfortable owning ambiguous reliability problems without a predefined root cause.
- Practical judgment about when to automate, simplify, or build minimal robust solutions.
- Comfortable using AI tools to increase speed, while relying on personal judgment to solve problems.
- Must be willing to work on-site in San Francisco, CA.
Skills
- Queueing systems
- Object storage
- Databases
- Observability tooling
- Cloud infrastructure
- Deployment automation
- AI-assisted development tools
Location
- San Francisco, CA
Work Type
- On-Site
- Remote
About the Company
- Hire Hangar connects top talent with vetted employers, competitive pay, and real growth opportunities.
- This is an early-stage AI video infrastructure startup. The company's first product turns prompts and documents into whiteboard-style explainer videos in seconds, working toward a bigger vision of making video generation as fast, cheap, and programmable as text is today.
