About the Role
We are building a Production Support function where the engineer owns the outcome. You'll be the primary owner of every product-reported ticket, investigating root causes across our API, SDKs, integrations, and vendor ecosystem, and resolving the majority of issues directly. This role is critical to FrankieOne's reliability and growth, requiring technical depth and operational discipline to diagnose distributed system issues, coordinate with Engineering, communicate clearly under pressure, and build the knowledge and automation that lets the team scale.
Responsibilities
- Own product-reported tickets from intake through closure, end-to-end.
- Triage tickets by severity and customer impact, setting clear expectations and timelines.
- Manage all customer communication for the life of the ticket, including status updates and technical explanations.
- Maintain a clean, well-triaged queue and report on volumes, trends, and recurring issues.
- Diagnose issues across the full stack (API, OneSDK, webhooks, Portal) using logs, traces, and metrics.
- Resolve the majority of tickets independently through configuration changes, data corrections, or known workarounds.
- Distinguish support-resolvable issues from genuine code defects and escalate bugs clearly to Product/Engineering.
- Coordinate Engineering fixes for genuine code defects, verifying delivery and closing the loop with customers.
- Act as incident commander, leading L2 incident response and owning resolution within SLA.
- Acknowledge on-call alerts with rapid analysis and next steps.
- Own status page updates, SLA monitoring, and post-incident reviews.
- Monitor vendor health and alerting, informing customers of vendor issues.
- Run and monitor batch processes and maintain operational health dashboards.
- Build runbooks, playbooks, and monitoring configurations to enable independent resolution.
- Own deep troubleshooting of hotspot surfaces.
- Maintain the technical Knowledge Base to ensure recurring issues are resolved efficiently.
- Identify opportunities to automate toil through Jira automation, scripting, or tooling improvements.
- Act as the technical liaison between customers, support teams, and Engineering.
- Respond to technical inquiries via Slack, Jira, and email.
- Reproduce reported issues against sandbox and production environments to validate fixes.
- Feed operational insights into Product and Engineering for roadmap prioritization.
Requirements
- 7+ years in production support, SRE, TechOps, DevOps, or similar operational engineering roles (ideally SaaS, fintech, or regulated industries).
- Proven ability to own tickets end-to-end and resolve the majority independently.
- Leverage AI tools to improve processes and ways of working.
- Strong troubleshooting across distributed systems, REST APIs, and webhook/event-driven integrations.
- Strong SQL and data investigation skills.
- Experience diagnosing third-party vendor integration issues and isolating fault.
- Scripting/automation skills (Python, Bash) for building tooling and automating runbooks.
- Familiarity with cloud infrastructure (AWS/GCP), message queues and async/event-driven systems (Kafka, SQS).
- Experience on an on-call rotation and leading incident response.
- Hands-on experience reading and interpreting logs, traces, and metrics from observability tools (Datadog, Grafana, ELK, etc.).
- Proven ability to review, triage and respond to production alerts.
- Sharp triage skills - quickly assessing severity and impact, classifying issues accurately, and routing to the right owner first time.
- Feature request triage - classifying requests, deduplicating demand, and negotiating workarounds.
- Excellent customer-facing written communication skills.
- Rigorous root cause analysis - distinguishing symptom fixes from actual fixes.
- Track record of reducing recurring ticket volume through automation, documentation, or process improvement.
Skills
- Debugging at depth
- Operational judgement
- Scripting for Operations (Python, Bash, JavaScript/TypeScript)
- SQL for investigation
- Clear communication under pressure
- Strong ownership mentality
- Go (Golang) code navigation (nice to have)
- Containers and CI/CD pipelines (nice to have)
- Supporting SDK-based or API-first products (nice to have)
- Handling sensitive PII in a regulated environment (nice to have)
- Working across time zones and managing async communication (nice to have)
Location
- Hybrid
Work Type
- Hybrid
Experience Level
- 7+ years
Benefits
- Competitive salary and benefits package aligned with market
- Growth pathway: Opportunity to mentor junior engineers and build Production Support capabilities as the team scales
About the Company
- At FrankieOne, our mission is to help fintechs and financial institutions scale by providing seamless access to the global ecosystem of identity and fraud solutions.
- Our customisable orchestration platform connects customers to all major identity verification, KYC/KYB, and fraud prevention vendors in one place- delivering unparalleled compliance and user experience.
- We work with a rapidly expanding global customer base, including some of the most recognisable fintech and financial brands.
- We've built a culture rooted in high performance, accountability, and being frank - a place where technical excellence and customer obsession go hand in hand.
