About the Role
Join a focused pod building two tightly-coupled capabilities for Humana's internal users: an Internal Agentic Search Engine and a Perplexity-Style Front End. You will build agentic pipelines, design retrieval and orchestration layers, and implement tool-use for accurate, cited, and trustworthy answers. You will also deliver a fast, intuitive, conversational search experience, partner across the stack to connect the backend to a responsive UI, and instrument usage and feedback signals.
Responsibilities
- Build hands-on, end to end, shipping agentic AI features into production.
- Engineer agentic systems, developing planning, retrieval, tool-use, and orchestration components.
- Integrate agents reliably and securely with numerous internal and enterprise data sources.
- Build, deploy, and operate solutions natively on Google Cloud (Vertex AI, GKE, BigQuery, Cloud Run, and related services).
- Contribute to the Perplexity-style front end for fast, grounded, well-cited answers.
- Own quality and evaluation by establishing evals, guardrails, observability, and feedback loops.
- Collaborate on-site in the NYC Mid-Town office Tuesday through Thursday, pairing closely with the pod and Humana stakeholders.
- Operate with urgency in a fast-moving engagement, iterating quickly and driving outcomes.
Requirements
- Strong, current software engineering fundamentals with clean, tested, production-quality code (Python strongly preferred).
- Demonstrated, hands-on experience building agentic AI systems (agents, tool-use, planning, multi-step reasoning, RAG/retrieval).
- Deep expertise with the Google Cloud technology stack (e.g., Vertex AI, GKE, BigQuery, Cloud Run, Cloud Storage, IAM).
- Experience integrating LLM applications with numerous, heterogeneous data sources (APIs, databases, document stores, search).
- Experience taking ML/GenAI systems to production, including deployment, scaling, monitoring, and reliability.
- Ability to work on-site in NYC Mid-Town 3 days per week (Tuesday–Thursday).
- Breadth of knowledge across frontier models (e.g., Gemini, and other leading LLM families) and when to use which.
- Hands-on experience with open-source agentic and LLM frameworks (e.g., LangChain, LangGraph, LlamaIndex, or similar).
- Front-end / full-stack exposure to help deliver a Perplexity-style user experience.
- Experience with evaluation frameworks, guardrails, and responsible-AI practices.
- Prior experience in regulated or enterprise environments (healthcare a plus).
Skills
- Python
- Agentic AI systems
- Tool-use
- Planning
- Multi-step reasoning
- RAG/retrieval
- Google Cloud
- Vertex AI
- GKE
- BigQuery
- Cloud Run
- Cloud Storage
- IAM
- LLM applications
- Data source integration
- ML/GenAI production deployment
- Scaling
- Monitoring
- Reliability
- Frontier models
- Gemini
- Open-source agentic frameworks
- LLM frameworks
- LangChain
- LangGraph
- LlamaIndex
- Front-end development
- Full-stack development
- Evaluation frameworks
- Guardrails
- Responsible-AI practices
Location
- Mid-Town NYC, USA
Work Type
- In-person 3 days/week
Experience Level
- 5+ Years
About the Company
- Quantiphi is an award-winning, AI-First global digital engineering company that helps Fortune 1000 organizations transform bold ideas into measurable business impact.
- We go beyond building innovative AI technologies—we solve the problems that matter most to our clients.
- Headquartered in Boston, with more than 4,000 professionals worldwide, we partner with global enterprises to deliver large-scale digital, cloud, and AI-driven transformation.
- We are an Elite and Premier partner to Google Cloud, AWS, NVIDIA, Snowflake, and other leading technology platforms.
- Quantiphi delivers First-in-class AI solutions across Life Sciences, Healthcare, Banking, Financial Services, CPG, Manufacturing, Energy, High-Tech, Telecommunications, etc., powered by cutting-edge Generative AI and Agentic AI accelerators.
- We are also proud to be certified as a Great Place to Work.
