About the Role
aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale. We are a fast-growing, VC-backed startup led by founders with a track record of successful exits, building the next generation of enterprise AI infrastructure.
Responsibilities
- Work directly at customer sites from factory floors to executive offices conducting discovery workshops and technical assessments to identify high-impact AI opportunities
- Design and architect end-to-end multimodal agent systems (voice + video + text) that leverage aion's distributed GPU infrastructure and managed services
- Build production-grade voice AI systems using STT, TTS APIs, and LLMs deployed on aion's platform
- Develop vision-enabled agents processing real-time video streams using computer vision pipelines on aion's infrastructure
- Implement sophisticated multi-agent orchestration with frameworks like LangChain or LlamaIndex
- Rapidly prototype POCs in 2-4 weeks, coding alongside client teams to validate concepts and iterate based on feedback
- Optimize for sub-500ms latency, natural conversation flow, turn detection, and interruption handling in real-time systems
- Integrate agents directly into customer codebases via REST/GraphQL/WebSocket APIs and custom SDKs (Python, TypeScript)
- Act as trusted technical advisor to customers, shaping AI strategy and guiding roadmap decisions from concept to production
- Design data architectures with efficient processing pipelines and ingestion workflows for training and inference on aion's platform
- Implement RAG systems with vector databases optimizing embedding strategies, chunk sizes, and retrieval methods
- Prepare and validate datasets for fine-tuning, evaluation, and synthetic data generation
- Work with other MLEs, MLOps, SREs to carry out model deployment and productionization
- Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure
- Instrument applications to trace LLM calls, retrieval operations, agent actions, and data flows
- Build evaluation frameworks with offline benchmarks (accuracy, relevance, safety metrics) and online monitoring (user feedback, drift detection)
Requirements
- 3-5+ years of hands-on experience building production AI/ML systems, with 1-2+ years deploying LLM applications to production
- Multimodal AI expertise practical experience building voice agents, vision systems, or conversational AI serving real users
- Strong LLM foundations hands-on with modern foundation models including fine-tuning, prompt engineering, and evaluation methodologies
- Agent framework proficiency production experience with LangChain, LlamaIndex, or similar orchestration frameworks
- Voice AI platform experience built real-time conversational systems with production STT/TTS integration
- Proficiency in Python (production-grade, async programming, type hints) and JavaScript/TypeScript (full-stack development)
- RAG implementation experience built retrieval-augmented generation systems with vector databases
- MLOps & deployment hands-on with Docker, Kubernetes, CI/CD pipelines, and infrastructure-as-code
- Cloud platforms experience with AWS, Azure, or GCP for ML workloads and infrastructure management
- Exceptional communication ability to explain complex AI concepts clearly to both technical and business stakeholders
- Customer-facing experience in Solutions Architecture, Technical Account Management, or Pre-Sales Engineering is highly desirable
- Computer vision experience working with video processing, object detection, or vision-language models is a plus
- Model fine-tuning practical experience with LoRA/QLoRA, supervised fine-tuning, or RLHF workflows is a plus
- Inference optimization experience with vLLM, TensorRT-LLM, Triton, or model quantization techniques is desirable
- Observability tooling practical experience with LLM monitoring, tracing, and evaluation frameworks is a strong plus
- Familiarity with WebRTC, real-time streaming protocols, and low-latency media processing
- Founder-level ownership and bias for action
- Strong strategic thinking and ability to connect technical decisions to business impact
- Excellent communication and mentoring skills
- Thrives in ambiguity, fast-paced environments, and early-stage startup culture
Skills
- AI Engineering
- Multimodal AI
- LLM Applications
- Production Code
- Technical Solutions Presentation
- AI System Debugging
- Voice Agents
- Video Processing Systems
- Conversational AI
- Business Requirements Translation
- Technical Solutions
- AI Deployment Lifecycle
- Use Case Discovery
- Solution Architecture
- Multimodal Agent Development
- MLOps Pipeline Implementation
- Production Optimization
- Agent Performance
- Observability
- Evaluation
- Voice AI Platforms
- RAG Systems
- LLM Orchestration Frameworks
- Communication Skills
- Customer Empathy
- Enterprise AI
- STT APIs
- TTS APIs
- LLMs
- Computer Vision Pipelines
- Multi-agent Orchestration
- LangChain
- LlamaIndex
- REST APIs
- GraphQL APIs
- WebSocket APIs
- Python SDKs
- TypeScript SDKs
- Data Architecture
- Data Processing Pipelines
- Data Ingestion Workflows
- Vector Databases
- Embedding Strategies
- Chunk Sizes
- Retrieval Methods
- Dataset Preparation
- Dataset Validation
- Fine-tuning
- Synthetic Data Generation
- Model Deployment
- Productionization
- LLM Observability
- Agent Monitoring
- Token Usage Tracking
- Latency Tracking
- Cost Tracking
- Quality Metrics
- Application Instrumentation
- LLM Call Tracing
- Retrieval Operation Tracing
- Agent Action Tracing
- Data Flow Tracing
- Offline Benchmarks
- Online Monitoring
- User Feedback Analysis
- Drift Detection
- Foundation Models
- Prompt Engineering
- Async Programming
- Type Hints
- Full-stack Development
- Docker
- Kubernetes
- CI/CD Pipelines
- Infrastructure-as-Code
- AWS
- Azure
- GCP
- ML Workloads
- Infrastructure Management
- Technical Stakeholder Communication
- Business Stakeholder Communication
- Solutions Architecture
- Technical Account Management
- Pre-Sales Engineering
- Video Processing
- Object Detection
- Vision-Language Models
- LoRA
- QLoRA
- Supervised Fine-tuning
- RLHF Workflows
- vLLM
- TensorRT-LLM
- Triton
- Model Quantization
- LLM Monitoring Tooling
- LLM Tracing Tooling
- LLM Evaluation Frameworks
- WebRTC
- Real-time Streaming Protocols
- Low-latency Media Processing
- Ownership
- Bias for Action
- Strategic Thinking
- Business Impact Analysis
- Mentoring
- Ambiguity Tolerance
- Fast-paced Environment Navigation
- Early-stage Startup Culture Adaptation
Experience Level
- 3-5+ years of experience
- 1-2+ years deploying LLM applications to production
Benefits
- Competitive compensation
- Flexible work options
- Wellness benefits
- Significant ownership and impact with equity reflective of your contributions
About the Company
- aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale.
- By abstracting away infrastructure complexity, aion unifies compute orchestration, training workflows, data pipelines, model versioning, and deployment into a streamlined enterprise experience.
- We forward-deploy AI for enterprises, enabling organizations to rapidly implement production-ready AI solutions.
- We're a fast-growing, VC-backed startup led by founders with a track record of successful exits.
- With teams across the US, UK, and India, we're building the next generation of enterprise AI infrastructure and are looking for exceptional people to help us scale.
