About the Role
aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale. We forward-deploy AI for enterprises, enabling organizations to rapidly implement production-ready AI solutions. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. We're building the next generation of enterprise AI infrastructure and are looking for exceptional people to help us scale.
Responsibilities
- Design and build aion's inference service platform the backbone for serving AI models at scale across diverse workloads
- Own and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and Autoscaler
- Design highly modular, scalable, and extensible low-level designs (LLDs) for inference infrastructure components
- Lead high-level design discussions, establish architectural patterns, and drive technical decision-making for the inference stack
- Understand and optimize the dynamics of model deployment, version upgrades, and rollback strategies
- Build robust deployment pipelines for seamless model updates with zero-downtime deployments
- Design intelligent routing systems for multi-model serving, A/B testing, and canary deployments
- Implement strategies for efficient GPU utilization and model cold-start optimization
- Implement highly performant and optimized software for low-latency, high-throughput inference serving
- Build and debug production-grade code in distributed systems handling real-time AI workloads
- Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization
- Design fault-tolerant systems with graceful degradation and automatic recovery mechanisms
- Build high-performance telemetry and observability stack for inference metrics, performance tracking, and debugging
- Implement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inference
- Conduct thorough code reviews to maintain code quality, performance standards, and architectural consistency
- Establish engineering best practices for testing, documentation, and production readiness.
Requirements
- Seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads
- Understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability
- Designed distributed systems that handle thousands of requests per second while maintaining sub-second response times and cost efficiency
- Experience with Golang is strongly preferred
- Exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus
- Take ownership of platform-level decisions
- Think strategically about performance vs. cost trade-offs
- Product-minded, understand how technical decisions impact developers using aion's platform and think about the end-to-end user experience
- Team player comfortable wearing multiple hats
- 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure
- Strong understanding of AI/ML inference systems and experience with inference engines (vLLM, TGI, TensorRT-LLM, or similar)
- Deep knowledge of distributed systems design, microservices architecture, and API gateway patterns
- Proficiency in Golang strongly preferred; Python, Rust, C++ for performance-critical components a plus
- Experience with container orchestration (Kubernetes, Docker) and infrastructure-as-code
- Solid understanding of autoscaling strategies, load balancing, and resource scheduling algorithms
- Experience building high-throughput, low-latency systems with sub-100ms response time requirements
- Familiarity with message queues (Kafka, RabbitMQ), databases (PostgreSQL, Redis), and event-driven architectures
- Knowledge of GPU computing, model serving optimizations (batching, quantization, multi-tenancy), and resource allocation
- Experience with observability tools (Prometheus, Grafana, OpenTelemetry) and distributed tracing
- Understanding of API design, rate limiting, authentication/authorization, and security best practices
- Exposure to AI model deployment workflows and model lifecycle management is highly desirable
- HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration
- Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks
- Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents)
- ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems
- Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installatiaons with emphasis on security, compliance, and operational simplicity
- High ownership, self driven and biased for action
- Strong strategic thinking and ability to connect technical decisions to business impact
- Excellent communication and mentoring skills
- Thrives in ambiguity, fast-paced environments, and early-stage startup culture.
Skills
- Golang
- Inference engines (vLLM, TGI, TensorRT)
- Containerization
- Distributed systems
- AI/ML inference systems
- Microservices architecture
- API gateway patterns
- Python
- Rust
- C++
- Kubernetes
- Docker
- Infrastructure-as-code
- Autoscaling strategies
- Load balancing
- Resource scheduling algorithms
- Message queues (Kafka, RabbitMQ)
- Databases (PostgreSQL, Redis)
- Event-driven architectures
- GPU computing
- Model serving optimizations (batching, quantization, multi-tenancy)
- Resource allocation
- Observability tools (Prometheus, Grafana, OpenTelemetry)
- Distributed tracing
- API design
- Rate limiting
- Authentication/authorization
- Security best practices
- AI model deployment workflows
- Model lifecycle management
- HPC & Cluster Management
- Kubernetes
- Slurm
- Job scheduling
- Resource allocation
- Workload orchestration
- Data Engineering
- Data pipelines
- ETL systems
- Large-scale data processing frameworks
- Systems-Level Programming
- Storage systems
- Kubernetes operators
- OS-level software development
- Daemon services (llm-d, system agents)
- ML Platform Engineering
- ML pipelines
- Batch job orchestration
- Model fine-tuning workflows
- Jupyter notebook orchestration systems
- Enterprise Deployment
- On-premises deployments
- Customer VPC installations
- Security
- Compliance
- Operational simplicity
Experience Level
- 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure
Benefits
- Competitive compensation
- Flexible work options
- Wellness benefits
About the Company
- aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale.
- By abstracting away infrastructure complexity, aion unifies compute orchestration, training workflows, data pipelines, model versioning, and deployment into a streamlined enterprise experience.
- We forward-deploy AI for enterprises, enabling organizations to rapidly implement production-ready AI solutions.
- We're a fast-growing, VC-backed startup led by founders with a track record of successful exits.
- With teams across the US, UK, and India, we're building the next generation of enterprise AI infrastructure and are looking for exceptional people to help us scale.
