About the Role
aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale. We forward-deploy AI for enterprises, enabling organizations to rapidly implement production-ready AI solutions. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits.
Responsibilities
- Build and maintain platform services across aion's Compute and Inference platforms.
- Implement features for multi-cloud orchestration, resource scheduling, model deployment pipelines, and autoscaling systems.
- Write well-maintained, production-grade code with proper abstractions, design patterns, and comprehensive test coverage.
- Contribute to low-level design (LLD) including service APIs, database schema design, data models, and component interactions.
- Collaborate with senior engineers on high-level design discussions, providing implementation perspectives and feasibility inputs.
- Develop RESTful APIs and gRPC services for platform control planes, resource management, and inference serving.
- Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics.
- Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components.
- Build event-driven architectures for asynchronous processing, job scheduling, and platform automation.
- Implement monitoring, logging, and alerting for platform services to ensure production reliability.
- Write comprehensive unit tests, integration tests, and end-to-end tests to ensure code reliability.
- Participate in code reviews, providing constructive feedback and learning from senior engineers' perspectives.
- Refactor existing code to improve maintainability, performance, and scalability.
- Document design decisions, API specifications, and operational runbooks for platform services.
- Debug production issues and contribute to incident response and post-mortems.
Requirements
- 2-4 years of experience in backend engineering, platform development, or distributed systems.
- Strong proficiency in Golang you write idiomatic Go code with proper error handling, concurrency patterns, and testing.
- Solid understanding of backend systems fundamentals: RESTful APIs, microservices architecture, and API design principles.
- Hands-on experience with databases (PostgreSQL, MySQL) including schema design, query optimization, and transactions.
- Familiarity with storage systems (object storage like S3, block storage, distributed file systems) and their use cases.
- Experience working with message queues (Kafka, RabbitMQ, NATS) and event-driven architectures.
- Understanding of distributed systems concepts: consensus, eventual consistency, fault tolerance, and retry mechanisms.
- Experience with containerization (Docker) and basic Kubernetes concepts.
- Knowledge of testing frameworks and practices (unit tests, integration tests, mocking).
- Familiarity with Git, CI/CD pipelines, and modern development workflows.
- High ownership, self driven and biased for action.
- Strong strategic thinking and ability to connect technical decisions to business impact.
- Excellent communication and mentoring skills.
- Thrives in ambiguity, fast-paced environments, and early-stage startup culture.
Skills
- Golang
- RESTful APIs
- gRPC
- PostgreSQL
- MySQL
- Redis
- Kafka
- RabbitMQ
- NATS
- Docker
- Kubernetes
- Git
- CI/CD
- HPC & Cluster Management
- Data Engineering
- Systems-Level Programming
- ML Platform Engineering
- Enterprise Deployment
Experience Level
- 2-4 years
Benefits
- Competitive compensation
- Flexible work options
- Wellness benefits
About the Company
- aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale.
- By abstracting away infrastructure complexity, aion unifies compute orchestration, training workflows, data pipelines, model versioning, and deployment into a streamlined enterprise experience.
- We're a fast-growing, VC-backed startup led by founders with a track record of successful exits.
- With teams across the US, UK, and India, we're building the next generation of enterprise AI infrastructure and are looking for exceptional people to help us scale.
