Senior Software Engineer, Inference Platform at AION | England, GB | Rezi

Senior Software Engineer, Inference Platform at AION

Senior Software Engineer, Inference Platform

AION · England, GB

1 months ago

Senior Software Engineer, Inference Platform

AION · England, GB

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale. We forward-deploy AI for enterprises, enabling organizations to rapidly implement production-ready AI solutions. We're a fast-growing, VC-backed startup led by founders with a track record of successful exits. We're building the next generation of enterprise AI infrastructure and are looking for exceptional people to help us scale.

Responsibilities

  • Design and build aion's inference service platform the backbone for serving AI models at scale across diverse workloads
  • Own and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and Autoscaler
  • Design highly modular, scalable, and extensible low-level designs (LLDs) for inference infrastructure components
  • Lead high-level design discussions, establish architectural patterns, and drive technical decision-making for the inference stack
  • Understand and optimize the dynamics of model deployment, version upgrades, and rollback strategies
  • Build robust deployment pipelines for seamless model updates with zero-downtime deployments
  • Design intelligent routing systems for multi-model serving, A/B testing, and canary deployments
  • Implement strategies for efficient GPU utilization and model cold-start optimization
  • Implement highly performant and optimized software for low-latency, high-throughput inference serving
  • Build and debug production-grade code in distributed systems handling real-time AI workloads
  • Optimize inference pipelines for latency, throughput, batching efficiency, and resource utilization
  • Design fault-tolerant systems with graceful degradation and automatic recovery mechanisms
  • Build high-performance telemetry and observability stack for inference metrics, performance tracking, and debugging
  • Implement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inference
  • Conduct thorough code reviews to maintain code quality, performance standards, and architectural consistency
  • Establish engineering best practices for testing, documentation, and production readiness.

Requirements

  • Seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads
  • Understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability
  • Designed distributed systems that handle thousands of requests per second while maintaining sub-second response times and cost efficiency
  • Experience with Golang is strongly preferred
  • Exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus
  • Take ownership of platform-level decisions
  • Think strategically about performance vs. cost trade-offs
  • Product-minded, understand how technical decisions impact developers using aion's platform and think about the end-to-end user experience
  • Team player comfortable wearing multiple hats
  • 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure
  • Strong understanding of AI/ML inference systems and experience with inference engines (vLLM, TGI, TensorRT-LLM, or similar)
  • Deep knowledge of distributed systems design, microservices architecture, and API gateway patterns
  • Proficiency in Golang strongly preferred; Python, Rust, C++ for performance-critical components a plus
  • Experience with container orchestration (Kubernetes, Docker) and infrastructure-as-code
  • Solid understanding of autoscaling strategies, load balancing, and resource scheduling algorithms
  • Experience building high-throughput, low-latency systems with sub-100ms response time requirements
  • Familiarity with message queues (Kafka, RabbitMQ), databases (PostgreSQL, Redis), and event-driven architectures
  • Knowledge of GPU computing, model serving optimizations (batching, quantization, multi-tenancy), and resource allocation
  • Experience with observability tools (Prometheus, Grafana, OpenTelemetry) and distributed tracing
  • Understanding of API design, rate limiting, authentication/authorization, and security best practices
  • Exposure to AI model deployment workflows and model lifecycle management is highly desirable
  • HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration
  • Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks
  • Systems-Level Programming: Experience with low-level systems programming such as storage systems, Kubernetes operators, OS-level software development, or daemon services (llm-d, system agents)
  • ML Platform Engineering: Experience productionizing ML pipelines, batch job orchestration, model fine-tuning workflows, and Jupyter notebook orchestration systems
  • Enterprise Deployment: Experience platformizing and packaging software for on-premises deployments or customer VPC installatiaons with emphasis on security, compliance, and operational simplicity
  • High ownership, self driven and biased for action
  • Strong strategic thinking and ability to connect technical decisions to business impact
  • Excellent communication and mentoring skills
  • Thrives in ambiguity, fast-paced environments, and early-stage startup culture.

Skills

  • Golang
  • Inference engines (vLLM, TGI, TensorRT)
  • Containerization
  • Distributed systems
  • AI/ML inference systems
  • Microservices architecture
  • API gateway patterns
  • Python
  • Rust
  • C++
  • Kubernetes
  • Docker
  • Infrastructure-as-code
  • Autoscaling strategies
  • Load balancing
  • Resource scheduling algorithms
  • Message queues (Kafka, RabbitMQ)
  • Databases (PostgreSQL, Redis)
  • Event-driven architectures
  • GPU computing
  • Model serving optimizations (batching, quantization, multi-tenancy)
  • Resource allocation
  • Observability tools (Prometheus, Grafana, OpenTelemetry)
  • Distributed tracing
  • API design
  • Rate limiting
  • Authentication/authorization
  • Security best practices
  • AI model deployment workflows
  • Model lifecycle management
  • HPC & Cluster Management
  • Kubernetes
  • Slurm
  • Job scheduling
  • Resource allocation
  • Workload orchestration
  • Data Engineering
  • Data pipelines
  • ETL systems
  • Large-scale data processing frameworks
  • Systems-Level Programming
  • Storage systems
  • Kubernetes operators
  • OS-level software development
  • Daemon services (llm-d, system agents)
  • ML Platform Engineering
  • ML pipelines
  • Batch job orchestration
  • Model fine-tuning workflows
  • Jupyter notebook orchestration systems
  • Enterprise Deployment
  • On-premises deployments
  • Customer VPC installations
  • Security
  • Compliance
  • Operational simplicity

Experience Level

  • 4+ years of experience building and scaling backend systems, distributed platforms, or inference infrastructure

Benefits

  • Competitive compensation
  • Flexible work options
  • Wellness benefits

About the Company

  • aion is the Enterprise AI Platform, a full-stack solution for building, fine-tuning, deploying, and forward-deploying AI for enterprises at scale.
  • By abstracting away infrastructure complexity, aion unifies compute orchestration, training workflows, data pipelines, model versioning, and deployment into a streamlined enterprise experience.
  • We forward-deploy AI for enterprises, enabling organizations to rapidly implement production-ready AI solutions.
  • We're a fast-growing, VC-backed startup led by founders with a track record of successful exits.
  • With teams across the US, UK, and India, we're building the next generation of enterprise AI infrastructure and are looking for exceptional people to help us scale.