About the Role
Nscale is seeking Senior / Staff AI Engineers to join its core AI team and develop the systems powering its GenAI cloud platform. This role focuses on designing and optimizing distributed systems for large-scale AI tasks, including training, post-training, evaluation, and inference, under strict performance and efficiency constraints. It's a hands-on position for engineers aiming to advance AI system development, optimization, and consumption.
Responsibilities
- Design, build, and optimize scalable AI platform systems.
- Drive inference performance and efficiency, including KV cache management, continuous batching, speculative decoding, and quantization.
- Build and improve post-training services, including fine-tuning (LoRA, QLoRA, adapters, full fine-tuning) and alignment (RLHF, DPO, reward modelling).
- Develop dataset curation and data processing workflows.
- Develop evaluation and benchmarking systems to measure model quality, safety, regression, system performance, and real-world behavior.
- Develop and optimize distributed systems for GPU/accelerator workloads, focusing on scalability, reliability, and efficiency.
- Conduct performance analysis and bottleneck investigations across multiple components and stacks.
- Collaborate with research, infrastructure, and product teams to build platform components based on customer demand and industry direction.
- Build developer-facing APIs, SDKs, and tooling for Nscale’s AI services.
Requirements
- 5+ years of experience building production systems in machine learning, distributed systems, or high-performance infrastructure.
- 4+ years of hands-on experience in at least one core area within large-scale, production AI environments (e.g., AI labs, hyperscalers), such as inference optimization, large-scale training/pre-training systems, post-training, or evaluation and benchmarking frameworks.
- Strong hands-on expertise in at least one core area, with working knowledge across others.
- Proven ability to design, optimize, and operate systems at scale, with a strong understanding of performance trade-offs.
- Deep understanding of transformer architectures, LLMs, and/or multimodal models.
- Strong proficiency in Python and PyTorch, with a track record of building production-grade ML systems.
- Experience with distributed compute and training paradigms (e.g., data/model parallelism, sharding, scheduling).
- Experience working close to the hardware/software boundary, such as GPU/accelerator optimization or memory management.
- Experience building or operating production inference or training systems at scale.
- Ability to design clean abstractions, APIs, and reusable systems for other engineers.
- Strong engineering fundamentals, with a track record of writing maintainable, well-tested, production-quality code.
Skills
- Python
- PyTorch
- Machine Learning
- Distributed Systems
- High-Performance Infrastructure
- Inference Optimization
- Large-Scale Training
- Post-Training (Fine-tuning, Alignment, Distillation)
- Evaluation and Benchmarking Frameworks
- Transformer Architectures
- LLMs
- Multimodal Models
- Distributed Compute
- Training Paradigms
- GPU/Accelerator Optimization
- CUDA
- ROCm
- Memory Management
- System-Level Performance Tuning
- API Design
- OpenAPI 3.0+
Experience Level
- Senior
- Staff
About the Company
- Nscale is building a vertically integrated GenAI cloud platform, owning data centers, software, and applications for the AI stack using sustainable technology.
- The company culture emphasizes relentless innovation, ownership, accountability, pride in work, excellence, urgency, openness, transparency, inspiration, collaboration, swiftness, respect, adaptability, and resilience.
Equal Opportunity
- Nscale is committed to fostering an inclusive, diverse, and equitable workplace.
- The company encourages applications from candidates of all backgrounds, experiences, and abilities, including people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
- Nscale will provide accommodations for specific situations upon request.
