About the Role
Architect and build the inference engine that powers virtual actors, bridging the gap between massive research models and the constraints of real-time interactive entertainment. You will optimize model architectures for speed on consumer hardware while maintaining character abilities.
Responsibilities
- Architect Low-Latency Runtimes: Build and maintain high-performance inference pipelines for Multimodal LLMs, TTS, and Vision models, targeting both server-side and consumer edge environments.
- State-of-the-Art Optimization: Implement advanced techniques like Speculative Decoding, KV-Cache quantization, PagedAttention, and Layer Pruning to minimize TTFT and TPOT, maximizing throughput.
- Model Compression: Lead efforts in post-training quantization and distillation to fit massive models into consumer VRAM budgets.
- Engine Integration: Collaborate with the game engineering team to ensure thread-safe, non-blocking asynchronous inference within the game loop.
- Custom Kernel Development: Write custom ops in CUDA, Triton, or Metal when off-the-shelf kernels aren't fast enough.
Requirements
- MSc or PhD in Computer Science, Machine Learning, or a related field (or equivalent industry experience)
- Strong experience with model optimization techniques (quantization, pruning, distillation, knowledge transfer)
- Experience with LLM-specific inference optimizations (KV-cache management, speculative decoding, attention mechanisms)
- Proficiency in C/C++
- Hands-on experience deploying ML models on-device or in latency-sensitive environments
- Proficiency in Python and deep learning frameworks (PyTorch, JAX, or TensorFlow)
- Experience with inference optimization tools and runtimes (TensorRT, ONNX Runtime, Core ML, or similar)
- Strong systems and engineering skills
- Excellent collaboration and communication skills
Skills
- Model Optimization
- Quantization
- Pruning
- Distillation
- Knowledge Transfer
- LLM Inference Optimization
- KV-Cache Management
- Speculative Decoding
- Attention Mechanisms
- C/C++
- Python
- PyTorch
- JAX
- TensorFlow
- TensorRT
- ONNX Runtime
- Core ML
- Systems Engineering
- CUDA
- Triton
- Metal
- ExecuTorch
- MLX
- AMD/ROCm
- DirectML
- Vulkan Compute
- Real-time Systems
- Game Engines
- Real-Time Rendering
Location
- London
Work Type
- Hybrid
Experience Level
- Foundational member
Education Level
- MSc or PhD in Computer Science, Machine Learning, or a related field (or equivalent industry experience)
Salary/Compensations
- Competitive salary and equity compensation
Benefits
- 25 days annual leave + bank holidays
- Private healthcare
- Inclusive & friendly company culture with socials and game breaks
About the Company
- Iconic is innovating at the intersection of AI, art, and storytelling, building virtual actors that perform.
- We are a company building toward something genuinely new.
