AI Research Engineer, Model Optimization and Inference at Iconic | GBR | Rezi

AI Research Engineer, Model Optimization and Inference at Iconic

AI Research Engineer, Model Optimization and Inference

Iconic · GBR

4 days ago

AI Research Engineer, Model Optimization and Inference

Iconic · GBR

4 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Architect and build the inference engine that powers virtual actors, bridging the gap between massive research models and the constraints of real-time interactive entertainment. You will optimize model architectures for speed on consumer hardware while maintaining character abilities.

Responsibilities

  • Architect Low-Latency Runtimes: Build and maintain high-performance inference pipelines for Multimodal LLMs, TTS, and Vision models, targeting both server-side and consumer edge environments.
  • State-of-the-Art Optimization: Implement advanced techniques like Speculative Decoding, KV-Cache quantization, PagedAttention, and Layer Pruning to minimize TTFT and TPOT, maximizing throughput.
  • Model Compression: Lead efforts in post-training quantization and distillation to fit massive models into consumer VRAM budgets.
  • Engine Integration: Collaborate with the game engineering team to ensure thread-safe, non-blocking asynchronous inference within the game loop.
  • Custom Kernel Development: Write custom ops in CUDA, Triton, or Metal when off-the-shelf kernels aren't fast enough.

Requirements

  • MSc or PhD in Computer Science, Machine Learning, or a related field (or equivalent industry experience)
  • Strong experience with model optimization techniques (quantization, pruning, distillation, knowledge transfer)
  • Experience with LLM-specific inference optimizations (KV-cache management, speculative decoding, attention mechanisms)
  • Proficiency in C/C++
  • Hands-on experience deploying ML models on-device or in latency-sensitive environments
  • Proficiency in Python and deep learning frameworks (PyTorch, JAX, or TensorFlow)
  • Experience with inference optimization tools and runtimes (TensorRT, ONNX Runtime, Core ML, or similar)
  • Strong systems and engineering skills
  • Excellent collaboration and communication skills

Skills

  • Model Optimization
  • Quantization
  • Pruning
  • Distillation
  • Knowledge Transfer
  • LLM Inference Optimization
  • KV-Cache Management
  • Speculative Decoding
  • Attention Mechanisms
  • C/C++
  • Python
  • PyTorch
  • JAX
  • TensorFlow
  • TensorRT
  • ONNX Runtime
  • Core ML
  • Systems Engineering
  • CUDA
  • Triton
  • Metal
  • ExecuTorch
  • MLX
  • AMD/ROCm
  • DirectML
  • Vulkan Compute
  • Real-time Systems
  • Game Engines
  • Real-Time Rendering

Location

  • London

Work Type

  • Hybrid

Experience Level

  • Foundational member

Education Level

  • MSc or PhD in Computer Science, Machine Learning, or a related field (or equivalent industry experience)

Salary/Compensations

  • Competitive salary and equity compensation

Benefits

  • 25 days annual leave + bank holidays
  • Private healthcare
  • Inclusive & friendly company culture with socials and game breaks

About the Company

  • Iconic is innovating at the intersection of AI, art, and storytelling, building virtual actors that perform.
  • We are a company building toward something genuinely new.