Senior Machine Learning Engineer at Connect | City of London, England, GB | Rezi

Senior Machine Learning Engineer at Connect

Senior Machine Learning Engineer

Connect · City of London, England, GB

1 months ago

Senior Machine Learning Engineer

Connect · City of London, England, GB

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Senior Machine Learning Engineer to design, deploy, and optimize next-generation Conversational AI and Data Analytics platforms, bridging the gap between AI research and production engineering. The primary focus will be optimizing and scaling core Speech (ASR, TTS) and Language Model (LLM, SLM) pipelines across hybrid cloud and local edge environments.

Responsibilities

  • Build production-grade, low-latency pipelines for ASR, TTS, and Small Language Models (SLMs).
  • Manage deployment topologies across multi-cloud environments and bare-metal local hardware.
  • Create high-performance, asynchronous REST and WebSocket APIs using FastAPI to serve real-time conversational agents.
  • Design automated machine learning pipelines for model testing, versioning, and continuous deployment.
  • Pack applications using Docker or Podman for consistent execution across dev, staging, and production.
  • Orchestrate cloud infrastructure across AWS, Azure, and GCP, optimizing for compute efficiency and cost.
  • Maximize hardware utilization for single-GPU and distributed multi-GPU environments.
  • Optimize Python code execution using Numba, NumPy, and specialized CUDA libraries.
  • Conduct rigorous load and stress testing to guarantee system stability under high concurrent traffic.

Requirements

  • Mastery of Python and its asynchronous ecosystem.
  • Deep expertise in PyTorch, Scikit-learn, and NumPy.
  • Experience accelerating Python code via Numba or Triton.
  • Hands-on experience deploying Automated Speech Recognition (ASR) and Text-to-Speech (TTS) models.
  • Familiarity with optimizing and serving Large Language Models (LLMs) and resource-efficient Small Language Models (SLMs).
  • Advanced knowledge of Docker, Podman, and container orchestration.
  • Practical experience managing AI workloads on AWS (EC2, SageMaker), Azure (Azure ML), and GCP (Vertex AI).
  • Experience with GitLab CI, GitHub Actions, Jenkins, or specialized MLOps platforms (e.g., Kubeflow, MLflow).
  • Experience with dialogue management, prompt engineering, and Retrieval-Augmented Generation (RAG).
  • Familiarity with real-time data streaming (e.g., Kafka) and vector databases (e.g., Pinecone, Milvus, Qdrant).
  • Experience with model compression techniques like quantization (INT8/FP4), pruning, and distillation for edge deployment.

Skills

  • Python
  • Asynchronous Programming
  • PyTorch
  • Scikit-learn
  • NumPy
  • Numba
  • Triton
  • ASR
  • TTS
  • LLMs
  • SLMs
  • Docker
  • Podman
  • Container Orchestration
  • AWS
  • Azure
  • GCP
  • GitLab CI
  • GitHub Actions
  • Jenkins
  • Kubeflow
  • MLflow
  • Dialogue Management
  • Prompt Engineering
  • RAG
  • Kafka
  • Vector Databases
  • Pinecone
  • Milvus
  • Qdrant
  • Quantization
  • Pruning
  • Distillation

Experience Level

  • Senior