Solution Architect - GPU & HPC at NexGen Cloud | London, England, GB | Rezi

Solution Architect - GPU & HPC at NexGen Cloud

Solution Architect - GPU & HPC

NexGen Cloud · London, England, GB

2 weeks ago

Solution Architect - GPU & HPC

NexGen Cloud · London, England, GB

14 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

The Solutions Architect sits at the intersection of sales, infrastructure, and the customer, translating complex workload requirements into technically sound, commercially viable solutions on the Hyperstack platform. You’ll be the primary technical authority through the sales cycle, engaging directly with prospective and existing customers, producing detailed solution designs, and ensuring proposed solutions are deliverable.

Responsibilities

  • Own the technical sales cycle end-to-end, acting as the primary technical authority for GPU cloud solution design.
  • Engage directly with prospective and existing customers to understand requirements and produce detailed solution designs.
  • Collaborate with internal teams to validate delivery feasibility before customer commitments.
  • Build and maintain a library of reference architectures and solution templates.
  • Develop high-quality technical proposals, RFP responses, and statements of work.
  • Define and maintain comprehensive Bills of Materials (BoMs) for all proposed solutions.
  • Feed back recurring customer requirements and competitive intelligence to engineering leadership.

Requirements

  • Proven experience in HPC or AI software stack design and delivery at scale.
  • Deep understanding of GPU software environments: CUDA, cuDNN, NCCL, driver stacks.
  • Hands-on experience optimising AI and HPC workloads across multi-GPU and multi-node configurations.
  • Strong working knowledge of containerisation and orchestration in HPC/AI contexts: Docker, Kubernetes, NVIDIA GPU Operator.
  • Background in an OEM, hyperscaler, neo-cloud, or enterprise/research HPC environment.
  • Ability to produce clear, professional technical documentation and architecture diagrams.
  • Confident engaging with customers, vendors, and internal engineering teams as a technical authority.
  • Experience with large-scale cluster performance benchmarking — NCCL tests, MLPerf, or equivalent.
  • Exposure to MLOps tooling and AI platform layers.
  • Familiarity with InfiniBand and high-performance networking.
  • Commercial awareness: experience contributing to BoMs, technical proposals, or RFP responses.

Skills

  • HPC
  • AI infrastructure
  • GPU cloud
  • Workload profiling
  • Scheduler configuration (SLURM, PBS, or equivalent)
  • MPI/NCCL tuning
  • Distributed training frameworks (PyTorch, JAX, DeepSpeed)
  • CUDA
  • cuDNN
  • NCCL
  • Driver stacks
  • AI training
  • Inference workloads
  • Containerisation
  • Orchestration
  • Docker
  • Kubernetes
  • NVIDIA GPU Operator
  • MLOps tooling
  • AI platform layers
  • MLflow
  • W&B
  • Model serving frameworks (Triton, vLLM)
  • Kubeflow
  • Airflow
  • InfiniBand
  • High-performance networking

Location

  • UK-based

Work Type

  • Customer-Site Travel Required
  • Flexible working arrangements

Experience Level

  • Proven experience
  • Deep understanding
  • Hands-on experience
  • Strong working knowledge
  • Background in
  • Demonstrable exposure
  • Confident engaging
  • Experience with
  • Exposure to
  • Familiarity with
  • Commercial awareness

Salary/Compensations

  • Competitive salary
  • Annual discretionary bonus scheme

Benefits

  • Employee wellbeing benefits
  • 25 days of holiday, plus public holidays

About the Company

  • NexGen Cloud is the company behind Hyperstack, a full-stack AI cloud serving tens of thousands of customers from AI researchers to enterprises running the world’s most compute-intensive workloads.
  • We deliver on-demand and private GPU infrastructure to teams who treat performance as a requirement, not a feature.
  • We’re a tight-knit, fast-moving team working at the cutting edge of AI cloud infrastructure.
  • We practise what we preach, equipping our people with AI at every level so we can solve harder problems, ship faster, and keep raising the bar for what enterprise GPU infrastructure looks like.