Staff Software Engineer, GPU Inference at Cerebras Systems | Canada | Rezi

Staff Software Engineer, GPU Inference at Cerebras Systems

Staff Software Engineer, GPU Inference

Cerebras Systems · Canada

4 days ago

Staff Software Engineer, GPU Inference

Cerebras Systems · Canada

4 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Cerebras is developing new disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine. This role involves productionizing and optimizing our GPU serving stack, working with custom inference APIs, the vLLM serving runtime, the AMD ROCm software stack, and rack-scale AMD GPU infrastructure to ensure reliability, numerical correctness, observability, and high performance. The position requires writing production code, establishing operational practices for a new accelerator fleet, and improving key performance metrics like time to first token, throughput, tail latency, and capacity efficiency. It is a hands-on role demanding deep debugging and optimization across application, runtime, distributed systems, and hardware layers.

Responsibilities

  • Productionize the GPU inference stack, including designing, building, deploying, and maintaining the complete GPU prefill path across API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Own GPU operational readiness by establishing deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet.
  • Build automation to ensure explicit and reproducible compatibility for drivers, firmware, runtimes, models, and containers.
  • Drive reliability in production by defining service-level indicators and objectives for GPU-backed inference.
  • Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation across the serving stack.
  • Improve inference performance by profiling and optimizing time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity under production workloads.
  • Optimize model-serving behavior by tuning scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
  • Debug complex failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
  • Ensure numerical correctness by building validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and compatibility across software and hardware releases.
  • Build performance and correctness infrastructure, including benchmarks, workload replay tools, profiling automation, release qualification, dashboards, and regression gates.
  • Transform one-off investigations into repeatable engineering systems.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
  • Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
  • Strong programming ability in C++ and Python, including experience with multithreading, concurrency, memory management, and performance-sensitive software.
  • Hands-on experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent internally developed system.
  • Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling methodology.
  • Experience debugging distributed systems across multiple layers rather than treating the serving framework or accelerator runtime as a black box.
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, CI/CD, and operating latency-sensitive services in production.
  • Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
  • Strong communication and technical leadership skills, with a demonstrated ability to drive ambiguous cross-functional projects to completion.
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Skills

  • C++
  • Python
  • Multithreading
  • Concurrency
  • Memory Management
  • Performance-sensitive software
  • vLLM
  • SGLang
  • TensorRT-LLM
  • Triton Inference Server
  • GPU execution
  • GPU performance
  • Asynchronous execution
  • Memory movement
  • Synchronization
  • Kernel launches
  • Communication overhead
  • Profiling methodology
  • Distributed systems debugging
  • Linux
  • Containers
  • Kubernetes
  • Observability
  • CI/CD
  • Latency-sensitive services
  • Benchmarking
  • Performance analysis
  • Bottleneck identification
  • AMD Instinct accelerators
  • ROCm ecosystem
  • HIP
  • RCCL
  • rocprofiler
  • AMD SMI
  • AITER
  • hipBLASLt
  • Composable Kernel
  • CUDA
  • PyTorch
  • Prefill-heavy inference architectures
  • Disaggregated prefill/decode inference architectures
  • KV-cache transfer
  • Prefix caching
  • Continuous batching
  • Chunked prefill
  • Request scheduling
  • Memory-aware admission control
  • Multi-GPU inference
  • Multi-node inference
  • Tensor parallelism
  • Pipeline parallelism
  • Expert parallelism
  • RDMA
  • Collective communication
  • Mixture-of-Experts models
  • Multimodal models
  • GPU kernel optimization
  • Operator fusion
  • Graph capture
  • Attention kernels
  • GEMM tuning
  • Communication/computation overlap
  • Reduced-precision inference
  • Quantization formats (BF16, FP8, FP4, INT8, INT4)
  • Numerical comparison systems
  • Determinism test systems
  • Model validation systems
  • Performance regression test systems

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • 8+ years of software engineering experience
  • Individual-contributor ownership of complex production systems

Education Level

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Simple, non-corporate work culture that respects individual beliefs.

About the Company

  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds.
  • Their architecture transforms the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
  • Cerebras collaborates with leading model labs, global enterprises, and cutting-edge AI-native startups.
  • OpenAI has partnered with Cerebras to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

Equal Opportunity

  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • We celebrate different backgrounds, perspectives, and skills.
  • We believe inclusive teams build better products and companies.
  • We strive to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.