About the Role
Cerebras is developing new disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine. This role involves productionizing and optimizing our GPU serving stack, working with custom inference APIs, the vLLM serving runtime, the AMD ROCm software stack, and rack-scale AMD GPU infrastructure to ensure reliability, numerical correctness, observability, and high performance. The position requires writing production code, establishing operational practices for a new accelerator fleet, and improving key performance metrics like time to first token, throughput, tail latency, and capacity efficiency. It is a hands-on role demanding deep debugging and optimization across application, runtime, distributed systems, and hardware layers.
Responsibilities
- Productionize the GPU inference stack, including designing, building, deploying, and maintaining the complete GPU prefill path across API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
- Own GPU operational readiness by establishing deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet.
- Build automation to ensure explicit and reproducible compatibility for drivers, firmware, runtimes, models, and containers.
- Drive reliability in production by defining service-level indicators and objectives for GPU-backed inference.
- Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation across the serving stack.
- Improve inference performance by profiling and optimizing time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity under production workloads.
- Optimize model-serving behavior by tuning scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
- Debug complex failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
- Ensure numerical correctness by building validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and compatibility across software and hardware releases.
- Build performance and correctness infrastructure, including benchmarks, workload replay tools, profiling automation, release qualification, dashboards, and regression gates.
- Transform one-off investigations into repeatable engineering systems.
Requirements
- 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
- Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
- Strong programming ability in C++ and Python, including experience with multithreading, concurrency, memory management, and performance-sensitive software.
- Hands-on experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent internally developed system.
- Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling methodology.
- Experience debugging distributed systems across multiple layers rather than treating the serving framework or accelerator runtime as a black box.
- Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, CI/CD, and operating latency-sensitive services in production.
- Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
- Strong communication and technical leadership skills, with a demonstrated ability to drive ambiguous cross-functional projects to completion.
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
Skills
- C++
- Python
- Multithreading
- Concurrency
- Memory Management
- Performance-sensitive software
- vLLM
- SGLang
- TensorRT-LLM
- Triton Inference Server
- GPU execution
- GPU performance
- Asynchronous execution
- Memory movement
- Synchronization
- Kernel launches
- Communication overhead
- Profiling methodology
- Distributed systems debugging
- Linux
- Containers
- Kubernetes
- Observability
- CI/CD
- Latency-sensitive services
- Benchmarking
- Performance analysis
- Bottleneck identification
- AMD Instinct accelerators
- ROCm ecosystem
- HIP
- RCCL
- rocprofiler
- AMD SMI
- AITER
- hipBLASLt
- Composable Kernel
- CUDA
- PyTorch
- Prefill-heavy inference architectures
- Disaggregated prefill/decode inference architectures
- KV-cache transfer
- Prefix caching
- Continuous batching
- Chunked prefill
- Request scheduling
- Memory-aware admission control
- Multi-GPU inference
- Multi-node inference
- Tensor parallelism
- Pipeline parallelism
- Expert parallelism
- RDMA
- Collective communication
- Mixture-of-Experts models
- Multimodal models
- GPU kernel optimization
- Operator fusion
- Graph capture
- Attention kernels
- GEMM tuning
- Communication/computation overlap
- Reduced-precision inference
- Quantization formats (BF16, FP8, FP4, INT8, INT4)
- Numerical comparison systems
- Determinism test systems
- Model validation systems
- Performance regression test systems
Location
- Remote
Work Type
- Full-time
Experience Level
- 8+ years of software engineering experience
- Individual-contributor ownership of complex production systems
Education Level
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
Benefits
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Simple, non-corporate work culture that respects individual beliefs.
About the Company
- Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds.
- Their architecture transforms the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
- Cerebras collaborates with leading model labs, global enterprises, and cutting-edge AI-native startups.
- OpenAI has partnered with Cerebras to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
Equal Opportunity
- Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
- We celebrate different backgrounds, perspectives, and skills.
- We believe inclusive teams build better products and companies.
- We strive to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
