Staff Software Engineer, Inference API at Cerebras Systems | CA | Rezi

Staff Software Engineer, Inference API at Cerebras Systems

Staff Software Engineer, Inference API

Cerebras Systems · CA

Today

Staff Software Engineer, Inference API

Cerebras Systems · CA

an hour ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Staff Software Engineer, Inference API role.

Rezi rewrites your resume against Cerebras Systems's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Staff Software Engineer, Inference API posting at Cerebras Systems — free, in seconds.

About the Role

Cerebras is developing a new generation of disaggregated AI inference systems. This role focuses on building and enhancing the ML API layer to ensure this heterogeneous serving system is accessible, reliable, and user-friendly. You will work across inference APIs, model integration, request routing, and the Cerebras inference platform to provide a consistent experience regardless of models or accelerator backends. This position is at the intersection of machine learning systems, API design, model serving, and distributed systems, enabling new model architectures and inference capabilities while ensuring production-quality features.

Responsibilities

  • Build production ML inference APIs for various capabilities including chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs.
  • Deliver a unified serving experience by creating consistent request and response semantics across heterogeneous inference backends like GPU prefill and Cerebras decode.
  • Enable new models and capabilities by integrating emerging foundation models, tokenizers, prompt formats, sampling methods, and other features into the serving platform.
  • Own API compatibility and evolution by maintaining compatibility with existing interfaces while designing Cerebras-specific extensions, establishing clear versioning, deprecation, validation, and backward compatibility practices.
  • Integrate custom inference services with various runtimes including vLLM, PyTorch, Hugging Face libraries, AMD ROCm stack, and Cerebras runtime components.
  • Support disaggregated inference by building control and data paths for coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management.
  • Improve serving performance by optimizing streaming, latency, throughput, batching, serialization, tokenization, scheduling, and inter-component communication.
  • Ensure functional and numerical correctness by building validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and compatibility across serving backends.
  • Strengthen reliability and observability by defining service indicators and building structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling.
  • Develop testing and qualification infrastructure, including conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates.
  • Improve developer experience by building intuitive configuration, SDKs, documentation, examples, debugging tools, and self-service workflows.
  • Collaborate with compiler, runtime, kernel, cloud, product, and solutions teams to translate model and customer requirements into scalable serving capabilities.

Requirements

  • 5+ years of software engineering experience with significant individual-contributor ownership of production software or distributed systems.
  • Strong programming ability in Python and Go.
  • Experience developing performance-sensitive or highly concurrent services in C++, Rust, or a similar systems language.
  • Experience building stable APIs with clear validation, error handling, observability, compatibility, and versioning practices.
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries.
  • Experience designing or maintaining OpenAI-compatible, gRPC, REST, or streaming inference APIs.
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production.
  • Ability to diagnose correctness, reliability, and performance issues across multiple components of a distributed serving system.
  • Strong communication and cross-functional execution skills.
  • Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Skills

  • Python
  • Go
  • C++
  • Rust
  • API Design
  • Distributed Systems
  • Machine Learning Systems
  • Model Serving
  • OpenAI-compatible APIs
  • gRPC
  • REST
  • Streaming APIs
  • Linux
  • Containers
  • Kubernetes
  • CI/CD
  • vLLM
  • SGLang
  • PyTorch
  • Hugging Face Transformers
  • Trition
  • TensorRT-LLM
  • Multi-model inference
  • Multi-tenant inference
  • Reduced-precision inference
  • Quantization formats (BF16, FP8, FP4, INT8, INT4)

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • 5+ years of software engineering experience

Education Level

  • Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

About the Company

  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs, delivering industry-leading training and inference speeds.
  • Their architecture transforms the user experience of AI applications, unlocking real-time iteration and increasing intelligence.
  • Cerebras collaborates with leading model labs, global enterprises, and AI-native startups, including a multi-year partnership with OpenAI.
  • The company offers a simple, non-corporate work culture that respects individual beliefs.
  • Employees can build a breakthrough AI platform beyond GPU constraints, publish and open source cutting-edge AI research, and work on one of the fastest AI supercomputers globally.
  • Cerebras provides job stability with startup vitality.

Equal Opportunity

  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • They celebrate different backgrounds, perspectives, and skills, believing inclusive teams build better products and companies.
  • The company strives to build a work environment that empowers people to do their best work through continuous learning, growth, and support.