Staff Software Engineer, Inference API at Cerebras Systems | CA | Rezi

Staff Software Engineer, Inference API at Cerebras Systems

Staff Software Engineer, Inference API

Cerebras Systems · CA

Yesterday

Staff Software Engineer, Inference API

Cerebras Systems · CA

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Staff Software Engineer, Inference API role.

Rezi rewrites your resume against Cerebras Systems's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Staff Software Engineer, Inference API posting at Cerebras Systems — free, in seconds.

About the Role

Cerebras is developing new disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine. This role involves building and evolving the ML API layer to make this system accessible, reliable, and user-friendly, working across inference APIs, model integration, request routing, and runtime components.

Responsibilities

  • Build production ML inference APIs for various capabilities like chat completions, text generation, streaming, and multimodal inputs.
  • Deliver a unified serving experience with consistent request and response semantics across heterogeneous inference backends.
  • Integrate emerging foundation models, tokenizers, and sampling methods into the serving platform.
  • Own API compatibility and evolution, maintaining compatibility with existing interfaces while designing extensions.
  • Integrate custom inference services with frameworks like vLLM, PyTorch, and Hugging Face libraries.
  • Support disaggregated inference by building control and data paths for coordinating GPU prefill with Cerebras decode.
  • Optimize serving performance, including streaming behavior, latency, and throughput.
  • Ensure functional and numerical correctness through validation systems.
  • Strengthen reliability and observability with structured logging, tracing, and metrics.
  • Develop testing and qualification infrastructure, including conformance tests and performance benchmarks.
  • Improve developer experience through intuitive configuration, SDKs, and documentation.
  • Collaborate with compiler, runtime, cloud, and product teams to translate requirements into scalable capabilities.

Requirements

  • 5+ years of software engineering experience with ownership of production software or distributed systems.
  • Strong programming ability in Python.
  • Experience developing performance-sensitive or highly concurrent services in C++, Go, or a similar systems language.
  • Hands-on experience with a model-serving framework (e.g., vLLM, SGLang, TensorRT-LLM, Triton).
  • Understanding of modern LLM inference concepts (tokenization, sampling, KV-cache management, etc.).
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries.
  • Experience building stable APIs with clear validation, error handling, and versioning practices.
  • Experience with Linux, containers, Kubernetes, CI/CD, and operating latency-sensitive services.
  • Ability to diagnose correctness, reliability, and performance issues in distributed systems.
  • Strong communication and cross-functional execution skills.
  • Bachelor’s degree in computer science, Computer Engineering, Electrical Engineering, or equivalent practical experience.

Skills

  • Python
  • C++
  • Go
  • vLLM
  • SGLang
  • TensorRT-LLM
  • Triton Inference Server
  • Hugging Face Text Generation Inference
  • Linux
  • Kubernetes
  • CI/CD
  • API Design
  • Distributed Systems
  • Machine Learning Systems
  • Model Serving
  • OpenAI-compatible APIs
  • gRPC
  • REST
  • Streaming APIs

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related discipline, or equivalent practical experience.

Benefits

  • Job stability with startup vitality
  • Simple, non-corporate work culture
  • Respect for individual beliefs
  • Continuous learning, growth, and support

About the Company

  • Cerebras Systems builds the world's largest AI chip, significantly larger than GPUs, enabling industry-leading training and inference speeds.
  • Their architecture transforms AI application user experiences, unlocking real-time iteration and increasing intelligence.
  • Cerebras collaborates with leading model labs, global enterprises, and AI startups, including a partnership with OpenAI to deploy 750 megawatts of scale for ultra high-speed inference.
  • The company offers a breakthrough AI platform beyond GPU constraints, opportunities to publish and open source AI research, and access to one of the fastest AI supercomputers globally.

Equal Opportunity

  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • They celebrate different backgrounds, perspectives, and skills, believing inclusive teams build better products and companies.
  • The company strives to build a work environment that empowers people to do their best work through continuous learning, growth, and support.