ML Systems Integration Engineer at Cerebras Systems | Sunnyvale, US | Rezi

ML Systems Integration Engineer at Cerebras Systems

ML Systems Integration Engineer

Cerebras Systems · Sunnyvale, US

2 weeks ago

ML Systems Integration Engineer

Cerebras Systems · Sunnyvale, US

18 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Cerebras Systems builds the world's largest AI chip, enabling industry-leading training and inference speeds that transform AI application user experiences and unlock real-time iteration and increased intelligence. We collaborate with leading AI organizations, including a multi-year partnership with OpenAI.

Responsibilities

  • Participate in bring-up of next-generation AI hardware systems and supporting software infrastructure.
  • Debug complex system-level issues spanning hardware and software interactions.
  • Investigate failures occurring during system bring-up and identify root causes using logs, telemetry, and diagnostic tools.
  • Build automation frameworks and internal tooling that improve system validation and debugging workflows.
  • Develop software used to test, validate, and stress distributed hardware systems during development and production cycles.
  • Collaborate closely with hardware engineers to isolate and resolve system integration issues.
  • Improve system observability by building tools that surface failures quickly and accelerate debugging.
  • Reproduce, triage, and diagnose difficult issues that arise during early hardware deployment.
  • Support validation and qualification of new hardware generations as systems move toward production readiness.
  • Continuously improve internal engineering workflows related to debugging, testing, and automation.

Requirements

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related technical field.
  • Strong programming skills in Python and/or C++.
  • Excellent debugging and problem-solving skills with ability to investigate complex technical issues methodically.
  • Solid understanding of operating systems fundamentals (processes, threads, memory management, concurrency, IPC).
  • Experience working in Linux development environments.
  • Understanding of computer architecture and interactions between hardware and software systems.
  • Strong analytical thinking and ability to break down complex system failures into actionable root causes.
  • Ability to work effectively across multiple engineering teams and collaborate in highly technical environments.
  • Strong communication skills and willingness to work on ambiguous technical problems.

Skills

  • Python
  • C++
  • Operating systems fundamentals
  • Linux development environments
  • Computer architecture
  • Hardware-software interactions
  • Automation frameworks
  • Internal tooling
  • Test infrastructure
  • Distributed systems concepts
  • Networking fundamentals
  • Performance analysis
  • System telemetry
  • Log analysis
  • Production systems validation
  • Infrastructure reliability engineering

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Mid-level

Education Level

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related technical field

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

About the Company

  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.
  • This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
  • This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
  • Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
  • OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

Equal Opportunity

  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • We celebrate different backgrounds, perspectives, and skills.
  • We believe inclusive teams build better products and companies.
  • We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.