Research Engineer, Audio and Speech at Decagon | CA, US | Rezi

Research Engineer, Audio and Speech at Decagon

Research Engineer, Audio and Speech

Decagon · CA, US

2 weeks ago

Research Engineer, Audio and Speech

Decagon · CA, US

19 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Research Engineer, Audio and Speech role.

Rezi rewrites your resume against Decagon's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Research Engineer, Audio and Speech posting at Decagon — free, in seconds.

About the Role

As a Research Engineer focused on Audio and Speech, you will build models and agent harnesses for Decagon’s real-time voice agents, taking them from idea to production. Your work will advance multimodal and full-duplex systems capable of natural, real-time listening, reasoning, speaking, and responding. We seek engineers eager to build the next generation of AI voice agents, who own their work end-to-end, ship improvements, and make high-impact technical decisions.

Responsibilities

  • Design and build next-generation agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction
  • Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages
  • Build evaluations and use production calls to ship measurable improvements in accuracy, latency, naturalness, and task outcomes
  • Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with Voice Platform and Infrastructure teams to deploy at scale

Requirements

  • 2+ years of experience in speech, audio ML, multimodal ML, or production machine learning
  • Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models
  • Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio
  • Fluency in Python and a modern deep-learning framework such as PyTorch
  • Strong foundations in machine learning and signal processing
  • A track record of taking research ideas from prototype to reliable, measurable production impact

Skills

  • Speech
  • Audio ML
  • Multimodal ML
  • Production machine learning
  • Autoregressive models
  • Diffusion models
  • Flow-matching models
  • Codec-based speech models
  • Streaming agent systems
  • Low-latency inference
  • Production model serving
  • Python
  • PyTorch
  • Machine learning
  • Signal processing
  • Speech-to-speech models
  • Full-duplex models
  • Telephony
  • Multilingual speech
  • Noisy-channel robustness
  • Speaker adaptation
  • Expressive speech generation

Location

  • In-office

Work Type

  • Full-time
  • In-office

Experience Level

  • 2+ years of experience

Salary/Compensations

  • $200K – $400K + Offers Equity

Benefits

  • Medical, Dental, and Vision benefits for you and your family
  • Life Insurance and Disability Benefits
  • Retirement Plan (e.g., 401K, pension)
  • Parental Leave
  • Fertility and family building benefits through Carrot
  • Monthly stipend to support your wellness, lifestyle, and work-life balance
  • Daily lunches and snacks in the office
  • Take what you need vacation policy

About the Company

  • Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.
  • Our technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel.
  • We’re building a future where customer experiences are being redefined from support tickets and hold music to faster resolutions, richer conversations, and deeper relationships.
  • We’re proud to be backed by world-class investors who share that vision, including a16z, Accel, Bain Capital Ventures, Coatue, and Index Ventures, along with many others.
  • We’re an in-office company, driven by a shared commitment to excellence and velocity.
  • Our values — Just Get It Done, Invent What Customers Want, Winner’s Mindset, and The Polymath Principle — shape how we work and grow as a team.