About the Role
As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy. We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.
Responsibilities
- Own a research question end to end
- Take one well-scoped problem from literature review through implementation, experimentation, and results
- Design ablations that isolate what actually caused an improvement
- Present your findings to the research team and defend the methodology
- Train and evaluate models on large-scale, real-world telephony audio
- Use our distributed GPU infrastructure
- Work with engineers to move research toward production
Requirements
- Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience
- Comfortable reading a paper and reimplementing it without hand-holding
- Experience with self-supervised, generative, or multimodal modeling
- Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning
- Strong intuition for audio quality and what makes synthetic speech sound wrong
- Fluent in PyTorch and comfortable in a real codebase
- Able to run your own experiments on GPU clusters without waiting to be unblocked
- Identify the single experiment that validates an idea in days, not months
- Measure everything and let data drive decisions
- Honest about negative results
- Obsessed with making voice agents sound truly human
- Use AI tools aggressively to amplify your own impact
Skills
- Speech-to-text
- Large language models
- Neural audio codecs
- Text-to-speech
- Self-supervised modeling
- Generative modeling
- Multimodal modeling
- Audio models
- Speech models
- TTS
- ASR
- Codecs
- Audio representation learning
- PyTorch
- GPU clusters
Location
- Levi's Plaza, SF
Work Type
- Internship
Experience Level
- Intern
Education Level
- MS or PhD in ML, CS, EE, or a related field, or equivalent research experience
Salary/Compensations
- Competitive intern compensation
Benefits
- Mentorship from researchers working on frontier voice AI
- Every tool you need to succeed
- Beautiful office in Levi's Plaza, SF with rooftop views
- A real shot at a return offer
About the Company
- Bland
