About the Role
As an intern on the Speech team at Spotify, you will collaborate with engineers and researchers to drive research in speech synthesis, speech recognition, and audio generation, helping define the next generation of audio interaction at Spotify.
Responsibilities
- Conduct research in speech technology areas such as text-to-speech synthesis, automatic speech recognition, neural audio codecs, or voice conversion.
- Develop and evaluate models using real-world data at scale, from problem formulation and data analysis through to experimentation and iteration.
- Collaborate with scientists, engineers, product managers, and designers across Spotify to translate research into product impact.
- Contribute to a long-term research roadmap while working on a focused internship project with clear deliverables.
- Publish work at relevant conferences.
Requirements
- Currently enrolled in or recently completed a PhD, MSc, or equivalent program in machine learning, signal processing, or a related area.
- Solid hands-on skills in sourcing, cleaning, manipulating, analysing, and modelling real data.
- Creative problem-solver passionate about digging into complex problems and devising new approaches to reach results.
- Experience with modern deep learning frameworks (e.g. PyTorch).
- Research experience in speech synthesis, speech recognition, audio generation, or related speech/audio domains.
- Familiarity with current speech and audio ML literature (e.g. text-to-speech, neural audio codecs, voice conversion).
Location
- London
- Stockholm
Work Type
- Hybrid
Experience Level
- Internship
Education Level
- PhD
- MSc
About the Company
- Spotify has more than 600M listeners in more than 180 markets around the world, who use our music, podcast, and audiobook services to find what delights, entertains, educates, and informs them.
- Personalization is a high impact organization that provides the technology to serve them what they expect to find, to help them explore and find new things to enjoy, and to suggest things they might not be aware of that they would like.
- The Speech team develops state-of-the-art technologies powering applications such as the AI DJ, generative advertisements, and other audio experiences across the platform.
- The team focuses on ensuring that the foundations of Spotify's speech technologies are at or above the state of the art.
- The team maintains strong ties both internally to product groups and externally to the research community, and actively encourages publishing at top venues.
Equal Opportunity
- Spotify is an equal opportunity employer. You are welcome at Spotify for who you are, no matter where you come from, what you look like, or what’s playing in your headphones. Our platform is for everyone, and so is our workplace. The more voices we have represented and amplified in our business, the more we will all thrive, contribute, and be forward-thinking! So bring us your personal experience, your perspectives, and your background. It’s in our differences that we will find the power to keep revolutionizing the way the world listens.
- At Spotify, we are passionate about inclusivity and making sure our entire recruitment process is accessible to everyone. We have ways to request reasonable accommodations during the interview process and help assist in what you need. If you need accommodations at any stage of the application or interview process, please let us know - we’re here to support you in any way we can.
