Voice artificial intelligence startup Modulate Inc. today announced it has raised $25 million in new funding to get its audio-native models in front of more developers.
Modulate offers AI models that work on the raw audio of a conversation. Emotion, tone and intent all register as signals, as do signs that a voice is a deepfake. The models can combine those readings to flag a fraud attempt or a caller losing patience with a voice agent, and some of that analysis runs live while the conversation is still happening.
The company got its start moderating voice chat in online games, and harassment and child grooming on social and gaming platforms are still among the things its models look for. Health care institutions use them to screen out callers impersonating staff with deepfake voices. A newer set of customers runs voice AI agents and uses the models to check how those agents are performing.
All of that runs through Velma, Modulate’s flagship platform. Underneath it, the Ensemble Listening Model architecture picks from more than 100 small specialized audio models for each job and blends their results. Modulate says that design is up to 1,000 times more efficient than handing the same audio to one large model.
More than 10 million hours of audio a month now run through the models, Modulate said, and the lifetime total recently passed 600 million hours. Two of its products have also taken first place on Hugging Face leaderboards this year. The transcription model topped the Open ASR Leaderboard in July, and Velma Deepfake Detect, which launched in March, sits at the top of the Speech Deepfake Arena. The company said the deepfake model is 98.9% accurate on public benchmark data.
“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript,” said co-founder and Chief Executive Carter Huffman.
Huffman said developers “shouldn’t have to rebuild the audio intelligence layer every time they create a new voice experience,” and part of the funding goes toward reaching them. New software development kits and application programming interfaces are in the works, as are models built for specific industries.
Hiring will pick up in research, engineering and developer relations, and the company is building out partner integrations and more ways for customers to deploy its models. Developers can already buy batch transcription through its existing API for three cents an hour.
Future Ventures led the round, with participation from returning investors Hyperplane and Lakestar. Hyperplane backed Modulate’s $2 million seed round, and Lakestar led its $30 million Series A in 2022. Steve Jurvetson, a co-founder of Future Ventures, said Modulate has “gained a significant technical lead in audio-native AI” and that the need for its technology is spreading well beyond where it started, into AI agents and security.
Huffman and fellow co-founder Mike Pappas are both Massachusetts Institute of Technology alumni and started the company in 2017. Counting the new round, backers have put $60 million into Modulate.
