Modulate Raises $25M to Scale Audio-Native AI for the Next Generation of Voice
Modulate has raised $25 million in new funding to expand its audio-native AI technology, which helps companies understand what is happening inside voice conversations beyond the words being spoken.
The round was led by Future Ventures, with participation from Hyperplane and Lakestar, bringing Modulate's total funding to $60 million.
The Boston-based company is focused on building the intelligence layer around voice AI rather than the voice agents themselves. Its technology is being used across fraud prevention, deepfake detection, customer experience, voice agent supervision, trust and safety, and other emerging audio applications.
Building an Intelligence Layer for Voice AI
As AI systems become increasingly capable of speaking naturally, understanding a conversation requires more than simply turning speech into text.
Modulate's flagship Velma platform analyzes audio directly to identify signals including emotion, tone, intent, emphasis, synthetic speech and conversational behavior.
These signals can then be combined to identify higher-level events, from fraud attempts and AI agent failures to customer dissatisfaction, harassment and policy violations.
Velma can also operate in real time, allowing applications to understand and respond to what is happening during a conversation rather than only analyzing it afterwards.
AI That Understands More Than Words
Traditional speech AI has largely focused on transcription.
Modulate is taking a broader approach by analyzing the characteristics and context of the audio itself.
Its technology can help voice applications understand when a customer is frustrated, when an AI agent is struggling with an interaction, or when a conversation contains potentially harmful or suspicious behavior.
The company says its models now analyze more than 10 million hours of audio every month, with more than 600 million hours processed in total.
Detecting Deepfakes and Synthetic Voices
Voice cloning is creating new challenges for organizations dealing with fraud and security.
Modulate's technology can identify synthetic and manipulated speech, helping organizations detect suspicious voices in high-risk conversations.
The company says its deepfake detection technology achieves 98.9% accuracy on public benchmark data and currently ranks first on Hugging Face's deepfake speech benchmark.
Its models are already being used to help protect healthcare institutions from deepfake attacks and identify suspicious voice activity.
Supervising Voice AI Agents
Another growing application is monitoring how voice agents actually perform.
Rather than building the agents themselves, Modulate provides audio intelligence that can help companies observe their performance and understand the quality of their interactions.
That includes identifying when an agent misunderstands a customer, when a conversation is becoming frustrating, or when an interaction does not follow expected policies.
The same technology can also help voice agents better understand emotion and empathy, giving them additional context for how they should respond.
More Than 100 Specialized Audio Models
Underneath Velma is Modulate's Ensemble Listening Model, or ELM, architecture.
Rather than relying on a single large model, ELM combines more than 100 specialized audio models, selecting and combining them for different types of audio understanding.
Modulate says this approach can deliver significantly greater efficiency than using a single large model, with Velma demonstrating up to 1,000x greater efficiency in certain comparisons.
The company says this architecture allows it to analyze audio at scale while reducing the compute, memory and cost required.
What the New Funding Will Fund
The $25 million investment will support Modulate's expansion across:
AI and machine learning research
Product and engineering
Developer relations
APIs and SDKs
Partner integrations
New industry-specific models
Additional deployment environments
The company is also expanding its developer and partner ecosystem as more businesses look to add audio intelligence to voice applications without having to build specialized models themselves.
The Bigger Picture
Voice is becoming an increasingly important interface for AI.
As businesses deploy more voice agents across customer service, communications, security and other applications, the ability to understand what happens during those interactions becomes just as important as generating the conversation itself.
Modulate's approach reflects that shift.
Rather than competing to build another voice agent, the company is building the intelligence layer that can sit underneath different voice experiences, helping organizations understand emotion, intent, synthetic speech, agent performance and conversational behavior.
The $25 million raise gives Modulate additional resources to expand that technology and its developer ecosystem as voice AI moves further into real-world applications.
The bigger question for the market is no longer just whether AI can talk.
It is whether AI can actually understand what is happening when people talk to it.
