Modulate Raises $25 Million to Hear What Your Hold Music Is Hiding
Somerville's Modulate raised $25 million for voice AI, expanding from game moderation into deepfake detection, fraud signals and AI agent supervision.
“That’s fine” is a remarkably ambitious sentence to hand a computer. In Greater Boston, it can mean agreement, resignation, or that your proposed route through Sullivan Square has permanently altered the friendship. A transcript records two words. The actual conversation may require an appeals process.
That gap is the business opportunity behind Modulate’s September 28 announcement of $25 million in new funding, led by Future Ventures with Hyperplane and Lakestar participating. Co-founders Carter Huffman and Mike Pappas say the round brings total funding to $60 million. Their pitch is to make software better at interpreting conversational audio, including signals that disappear when speech becomes text.
The local connection is satisfyingly literal. Modulate’s published contact address is One Davis Square in Somerville. This is Massachusetts technology with an actual Massachusetts mailing address, rather than a company claiming regional citizenship because someone once changed planes at Logan.
My judgment: a serious, promising enterprise-AI bet, strengthened by an existing customer foothold. Listening deserves investment. Whether a machine’s interpretation deserves authority is the more interesting question.
The transcript has omitted the entire vibe
Modulate describes its approach as an Ensemble Listening Model. Think of specialized models examining different properties of audio rather than asking one general-purpose system to infer everything from a typed script. The funding announcement reports more than 100 specialized models and more than 600 million hours of audio processed. Those are company-reported scale figures, not independent measurements of accuracy.
Its Velma platform description lays out the practical workflow: ingest recordings or live streams, examine acoustic and conversational signals, and deliver alerts into dashboards or other software. Tone, overlap and stress can contribute information alongside the words. An API lets another application request that analysis; a webhook lets the analysis notify that application when something relevant happens.
Consider a hypothetical customer who calmly repeats that a promised refund never arrived. A cheerful service agent can produce an immaculate transcript of apologies while accomplishing absolutely nothing. Useful supervision would recognize the unresolved problem and prompt a handoff. Counting courteous phrases would merely award the building a certificate for excellent wallpaper.
This is a different infrastructure problem from Cambridge-based Subconscious’s effort to reduce agent inference costs. One asks how to make the machine’s work economical. The other asks what evidence the machine needs to understand the work. Both become more valuable when a demo graduates into a queue of actual customers.
Call of Duty supplied the unusually loud references
There is customer evidence here beyond logos decorating a pitch deck. In its November 2025 player-safety report, Activision explicitly identifies Modulate’s ToxMod as part of its voice-chat moderation system. Activision also describes a broader system involving community reports and its own enforcement technology.
That distinction matters. Modulate contributes detection; it does not single-handedly operate every part of Call of Duty’s moderation machinery. Nor does success in a gaming environment automatically validate a banking deployment. Different decisions carry different consequences.
Still, gaming is a credible place to encounter the audio problems that polished product videos prefer to leave outside: interruptions, overlapping speakers and unpredictable context. Building for conversations that decline to behave is useful experience. The enterprise market can provide nicer headsets, but it cannot promise nicer inputs.
The new financing is therefore an expansion story. It is not the first appearance of a product that suddenly learned to listen on September 28. The chronology deserves better than a venture-capital jump cut.
A synthetic voice is a clue, not a conviction
Modulate announced Velma Deepfake Detect on March 31. Its described capabilities include batch and streaming analysis, probability scores and segment-level inspection. The company says results can feed escalation, secondary verification or review workflows. Those are existing product claims, not features newly launched with this financing.
The useful design choice is what happens after a score appears. A synthetic voice can belong to a legitimate automated service. A human voice can belong to a scammer. Detecting the first property does not settle the second question.
Our look at IPID’s payee-verification business encountered the same limit from the payment side: checking one attribute does not establish that the whole transaction is trustworthy. Voice analysis can supply another piece of evidence. It should not become an electronic courtroom where the waveform is also the judge.
A sensible buyer would test recordings representative of its own callers, devices and connection quality, then measure missed attacks and false alarms separately. It should also ask how quickly a reviewer can understand and correct a mistake. An impressive benchmark is useful evidence; the customer-support queue is the operating environment.
Please do not give the dashboard a psychology license
The Velma Triage product page illustrates outputs such as behavior flags, speaker emotions, summaries and confidence scores. That makes the intended enterprise experience tangible. It also makes the interpretation problem unavoidable: an inferred emotion is an estimate, not direct access to somebody’s interior life.
Stress could indicate urgency, frustration, a poor connection or a perfectly understandable dislike of having to explain the same invoice again. A useful implementation would preserve the evidence and allow context to change the decision. A bad one would turn “the computer thinks you sound anxious” into a new obstacle between a person and a refund.
The operational question echoes our Onpipeline coverage: does automation reduce the work after verification and corrections are included? Here, that means fewer unresolved calls and actionable alerts, rather than a beautifully colored dashboard accumulating suspicions.
Davis Square has a worthwhile listening assignment
Readers outside Massachusetts should care because this is a portable problem. Any organization adding voice agents must decide how to notice failure while a conversation is still recoverable. That creates room for specialized monitoring alongside the systems doing the talking.
Modulate has a concrete funding event, an established customer reference and products addressing that need. The next persuasive evidence would be named enterprise deployments with measured outcomes, including the cost of false alarms and human review. Capital and processed hours alone cannot provide those answers.
For now, this looks like a worthwhile Boston-area technical bet: use experience from difficult conversations to make more conversations work. Teaching software to hear the hesitation before “that’s fine” would be useful. Teaching it when to ask a human would be even better.