We Started by Asking What Machines Could Hear. Now We're Building What Comes Next.

September 28, 2026
Carter Huffman
(HE/HIM/HIS)
Mike Pappas
(HE/HIM/HIS)

By Carter Huffman and Mike Pappas, Co-founders of Modulate

‍

Nearly ten years ago, we started Modulate around a question.

What if computers could actually understand voice?

Not simply recognize the words someone said. Understand the conversation itself: the tone, emotion, intent, emphasis, behavior and hundreds of other signals human beings process almost instinctively when we listen to one another.

At the time, that felt like a very different question than it does today.

Voice wasn't yet becoming the interface for AI. Deepfakes weren't showing up in fraud attempts. Companies weren't deploying fleets of AI agents that could hold natural conversations with customers. And most developers weren't thinking about how machines would understand the increasingly complicated world of human and synthetic voices interacting with one another.

We just thought audio was an incredibly rich source of information that software didn't understand very well.

So we started building.

Today, we're announcing $25 million in new funding led by Future Ventures, with participation from Hyperplane and Lakestar, bringing Modulate's total funding to $60 million.

We're incredibly excited about what this investment allows us to do.

But to explain why we're raising this capital now — and what we intend to do with it — it's worth explaining how we got here.

‍The problem kept getting bigger

Some of Modulate's earliest work was in online games and social platforms.

That turned out to be an extraordinary environment in which to learn about audio.

Real conversations are messy. People interrupt each other. They yell. They whisper. They joke. They switch languages. Their microphones are terrible. Music is playing in the background. Someone's dog is barking. Meaning depends on tone and context as much as it depends on the literal words being spoken.

And the problems our customers wanted to solve couldn't be answered simply by asking, “What words were said?”

They needed to understand what was actually happening. Was someone joking with a friend or threatening a stranger? Was a conversation escalating? Was someone attempting to manipulate or groom another person?

Those are fundamentally different problems from transcription. Solving them pushed us toward an architecture built specifically to understand audio itself.

Over time, that became Velma and the Ensemble Listening Model architecture underneath it: a system of more than 100 specialized models capable of identifying and combining different signals from audio to understand higher-level events and behaviors.

And then something happened that dramatically expanded the importance of that work. AI started talking.

The world taught machines to speak. Now they need to listen.

The progress in generative AI over the last several years has been remarkable.

AI systems can increasingly speak naturally, fluidly and convincingly. Voice agents are moving from demos into real customer interactions. Synthetic voices can sound almost indistinguishable from real people.

That creates enormous opportunity. It also creates an entirely new class of problems. If an AI agent is speaking with a customer, how does it know that person is becoming frustrated before they explicitly say so? How does a company know whether its voice agents are behaving correctly across millions of conversations? How does a financial institution know whether the executive supposedly authorizing a transaction is actually that person? How does a platform recognize manipulation, harassment or grooming happening through voice?

And as humans and AI increasingly talk to one another, how do applications understand everything happening in those interactions that isn't captured by the words alone?

A transcript is useful. But a transcript is not a conversation.

That distinction has become the foundation of what we're building at Modulate.

Nearly a decade of work — and we're still early

Our models now analyze more than 10 million hours of audio every month and they've processed more than 600 million hours in total. Along the way, we've learned an enormous amount about what it takes to build audio AI that works outside of controlled demos and research environments. We've also watched the number of problems this technology can solve expand dramatically.

Today, Modulate's technology is being applied across deepfake and fraud detection, Voice AI agent supervision, customer experience, trust and safety, transcription and other emerging voice applications. And increasingly, developers come to us with problems we hadn't imagined ourselves. 

That's one of the things that excites us most about this next chapter. We don't want every company building a voice product to also need to become an audio AI research lab.

We want to build the intelligence layer underneath those products.

What this funding changes

This investment gives us the resources to move faster across the areas we believe matter most.

We're expanding our AI and machine learning research. We're investing further in product and engineering. We're broadening our APIs, models, SDKs and deployment options. We're growing our developer ecosystem so more builders can experiment with audio-native intelligence without having to build these capabilities from scratch. And we're expanding our partnerships so Modulate's technology can become part of the platforms and products where voice interactions are already happening.

The goal isn't simply to build one great model or solve one category of problems. It's to make sophisticated audio understanding available anywhere developers are building with voice.

The part that doesn't fit in a press release

Funding announcements naturally focus on numbers. 

$25 million raised.

$60 million in total funding.

600+ million hours of audio processed.

100+ specialized models.

Those numbers matter. But they aren't really what today means to us. Building Modulate has taken nearly a decade. During that time, the company has evolved. The technology has evolved. The market has evolved. We've been right about some things, wrong about others, and surprised more times than we can count.

What has remained remarkably consistent is the underlying belief that brought us here in the first place: Voice contains an extraordinary amount of information and machines should be capable of understanding it.

We also wouldn't be here without the people who decided that problem was worth spending a meaningful part of their lives solving alongside us. Our team has tackled extraordinarily difficult technical challenges. Our customers trusted us with some of the hardest real-world audio environments imaginable. Our partners helped us see applications for this technology that we hadn't yet considered. And our investors have backed a vision that has always been much bigger than any single product or market. We're deeply grateful to all of them.

Nearly ten years ago, we started Modulate asking what machines could learn to hear. 

Today, that question feels more relevant than ever.

The next generation of AI won't just need to know how to talk.

It will need to know how to listen.

We're building that layer.

And we couldn't be more excited about what comes next.

‍

— Carter & Mike

‍