Pay by the hour

Modulate API Pricing

Industry-leading voice AI models at a fraction of the cost.
Self-serve, pay-as-you-go pricing available for every model
API

Pay for what you need,
when you need it

Self-serve, usage-based pricing per hour of audio processed.
Every model runs as a batch REST API; most also stream in real time over WebSocket
Transcribe
Batch
REST · file
Streaming
Real-time WS
English Fast Speech-to-Text
Leading english-only transcription
$0.025
$0.05
English-only, ultra fast. 70× real-time throughput. #1 ranked on Hugging Face
Multilingual Speech-to-Text
Leading multilingual transcription
$0.03
$0.06
70+ languages with diarization
Optional add-ons: Emotion ($.02), Accent ($.01) and PII/PHI tagging ($.02)
Multilingual Fast Speech-to-Text
Multilingual transcription tuned for speed
$0.03
Ultra fast, 70+ languages
Detect
Batch
REST · file
Streaming
Real-time WS
Deepfake Detection
Flags AI-generated & cloned synthetic voices
$0.25
$0.25
98.9% accuracy and a reliable verdict from 3 seconds of audio. #1 ranked on Hugging Face
Emotion Detection
Identify emotion tone & affect in voice
$0.02
20+ emotions, time-aligned per speaker
AI Music & Singing Detection
Identify AI-generated music & synthetic vocals
$0.07
$0.07
Independent vocal + instrumental scoring per 4-second window
Music vs Speech Detection
Separate speech from music
$0.01
$0.01
Speech vs. music, noise & silence. Frame-level (~192 ms)
Accent Identification
Identify speaker accent in voice
$0.01
Per-speaker accent labels
Language Detection
Identify which language is being spoken
$0.01
100 languages from the first 30 seconds of audio
Triage
Batch
REST · file
Streaming
Real-time WS
Velma - up to 10 behaviors
The most capable and complete ensemble of models
$0.75
$0.75
Complete conversation understanding with 150+ configurable behaviors, diarization, speaker roles, topics & sentiment, and a summary of each call
Velma - up to 25 behaviors
$1.25
$1.25
Redact
Batch
REST · file
Streaming
Real-time WS
PII / PHI Redaction
Remove sensitive data from audio and transcripts
$0.05
$0.08
Redacted transcript + clean audio across 100+ entity types
Prices are per hour of audio processed; 100 credits = $1.00 at pay-as-you-go. Batch runs as a REST API over uploaded files; streaming runs in real time over WebSocket. Every call has a minimum charge and a minimum billable clip length.
Ways to buy

Pay as you go, or commit to a credit bucket

Start with no commitment, or lock in an annual credit bucket for a lower effective rate.
Usage above your included block bills at your tier's rate with no surprise overage.
100 credits = $1.00 at pay-as-you-go
INCLUDED WITH EVERY PLAN:      Console Access         Auth/SAML SSO        API + Webhooks
How Modulate compares

Best-in-class accuracy,
a fraction of the cost

See how our pricing and accuracy stacks up against the field
How our Speech-to-Text compares
Feature
Modulate
Deepgram
AssemblyAI
ElevenLabs
Batch cost (multilingual)
$0.03/hr
$0.31/hr
$0.21/hr
$0.39/hr
Streaming cost (multilingual)
$0.06/hr
$0.35/hr
$0.45/hr
$0.39/hr
Full STT bundle
$0.09/hr
$0.41/hr
n/a
n/a
Speaker diarization
Included
$0.12/hr
$0.02/hr
n/a
Emotion detection
$0.02/hr
Pay for sentiment
$0.02/hr
n/a
PII / PHI tagging
$0.02/hr
$0.12/hr
$0.08/hr
Unavailable
Keyterm prompting
$0.02/hr
$0.08/hr
n/a
n/a
Credits included
Up to 400 hrs
Up to 360 hrs
Up to 300 hrs
2.5 hrs
How our Deepfake Detection compares
Feature
Modulate
Resemble AI
Reality Defender
Sensity AI
Pricing
$0.25/hr
$144/hr → $28/hr
Not listed
Not listed
Self-serve API
Yes
Yes
Yes
No — talk to sales
Accuracy
98.9%
97.4%
Unavailable
95–98%
Equal Error Rate
1.1%
2.57%
Unavailable
Unavailable
Audio required
3 seconds
Not listed
6 seconds
30 seconds
Free credits
Up to 40 hrs
None
50 scans/mo
None

Frequently
Asked Questions

Straight answers on how the API works and how to get started
What is a credit?
Credits are a simple unit that maps to audio processing usage across workflows and API capabilities. Different workflows consume credits at different rates.
Can I split spend between the API and the Platform?
Yes. Each tier includes a volume discount and a set number of credits. Those credits can be applied to your Included Calls or used directly within the API.
Do you offer volume discounts?
Yes. Larger platform tiers include bigger credit bundles, and API pricing supports volume-based discounts. Talk to sales to learn more.
Can we start small and scale?
Yes. Most customers start by paying as they go and expand as usage grows. Upgrading to a paid tier with a volume discount is quick and painless.
How do you integrate with our stack?
Use Modulate's enterprise platform and integrations, or access Velma programmatically through the API.
Is my data secure?
Modulate maintains ISO 27001 certification as part of its organization-wide security program, and will not sell or rent your data, ever.

Get pricing tailored to your conversation volume

Talk with our team to estimate usage, compare plans, and design something that fits your workflows and scale.