Best Speech to Text Models in 2026: 10 Leading Models Compared

September 30, 2026

Your transcript can only be as good as the model powering it.

The best speech to text models in 2026 are judged above all on word error rate on real, messy audio, not clean studio recordings. The speech-to-text API market reached $4.66 billion in 2025 and is projected to hit $25.28 billion by 2034 (Fortune Business Insights), and the models below are what that spend is buying.

Speech to text now runs as backend infrastructure for everything from call center analytics to virtual AI voice agents. But not all speech to text models are built the same. Some are engineered for ultra-low latency to power live conversations. Others focus on real-world accuracy over clean studio recordings. Some go beyond plain transcription to add emotion detection, custom vocabulary or understanding of the full conversation.

This guide compares 10 leading speech to text models: Modulate Transcribe, ElevenLabs Scribe v2, Google Chirp 3, Mistral Voxtral, Qwen3-ASR, Microsoft MAI-Transcribe-1, AssemblyAI Universal-3, OpenAI Whisper Large v3, Deepgram Nova-3 and Speechmatics Ursa-2, scored on accuracy, latency, language coverage, diarization and price so you can match a model to your use case.

In this guide

Speech to Text Models Compared at a Glance

We compare models on average Word Error Rate (WER) across Earnings-22 and VoxPopuli, via the Modulate benchmark matrix. We use this benchmark because both datasets are independent, publicly available, and widely used elsewhere in the industry. Earnings-22 stress-tests noisy, accented real-world audio; the dataset contains 119 hours of English-language earnings calls from global companies. VoxPopuli tests multilingual consistency with a dataset containing 400,000 hours of unlabeled speech data in 23 languages from European Parliament event recordings. Neither is built or curated by a single vendor, so results are reproducible outside this matrix. Modulate, which publishes this matrix, also makes Modulate Transcribe. 

Word error rate comparison showing Modulate Transcribe at 8%, ElevenLabs Scribe v2 at 9%, AssemblyAI Universal-3 Pro at 11%, Speechmatics Enhanced at 12%, and Deepgram Nova-3, Google Chirp, and OpenAI Whisper Large v3 at 13%.

Average WER across Earnings-22 and VoxPopuli. Source: Modulate benchmark matrix.

Accuracy figures are drawn from the Modulate benchmark matrix, which reports average Word Error Rate across the Earnings-22 and VoxPopuli datasets. Vendor-only claims are labeled as such. Language counts, latency and pricing are vendor-reported; verify current pricing with each provider before purchase.

Speech-to-Text Model Comparison — Modulate
Modulate Speech-to-Text Model Comparison · 2026
Model ▲▼ Best for WER (real-world) ▲▼ Latency Languages ▲▼ Diarization Open source Pricing
WER shown is average Word Error Rate across the Earnings-22 and VoxPopuli benchmarks, lower is better, per the Modulate benchmark matrix. "Not tested here" means the model was not independently scored on those datasets. Language counts, latency and pricing are vendor-reported. Verify current pricing with each provider before purchase.

‍The Most Accurate Speech to Text Model

The most accurate speech to text model on real-world conversational audio is Modulate Transcribe, which records 7 to 8 percent word error rate across the IHM and Earnings-22 benchmarks, roughly half the 14 to 15 percent WER of other leading providers. IHM (Individual Headset Microphone) isolates each speaker's audio to measure raw transcription accuracy in multi-speaker settings, while Earnings-22 tests performance on noisy, accented, real-world financial calls. Together they cover two of the hardest conditions for production speech-to-text: distinguishing overlapping speakers and handling messy, non-studio audio.

On the AMI Meeting Corpus, a separate benchmark built from overlapping speakers and noisy multi-party calls, Modulate Transcribe reduces errors by over 40 percent versus ElevenLabs Scribe and over 70 percent versus OpenAI GPT-4o-transcribe. Accuracy rankings flip depending on the dataset, so always benchmark a speech to text model against audio that matches your real workload.

Modulate Transcribe

Modulate Transcribe

Overview

Modulate Transcribe, Modulate's speech to text API, is available in three modes: Batch, Streaming and Fast.

Modulate Transcribe Batch is designed for large-scale transcription of pre-recorded audio (call archives, meeting recordings, QA libraries) where throughput and maximum accuracy matter more than immediate results. Modulate Transcribe Streaming feeds the same underlying model over a real-time streaming endpoint with sub-second latency, returning partial transcripts while audio is still being processed. That makes Streaming ideal for live agent assist, voice AI pipelines and real-time transcript dashboards.

Modulate Transcribe Fast is a smaller, latency-optimized version of our core model, tuned for situations where low latency matters more than squeezing out every last bit of accuracy. Transcribe Fast suits high-volume automated workflows that need a transcript returned as quickly as possible. Batch and Streaming have identical accuracy scores because they are powered by the same model; only the method of delivery differs.

Key Features

High Accuracy on Messy Speech: As measured on the AMI Meeting Corpus (an industry-standard dataset defined by overlapping speakers and noisy, multi-party conversations) Modulate Transcribe reduces errors by over 40% versus ElevenLabs and by over 70% versus OpenAI's GPT-4o-transcribe, keeping transcripts accurate even when speakers interrupt each other or audio quality is poor.

Industry-Leading Accuracy Across Conversational Benchmarks: Beyond AMI, Modulate Transcribe averages 7-8% WER across datasets like AMI-IHM (headset-microphone recordings of real multi-speaker meetings) and Earnings-22 (noisy, accented real-world earnings calls). That is roughly 50% lower than other leading providers, which average 14-15% WER on these datasets.

Conversation-First Transcription: Designed for conversation, Transcribe performs strongly with overlaps, interruptions and colloquial speech rather than just clean, read-from-a-script studio audio.

Low-Latency Streaming: Real-time streaming transcription built to power live UIs, agent pipelines and systems that need results immediately.

Free Speaker Diarization: Includes real-time and batch speaker diarization at no additional cost, something many competitors charge extra for.

PII/PHI Redaction: All models detect 94 types of PII and PHI out of the box (contact info, IDs, financial, health, employment, digital and security-related info). You can redact that sensitive audio and transcript in real time or tag it for reference without full redaction.

Emotion and Accent Detection: Modulate Transcribe detects over 20 emotions and 20 accents on top of the standard transcript, surfacing insights a basic transcription API would not.

Broad Language Support: Transcribe supports 57 languages, including major global languages and less widely supported ones like Welsh, Maori and Galician, as well as dialects.

Trained on Real Conversations: Unlike most competitors, Modulate trained Transcribe on over 500 million hours of real-world conversational audio rather than heavily curated or synthetic datasets. That is the primary basis for Modulate Transcribe's accuracy on real-world, messy conversations.

Pros

  • Reduces errors by over 40% versus ElevenLabs and over 70% versus OpenAI's GPT-4o-transcribe on the AMI Meeting Corpus, the industry-standard conversational-speech benchmark.
  • Averages roughly 50% of the WER of Deepgram, Google and NVIDIA on other conversational benchmarks.
  • Rated Robust on noise and accent/multilingual robustness criteria where several competitors see rising WER.
  • Both real-time and batch speaker diarization included free, not as a paid add-on.
  • Integrated PII/PHI redaction and tagging across 94 unique data types.
  • 10x cost savings versus other market leaders while delivering superior accuracy for real-world conversations.
  • Three deployment modes (Batch, Streaming, Fast) let teams match the model to their latency and cost needs.

Cons

  • Transcribe's tooling is newer than some long-standing providers, so you may find fewer community examples or SDKs initially.
  • Fewer third-party integrations than more established providers, so some workflows may require custom development. 

Get Started with Free Credits

Scribe v2 (ElevenLabs)

Scribe v2 (ElevenLabs)

Overview

ElevenLabs' core speech-to-text model is Scribe v2, designed to transcribe pre-recorded audio and video with high accuracy, even with different accents and poor recording conditions.

Scribe v2 Realtime, built on ElevenLabs' streaming-first tech, can transcribe speech in under 150 milliseconds, aimed at voice agents and other applications that need real-time transcription rather than post-production.

Scribe v2 and Scribe v2 Realtime are accessible through ElevenLabs' API. Scribe v2 is also integrated into ElevenLabs Studio for captions, subtitles and editable transcripts for podcasts, videos and interviews.

Key Features

High-Accuracy Batch Transcription: Clean, editable text from noisy or accented audio. Whether you are transcribing podcasts, interviews or meetings, Scribe v2 produces accurate, searchable text.

Sub-150ms Real-Time Transcription: Scribe v2 Realtime delivers true sub-150ms live transcription, ideal for agents and live meetings that need to understand humans instantly.

Voice Activity Detection: Automatically detects when someone starts and stops speaking, segmenting speech accurately for live use.

Keyterm Prompting: Highlight up to 1,000 words or phrases for Scribe to recognize based on context, useful for disambiguating names, jargon or brand words.

Dynamic Audio Tagging: Automatic tagging of non-verbal sound events like laughter, applause and footsteps.

Speaker and Entity Detection: Detects and tags every speaker, provides start/end offsets for entities and can redact PII.

Multi-Language Support: Supports over 90 languages, with accuracy varying by language.

Strong Multilingual Benchmark Performance: Scribe v2 performs well on multilingual, formal-recording benchmarks such as VoxPopuli (European parliamentary speech recordings), scoring 6.10% WER. 

Enterprise Security and Deployment Options: SOC 2, HIPAA and GDPR compliant, with EU Data Residency and Zero Retention modes. Self-hosted and VPC deployment available for enterprise.

Pros

  • Sub-150ms real-time latency, among the fastest here.
  • Support for over 90 languages and variants.
  • Keyterm prompting enables custom vocabularies for names, jargon and brand terminology.
  • Flexible deployment across SaaS, self-hosted and VPC.
  • Built-in speaker diarization, entity detection and sensitive-information redaction.
  • Performs well on benchmark metrics (6.10% WER on VoxPopuli, a multilingual dataset of formal parliamentary speech recordings).

Cons

  • WER and diarization accuracy affected by noisy audio.
  • Higher WER on accented and multi-language conversations than the top performers here.
  • Struggles with domain-specific jargon and proper nouns in practice.
  • Does not support keyword weighting (many competitors do, alongside custom vocab building).
  • Accuracy differs by language, with some languages in the Moderate (20%+ WER) tier.
  • Starting price is $0.40/hour, higher than several lower-cost alternatives here.

Chirp 3 (Google)

Chirp 3 (Google)

Overview

Chirp 3 is Google's newest generation of multilingual Automatic Speech Recognition (ASR) models, trained with accuracy and speed improvements over previous Chirp models plus diarization and language detection.

Chirp 3 is accessed only through the Speech-to-Text API V2. It is meant for teams already on Google Cloud, and supports synchronous recognition of short clips, batch recognition for larger files and streaming recognition for real-time audio.

Key Features

Automatic Language Detection: Chirp 3 detects and transcribes the predominant language in a file. You can also restrict it to a set of expected locales to improve accuracy.

Speaker Diarization: Chirp 3 determines speaker changes within a single channel. This is available only via batch recognition, not live streaming.

Adjustable Endpointing Sensitivity: Configure how soon the model produces a final result. Streaming clients can switch between accuracy-first behavior and low-latency modes for brief commands or single-word responses.

Noise Reduction: Optionally enable a denoiser to remove background sounds like music, rain or traffic. Background voices cannot be removed.

Speech Adaptation (Biasing): Provide up to 1,000 phrases or words to bias the model toward domain-specific terminology or proper nouns.

Custom Prompting: As a preview feature, Chirp 3 accepts custom formatting directives, such as how dates or acronyms should appear in the output.

Extensive Language Support: Transcribe speech in 125 languages and regions, including generally available and preview languages.

Strong Multilingual Benchmark Performance: Chirp 3 performs well on multilingual, formal-recording benchmarks such as VoxPopuli (European parliamentary speech recordings), scoring 6.40% WER. 

Pros

  • Broadest language coverage here at 125 supported languages.
  • Strong accuracy on the VoxPopuli benchmark (6.40% WER), a multilingual dataset of formal parliamentary speech recordings.
  • Optional built-in denoiser for noisy environments.
  • Speech adaptation and custom vocabulary for domain-specific terms.
  • Flexible endpointing to tune latency versus accuracy in streaming.
  • Integrated into Google Cloud’s infrastructure, security, and support ecosystem, a strong fit if you’re already on GCP, though it does mean deeper reliance on Google’s platform. 

Cons

  • WER issues reported in noisy conditions and with accented or multilingual speech.
  • Diarization available only with batch recognition, not live streaming.
  • No word-level timestamps, only utterance-level, and only with streaming recognition.
  • Automatic punctuation and formatting is more limited than competitors offering full formatting.
  • SaaS only, with no self-hosted or in-VPC deployment. Audio must be sent to and from Google Cloud, adding data ingress/egress costs and latency on top of the data-residency limitations for teams with stricter compliance needs.
  • Higher out-of-the-box WER (11.30%) on the Earnings-22 benchmark (noisy, accented real-world earnings calls) than the best performers here, which range from roughly 7-8% to 12% on this benchmark.

Voxtral Small (Mistral)

Voxtral Small (Mistral)

Overview

Voxtral is Mistral AI's family of open speech understanding models, in two sizes: a 24B variant for production-scale deployments and a smaller 3B variant for on-device and edge use. Both are Apache 2.0 licensed and available through Mistral's API, which runs requests on a transcribe-optimized variant of Voxtral Mini tuned for cost and latency.

Compared to Voxtral Mini, Voxtral Small adds capabilities that make it a speech understanding model rather than strictly a transcription model. It can answer questions about audio, summarize audio and take backend actions from textual intent, combining ASR and language understanding into a single integrated model rather than requiring developers to chain two separate ones. 

Key Features

Integrated LLM Framework: Beyond transcription, Voxtral pairs its ASR output with a built-in language model that can answer questions about audio or produce structured summaries, without requiring developers to chain separate ASR and LLM models themselves. 

Long-Form Context: A 32k token context supports up to 30 minutes of audio for transcription or 40 minutes for audio understanding in a single pass.

Function-Calling from Voice: Call backend functions, workflows and APIs directly from voice input without a separate parsing layer.

Multi-Lingual Support: Automatic language detection with robust support for popular languages including English, Spanish, French, Portuguese, Hindi, German, Dutch and Italian.

Retained Text Capabilities: Based on Mistral Small 3.1, Voxtral inherits Mistral's text understanding, so it can serve as both a speech model and a text model.

Flexible Deployment Options: Run Voxtral on-device, via Mistral's API or privately behind the firewall for regulated or internal data.

Pros

  • Apache 2.0 licensed open-weight model teams can self-host for full control over customization and deployment.
  • Native Q&A, summarization and function calling alongside transcription, removing a separate proprietary LLM step.
  • Competitive VoxPopuli accuracy (6.40% WER) among the open-source and API models here. VoxPopuli is a multilingual dataset of formal parliamentary speech recordings.
  • Supports keyword weighting and custom vocabulary, where some competitors do not.
  • Deployable via SaaS API or edge, beyond SaaS-only competitors.
  • Token-based pricing that Mistral claims is less than half the cost of comparable proprietary APIs.

Cons

  • Rated only reasonably robust on noise and accent, trailing the top performers here.
  • Supports 13 languages, while several competitors support 50+.
  • Speaker diarization is batch only, not real-time.
  • Does not yet support word-level timestamps. 
  • Higher (worse) WER on noisier, real-world benchmarks such as Earnings-22 (11.30% WER) than top performers, which range from roughly 7-8% to 12% on this benchmark. Earnings-22 is a dataset of noisy, accented real-world earnings calls.
  • Token-based pricing (input per minute, output per token) can be less predictable at scale than flat per-minute tiers.

Qwen3-ASR (1.7B and 0.6B)

Qwen3-ASR (1.7B and 0.6B)

Overview

Qwen3-ASR is a family of open-source speech recognition models from Alibaba's Qwen team, based on the Qwen3-Omni architecture and trained on large speech datasets. It ships in two sizes: Qwen3-ASR-1.7B, the flagship, and Qwen3-ASR-0.6B, which is significantly smaller and faster with a modest accuracy loss. Both handle language identification and transcription across 52 languages and dialects.

The models run offline ASR and streaming ASR from a single unified model rather than separate models per task. The release also includes Qwen3-ForcedAligner-0.6B, a companion model purpose-built for accurate word- and character-level timestamps.

Key Features

All-in-One Language Coverage: Native language ID and speech recognition for 30 languages and 22 Chinese dialects, plus English accents from many regions. There is no separate language-detection step.

Unified Streaming and Offline Inference: Each model supports streaming and offline inference from one model and can transcribe long-form audio, removing the need for separate streaming and batch variants.

High-Throughput Small Model: Qwen3-0.6B is optimized for the accuracy-throughput trade-off, reaching roughly 2,000x higher throughput at a concurrency of 128, ideal for large-scale, cost-sensitive scenarios.

Companion Forced Aligner: Qwen3-ForcedAligner-0.6B predicts timestamps for units up to five minutes long across 11 languages, with higher timestamp accuracy than end-to-end forced-alignment models in Alibaba's tests.

Music and Song Transcription: Trained on transcribing songs and singing voice with music, a use case most transcription models are not trained for out of the box.

Robust Acoustic Performance: Built to maintain quality in noisy, acoustically challenging environments and to recognize heavy dialect.

Open-Weight, Self-Hostable: Apache 2.0 licensed, downloadable and self-hostable via Hugging Face or ModelScope, with a full inference toolkit supporting vLLM batch inference and async serving.

Pros

  • Open source under Apache 2.0, free to use and self-host with no per-minute or per-token fees.
  • Robust noise and accent performance.
  • Accepts keyword weighting and custom vocab, where many closed-source competitors do not.
  • 1.7B model offers accuracy on par with top commercial APIs while staying open-weight.
  • 0.6B model provides a genuinely low-latency (~100ms) option for high-throughput scenarios.
  • Singing and music transcription capabilities.

Cons

  • Self-host only; no managed SaaS. You own infrastructure, configuration and scaling.
  • No built-in speaker diarization, real-time or batch.
  • Requires significant GPU infrastructure and ML expertise to deploy and maintain, unlike plug-and-play API competitors (the 0.6B model needs at least an RTX 4090 (24GB) to run slowly, while the 1.7B model requires 4x H100s at minimum).
  • Not independently benchmarked on AMI-IHM, Earnings-22 or VoxPopul, the real-world conversational and multilingual benchmarks used to compare accuracy across models in this article.
  • As a newer community project, model support and SLAs are not guaranteed by a commercial vendor.

MAI-Transcribe-1 (Microsoft)

MAI-Transcribe-1 (Microsoft)

Overview

MAI-Transcribe-1 is Microsoft's multilingual speech-to-text model, built to work consistently across languages, accents and real-world production environments such as conference rooms, phone lines and crowded streets.

Microsoft claims it is the most accurate transcription model for each of its 25 languages when benchmarked against competitors like Scribe v2, Whisper Large v3, GPT-Transcribe and Gemini 3.1 Flash-Lite on FLEURS, a multilingual speech dataset covering over 100 languages commonly used to benchmark ASR accuracy. It is available via Microsoft Foundry and is being deployed inside Microsoft products like Copilot voice mode and Microsoft Teams.

Key Features

Top FLEURS Accuracy: MAI-Transcribe-1 has the lowest WER of any speech-to-text model on the FLEURS benchmark across all 25 languages it supports. FLEURS spans over 100 languages and is a standard benchmark for evaluating multilingual speech recognition.

Robust Multilingual Accuracy: Trained for competitive accuracy across 25 languages for global products, holding up against accents and unusual speaking patterns.

Fast Batch Transcription: According to Microsoft, batch transcription is 2.5x faster than its prior Azure Fast model, built for high-volume production.

Tailored for Real-World Audio: Trained on noisy, low-quality audio with background noise rather than clean studio audio.

Voice Agent Stack Integration: Microsoft sells MAI-Transcribe-1 as the transcription layer in a voice agent stack, used alongside MAI-Voice-1 (text-to-speech) and an LLM of your choice.

Broad Use Cases: Supports offline uses (subtitle generation, podcast transcription, compliance recording, legal discovery, call center QA) and online uses (meeting transcription, live captioning, dictation) with a single model.

Competitive Cost: Priced at $0.36 per hour of audio, which Microsoft positions as the best price-to-performance ratio among large cloud providers.

Pros

  • Lowest WER of any speech to text model on the FLEURS benchmark (a standard benchmark for evaluating multilingual speech recognition) across 25 languages.
  • Supports keyword weighting and custom vocabulary.
  • Built-in real-time and batch speaker diarization.
  • Hourly pricing ($0.36/hr) is on par with other large cloud providers here.
  • Backed by Microsoft's platform and integration with Foundry and Microsoft 365.
  • Designed for noisy, real-world audio rather than clean benchmark audio.

Cons

  • Real-world results from independent testing indicate higher WER on noisy audio and accented/multilingual speech.
  • Slowest latency here at ~700ms, which may limit latency-sensitive real-time use.
  • Narrower language coverage (25 languages) than competitors supporting 50+ or 90+.
  • SaaS only; no self-hosting or VPC option for strict data-residency needs.
  • No independent benchmarking on AMI-IHM, Earnings-22 or VoxPopuli (the real-world conversational and multilingual benchmarks used to compare accuracy across models in this article), so its accuracy claims rest mainly on Microsoft's own FLEURS results.
  • Microsoft positions MAI-Transcribe-1 as one tier in a fast-moving lineup, so confirm it is the current recommended transcription model in Foundry before you build on it.

Universal-3 (AssemblyAI)

Universal-3 (AssemblyAI)

Overview

Universal-3 Pro is AssemblyAI's most advanced speech-to-text model, which they call the world's first promptable Speech Language Model. Instead of transcribing audio then cleaning it up downstream, Universal-3 Pro lets developers describe the context up front in natural language: the domain, specific terminology, speaker roles and formatting preferences, before transcription, which can improve accuracy from the start.

Use cases include medical transcription, legal transcription, contact center and conversation intelligence. Universal-3 Pro is part of AssemblyAI's Voice AI Platform, with a cloud-based API and on-premise options. Batch transcription works in 99 languages; real-time streaming supports 6: English, Spanish, French, German, Italian and Portuguese (beta).

Key Features

Natural-Language Prompting: Provide context about the audio (for example, that it is a medical appointment) to bias transcription toward that use case before it runs, rather than fixing it later.

Custom Vocabulary via Keyterms: Provide up to 1,000 domain-specific words as prompts, which AssemblyAI reports can substantially reduce errors on specialized vocabulary.

Audio Event Tagging: Tag non-speech events such as beeps or hold music with a prompt, retaining acoustic context a bare transcript loses.

Speaker Role Labeling: Label speakers by role (Nurse, Patient) instead of Speaker A, Speaker B when prompted.

Verbatim and Clean Dictation Modes: Switch modes in the same model. Verbatim keeps fillers and false starts for legally defensible transcripts; Clean filters conversational noise for readable summaries.

Single File Multilingual Transcription Support: Maintains accurate transcription across multiple  languages within one file, for example English to Spanish mid-conversation.

Strong Multilingual Benchmark Performance: Universal-3 Pro performs well on VoxPopuli, a benchmark of formal, multilingual European parliamentary speech recordings, scoring 6.80% WER.

Flexible Deployment Options: Cloud API plus a self-hosted option for teams with stricter data-residency requirements or existing infrastructure.

Pros

  • Strong VoxPopuli (formal, multilingual parliamentary speech) performance (6.80% WER).
  • Highest word accuracy (94.07%) and lowest missed-entity rates across categories on AssemblyAI's own tests against ElevenLabs, OpenAI, Speechmatics, Microsoft, Amazon and Deepgram.
  • Prompt-based customization without training separate models per domain.
  • Support for weighted keywords and custom vocabulary.
  • SaaS, self-hosted and VPC deployment.
  • Integrated real-time and batch speaker diarization.

Cons

  • Higher (worse) WER on both noisy audio and accented or multilingual speech.
  • Streaming supports only 6 languages (batch supports 99).
  • Higher (worse) WER (11.90%) than top performers (which range from roughly 7-8% to 12%) on Earnings-22, a noisier, real-world benchmark of earnings calls.
  • The promptable framework may need extra learning compared to plug-and-play APIs.
  • AssemblyAI's benchmark data is self-reported, not third-party tested.
  • Per-hour pricing is harder to compare directly against APIs priced per minute.
  • Billing can get complicated and add up quickly: multichannel audio bills each channel as a separate stream (a 1-hour, 3-channel call bills as 3 hours), and add-on features are billed separately, either per hour or per token (see current pricing).

Whisper Large v3 (OpenAI)

Whisper Large v3 (OpenAI)

Overview

Whisper large-v3 is OpenAI's flagship open-source ASR and speech-translation model, part of the Whisper family first released in 2022. It is a general-purpose model trained on diverse audio, capable of multilingual speech recognition, speech translation and language identification in one system.

large-v3 differs slightly from large and large-v2: it uses a spectrogram with 128 Mel frequency bins (versus 80 previously), which translates to better differentiating timbre, helpful when interpreting non-vocal audio and tonal languages. That said, the typical sweet spot for optimized vocal speech recognition is considered to be 40-80 Mel bins, so the increase to 128 is a trade-off: it broadens the model's ability to differentiate a wider array of audio types at some cost to vocal-speech optimization. It also adds a language token for Cantonese. Whisper's code and weights are open source, so it can be self-hosted or run through community tooling like whisper.cpp. It's also available via OpenAI's hosted API for teams who'd rather not manage infrastructure.

Key Features

Open Model Weights: Both the codebase and weights are public, so developers can run the model locally, fine-tune it or simply use OpenAI's API.

Broad Language Coverage: Supports 99 languages, among the widest coverage here.

Multitask Architecture: Multilingual recognition, translation and language identification are handled by a single Transformer sequence-to-sequence model instead of a multi-stage pipeline.

Reduced Error Rates Over Predecessor: According to OpenAI, large-v3 reduces the error rate across a broad range of languages compared with Whisper large-v2.

Massive Training Scale: Trained on a very large mixture of weakly labeled and pseudo-labeled audio, giving strong generalization across domains without fine-tuning for most uses.

Community Tooling: As open source, Whisper has many third-party tools, including optimized ports like whisper.cpp that run efficiently on CPUs and Apple Silicon.

Long-Form Transcription Support: Transcribes audio longer than its 30-second window using sequential or chunked algorithms that trade speed for accuracy.

Pros

  • Open weights let you self-host, fine-tune or run offline (most direct competitors here are closed source).
  • Supports 99 languages, the most of any model here.
  • No licensing fees for self-hosted use, since the model is free and openly available.
  • Large open-source community with many ports, tools and documentation.

Cons

  • Highest WER of the models compared in this article on AMI-IHM (headset-mic recordings of real multi-speaker meetings) at 15.95%, indicating weaker results on meeting-style audio.
  • Reportedly higher (worse) WER in noisy audio and accented/multilingual speech.
  • No keyword weighting or custom vocabulary, limiting domain-specific tuning.
  • Lacks text-normalization support, unlike nearly every other model here.
  • No native speaker diarization; requires extra tools like pyannote.audio.
  • Not optimized for real-time streaming; needs additional tooling for near-real-time use.
  • OpenAI's hosted API uses per-token pricing, which is harder to estimate than flat per-minute or per-hour billing.

Nova-3 (Deepgram)

Nova-3 (Deepgram)

Overview

Nova-3 is Deepgram's core speech-to-text model, built to lead on accuracy in harsh real-world use cases such as drive-thrus, call centers and air traffic control. Deepgram advertises Nova-3 as one of the first voice AI models to support real-time multilingual transcription (accurately transcribing a speaker who switches languages mid-sentence rather than relying on multiple language-specific models). This is the same underlying capability AssemblyAI supports with single-file multilingual transcription, applied here in real time.

Nova-3 is also Deepgram's first model with self-serve customization via keyterm prompting, which adapts the model to specific vocabulary at inference time without retraining. It is available through Deepgram's cloud API and can be self-hosted.

Key Features

Real-Time Multilingual Transcription: Nova-3 follows a conversation as the speaker switches languages, without separate language-specific components, handling real-time code-switching in 10 languages.

Keyterm Prompting: Tune up to 100 key terms critical to your use case at inference time, without retraining.

Robustness in Challenging Acoustic Conditions: Accuracy is minimally affected by speaker-to-microphone distance, cross-talk or background noise, even in drive-thrus and call centers.

Real-Time Entity Redaction: Redact up to 50 types of PII from live conversations, aimed at compliance-driven use in finance and customer support.

Enhanced Numeric and Formatting Accuracy: Improved transcription of numeric sequences and entities, plus better punctuation and paragraphing over Nova-2. Deepgram does not supply a benchmark or percentage for this feature.

Fast Inference: Deepgram reports one of the lowest median inference times in the market.

Flexible Deployment: Cloud API plus self-hosted and VPC options for stricter infrastructure requirements.

Pros

  • Deepgram's own benchmarking reports a median streaming WER of 6.84% and batch WER of 5.26%, based on Deepgram's proprietary internal test sets. The methodology isn't publicly reproducible, so these figures can't be independently verified.
  • Real-time multilingual transcription with code-switching.
  • Keyterm customization without retraining.
  • Real-time redaction of sensitive information for regulated industries.
  • Flexible SaaS, self-hosted and VPC deployment.

Cons

  • Increased WER in noisy audio and accented/multilingual speech.
  • Higher (worse) WER on noisier, real-world benchmarks such as Earnings-22 (15.70% WER) than most models here.Top performers on this dataset range from roughly 7-8% to 12% on this benchmark.
  • Headline accuracy numbers come from Deepgram's internal benchmarks, not a third-party source.
  • Supports fewer languages (52) than some competitors here.
  • Credit-based or pay-as-you-go pricing can be harder to predict than per-minute or per-hour pricing.

Ursa-2 (Speechmatics)

Ursa-2 (Speechmatics)

Overview

Ursa-2 is Speechmatics' latest speech recognition model. Speechmatics uses its own self-supervised learning (SSL) to pre-train on large amounts of unlabeled audio before targeted, language-specific training. The company advertises an 18% Word Error Rate improvement across 50+ languages over the previous Ursa generation.

Speechmatics also expanded coverage to include Irish and Maltese, covering all 24 major European languages. It positions Ursa-2's strength around consistency: strong performance across a very wide range of languages rather than the absolute lowest WER on any single benchmark.

Key Features

Self-Supervised Pre-Training: Ursa-2 is pre-trained extensively on unlabeled audio before language-specific supervised training, which Speechmatics says helps it perform well even for languages with limited labeled data.

Expanded Model Scale: Speechmatics increased pre-training data and overall model size, and credits a move to GPU inference for accuracy gains.

Sub-Second Real-Time Transcription: Ursa-2 lowers real-time latency to under one second without sacrificing accuracy.

Enhanced Arabic Language Pack: A major upgrade to the global Arabic pack, which Speechmatics claims surpasses competitor accuracy for Egyptian, Gulf, Levant and Modern Standard Arabic.

Broad, Consistent Language Coverage: Ursa-2 supports 55 languages, and Speechmatics claims top accuracy in the majority of them.

Enterprise Compliance Certifications: ISO 27001, SOC 2, GDPR and HIPAA compliance for regulated markets like healthcare and legal.

Flexible Deployment: SaaS or on-premises deployment for greater data-residency options than SaaS-only competitors.

Pros

  • Keyword weighting and custom vocabulary for specialized terminology.
  • Supports 55 languages, among the higher counts here.
  • Enhanced Arabic language pack covering Egyptian, Gulf, Levant and Modern Standard Arabic, a dialect depth many competitors lack.
  • Real-time and batch speaker diarization.
  • Speechmatics reports top or near-top accuracy on its own FLEURS benchmark for most supported languages.
  • SaaS and on-premises deployment for data-residency needs.
  • Some built-in handling for noisy audio.

Cons

  • Not independently benchmarked on the noisy, real-world conversational and multilingual datasets used to compare other models here (AMI-IHM, Earnings-22 and VoxPopuli), making direct accuracy comparison difficult.
  • Rated as having increased WER on accented and multilingual speech.
  • Tied for the highest latency here at roughly 700ms, which may limit real-time or latency-sensitive use.
  • Published leaderboard and accuracy figures come from internal FLEURS benchmarking (a widely used, publicly available multilingual speech dataset); Speechmatics notes the score may be inflated because the public dataset can appear in training data.
  • Charges by the hour, making direct comparison to per-minute providers harder.

Best Speech to Text Models Comparison Chart

Metric Velma Transcribe Batch Velma Transcribe Streaming Velma Transcribe Fast Scribev2 Chirp-3 Voxtral Small Qwen3-ASR (1.7B & 0.6B) MAI-Transcribe-1 Universal-3 Whisper Large v3 Nova-3 Ursa-2
Provider Modulate Modulate Modulate Eleven Labs Google Mistral Alibaba Microsoft AssemblyAI OpenAI Deepgram Speechmatics
Accuracy (Word Error Rate) - AMI-IHM 7.92% 7.92% N/A N/A N/A N/A N/A N/A N/A 15.95% N/A N/A
Accuracy (Word Error Rate) - Earnings22 7.55% 7.55% N/A 9.70% 11.30% 11.30% N/A N/A 11.90% 11.29% 15.70% N/A
Accuracy (Word Error Rate) - VoxPopuli 8.04% 8.04% N/A 6.10% 6.40% 6.40% N/A N/A 6.80% 9.54% 8.20% N/A
Supported Languages 70+ 70+ 70+ 90+ Languages 125 Languages 13 Languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, and Dutch 21 Languages 25 Languages: English, French, German, Italian, Spanish, Hindi, Portuguese, Czech, Danish, Finnish, Hungarian, Dutch, Polish, Romanian, Swedish, Japanese, Korean, Chinese, Arabic, Indonesian, Russian, Thai, Turkish, and Vietnamese. 6 Languages 99 Languages 52 Languages 55 Languages
Latency Sub-second Sub-second Sub-second ~150 ms 200 ms ~240 ms ~100 ms 700 ms 300 ms N/A 200-300 ms 700 ms
Noise Performance Robust Robust Robust Increased WER, Diarization Issues Some WER Issues Reasonably Robust Robust Increase WER Increase WER WER Increase WER Increase Some Options to mitigate
Accent/Multilingual Performance Robust Robust Robust Increased WER Known issues Reasonably Robust Robust Increase WER Increase WER WER Increase WER Increase WER Increase
Word or Segment Level Timestamps Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes
Real-Time Speaker Diarization Yes Yes Yes Yes Yes No No Yes Yes Yes Yes Yes
Batch Speaker Diarization Yes Yes Yes Yes Yes Yes No Yes Yes Yes Yes Yes
Automatic Punctuation and Formatting Full Formatting Full Formatting Full Formatting Full Formatting Limited Full Formatting Full Formatting Full Formatting Full Formatting Full Formatting Full Formatting Full Formating
Keyword Weighting No No No No Yes Yes Yes Yes Yes No Yes Yes
Custom Vocabulary and Language Controls (Advanced Metadata Specification) No No No Yes Yes Yes Yes Yes Yes No Yes Yes
Text Normalization Yes Yes Yes Yes Yes Yes Yes Yes Yes No Yes Yes
Deployment SaaS SaaS SaaS SaaS, Self-host, VPC SaaS SaaS, Edge Deploy Self-host SaaS SaaS, Self-host, VPC SaaS SaaS, Self-hosted, VPC Saas, On-prem
Pricing Model Price Per Usage (Minute) Price Per Usage (Minute) Price Per Usage (Minute) Price Per Usage (Minute) Price Per Usage (Minute) Token Base (Input per min, output per token) Open Source Price Per Usage (Hour) Price Per Usage (Hour) Per Token Pricing Credit based Plans or Pay as you Go Per minute Price Per Usage (Hour)
Best For High-accuracy offline processing of recorded audio at scale (call archives, compliance recordings, QA libraries) Real-time voice agents and live dashboards needing top-tier accuracy without sacrificing speed High-volume, latency-sensitive workflows where turnaround speed matters more than peak accuracy Podcast, video, and content-creation teams needing captions/subtitles plus keyterm customization Teams already on Google Cloud needing the broadest language coverage (125 languages) Teams wanting an open-weight model that goes beyond transcription (Q&A, summarization, function-calling) Cost-sensitive, self-hosted deployments needing free, high-throughput multilingual ASR Teams embedded in the Microsoft ecosystem (Teams, Copilot, Azure/Foundry) Domain-specific use cases (medical, legal, contact centers) needing prompt-based customization without model retraining Budget-conscious or research projects needing a free, self-hostable model with wide language support Real-time multilingual voice agents and contact centers needing low latency and self-serve vocabulary tuning Enterprises needing consistently high accuracy across a very broad, non-English-heavy language set

What Is a Speech to Text Model?

A speech-to-text model, sometimes called an Automatic Speech Recognition (ASR) model, is an artificial intelligence model trained to transcribe audio into written text. It can take raw audio such as a phone conversation, a recording of a meeting or a live voice command and produce a transcript. Ideally, speech to text transcripts include not just the words but punctuation, speaker labels, pauses and formatting.

Modern speech to text models generally use deep learning, often Transformer-based sequence-to-sequence models. Instead of matching sounds to a predefined dictionary like older systems, modern models are trained on large datasets of audio paired with human transcriptions. From that data they learn the patterns in audio that occur across different accents, dialects, background noises and speaking styles.

Some of the most recent models even allow general speech understanding, where the model can summarize, answer questions about or take actions on audio rather than simply transcribing it.

What's the Difference Between a Speech to Text Model and a Speech to Text API?

A speech-to-text model and a speech-to-text API are two layers of the same service: the model is the brain and the API is the delivery mechanism that makes that brain usable.

Speech to Text Models

A speech-to-text model is the actual AI system that does the transcription work. Think of it as the trained neural network itself: the weights, architecture and learned patterns that connect input audio to text output. Each model has different accuracy and language support based on how it was trained. The model is literally a file (or set of files) of parameters; something needs to send it audio and collect its output.

Speech to Text APIs

A speech to text API is the software that packages up one or more models and exposes them as a consumable service. It lets you upload an audio file or stream and receive text back, handling authentication, request and response formats, streaming versus batch endpoints, JSON and rate limits. The transcription capabilities available through a speech to text API, such as speaker diarization, timestamps, custom vocabulary and formatting, come from the model underneath, not the API layer itself. For a delivery-layer view, see our guide to the best speech to text APIs.

How They Relate

Most developers interact with a model only through its API.  Modulate Transcribe is the API that exposes it as Batch, Streaming or Fast endpoints. OpenAI’s Whisper Large v3 is the model; it can be used via OpenAI’s hosted API or self-hosted directly, since it is open-weight. Endpoints can even route to multiple models for cost, latency or language coverage without the developer knowing which one.

The practical difference for buyers: when you evaluate a model, you are comparing pure engine functionality (accuracy, language coverage, noise and accent robustness). When you evaluate an API, you are also comparing how that model is packaged into a product: documentation, SDKs, pricing model, latency under load, uptime guarantees, compliance certifications and value-added features built on top of transcription.

Speech to Text vs. Voice Intelligence

There is a distinction between speech to text transcription and a broader category sometimes called voice intelligence.

Speech-to-text answers one question: what was said. Given sound, it produces text. Everything about how it sounded (the tone, speed, pauses, emotional range) is lost when you transcribe speech. That is true of even the highest-accuracy transcription models. An accurate transcript of the phrase "I’m fine" could have been spoken gently, harshly or sarcastically, but the text is identical either way.

Modulate’s founder frames the limit of transcription this way:

"The emotion, the pauses, even the timbre of my voice tell you really meaningful things."
Mike Pappas, CEO and Co-Founder, Modulate. CTO Magazine, April 2026.

That gap between what was said and how it was said is exactly where voice intelligence picks up. Voice intelligence platforms (like Modulate’s Velma voice intelligence platform) analyze the audio signal itself (pitch, cadence, pauses, emphasis) to reveal what text cannot, such as emotional sentiment, cues of deception or rising agitation during a live call.

Velma-2, Modulate’s recent voice intelligence model, is a voice-native AI, built on the Ensemble Listening Model (ELM) architecture, and is designed to understand full conversations with nuance rather than simply generating text. Traditional large language models listen to a call, transcribe it to text and then reason over text. Velma instead operates hundreds of detector models in tandem, each specialized in its own area of analysis and working together under one orchestration layer.

If your use case is captioning, searchable meeting notes, compliance transcripts or basic call logging, a good transcription model is all you need, and the comparisons above will help you find one. If you are after fraud detection, sentiment analysis or understanding how a conversation progressed rather than just what was said, transcription accuracy will not tell the whole story. Consider whether a model offers audio-native analysis versus text-only output. For teams running contact centers, the same signal powers call center analytics well beyond a raw transcript.

How Speech to Text Models Work

Most speech to text models follow a similar pipeline:

  1. Audio preprocessing. The raw waveform is filtered (noise removal if needed) and transformed into something numerical the model can use, typically a spectrogram that encodes frequency over time.
  2. Acoustic modeling. The model analyzes the spectrogram, slices it into chunks and predicts which phonetic or sub-word unit is likely in each. Phonemes are the smallest units of sound that distinguish one word from another.
  3. Language modeling. The model predicts the most likely words given the acoustic data. Language context resolves ambiguous sounds (like "there," "their" and "they’re"). This is often a network layer trained on top of the acoustic model within the same end-to-end package.
  4. Post-processing. The raw transcript is cleaned up with punctuation, capitalization, formatting and speaker labels or timestamps to produce a readable result.

Batch vs. Streaming Transcription

Speech to text models are generally deployed in one of two modes, and the distinction matters for choosing the right one:

  • Batch (offline) transcription processes an entire audio file that has already been recorded. Because it can prioritize accuracy over speed, batch suits podcast transcription, call center QA, legal recordings and compliance archiving.
  • Streaming (real-time) transcription analyzes audio as it happens and returns partial results within milliseconds, for live captioning, voice assistants and AI voice agents that cannot wait for a recording to finish.

Some providers offer both modes from a single underlying model; others offer variants trained specifically for each mode.

How Much Do Speech to Text Models Cost?

Speech to text pricing follows four models: per minute, per hour, per token and free for open-weight models you self-host. Among the APIs compared here, Microsoft MAI-Transcribe-1 lists $0.36 per hour and ElevenLabs Scribe v2 starts at $0.40 per hour, while open-weight models like OpenAI Whisper Large v3, Mistral Voxtral and Qwen3-ASR carry no licensing fee when you self-host on your own GPUs. 

Streaming typically costs more than batch on the same model, and add-ons like speaker diarization or PII redaction cost extra on some platforms while Modulate Transcribe includes them at no charge. Run a real cost estimate with your expected volume and audio type before committing, because per-token billing can swing widely at scale.

What to Look For When Comparing Speech to Text Models

Different models have different strengths, and the ideal one depends on your use case. Features to compare:

  • Word Error Rate (WER). The common accuracy metric: the percentage of words a model transcribes incorrectly versus a human-reviewed reference.
  • Language and accent support. How many languages and dialects are supported, and how accurate the model is in each.
  • Speaker diarization. Whether the model can identify individual speakers within the same audio stream.
  • Real-world accuracy. How well the model performs with background noise, overlapping voices, children and non-native accents versus clean single-speaker studio audio.
  • Custom vocabulary and keyword weighting. Whether you can tune the model for domain-specific words, names or jargon.
  • Deployment flexibility. Cloud-hosted API only, or self-hosted for teams with data-residency or compliance requirements.

Which Models Hold Up on Noisy, Real-World Audio

On noisy, real-world conversational audio with overlapping speakers, model accuracy separates sharply from clean-studio benchmarks. Models trained mostly on curated or read speech, such as OpenAI Whisper-Large-v3, degrade severely in real world conditions such as when speakers interrupt each other or audio quality is poor. This is reflected in poor WER such as Whispers 15.95 percent on AMI-IHM (headset-microphone recordings of real multi-speaker meetings). In contrast, Modulate Transcribe, which is trained on over 500 million hours of real conversational audio, holds 7 to 8 percent WER on conversational benchmarks and is rated robust on both noise and accent, where several competitors show rising word error rate.

Frequently Asked Questions

What is the difference between a speech to text model and a speech to text API?

The model is the AI that does the transcription: the trained weights and architecture that transform audio into text. An API exposes that model through a usable interface, adding business logic for authentication, request and response formatting and batch versus streaming delivery.

What is word error rate (WER), and how should I interpret it?

WER is a common metric for speech recognition performance, calculated by comparing a model’s transcription against human-verified words. It counts substitutions, insertions and deletions as errors. Lower is better, but consider where the number comes from: a 7-8% WER on a difficult real-world benchmark like AMI Meeting Corpus (real, overlapping-speaker meeting audio) is far more meaningful than the same score on clean, read audio such as VoxPopuli (formal, multilingual parliamentary speech recordings).

Do all speech to text models support speaker diarization?

No. Some models advertised or assumed to have it do not. Whisper’s standard model has no native diarization (though it can be paired with a separate tool), nor does Qwen3-ASR. Some models support it only in batch mode, not while streaming, like Chirp 3.

Are accents and noisy audio handled equally well across models?

No. Accuracy degrades differently across models, and this is one of the starkest differences between them. Many models here have measurably higher WER on accented speech or noisy audio, even when the model claims to handle noise well. Test any model with your own real-world audio before deciding.

The gap is well documented. A peer-reviewed study of five major commercial systems found:

"All five ASR systems exhibited substantial racial disparities, with an average word error rate (WER) of 0.35 for black speakers compared with 0.19 for white speakers."
Koenecke et al., "Racial disparities in automated speech recognition," PNAS, 2020.

The lesson holds today: test any model on audio that matches the speakers and conditions you actually serve.

How do I add my own vocabulary or domain-specific terms?

Many models support custom vocabulary lists or keyword weighting, useful for product names, jargon and brand terms. Support and real-world effectiveness vary, and some models that claim the feature still miss domain terms in practice, so test with your own vocabulary.

Can I believe a vendor’s own benchmark claims?

Treat them as directional, not definitive. Most companies report favorable comparisons using their own methodology, datasets and competitor selection, which tends to favor their own model. It does not mean the numbers are wrong; third-party testing simply is not common in this industry. Benchmark any model yourself with your audio before purchase.

How much do these models cost to run?

Pricing ranges across per-minute, per-hour, per-token and credit-based models, or free for open-weight self-hosted models. Streaming usually costs more than batch for the same model. Features like diarization, redaction or custom vocabulary may cost extra on some platforms and be free on others. Run a real cost estimate with your expected volume and audio type when comparing providers.