Resources  /  Glossary

Voice AI Glossary

Plain-language definitions for the terms behind speech recognition, transcription accuracy, voice intelligence, deepfake detection, voice moderation and contact center analytics. Each entry links to related terms and to deeper reading on modulate.ai.

286terms defined
10topic areas
286 terms

0-9

1 termsBack to top
Voice intelligence#

100% Coverage

100% coverage means analyzing all eligible conversations rather than reviewing a sample. Manual quality programs typically evaluate only a portion of interactions, so issues outside the sample may go undetected. Automated conversation analysis makes it practical to evaluate interactions at scale for purposes such as quality assurance, compliance monitoring and coaching.

Real-Time ScoringQuality Assurance (QA)Sampling

Call center monitoring statistics

A

28 termsBack to top
Deepfake & security#

Account Takeover (ATO)

Account takeover occurs when a fraudster gains control of a legitimate customer account, often by passing authentication checks and then changing account details, accessing data or moving money. Voice channels can be a target because some authentication methods rely on information that attackers can obtain through social engineering or data breaches.

Caller AuthenticationSocial EngineeringVoice Fraud

Contact centers as a fraud gateway
Accuracy#

Accuracy

Accuracy measures the proportion of predictions that are correct. In speech recognition, performance is commonly measured using word error rate (WER), and word accuracy is sometimes calculated as 100 minus WER, though the two are not strictly equivalent. For classifiers such as fraud or toxicity detectors, accuracy can be misleading when one class is much rarer than another, making metrics such as precision and recall important for evaluating performance.

Word Error Rate (WER)Precision and Recall

Audio#

Acoustic Features

Acoustic features are the numeric representations of characteristics in an audio signal, such as spectrogram values, Mel-frequency cepstral coefficients (MFCCs), pitch and energy. Traditional speech recognition systems relied heavily on hand-designed features, while modern neural models can learn useful representations directly from waveforms or spectrograms. These features and learned representations capture information the model uses to recognize speech.

Mel Spectrogram and MFCCsAudio Embedding

Transcription#

Acoustic Model

An acoustic model scores how well a slice of audio matches units of speech such as phonemes or characters. In classic ASR pipelines it worked alongside a separate language model and pronunciation dictionary. End-to-end ASR models typically learn these relationships jointly within a single neural network, though some systems still incorporate external language models during decoding.

Language ModelEnd-to-End ASRPhoneme

Deepfake & security#

Active Authentication

Active authentication requires a person to complete a deliberate verification step, such as repeating a phrase or answering a question. It provides an explicit and auditable check but can add friction to calls and may be vulnerable to replayed recordings or compromised answers.

Passive Voice VerificationKnowledge-Based Authentication (KBA)Replay Attack

Deepfake & security#

Adversarial Attack

An adversarial attack intentionally modifies an input to cause a machine learning model to make an incorrect decision. In voice security, this could include adding carefully designed audio changes intended to fool a deepfake detector or speaker verification system.

Deepfake DetectionOut-of-Distribution Detection

Contact center#

After-Call Work (ACW)

After-call work is the time an agent spends completing tasks after an interaction ends, such as writing notes, selecting disposition codes or completing follow-up actions. Automation such as summarization and disposition prediction can reduce the amount of manual work required.

SummarizationAverage Handle Time (AHT)Call Disposition

Trust & safety#

Age Signals

Age signals are indicators that a person may belong to a particular age group, based on factors such as voice characteristics, language patterns and conversation context. Platforms may use these signals to apply additional safety protections, but they are probabilistic and should be treated as one input among multiple safety signals rather than as identity verification.

Child SafetyGrooming

Contact center#

Agent Well-Being

Agent well-being refers to the physical and emotional impact of contact center work, including stress, fatigue and exposure to difficult customer interactions. Voice intelligence can help identify patterns such as abusive interactions or high-pressure situations and support programs designed to improve the agent experience.

Customer Abuse DetectionOccupancy RateVocal Stress Indicators

Deepfake & security#

AI Fraud Detection

AI fraud detection uses machine learning models to identify patterns associated with fraudulent activity. In voice channels, it can combine signals such as authentication results, behavioral patterns, conversation context and audio characteristics to assess risk and support fraud prevention.

Voice FraudReal-Time ScoringDeepfake Detection

AI fraud detection
Voice agents#

AI Guardrails

AI guardrails are controls that keep AI systems operating within defined policies and safety requirements. They can block prohibited content, detect manipulation attempts, enforce required disclosures and route uncertain cases to human review. For voice agents, guardrails can also monitor generated responses and actions before or during an interaction.

Voice AgentHuman-in-the-LoopPolicy Enforcement

AI guardrails
Voice intelligence#

AI Monitoring

AI monitoring uses machine learning models to analyze interactions such as customer calls, voice chats or agent behavior and identify events or behaviors that warrant attention. It can operate in real time or analyze recorded interactions after they occur. Compared with manual review, automated monitoring can evaluate interactions at greater scale, apply defined criteria consistently and surface relevant interactions for further review.

Real-Time ScoringCall Center Monitoring100% Coverage

AI monitoring guide
Transcription#

Alphanumerics

Alphanumerics are sequences containing letters and/or digits, such as order numbers, policy IDs and postcodes. They can be challenging for speech recognition because they often provide little linguistic context and may be spoken character by character. Formatting rules and custom vocabulary help.

Custom VocabularyInverse Text Normalization (ITN)Named Entity Recognition (NER)

Privacy & compliance#

Anonymization

Pseudonymization

Anonymization removes identifying information so data can no longer be linked to an individual. Pseudonymization replaces direct identifiers with a coded value while keeping the possibility of re-identification with additional information. Voice data is difficult to fully anonymize because a person’s voice may itself be identifying, so organizations often combine redaction with access controls.

PII RedactionVoiceprintGDPR

Deepfake & security#

Anti-Spoofing

Anti-spoofing refers to techniques that detect or prevent attempts to fool a voice security system with fake audio. Methods can include detecting replayed recordings, synthetic speech, converted voices, liveness signals and other indicators that the audio may not come from a genuine speaker.

Spoofing AttackLiveness DetectionDeepfake Detection

AI & ML#

Artificial Intelligence (AI)

Artificial intelligence is the broad field of creating systems that perform tasks associated with human intelligence, such as understanding language, recognizing patterns, making predictions and supporting decisions. Modern voice AI systems are primarily built using machine learning models trained on data rather than manually programmed rules.

Machine LearningDeep LearningGenerative AI

Transcription#

ASR Hallucination

Hallucination (speech recognition)

An ASR hallucination is when a recognizer outputs words that were never spoken, often during silence, music or background noise. Unlike an ordinary recognition error, where the system mishears actual speech, a hallucination produces text that is unsupported by the audio. Some end-to-end models can generate plausible or fluent text in these cases. Voice activity detection, confidence thresholds and other filtering techniques reduce the problem.

Voice Activity Detection (VAD)Confidence ScoreLLM Hallucination

AI & ML#

Attention Mechanism

An attention mechanism allows a model to weigh different parts of its input based on their relevance to the current task. In speech systems, it helps models identify which audio frames or tokens are most useful when generating an output. Attention is a core component of transformer architectures.

TransformerEncoder-Decoder

Audio#

Audio Embedding

An audio embedding is a fixed-length numerical vector produced by a neural network that represents characteristics of an audio clip. Audio that is similar according to the characteristics learned by the model tends to have similar embeddings. These representations are used for tasks such as speaker recognition, sound classification and audio search.

Speaker EmbeddingEmbeddingAcoustic Features

Voice intelligence#

Audio Event Detection

Sound event detection

Audio event detection identifies sounds or acoustic events in recordings, such as laughter, crying, keyboard clicks, alarms, gunshots or doors slamming. Depending on the application, it can help moderation systems identify audio associated with potentially concerning events or help contact centers detect background sounds that provide context about the caller’s environment.

Non-Speech AudioDynamic RangeVoice Moderation

Modulate's audio event detection model
Audio#

Audio File Formats

WAV, FLAC, MP3, Opus

Audio can be stored or encoded in formats such as WAV, FLAC, MP3, AAC and Opus. WAV commonly contains uncompressed pulse-code modulation (PCM) audio, while FLAC uses lossless compression and formats such as MP3 and AAC use lossy compression. Telephony audio may use codecs such as G.711 μ-law. Lossless encoding preserves the original digital audio samples, while lossy compression discards some audio information to reduce file size and can affect speech recognition accuracy, particularly at low bitrates.

CodecPulse-Code Modulation (PCM)Lossy and Lossless Compression

Audio#

Audio Frame

Chunk

An audio frame is a short window of audio samples that a speech or audio system processes as a unit, often spanning 10 to 25 milliseconds. Streaming systems process audio frames as they arrive. In frame-based spectral analysis, shorter frames provide better time resolution, while longer frames provide better frequency resolution.

SpectrogramStreaming Transcription

Voice intelligence#

Audio Intelligence

Audio intelligence is a broad term for technologies that analyze audio to extract information beyond basic speech-to-text. This can include sentiment, topics, summaries, speaker characteristics and sound events, using the transcript, the audio signal or both. These capabilities may be built into speech-to-text platforms or provided as a separate analysis layer for applications such as contact center analytics, moderation and media analysis.

Voice IntelligenceSentiment AnalysisAudio Event DetectionSummarization

Deepfake & security#

Audio Watermarking

Audio watermarking embeds a hidden signal into generated audio so that its origin can be verified later. It can help identify content created by systems that support watermarking, but it depends on the watermark being preserved through processing, compression and recording.

Content ProvenanceDeepfake DetectionSynthetic Media

Privacy & compliance#

Audit Trail

An audit trail is a record of system activity that shows what happened, when it occurred, what evidence supported a decision and what action followed. In voice intelligence, audit trails can link detections or decisions to specific interaction data so organizations can review outcomes for compliance, quality or security purposes.

ExplainabilityCompliance MonitoringReal-Time Audio Logging

Contact center#

Automatic Call Distributor (ACD)

An automatic call distributor routes incoming calls to the appropriate agents based on factors such as skills, availability, priority and queue time. It is a core contact center system that connects customer interactions with routing, workforce management and analytics processes.

Interactive Voice Response (IVR)Workforce Management (WFM)CCaaS

Transcription#

Automatic Speech Recognition (ASR)

Speech recognition

Automatic speech recognition is the technology that converts spoken audio into written text. ASR systems analyze audio signals to determine the most likely words being spoken, drawing on patterns in speech sounds and likely word sequences. Modern ASR systems typically use neural networks trained on large volumes of paired audio and transcripts. They power dictation, captions, call transcription and the speech-recognition stage of most voice assistants.

Speech-to-Text (STT)Acoustic ModelEnd-to-End ASR

Contact center#

Average Handle Time (AHT)

Average handle time is the average duration of a customer interaction, including talk time, hold time and after-call work. It is commonly used as a contact center efficiency metric, but it should be evaluated alongside measures such as resolution quality and customer satisfaction.

After-Call Work (ACW)Dead AirKey Performance Indicators (KPIs)

B

8 termsBack to top
Voice agents#

Backchannel

Backchannels are brief listener responses that signal attention or understanding without taking control of the conversation, such as “mm-hm,” “right” or “okay.” Voice systems must distinguish backchannels from full user turns so they do not incorrectly interrupt or respond to simple acknowledgments. Voice agents can also generate backchannels to create more natural conversations.

Turn DetectionBarge-In

Voice agents#

Barge-In

Barge-in occurs when a user begins speaking while a voice agent or IVR system is still playing audio. A capable system detects the interruption, stops its own speech and responds to the new input. Effective barge-in handling requires accurate turn detection, low latency and technologies such as echo cancellation.

Turn DetectionEcho CancellationVoice Agent

Transcription#

Batch Transcription

Asynchronous transcription, Pre-recorded transcription

Batch transcription processes recorded audio after the recording is complete rather than transcribing it in real time. With access to the whole recording and fewer latency constraints, batch systems can use additional context and more computationally intensive processing, potentially improving accuracy and supporting features such as multi-pass diarization. Batch transcription is well-suited for call recordings, meetings and media archives.

Streaming TranscriptionSpeaker DiarizationReal-Time Factor (RTF)

Accuracy#

Benchmark

A benchmark is a standard dataset or task, together with an evaluation method, used to measure and compare model performance. Public speech benchmarks include LibriSpeech, Common Voice and telephony sets such as Switchboard. Results on clean read speech rarely predict performance on real customer calls, so many teams run their own.

Word Error Rate (WER)Test SetText Normalization

Modulate benchmarks
Privacy & compliance#

Biometric Privacy Laws

BIPA

Biometric privacy laws regulate the collection, storage and use of biometric identifiers such as voiceprints. Requirements vary by jurisdiction and may include consent, disclosure, retention policies and security obligations. These laws affect how organizations deploy voice authentication and other voice-based biometric technologies.

VoiceprintVoice BiometricsConsent Management

Audio#

Bit Depth

Bit depth is the number of bits used to store each audio sample, determining the number of possible amplitude levels and the audio’s theoretical dynamic range. 16-bit audio is standard for CDs and is typically sufficient for speech recognition. Lower bit depths increase quantization noise, while higher bit depths such as 24-bit provide greater dynamic range but typically offer little additional benefit for speech recognition.

Sample RateDynamic Range

Deepfake & security#

Bona Fide Speech

Bona fide speech is the term used in anti-spoofing research for genuine human speech, as opposed to spoofed, replayed or synthetic speech. Speech anti-spoofing benchmarks typically measure how well systems distinguish bona fide samples from attacks.

Synthetic SpeechAnti-SpoofingEqual Error Rate (EER)

C

37 termsBack to top
Accuracy#

Calibration

A model is calibrated when its confidence matches its accuracy: among predictions made at 90% confidence, about 90% should be correct. Well-calibrated scores let downstream systems set meaningful thresholds and decide when to involve a human.

Confidence ScoreHuman-in-the-Loop

Contact center#

Call Center Analytics

Call center analytics uses interaction data to identify patterns and insights about customer needs, agent performance, operational efficiency and compliance. Speech and voice analytics add conversational information to traditional metrics such as handle time, queue volume and resolution rates.

Speech AnalyticsCall Center MonitoringKey Performance Indicators (KPIs)

Call center analytics
Contact center#

Call Center Monitoring

Call center monitoring is the practice of evaluating customer interactions for quality, compliance and customer experience. Traditional monitoring relies on supervisors reviewing a sample of calls, while AI monitoring can analyze interactions at larger scale and provide automated insights.

AI MonitoringQuality Assurance (QA)Call Scoring

Call center monitoring
Contact center#

Call Disposition

A call disposition is the outcome or classification assigned to an interaction after it ends, such as resolved, callback scheduled or sale completed. Accurate dispositions support reporting and analysis, while automated prediction can help improve consistency and reduce manual effort.

After-Call Work (ACW)Topic Detection

Contact center#

Call Recording

Call recording captures audio from customer interactions for purposes such as quality assurance, training, dispute resolution and compliance. Because recordings may contain sensitive information, organizations must follow applicable requirements for notification, consent, storage and data protection.

Call Recording ConsentData RetentionPCI DSS

Contact center#

Call Scoring

Call scoring evaluates an interaction against defined criteria and assigns a grade or rating. Automated call scoring can apply consistent standards across more conversations and identify patterns that may be difficult to measure through manual review alone.

Quality Assurance (QA)ScorecardReal-Time Scoring

Transcription#

Call Transcription

Call transcription converts contact center or sales calls into text, usually with speaker separation for agent and customer. Transcripts can support quality assurance, compliance review, coaching and analytics. Real-time call transcription can also enable live agent assistance by providing text that downstream systems analyze during the conversation.

Telephony TranscriptionSpeech AnalyticsCall Center Analytics

Call center analytics
Deepfake & security#

Caller Authentication

Caller authentication is the process of verifying that a caller is authorized to access an account or service. Methods can include passwords, knowledge-based questions, one-time codes, device signals, phone-number signals and voice biometrics combined with anti-spoofing protections.

Knowledge-Based Authentication (KBA)Voice BiometricsAccount Takeover (ATO)

Deepfake & security#

Caller ID Spoofing

Caller ID spoofing changes the phone number displayed to the recipient so a call appears to come from a different source, such as a bank, government agency or local number. Because caller ID can be manipulated, the displayed number alone should not be treated as proof of identity.

VishingRobocallCaller Authentication

Transcription#

Captions and Subtitles

SRT, WebVTT

Captions and subtitles display timed text alongside video. Captions typically represent spoken dialogue as well as relevant non-speech audio, such as sound effects and speaker identification, while subtitles primarily represent dialogue and may translate it into another language. Common file formats are SRT and WebVTT. High-quality captions and subtitles depend on accurate timing, sensible line breaks and appropriate reading speeds.

Word TimestampsForced Alignment

Voice agents#

Cascaded Pipeline

A cascaded pipeline builds a voice system from separate components, typically speech recognition, a language model or dialogue system and text-to-speech. Each stage can be selected, evaluated and replaced independently. The trade-off is additional processing steps, which can increase latency and may lose some information carried in the original audio signal.

Speech-to-Speech ModelVoice AgentLatency

Cascade vs speech-to-speech for enterprise voice agents
Contact center#

CCaaS

Contact Center as a Service (CCaaS) is cloud-delivered contact center software that provides capabilities such as call routing, IVR, recording, agent desktops, reporting and analytics. CCaaS platforms allow organizations to manage customer interactions across channels without maintaining on-premises contact center infrastructure. Many providers offer APIs and integrations that enable voice intelligence tools to analyze interaction data.

Contact CenterSession Initiation Protocol (SIP)Call Recording

Privacy & compliance#

CCPA

The California Consumer Privacy Act (CCPA), as amended by the California Privacy Rights Act (CPRA), gives California residents rights over their personal information. These rights include knowing what data is collected, requesting deletion in certain circumstances and understanding how personal information is used and shared.

GDPRBiometric Privacy Laws

Accuracy#

Character Error Rate (CER)

Character error rate applies the same formula as WER at the character level. It is preferred for languages without clear word boundaries, such as Mandarin and Japanese. It is also more forgiving of small spelling variants in names and alphanumerics.

Word Error Rate (WER)

Trust & safety#

Child Safety

Child safety refers to protecting minors from risks such as grooming, exploitation, exposure to inappropriate content and harassment. Online platforms use policies, detection systems and moderation processes to reduce these risks. Voice chat presents unique challenges because conversations happen in real time and may involve limited context.

GroomingOnline Safety RegulationAge Signals

Raising expectations for child safety online
Contact center#

Churn Risk

Churn risk is the likelihood that a customer may stop using a company’s products or services. Conversation signals such as cancellation requests, repeated issues or negative interaction patterns can help organizations identify customers who may need additional support.

Speech Emotion Recognition (SER)Net Promoter Score (NPS)Escalation Detection

AI & ML#

Cloud Deployment

Cloud deployment runs AI models and applications on infrastructure managed by a cloud provider or vendor rather than on an organization's own hardware. It enables scalable access through APIs but requires consideration of factors such as security controls, data handling, compliance and vendor trust.

On-Premise DeploymentREST APIISO 27001

Transcription#

Code-Switching

Code-switching is when a speaker alternates between two or more languages within a conversation or even within a sentence. It is common in multilingual communities and contact centers. Speech recognition models trained on one language at a time often struggle with code-switched speech, while multilingual models trained to recognize language switching can handle it more effectively.

Language Identification (LID)Multilingual Speech Recognition

Audio#

Codec

A codec encodes and decodes digital audio for storage or transmission, often using compression to reduce the amount of data required. Telephony commonly uses codecs such as G.711 and G.729, while Opus is widely used for internet-based voice communication and MP3 and AAC are typically used for media. Codecs can alter the audio signal through compression, bandwidth limitations and other processing, which can affect speech recognition accuracy. Models trained on similar audio conditions are generally better equipped to handle these effects.

Audio File FormatsMu-law (G.711)Lossy and Lossless Compression

Trust & safety#

Community Guidelines

Code of conduct

Community guidelines are the published rules that define acceptable behavior on a platform. They describe prohibited activities, expected conduct and potential enforcement actions. Moderation systems are typically designed around each platform’s specific guidelines rather than a universal definition of harmful behavior.

Policy EnforcementToxicityTrust and Safety

Contact center#

Compliance Monitoring

Compliance monitoring evaluates customer interactions against regulatory requirements and internal policies, such as required disclosures, consent procedures and rules for handling sensitive information. Automated monitoring can review more interactions and create evidence for audits and quality reviews.

Script AdherenceCall Recording ConsentPCI DSSAudit Trail

AI & ML#

Concurrency

Concurrency refers to the number of requests, users or data streams a system can handle at the same time. In voice applications, concurrency often represents the number of simultaneous calls or audio streams being processed and is an important factor in capacity planning and pricing.

ThroughputRate LimitStreaming Transcription

Transcription#

Confidence Score

A confidence score is the recognizer's own estimate of how likely a word or segment is correct. It is often expressed as a value between 0 and 1 and can be used to flag uncertain results for review or trigger human intervention. A confidence score is not a guarantee of accuracy, and its reliability depends on how well the model’s scores are calibrated.

CalibrationWord Error Rate (WER)

Accuracy#

Confusion Matrix

A confusion matrix is a table of predicted versus actual classes. For a binary detector, a confusion matrix has four cells: true positives, false positives, false negatives and true negatives. This makes it easy to see the types of errors a model makes rather than looking only at its overall accuracy.

Precision and RecallFalse Positive and False Negative

Transcription#

Connectionist Temporal Classification (CTC)

CTC is a training method that lets a model learn to map audio frames to characters, subwords or other output units without requiring frame-level alignment in the training data. It uses a special blank token and allows repeated outputs, which are collapsed during decoding to produce the final transcription. CTC supports efficient decoding and is commonly used in ASR systems, including models designed for streaming.

End-to-End ASRTransducer (RNN-T)Decoding

Contact center#

Contact Center

A contact center manages customer interactions across multiple channels, including phone, chat, email, messaging and social platforms. A call center focuses specifically on voice interactions, while a modern contact center typically combines multiple channels with routing, customer data and analytics capabilities.

Call Center AnalyticsCCaaSOmnichannel

Call center guide
Trust & safety#

Content Moderation

Content moderation evaluates user-generated content against platform rules and takes action when violations are identified. It can be reactive, based on user reports, or proactive, using automated systems to detect potential issues before they are reported. Voice moderation adds complexity because speech requires understanding timing, tone and context.

Voice ModerationProactive Voice ModerationReactive Moderation

Deepfake & security#

Content Provenance

C2PA

Content provenance records information about where media came from and how it was created or modified, often through signed metadata standards such as C2PA. It can help verify content when provenance information remains intact but cannot authenticate media that has no trusted origin data.

Audio WatermarkingSynthetic Media

Transcription#

Context Biasing

Context biasing nudges a recognizer toward expected words or phrases at inference time by feeding it context such as a contact list, a product catalog or the previous turn of a conversation. It can improve recognition of names, specialized terminology and other contextually relevant words and is commonly used to support custom vocabulary features.

Custom VocabularyLanguage Model

Trust & safety#

Contextual Analysis

Contextual analysis evaluates content based on surrounding information, such as who is speaking, the relationship between participants, the tone of delivery, previous conversation and the situation in which it occurs. In voice moderation, context helps distinguish playful conversation from harmful behavior that may use similar words.

Trash Talk vs ToxicityParalinguisticsVoice-Native AI

Voice intelligence#

Conversation Intelligence

Conversation intelligence uses AI to record, transcribe and analyze customer interactions across channels such as calls, meetings and chats. It can identify trends, coaching opportunities, customer concerns and sales risks from conversation data. Commonly associated with sales technology, conversation intelligence increasingly overlaps with speech analytics and customer experience platforms used to analyze interactions at scale.

Speech AnalyticsSummarizationCall Center Analytics

Voice agents#

Conversational AI

Conversational AI refers to systems designed to interact with people through natural language in text or speech. Examples include chatbots, virtual assistants and voice agents. Voice-based systems add challenges such as speech recognition, speech synthesis, turn-taking and handling interruptions that do not exist in text-only conversations.

Voice AgentNatural Language Understanding (NLU)Large Language Model (LLM)

Transcription#

Custom Vocabulary

Keyword boosting, Word boost

Custom vocabulary lets a user supply product names, jargon, acronyms, people's names or other domain-specific terms so the recognizer is more likely to select them when the audio is ambiguous. Related features go by names such as keyword boosting, word boost and phrase lists. It can improve recognition of domain-specific terms without retraining the underlying model.

Context BiasingFine-TuningAlphanumerics

Contact center#

Customer Abuse Detection

Customer abuse detection identifies interactions where a customer may be threatening, harassing or verbally abusive toward an agent. It uses signals such as language, conversation context and vocal characteristics to help organizations protect agents, support intervention and improve workplace safety.

Agent Well-BeingEscalation DetectionVoice Moderation

Contact center#

Customer Satisfaction (CSAT)

Customer satisfaction score (CSAT) measures how satisfied customers are with a specific interaction, typically through a post-contact survey. Because survey response rates can be limited, organizations may combine CSAT results with conversation analytics to identify satisfaction patterns across more interactions.

Net Promoter Score (NPS)Sentiment AnalysisKey Performance Indicators (KPIs)

Customer satisfaction

D

13 termsBack to top
AI & ML#

Data Labeling

Annotation

Data labeling is the process of adding meaningful information to raw data so a model can learn from it. In speech AI, this may include creating transcripts, marking speaker turns, identifying emotions or labeling policy violations. Label quality directly affects model performance.

Training DataGround TruthSupervised Learning

Privacy & compliance#

Data Residency

Data residency refers to requirements or policies governing where data is stored and processed geographically. For voice data, residency requirements may influence decisions about cloud regions, infrastructure location and deployment models based on privacy, regulatory or organizational needs.

Data RetentionOn-Premise DeploymentGDPR

Privacy & compliance#

Data Retention

Data retention defines how long organizations keep recordings, transcripts, metadata and analysis results before deletion or anonymization. Retention periods balance operational needs, compliance requirements and privacy risks. Some organizations retain derived insights longer than raw audio to reduce exposure.

Data ResidencyReal-Time Audio LoggingGDPR

Voice intelligence#

Dead Air

Dead air is a period during a call when neither participant is speaking. It can occur while an agent searches for information, navigates a system or waits for a response. Excessive or unexplained dead air can negatively affect the customer experience. Speech analytics systems can automatically measure periods of silence and use them as a metric for evaluating call handling and agent performance.

Speaker DynamicsVoice Activity Detection (VAD)Average Handle Time (AHT)

Transcription#

Decoding

Decoding is the process of turning a speech model’s output probabilities into a final sequence of words or tokens. Greedy decoding selects the most likely output at each step. Beam search considers several candidate sequences at once and can improve transcription accuracy compared with greedy decoding, but requires more computation.

Beam SearchLanguage Model

AI & ML#

Deep Learning

Deep learning is a type of machine learning that uses neural networks with multiple layers to learn complex patterns from data. It transformed speech technology by enabling advances such as end-to-end speech recognition, neural text-to-speech and realistic voice generation.

Neural NetworkTransformerMachine Learning

Deepfake & security#

Deepfake

A deepfake is synthetic or manipulated media created using AI to make audio, video or images appear to show something that did not occur. The term is commonly used for content that imitates a real person, such as a synthetic voice that appears to come from someone who did not say the words.

Voice DeepfakeSynthetic MediaDeepfake Detection

Deepfake detection at Modulate
Deepfake & security#

Deepfake Detection

Deepfake detection uses machine learning models to distinguish genuine recordings from synthetic or manipulated media. Audio detectors analyze patterns that may indicate generated or altered speech, such as inconsistencies in spectral features, timing or other acoustic characteristics. Performance is typically measured with metrics such as equal error rate and must account for new, previously unseen generators.

Equal Error Rate (EER)Anti-SpoofingOut-of-Distribution DetectionVoice Deepfake

Deepfake detection API
Accuracy#

Diarization Error Rate (DER)

cpWER

Diarization error rate measures how much of a recording is attributed to the wrong speaker. It accounts for missed speech, speech detected where none occurred and speech attributed to the wrong speaker, relative to the total amount of reference speech. Related metrics such as cpWER combine speaker attribution with word accuracy.

Speaker DiarizationWord Error Rate (WER)

Transcription#

Dictation

Dictation is speech recognition used to compose text directly, such as notes, messages or clinical documentation. It favors a single cooperative speaker, immediate feedback and voice commands for editing. Recognition of domain-specific vocabulary is particularly important in specialized applications such as clinical or legal dictation.

Speech-to-Text (STT)Custom Vocabulary

Deepfake & security#

Diffusion Model

A diffusion model is a type of generative AI model that learns to create data by reversing a gradual process of adding noise. Diffusion models are used in applications such as image, audio and speech generation. Because new generation methods can introduce different artifacts, detection systems must continually adapt.

Generative Adversarial Network (GAN)Generative AIModel Drift

Transcription#

Disfluency

Disfluencies are the ums, ahs, repetitions and self-corrections of natural speech. Recognizers may transcribe them or filter them out depending on settings. In voice intelligence, disfluency patterns can provide useful conversational signals, such as where a speaker hesitates or reformulates what they are saying.

Verbatim TranscriptionParalinguistics

Audio#

Dynamic Range

Dynamic range is the difference between the lowest usable signal level and the highest level a system can capture without distortion or clipping, typically measured in decibels (dB). Limited dynamic range can make very quiet speech harder to distinguish from background noise or cause loud speech to clip. Variations in speech level can also provide acoustic information for applications such as emotion and behavioral analysis.

Bit DepthAudio Event Detection

E

14 termsBack to top
Audio#

Echo Cancellation

Acoustic echo cancellation reduces the far-end audio that plays through a speaker and is picked up again by the microphone. Without effective echo cancellation, a voice agent may hear and transcribe its own synthesized speech. Echo cancellation is a standard component of many telephony, conferencing and voice communication systems.

Noise SuppressionBarge-In

AI & ML#

Embedding

An embedding is a numerical representation of data, such as text, audio, a speaker's voice or a conversation, in a form that models can compare mathematically. Similar items are represented by nearby points in an embedding space, enabling search, classification and clustering.

Audio EmbeddingSpeaker EmbeddingTokenization

Voice intelligence#

Emotion AI

Affective computing

Emotion AI refers to technologies designed to detect, interpret or respond to emotional cues from signals such as voice, facial expressions, text or physiological data. It falls within the broader field of affective computing and, for voice applications, overlaps with speech emotion recognition. Because emotional states cannot always be reliably inferred from observable signals, results should be interpreted cautiously and with appropriate human oversight.

Speech Emotion Recognition (SER)Sentiment Analysis

AI & ML#

Encoder-Decoder

An encoder-decoder model uses one component to transform input data into a learned representation and another to generate an output from that representation. In speech recognition, the encoder processes audio while the decoder generates text. In speech translation, the decoder produces the translated output.

Attention MechanismEnd-to-End ASRSpeech Translation

Privacy & compliance#

Encryption

Encryption protects data by converting it into a form that can only be accessed with the appropriate key. Encryption in transit protects data while it moves between systems, while encryption at rest protects stored data such as recordings and transcripts. It is a core security control for voice data.

ISO 27001SOC 2Data Retention

Transcription#

End-to-End ASR

End-to-end ASR maps audio directly to text using a jointly trained model rather than separate acoustic, pronunciation and language components. Common approaches include CTC, attention-based encoder-decoders and transducers. Most modern speech-to-text APIs use end-to-end models.

Connectionist Temporal Classification (CTC)Transducer (RNN-T)Transformer

Transcription#

Endpointing

Endpoint detection, End-of-speech detection

Endpointing is how a speech system decides a speaker has finished an utterance or conversational turn. Simple approaches may wait for a fixed period of silence, while more advanced systems can use acoustic clues, linguistic context or both to determine whether the speaker is likely finished. Faster, more accurate endpointing can reduce awkward pauses in voice agents.

Voice Activity Detection (VAD)Turn DetectionUtterance

Voice intelligence#

Ensemble Listening Model (ELM)

An Ensemble Listening Model is Modulate’s architecture for voice intelligence, combining many specialized models that analyze different aspects of audio, such as emotion, stress, speaker dynamics and synthetic speech. A shared orchestration layer combines these signals into a coherent interpretation of the conversation. Modulate’s Velma platform is built on the ELM architecture.

Voice-Native AIEnsemble ModelVoice Intelligence

Introducing Ensemble Listening Models
AI & ML#

Ensemble Model

An ensemble model combines the outputs of multiple models to improve accuracy, reliability or robustness compared with a single model. Different models may specialize in different signals or tasks, with their outputs combined into a final prediction.

Ensemble Listening Model (ELM)Neural Network

Accuracy#

Equal Error Rate (EER)

Equal error rate is the operating point where the false accept rate equals the false reject rate. It is commonly used as a single-number performance measure for speaker verification and deepfake detection. Lower EER indicates better performance. For example, a 1% EER means that at the equal-error threshold, approximately 1% of genuine trials are falsely rejected and 1% of impostor or spoof trials are falsely accepted.

False Positive and False NegativeDeepfake DetectionSpeaker Verification

Modulate deepfake detection
Trust & safety#

Escalation

In trust and safety, escalation can refer either to harmful behavior becoming more severe or to a flagged case being transferred to a higher level of review. Detection systems can help identify increasing risk so teams can intervene before an incident requires additional enforcement.

Escalation DetectionSeverity ScoringTriage

Voice intelligence#

Escalation Detection

Escalation detection identifies interactions that show signs of increasing conflict, dissatisfaction or risk, using signals such as keywords, interruptions, speaking patterns and acoustic cues. In real-time systems, it can alert supervisors when an interaction may require intervention, such as when a customer appears likely to complain, cancel or become abusive.

Speech Emotion Recognition (SER)Real-Time Agent AssistCustomer Abuse Detection

Privacy & compliance#

EU AI Act

The EU AI Act is a European Union regulation that establishes a risk-based framework for artificial intelligence systems. It creates different requirements depending on the level of risk, including obligations related to transparency, high-risk AI systems and certain uses of biometric or emotion-related technologies.

Emotion AIVoice BiometricsDeepfakeResponsible AI

AI & ML#

Explainability

Explainability is the ability to understand and communicate why a model produced a particular result. In voice intelligence, this may include identifying the audio segment, detected signal, confidence level or reasoning factors behind a classification or alert.

Audit TrailHuman-in-the-LoopResponsible AI

F

9 termsBack to top
Accuracy#

F1 Score

The F1 score is the harmonic mean of precision and recall, combining them into a single measure that gives a lower score when either is weak. It is commonly used to evaluate classification and detection tasks, particularly when the classes are imbalanced, such as toxicity detection, audio event detection and speaker-change detection.

Precision and Recall

Accuracy#

False Positive and False Negative

A false positive is an alert on something harmless, such as flagging friendly trash talk as harassment. A false negative is a miss, such as letting a cloned voice through authentication. Detection systems often face a trade-off between false positives and false negatives, and the appropriate threshold depends on the consequences of each type of error.

Precision and RecallEqual Error Rate (EER)Trash Talk vs Toxicity

AI & ML#

Fine-Tuning

Fine-tuning adapts a pre-trained model by continuing training it on a smaller, task-specific dataset. It can improve performance for specialized domains such as medical speech, industry terminology, product names or specific accents without requiring training a model from scratch.

Foundation ModelCustom VocabularyTransfer Learning

Contact center#

First Call Resolution (FCR)

First call resolution measures the percentage of customer issues resolved during the first interaction without requiring additional follow-up. It is commonly used as a contact center performance metric because it can affect customer experience, operational costs and agent efficiency.

Key Performance Indicators (KPIs)Customer Satisfaction (CSAT)

Transcription#

Forced Alignment

Forced alignment takes an audio file plus a known transcript and determines when each word or phoneme occurs. It is used to add timings to human transcripts, to build training data and to line up lyrics or scripts with recordings.

Word TimestampsPhoneme

Audio#

Formant

Formants are resonant frequencies of the vocal tract that appear as concentrations of energy in the speech spectrum. The first two formants, F1 and F2, are particularly important for distinguishing vowel sounds. Formant patterns also vary with vocal-tract characteristics and articulation, providing acoustic information that can contribute to speaker analysis and synthetic-speech detection.

PhonemeSpectrogramSpeaker Embedding

AI & ML#

Foundation Model

A foundation model is a large AI model trained on broad datasets that can be adapted for many downstream tasks. In speech AI, foundation models learn general patterns from large amounts of audio and can be fine-tuned for tasks such as recognition, speaker analysis, classification or detection.

Self-Supervised LearningFine-TuningLarge Language Model (LLM)

Audio#

Fourier Transform

The Fourier transform represents a signal in terms of the frequencies it contains. Applied to successive short, often overlapping windows of audio, a form called the short-time Fourier transform (STFT) can be used to create a spectrogram showing how frequency content changes over time. The fast Fourier transform (FFT) is an efficient way to compute the discrete Fourier transform, making frequency analysis practical for real-time audio processing.

SpectrogramAcoustic Features

Audio#

Frequency and Pitch

Fundamental frequency (F0)

Frequency describes how rapidly a periodic signal repeats and is measured in hertz (Hz). In voiced speech, the fundamental frequency (F0) corresponds to the rate of vocal-fold vibration and is a major acoustic correlate of perceived pitch. Pitch is the perception of how high or low a sound is. Changes in pitch across an utterance contribute to intonation and can convey information such as emphasis, questions and emotional cues.

ProsodyFormantSpeech Emotion Recognition (SER)

G

7 termsBack to top
Privacy & compliance#

GDPR

The General Data Protection Regulation (GDPR) is a European Union data protection law that governs how organizations collect, process and protect personal data. It treats voice recordings as personal data and may classify voiceprints as biometric data requiring additional protections. It also provides rights such as access, correction and deletion.

Biometric Privacy LawsData RetentionConsent Management

AI & ML#

Generalization

Generalization is a model’s ability to perform accurately on new data that differs from its training examples. In voice AI, this includes handling new accents, environments, devices, codecs, attack methods or voice generators. Strong generalization is essential for reliable production performance.

OverfittingOut-of-Distribution DetectionModel Drift

Deepfake & security#

Generative Adversarial Network (GAN)

A generative adversarial network (GAN) is a machine learning architecture with two competing models: a generator that creates synthetic data and a discriminator that evaluates whether data is real or generated. GANs were widely used in early deepfake systems and remain important in the history of synthetic media generation.

Diffusion ModelVocoderDeepfake

AI & ML#

Generative AI

Generative AI refers to systems that create new content based on patterns learned from existing data. It can generate text, speech, music, images and video. In voice applications, generative AI powers technologies such as voice agents and synthetic speech, while also enabling voice deepfakes.

Large Language Model (LLM)Diffusion ModelSynthetic Media

Audio#

Grapheme-to-Phoneme (G2P)

Grapheme-to-phoneme conversion maps written words or characters to the phonemes that represent their pronunciation. Text-to-speech systems can use G2P to determine pronunciations for unfamiliar names and new words. Speech recognition systems that rely on pronunciation lexicons can use it to generate pronunciations for words that are not already in the lexicon, including custom vocabulary.

PhonemeText-to-Speech (TTS)Custom Vocabulary

Trust & safety#

Grooming

Grooming is a process in which an adult builds trust or an emotional connection with a minor to facilitate exploitation or abuse. It may involve tactics such as excessive attention, secrecy, isolation or gradual boundary violations. Detecting grooming requires analyzing patterns across conversations and behavior over time rather than relying on a single message.

Child SafetyProactive Voice ModerationSeverity Scoring

The duty to protect children online
Accuracy#

Ground Truth

Ground truth is the verified correct answer used to train or evaluate a model, such as a checked transcript, a confirmed fraud outcome or a moderator's decision. Creating reliable ground truth can be costly and time-consuming, particularly when expert review is required. Disagreement between human annotators can reveal ambiguity or uncertainty in the labels and make model performance harder to measure reliably.

Reference TranscriptData LabelingTraining Data

H

4 termsBack to top
Trust & safety#

Harassment

Harassment is unwanted behavior directed at a person that causes harm, distress or a hostile environment. It can include repeated insults, threats, stalking, sexual harassment or targeted abuse. In voice environments, detection often requires evaluating patterns across an interaction rather than judging isolated words alone.

ToxicitySeverity ScoringPlayer Risk Categories

Trust & safety#

Hate Speech

Hate speech refers to attacks, threats or dehumanizing content directed at people based on protected characteristics such as race, religion, gender, sexual orientation or disability. Platforms define prohibited hate speech through their own policies and applicable laws. Voice detection requires understanding context, coded language and the difference between harmful use and discussion of such terms.

ToxicityPolicy EnforcementOnline Safety Regulation

Privacy & compliance#

HIPAA

The Health Insurance Portability and Accountability Act (HIPAA) establishes US requirements for protecting certain health information. Organizations covered by HIPAA and their business associates must implement privacy and security safeguards when handling protected health information (PHI). Vendors that process PHI on behalf of covered entities typically require a business associate agreement.

Protected Health Information (PHI)PII Redaction

AI & ML#

Human-in-the-Loop

Human-in-the-loop systems combine automated model outputs with human review or decision-making. AI systems may identify, rank or recommend actions, while people handle confirmation, judgment or intervention. This approach is commonly used in areas such as moderation and fraud detection where mistakes can have significant consequences.

TriageExplainabilityCalibration

I

6 termsBack to top
AI & ML#

Inference

Inference is the process of using a trained model to analyze new data and produce a prediction or output. Unlike training, which creates the model, inference applies the model in production. Inference speed and cost affect whether systems can operate in real time or require batch processing.

LatencyThroughputModel Distillation

Deepfake & security#

Injection Attack

An injection attack introduces synthetic or recorded audio directly into a communication stream instead of using a live microphone input. In voice security, this can bypass checks that rely on environmental signals, device characteristics or the assumption that audio is coming from a person speaking in real time.

Spoofing AttackDeepfake DetectionCaller Authentication

Voice intelligence#

Intent Detection

Intent detection identifies what a speaker is trying to accomplish, such as canceling a subscription, disputing a charge or asking for a manager. IVRs and voice agents can use detected intent to route interactions, trigger actions or determine an appropriate response. In conversation analytics, intent detection helps classify contact reasons and identify patterns in why customers reach out.

Natural Language Understanding (NLU)Interactive Voice Response (IVR)Topic Detection

Contact center#

Interactive Voice Response (IVR)

An interactive voice response (IVR) system is an automated phone system that collects information from callers and routes them to the appropriate destination using speech input, keypad selections or both. Modern conversational IVRs use speech recognition and intent detection to handle more natural requests.

Intent DetectionVoice AgentAutomatic Call Distributor (ACD)

Transcription#

Inverse Text Normalization (ITN)

Inverse text normalization converts spoken-form output into written form: 'twenty twenty six' becomes '2026' and 'five dollars' becomes '$5'. It is part of the formatting or post-processing stage that makes raw recognizer output easier to read and use.

Smart FormattingAlphanumerics

Privacy & compliance#

ISO 27001

ISO 27001 is an international standard for establishing, implementing and improving an information security management system (ISMS). Certification involves an independent audit of an organization’s security processes and controls against the standard’s requirements.

SOC 2EncryptionCloud Deployment

Modulate compliance

K

3 termsBack to top
Contact center#

Key Performance Indicators (KPIs)

Contact center key performance indicators (KPIs) are measurable metrics used to evaluate operational performance. Common KPIs include first call resolution, average handle time, customer satisfaction, net promoter score, occupancy and service level. Voice intelligence can add conversational measures such as compliance, silence and sentiment trends.

First Call Resolution (FCR)Average Handle Time (AHT)Customer Satisfaction (CSAT)

Call center statistics
Transcription#

Keyword Spotting

Keyword spotting detects specific words or phrases in audio, often without requiring a full transcription. It is used for applications such as wake-word detection and identifying predefined terms or phrases in recorded or live conversations. In contact centers, it can help flag interactions containing specific language for compliance review or analysis.

Wake WordCompliance Monitoring

Deepfake & security#

Knowledge-Based Authentication (KBA)

Knowledge-based authentication verifies identity by asking questions based on personal information, such as a date of birth, address or recent transaction details. Because this information can often be exposed through data breaches or public sources, many organizations supplement or replace KBA with stronger authentication signals.

Caller AuthenticationPassive Voice VerificationSocial Engineering

L

7 termsBack to top
Transcription#

Language Identification (LID)

Language detection, Spoken language identification

Language identification detects which language is being spoken so the appropriate recognition model, settings or downstream processing can be applied. It can run once at the start of a file or continuously to catch language changes mid-conversation.

Code-SwitchingMultilingual Speech Recognition

Transcription#

Language Model

In speech recognition, a language model estimates how likely a sequence of words is, helping the decoder choose between acoustically similar possibilities such as 'their' and 'there'. Traditional ASR systems use a separate language model during decoding, while end-to-end systems may learn linguistic patterns within the model or incorporate an external language model. Large language models (LLMs) are a more recent type of language model designed to understand and generate text.

Acoustic ModelLarge Language Model (LLM)Beam Search

AI & ML#

Large Language Model (LLM)

A large language model (LLM) is an AI model trained on large amounts of text to understand and generate language. In voice systems, LLMs can analyze transcripts, identify intent, summarize conversations and generate responses for voice agents. An LLM can only use information available in its input, so missing transcript information cannot be recovered.

TransformerTranscript-Based AnalysisLLM HallucinationGenerative AI

Transcription#

Latency

Latency is the delay between input and output. For streaming transcription it is the time from a word being spoken to its text appearing. For voice agents the total latency budget spans recognition, reasoning and speech synthesis. Lower latency helps conversations feel more responsive and natural.

Streaming TranscriptionReal-Time Factor (RTF)Voice Agent

Deepfake & security#

Liveness Detection

Liveness detection determines whether a voice sample comes from a live speaker rather than a replayed recording or synthetic audio. Methods can include challenge-response prompts, analysis of acoustic characteristics and models trained to identify signs of spoofed speech.

Anti-SpoofingReplay AttackSpeaker Verification

AI & ML#

LLM Hallucination

An LLM hallucination occurs when a language model generates information that is incorrect, unsupported or invented while presenting it confidently. In voice applications, this could include an inaccurate call summary or an unsupported response from a voice agent. Grounding outputs in source data and applying validation steps can reduce hallucinations.

Large Language Model (LLM)ASR HallucinationExplainability

Audio#

Lossy and Lossless Compression

Lossless compression reduces file size while preserving all of the original digital audio data, allowing it to be reconstructed exactly. Lossy compression permanently removes some audio information to achieve smaller file sizes, typically prioritizing information that is less perceptible to human listeners. Lossy compression can affect speech recognition accuracy, particularly at low bitrates, so retaining a lossless copy is preferable when audio will be archived or reused for analysis.

CodecAudio File Formats

M

15 termsBack to top
AI & ML#

Machine Learning

Machine learning is an approach to building AI systems that learn patterns from examples rather than relying only on explicitly programmed rules. Speech recognition, deepfake detection, classification and prediction systems are common machine learning applications. Performance depends heavily on the quality, diversity and relevance of the training data.

Deep LearningTraining DataSupervised Learning

Accuracy#

Mean Opinion Score (MOS)

Mean opinion score is the average of listener ratings, usually on a 1 to 5 scale, used to judge the naturalness of synthetic speech or the quality of a phone line. Human MOS testing can be time-consuming and costly, so automated MOS prediction models are increasingly used to estimate perceived quality.

Text-to-Speech (TTS)Voice Cloning

Transcription#

Meeting Transcription

Meeting transcription converts multi-party conversations from video calls or conference rooms into text, often with speaker attribution. Transcripts can also support downstream features such as summaries and action items. Overlapping speech, remote audio quality and many voices can make meeting transcription more challenging than single-speaker dictation.

Speaker DiarizationSummarization

Audio#

Mel Spectrogram and MFCCs

A mel spectrogram represents the frequency content of audio using the mel scale, which approximates aspects of human auditory perception and provides less resolution at higher frequencies. Mel-frequency cepstral coefficients (MFCCs) transform mel-scaled spectral information into a compact set of numerical features for each audio frame. Both are widely used representations in speech and speaker processing, although modern neural models may also learn features directly from spectrograms or raw audio.

SpectrogramAcoustic Features

AI & ML#

Model Bias

Model bias is a systematic difference between a model’s performance and the desired outcome, which can affect some groups or situations more than others. In voice AI, this may appear as different error rates across accents, languages, speakers or environments. Evaluating performance across relevant groups helps identify and reduce bias.

Training DataResponsible AIWord Error Rate (WER)

AI & ML#

Model Distillation

Model distillation is a technique where a smaller model is trained to reproduce the behavior of a larger model while using fewer resources. It can reduce latency and compute costs while maintaining much of the original model’s performance, making deployment in real-time applications more practical.

InferenceLatencyReal-Time Processing

Distillation and efficiency at Modulate
Accuracy#

Model Drift

Model drift is the decline in a model’s performance as real-world data moves away from what the model was trained on. New slang, new products, new fraud tactics and new deepfake generators all cause model drift. Monitoring live accuracy and retraining on fresh data can help minimize drift.

Out-of-Distribution DetectionTraining Data

Privacy & compliance#

Model Training Opt-Out

A model training opt-out is a policy or contractual commitment that customer data will not be used to train or improve a vendor’s AI models. It is particularly relevant for voice data because recordings and transcripts may contain customer, employee and other personal information.

Data RetentionConsent Management

Modulate privacy and data practices
Trust & safety#

Moderation Queue

A moderation queue is a collection of flagged items waiting for review or action. Queue size, age and resolution time are important operational metrics. Effective moderation tools provide reviewers with relevant evidence, such as audio clips, transcripts, context and user history, to support decisions.

TriageSeverity ScoringReal-Time Audio Logging

Audio#

Mono and Stereo

Mono audio contains one channel, while stereo audio contains two. For conversation analysis, two-channel audio is particularly useful when each participant is recorded on a separate channel, such as an agent on one channel and a customer on the other. This allows the system to attribute speech by channel without relying on speaker diarization. In a mixed mono recording, speaker diarization may be needed to determine who spoke when.

Multichannel AudioSpeaker Diarization

Audio#

Mu-law (G.711)

A-law

Mu-law (μ-law) is an 8-bit companding method defined by the G.711 standard and widely used in North American and Japanese telephony. G.711 μ-law typically encodes audio sampled at 8 kHz at a bit rate of 64 kbps. A-law is a related G.711 companding method widely used in Europe and other regions. Both are commonly encountered when processing traditional telephony audio.

Telephony TranscriptionNarrowband and Wideband AudioCodec

Transcription#

Multichannel Audio

Channel separation, Dual-channel audio

Multichannel audio records audio sources on separate channels, such as an agent on one channel and the customer on another. Transcribing channels independently can provide reliable speaker attribution and avoid many of the speaker diarization errors that occur with mixed mono audio.

Speaker DiarizationSpeaker LabelsMono and Stereo

Transcription#

Multilingual Speech Recognition

Multilingual speech recognition enables a system to recognize speech across multiple languages, often using a single model trained on multilingual data rather than a separate model for each language. This can simplify deployment, improve handling of code-switching and boost performance for some low-resource languages through knowledge shared across languages,

Language Identification (LID)Code-SwitchingSpeech Translation

AI & ML#

Multimodal AI

Multimodal AI processes and combines information from multiple types of input, such as audio, text, images or video. In voice applications, multimodal systems can combine what was said with information from the audio signal, enabling richer analysis than text alone.

Voice-Native AISpeech-to-Speech ModelFoundation Model

Voice intelligence#

Music Detection

AI music detection

Music detection identifies whether an audio recording contains music and can distinguish musical segments from speech or other sounds. It is used in applications such as content analysis, moderation and media processing. A related but distinct capability, AI-generated music detection, analyzes music for signals that may indicate it was created or modified using generative AI.

Non-Speech AudioAudio Event DetectionSynthetic Media

AI music detection API

N

8 termsBack to top
Transcription#

Named Entity Recognition (NER)

Named entity recognition identifies names, organizations, dates, amounts and other structured items in text or transcripts. In many business applications, errors involving important entities can have greater consequences than other transcription errors, such as when a system incorrectly transcribes an account number or drug name.

AlphanumericsWord Error Rate (WER)PII Redaction

Modulate's public entity transcription benchmark
Audio#

Narrowband and Wideband Audio

Narrowband speech audio typically covers frequencies from about 300 Hz to 3.4 kHz, as in traditional telephone calls, and is commonly sampled at 8 kHz. Wideband speech extends to about 7 kHz and is commonly sampled at 16 kHz, capturing more speech information and generally sounding clearer. Speech recognition models trained primarily on wideband audio may perform less accurately on narrowband telephone audio unless their training also represents telephony conditions.

Sample RateTelephony TranscriptionMu-law (G.711)

Voice intelligence#

Natural Language Processing (NLP)

Natural language processing is a field of AI focused on analyzing, understanding and generating human language. In many voice systems, NLP processes speech after it has been transcribed, supporting tasks such as intent detection, named entity recognition, sentiment analysis and summarization. Modern multimodal systems may also process speech and language jointly rather than relying on a strictly separate transcription stage.

Natural Language Understanding (NLU)Large Language Model (LLM)Transcript-Based Analysis

Voice intelligence#

Natural Language Understanding (NLU)

Natural language understanding is the area of NLP focused on interpreting the meaning of human language. It includes tasks such as identifying intents, entities and relationships. In a voice agent, NLU might interpret “I need to move my appointment to Thursday” as a rescheduling request and identify Thursday as the requested date, allowing the system to take the appropriate action.

Natural Language Processing (NLP)Intent DetectionNamed Entity Recognition (NER)

Contact center#

Net Promoter Score (NPS)

Net promoter score measures customer loyalty by asking how likely a customer is to recommend a company on a scale from 0 to 10. The score is calculated by comparing promoters, passives and detractors. Unlike CSAT, NPS measures overall relationship sentiment rather than satisfaction with a single interaction.

Customer Satisfaction (CSAT)Churn Risk

AI & ML#

Neural Network

A neural network is a machine learning model made of interconnected layers that transform input data into outputs using learned patterns. Different architectures, including convolutional networks, recurrent networks and transformers, are designed to handle different types of data such as audio, text and images.

Deep LearningTransformerInference

Audio#

Noise Suppression

Denoising

Noise suppression reduces unwanted background sounds such as traffic, keyboard noise or crowd noise before or during speech processing. Modern approaches often use neural networks to distinguish speech from background noise and enhance the desired speech signal. Excessive noise suppression can distort speech or remove useful speech information, potentially reducing recognition accuracy.

Signal-to-Noise Ratio (SNR)Echo CancellationSpeech Enhancement

Voice intelligence#

Non-Speech Audio

Non-speech audio includes sounds in a recording that are not spoken language, such as music, environmental noise, sound effects and vocalizations like laughter or sighs. Standard speech-to-text typically does not capture this information, although some systems can detect or label non-speech events. These sounds can provide useful context for understanding an interaction or analyzing the environment in which it occurred.

Audio Event DetectionMusic Detection

O

9 termsBack to top
Contact center#

Occupancy Rate

Occupancy rate measures the percentage of an agent’s logged-in time spent handling customer interactions or completing related work such as after-call tasks. Very high occupancy can contribute to fatigue, while very low occupancy may indicate unused capacity. Workforce teams use the metric to balance staffing and workload.

Workforce Management (WFM)Agent Well-Being

Contact center#

Omnichannel

Omnichannel describes a customer experience where interactions across channels such as phone, chat, email and messaging are connected with shared context. In analytics, an omnichannel approach allows organizations to understand the complete customer journey rather than viewing each interaction separately.

Contact CenterCCaaS

AI & ML#

On-Device Processing

Edge processing

On-device processing runs AI models locally on a device such as a phone, headset or gaming console instead of sending data to the cloud. It can reduce latency and improve privacy but is limited by available processing power and memory, making it most suitable for efficient models.

Wake WordOn-Premise DeploymentLatency

AI & ML#

On-Premise Deployment

On-premise deployment runs AI models and infrastructure within an organization’s own environment rather than through a vendor-managed cloud service. Organizations may choose this approach for requirements related to data control, compliance, security or latency, while taking responsibility for maintenance and scaling.

On-Device ProcessingData ResidencyCloud Deployment

Trust & safety#

Online Safety Regulation

Online safety regulation refers to laws and policies that require platforms to assess and reduce risks to users. These regulations may address issues such as child safety, harmful content, transparency and risk management. Requirements vary by jurisdiction and can affect how platforms design moderation systems and reporting processes.

Trust and SafetyChild SafetyPolicy Enforcement

New regulations for games industry trust and safety
Transcription#

Open-Source Speech Recognition

Open-source speech recognition refers to ASR models, frameworks and toolkits that make their code, model weights or other components freely available for users to run and modify under their respective licenses. Examples include Whisper, wav2vec 2.0, Kaldi and NVIDIA NeMo. Self-hosting can avoid usage-based API fees but shifts infrastructure, tuning and maintenance responsibilities to the user.

WhisperFine-TuningOn-Premise Deployment

Deepfake & security#

Out-of-Distribution Detection

Out-of-distribution detection identifies inputs that differ significantly from the data a model was trained on. In deepfake detection, it helps determine whether a system can handle new generation techniques rather than only recognizing known attack types. Strong generalization is essential as synthetic speech technology changes.

Deepfake DetectionModel DriftGeneralization

AI & ML#

Overfitting

Overfitting occurs when a model learns the details of its training data too closely and performs poorly on new examples. A deepfake detector that recognizes only the characteristics of known generators rather than general patterns of synthetic speech is an example of overfitting.

GeneralizationTest SetOut-of-Distribution Detection

Voice intelligence#

Overtalk

Crosstalk, Interruptions

Overtalk occurs when two or more speakers talk at the same time. It may result from interruptions, but overlapping speech can also occur naturally during conversation. Overtalk can provide useful information about conversational dynamics and can make transcription and speaker diarization more difficult, particularly when multiple voices are mixed into the same audio channel.

Speaker DynamicsBarge-InSpeaker Diarization

P

16 termsBack to top
Voice intelligence#

Paralinguistics

Paralinguistics refers to aspects of spoken communication beyond the words themselves, including pitch, loudness, speaking rate, pauses, laughter and sighs. These cues can provide information about how something is said, including signals associated with emotion, attitude and conversational intent. Paralinguistic analysis allows voice intelligence systems to use information from the audio signal that is typically absent from plain text transcripts.

ProsodySpeech Emotion Recognition (SER)Disfluency

Transcription#

Partial Results

Interim results, Final results

Partial results are provisional transcript fragments a streaming recognizer emits as speech is being processed. They can change as more audio arrives and additional context becomes available. Once the model commits to a portion of the transcript, it returns a final result that is no longer expected to change. Applications can display partial results for responsiveness while using final results as the stable transcript.

Streaming TranscriptionEndpointingUtterance

Deepfake & security#

Passive Voice Verification

Passive voice verification authenticates a caller using their natural speech during a normal conversation, without requiring a separate password phrase or security questions. It can reduce friction compared with active verification methods but is typically paired with anti-spoofing protections to detect replayed or synthetic voices.

Speaker VerificationActive AuthenticationKnowledge-Based Authentication (KBA)

Passive voice verification in insurance
Privacy & compliance#

PCI DSS

The Payment Card Industry Data Security Standard (PCI DSS) is a security standard for protecting payment card information. In contact centers, organizations must prevent sensitive payment data such as card numbers and security codes from being stored or exposed in recordings, transcripts or systems that do not require access.

PII RedactionCall RecordingCompliance Monitoring

Privacy & compliance#

Personally Identifiable Information (PII)

Personally identifiable information (PII) is information that can identify an individual, either directly or when combined with other data. Examples include names, addresses, phone numbers, account details and dates of birth. Voice recordings and transcripts often contain PII, requiring appropriate protection, storage and access controls.

PII RedactionProtected Health Information (PHI)Data Retention

Audio#

Phoneme

A phoneme is the smallest sound unit that can distinguish one word from another in a language, such as /b/ and /p/ in “bat” and “pat.” English is often described as having about 44 phonemes, although the number varies by dialect and linguistic analysis. Pronunciation dictionaries and some speech recognition systems represent words as sequences of phonemes.

Acoustic ModelGrapheme-to-Phoneme (G2P)Forced Alignment

Privacy & compliance#

PII Redaction

PII redaction removes, masks or replaces personal identifiers in data such as transcripts and audio recordings. For example, an account number may be replaced with a placeholder in text or removed from audio. Automated systems can use entity recognition and timestamps to identify and redact sensitive information at the correct location.

Personally Identifiable Information (PII)Named Entity Recognition (NER)Word TimestampsPCI DSS

Modulate's PII and PHI redaction model
Trust & safety#

Player Risk Categories

Player risk categories group users based on patterns of behavior over time, helping platforms apply proportionate responses. Categories may consider factors such as the frequency, severity and type of violations. This allows systems to distinguish repeated harmful behavior from isolated incidents.

Severity ScoringPolicy EnforcementHarassment

Player risk categories
Trust & safety#

Policy Enforcement

Policy enforcement is the process of applying actions when users violate platform rules. Actions may include warnings, content removal, mutes, temporary restrictions, suspensions or bans. Effective enforcement relies on clear policies, consistent decisions and sufficient evidence, including relevant context for voice interactions.

Community GuidelinesPlayer Risk CategoriesTrust and Safety

Policy enforcement
Accuracy#

Precision and Recall

Precision is the share of flagged items that were truly positive. Recall is the share of true positives the system correctly identifies. A fraud detector with high precision rarely raises false alarms; one with high recall rarely misses a real attack. In many detection systems, adjusting the decision threshold creates a trade-off between precision and recall.

F1 ScoreFalse Positive and False NegativeEqual Error Rate (EER)

Trust & safety#

Proactive Voice Moderation

Proactive voice moderation analyzes voice conversations to identify potential policy violations before they are reported by users. It can help surface harmful behavior earlier by detecting signals such as harassment, threats or other rule violations during or shortly after an interaction.

Voice ModerationReactive ModerationUser Report Correlation

What is proactive voice moderation
Trust & safety#

Prosocial Behavior

Prosocial behavior is conduct that supports a healthy online community, such as welcoming others, helping teammates, encouraging positive interactions or de-escalating conflict. Measuring and promoting prosocial behavior gives platforms a way to improve community health alongside enforcing rules against harmful behavior.

ToxicityCommunity Guidelines

What is online prosocial behavior
Audio#

Prosody

Prosody is the pattern of rhythm, stress, intonation and timing in speech. It provides information beyond the words themselves, including cues about emphasis, questions, hesitation, urgency and other aspects of how something is said. Much of this information is absent from a plain text transcript, while audio-based analysis can use prosodic features directly.

ParalinguisticsFrequency and PitchVoice-Native AI

Privacy & compliance#

Protected Health Information (PHI)

Protected health information (PHI) is individually identifiable health information protected under HIPAA in the United States. It can include medical conditions, treatments, insurance details, appointment information and other health-related data connected to a person. Organizations handling PHI must apply appropriate privacy and security safeguards.

HIPAAPII RedactionPersonally Identifiable Information (PII)

Trust & safety#

Proximity Voice Chat

Proximity voice chat allows players to hear and communicate with other players based on their location within a virtual environment. Unlike fixed party chat, it creates conversations that change based on distance and movement. This increases immersion but makes moderation more complex because participants and conversation context can change continuously.

Voice over IP (VoIP)Voice Moderation

Proximity voice chat
Audio#

Pulse-Code Modulation (PCM)

Pulse-code modulation represents digital audio as a sequence of numerical sample values. Linear PCM (LPCM) stores these samples without audio compression and is commonly used in WAV files. Speech recognition APIs often accept PCM audio directly for streaming because it requires no decompression, although it uses more bandwidth than compressed audio formats.

Audio File FormatsCodecSample Rate

Q

1 termsBack to top
Contact center#

Quality Assurance (QA)

Quality management

Quality assurance in a contact center evaluates agent interactions against defined criteria such as accuracy, compliance, communication skills, empathy and resolution quality. Traditional QA programs review a sample of interactions, while automated QA can evaluate more conversations and identify patterns for coaching and process improvement.

Call ScoringScorecard100% CoverageSampling

R

14 termsBack to top
AI & ML#

Rate Limit

A rate limit restricts the number of requests, operations or connections a system allows within a specific period. Rate limits help maintain service reliability and ensure resources are shared fairly across users and applications.

ConcurrencyREST API

Trust & safety#

Reactive Moderation

Reactive moderation takes action after a problem has been reported or identified by a user. It can be simpler to implement and may reduce unnecessary monitoring, but it depends on users reporting harmful behavior. Many incidents go unreported, which can leave some violations undetected.

Proactive Voice ModerationUser Report Correlation

Contact center#

Real-Time Agent Assist

Agent assist

Real-time agent assist analyzes a live customer interaction and provides guidance to an agent while the conversation is happening. It can surface relevant knowledge, suggest next steps, remind agents about compliance requirements or highlight potential issues. It depends on low-latency transcription and conversation analysis.

Streaming TranscriptionEscalation DetectionScript Adherence

Trust & safety#

Real-Time Audio Logging

Real-time audio logging captures relevant portions of voice conversations associated with a moderation event, often with a limited amount of surrounding context. It provides reviewers with evidence for decisions while reducing the need to store complete conversations, helping balance enforcement needs with privacy and storage considerations.

Moderation QueueData RetentionAudit Trail

Real-time audio logging
Transcription#

Real-Time Factor (RTF)

Real-time factor is processing time divided by audio duration. An RTF of 0.1 means one hour of audio can be processed in six minutes, while an RTF of 1 means processing takes as long as the audio itself. Values below 1 indicate faster-than-real-time processing, meaning the system can process prerecorded audio faster than the audio would take to play at normal speed. Lower RTF values indicate greater processing throughput, although they do not directly measure latency or cost.

LatencyBatch TranscriptionThroughput

AI & ML#

Real-Time Processing

Real-time processing analyzes data as it arrives and produces results quickly enough to support immediate action. In voice applications, this can include live transcription, risk scoring or alerts during an ongoing conversation rather than after the interaction ends.

Streaming TranscriptionReal-Time ScoringLatency

Voice intelligence#

Real-Time Scoring

Real-time scoring evaluates an interaction while it is occurring rather than only analyzing it after completion. Scores can assess factors such as quality, compliance, fraud risk or safety. When combined with alerts or automated actions, real-time scoring can help supervisors, agents or other systems respond to potential issues during the live interaction rather than waiting for retrospective review.

AI MonitoringCall Scoring100% Coverage

AI monitoring explained
Accuracy#

Reference Transcript

A reference transcript is the human-verified text that a model's output is scored against. Its quality caps the quality of any benchmark: errors, inconsistent formatting or missing words in the reference show up as errors that can be incorrectly attributed to the model.

Word Error Rate (WER)Ground TruthText Normalization

Deepfake & security#

Replay Attack

A replay attack uses a previously recorded sample of a legitimate speaker’s voice to attempt to pass an authentication check. Unlike AI-generated attacks, replay attacks do not require voice synthesis, which is why systems often combine replay detection, liveness checks and anti-spoofing methods.

Spoofing AttackLiveness DetectionSpeaker Verification

AI & ML#

Responsible AI

Responsible AI is the practice of designing, developing and deploying AI systems with attention to factors such as fairness, transparency, privacy, safety and accountability. In voice AI, this can include evaluating performance across different speakers, protecting sensitive data, explaining model outputs and maintaining appropriate human oversight.

Model BiasExplainabilityHuman-in-the-Loop

Modulate's ethics commitments
AI & ML#

REST API

A REST API is a way for software systems to communicate through standard HTTP requests. Speech and AI services commonly use REST APIs for batch workflows, such as submitting audio files, retrieving results or receiving completed processing outputs. Real-time audio applications often use streaming protocols instead.

WebSocketWebhookTranscription API

Modulate API overview
Audio#

Reverberation

Reverberation is the persistence of sound caused by reflections from surfaces in an enclosed space. These reflections overlap with subsequent speech sounds, reducing clarity and potentially increasing speech recognition errors, particularly with distant microphones in spaces such as conference rooms. Dereverberation processing and placing microphones closer to speakers can reduce its effects.

Speech EnhancementSignal-to-Noise Ratio (SNR)

Deepfake & security#

Robocall

A robocall is an automated phone call that delivers a recorded or synthesized message rather than relying on a live caller. Robocalls can serve legitimate purposes, such as reminders or notifications, but are also commonly used for large-scale scams and unwanted marketing. AI-generated voices have made some robocalls more realistic and personalized.

Caller ID SpoofingVishingSynthetic Speech

Accuracy#

ROC Curve and AUC

A receiver operating characteristic (ROC) curve plots true positive rate against false positive rate across every threshold. The area under the curve summarizes performance across the thresholds: 1.0 is perfect, 0.5 is guessing. AUC can be used to compare detectors without selecting a single operating threshold.

Equal Error Rate (EER)Precision and Recall

S

36 termsBack to top
Audio#

Sample Rate

Sample rate is the number of times per second an analog audio signal is measured, expressed in hertz (Hz). Traditional telephone audio is commonly sampled at 8 kHz, wideband speech at 16 kHz and higher-fidelity audio at 44.1 or 48 kHz. Higher sample rates can capture higher frequencies, preserving more acoustic detail for speech recognition.

Narrowband and Wideband AudioBit DepthWaveform

Contact center#

Sampling

Sampling is the practice of reviewing a subset of interactions to estimate the quality or performance of a larger group. Manual quality programs often rely on sampled calls, which can leave some compliance issues or coaching opportunities undiscovered. Automated analysis enables broader interaction coverage.

100% CoverageQuality Assurance (QA)

Contact center#

Scorecard

A scorecard is a structured framework used to evaluate an interaction against specific criteria. It may include required behaviors, compliance checks and quality measures, with each item assigned a rating or weight. Effective scorecards distinguish objective requirements from more subjective measures such as communication quality.

Call ScoringQuality Assurance (QA)Script Adherence

Contact center#

Script Adherence

Script adherence measures whether agents follow required language, processes or disclosures during customer interactions. It can include regulatory statements, identity verification steps and required explanations. Automated detection can evaluate adherence across more interactions than manual sampling alone.

Compliance MonitoringKeyword SpottingScorecard

AI & ML#

Self-Supervised Learning

Self-supervised learning trains models using large amounts of unlabeled data by creating learning tasks from the data itself. In speech AI, a model may learn from untranscribed audio before being fine-tuned with smaller labeled datasets. This approach helps models learn broad patterns from large audio collections.

Supervised LearningFoundation ModelFine-Tuning

Voice intelligence#

Sentiment Analysis

Sentiment analysis identifies attitudes or opinions expressed in content, commonly classifying them as positive, negative or neutral. Text-based analysis uses the words themselves, while speech sentiment analysis can also incorporate acoustic cues such as pitch, intensity and speaking rate. Combining linguistic and acoustic information can reveal differences between what someone says and how they say it. Sentiment can also be tracked across customer interactions to identify trends.

Speech Emotion Recognition (SER)Customer Satisfaction (CSAT)

Contact center#

Session Initiation Protocol (SIP)

SIPREC

Session Initiation Protocol (SIP) is a communication protocol used to establish, manage and end voice-over-IP calls. SIPREC is an extension that allows a copy of call audio and metadata to be sent to recording or analysis systems, enabling real-time monitoring and voice intelligence applications.

Voice over IP (VoIP)Call RecordingCCaaS

Trust & safety#

Severity Scoring

Severity scoring ranks detected incidents based on factors such as the type of violation, potential harm, context, user history and confidence in the detection. It helps moderation teams prioritize the cases that require the most urgent review or strongest response.

TriagePlayer Risk CategoriesEscalation

Audio#

Signal-to-Noise Ratio (SNR)

Signal-to-noise ratio compares the level of a desired signal, such as speech, with the level of background noise and is typically expressed in decibels (dB). A higher SNR means the speech is more prominent relative to the noise. Lower SNRs can make speech recognition more difficult, particularly in noisy environments such as cars, offices and busy public spaces.

Noise SuppressionVoice Activity Detection (VAD)

Transcription#

Smart Formatting

Punctuation and capitalization

Smart formatting converts raw speech recognition output into more readable written text by adding features such as punctuation, capitalization, numerals and paragraph breaks. While raw or traditional ASR output may contain little formatting, modern systems can predict formatting as part of the recognition process or add it during post-processing so the transcript reads like written text.

Inverse Text Normalization (ITN)Verbatim Transcription

Privacy & compliance#

SOC 2

SOC 2 is an audit framework developed by the American Institute of Certified Public Accountants (AICPA) for evaluating service providers’ controls related to security, availability, processing integrity, confidentiality and privacy. A Type II report evaluates whether those controls operated effectively over a defined period.

ISO 27001Encryption

Deepfake & security#

Social Engineering

Social engineering uses psychological manipulation to persuade people to reveal information, provide access or take an action that benefits an attacker. Common tactics include creating urgency, impersonating trusted individuals and exploiting helpfulness. In contact centers, prevention requires both conversation analysis and strong identity verification processes.

VishingKnowledge-Based Authentication (KBA)Account Takeover (ATO)

AI & ML#

Software Development Kit (SDK)

A software development kit (SDK) is a collection of tools, libraries and documentation that helps developers integrate with a service or platform. An SDK can simplify tasks such as authentication, sending requests, handling responses and managing real-time connections compared with building directly from an API.

REST APITranscription API

Transcription#

Speaker Diarization

Speaker diarization determines who spoke when by dividing audio into speaker-specific segments and assigning labels such as Speaker 1 and Speaker 2, typically without identifying the speakers by name. It is useful for meeting notes, mixed-audio call transcripts and any analysis that needs to distinguish between speakers.

Speaker LabelsDiarization Error Rate (DER)Speaker EmbeddingSpeaker Identification

Voice intelligence#

Speaker Dynamics

Speaker dynamics describe how participants interact over the course of a conversation, including talk time, interruptions, overlapping speech, pauses and turn-taking patterns. These measures can provide useful signals about conversational flow and interaction quality when interpreted in context. Measuring speaker dynamics requires reliable speaker attribution, which may come from speaker diarization or separate audio channels for each participant.

Talk-to-Listen RatioOvertalkDead AirSpeaker Diarization

Deepfake & security#

Speaker Embedding

x-vector

A speaker embedding is a numerical representation of the characteristics of a person’s voice that can be used for speaker-related tasks. Embeddings from the same speaker tend to be closer together in the embedding space. They are used in applications such as speaker verification, identification and diarization. Common approaches include x-vectors and ECAPA embeddings.

VoiceprintSpeaker DiarizationAudio Embedding

Deepfake & security#

Speaker Identification

Speaker identification determines which known speaker is talking by comparing an audio sample against a database of enrolled speakers. It is a one-to-many search problem and can be used to recognize repeat callers, identify participants in recordings or attribute speech when speaker identities are known.

Speaker VerificationSpeaker DiarizationSpeaker Embedding

Transcription#

Speaker Labels

Speaker labels are the tags a diarization system attaches to each transcript segment to show who was speaking. Diarization systems typically assign anonymous labels, while additional information can map speakers to known names or roles. For example, separate audio channels may identify an agent and customer, while speaker recognition can associate a voice with a known individual.

Speaker DiarizationMultichannel Audio

Deepfake & security#

Speaker Verification

Speaker verification determines whether a voice belongs to a claimed identity. It is a one-to-one comparison between an incoming voice sample and an enrolled voice representation, producing a score that is evaluated against an acceptance threshold. Performance is commonly measured using metrics such as equal error rate.

Speaker IdentificationVoiceprintEqual Error Rate (EER)Voice Biometrics

Audio#

Spectrogram

A spectrogram is a visual representation of sound showing time on one axis, frequency on the other and signal intensity through color or brightness. Speech characteristics such as formants and harmonic patterns related to pitch can be seen in a spectrogram. Many speech recognition and synthetic-speech detection models use spectrogram-based representations as input instead of raw audio waveforms.

Mel Spectrogram and MFCCsWaveformFourier Transform

Voice intelligence#

Speech Analytics

Speech analytics uses software to analyze recorded or live conversations for information such as topics, compliance language, sentiment, silence, talk ratios and interaction outcomes. Many systems transcribe speech and analyze the resulting text, while others also use acoustic and conversational features from the audio itself. Combining these signals can provide additional information about how participants speak and interact, not just the words they use.

Call Center AnalyticsConversation IntelligenceVoice Intelligence

Call center analytics guide
Voice intelligence#

Speech Emotion Recognition (SER)

Speech emotion recognition analyzes speech for patterns associated with emotion using acoustic features such as pitch, intensity and speaking rate, sometimes combined with linguistic information. Systems may classify speech into categories such as frustration or calmness or represent it along dimensions such as arousal and valence. SER can support applications such as identifying potentially escalating interactions and analyzing patterns across customer conversations.

Emotion AIParalinguisticsSentiment AnalysisEscalation Detection

Audio#

Speech Enhancement

Speech enhancement improves the quality or intelligibility of speech using techniques such as noise suppression, dereverberation and bandwidth extension. It can make speech easier for people to understand and improve the input to downstream systems such as speech recognition. However, processing designed to improve perceived audio quality does not always improve recognition accuracy and can sometimes remove useful speech information.

Noise SuppressionReverberation

Voice agents#

Speech Synthesis Markup Language (SSML)

Speech Synthesis Markup Language (SSML) is an XML-based markup language that gives text-to-speech systems instructions about how to speak text. It can control elements such as pronunciation, pauses, emphasis, pitch, rate and the handling of numbers or abbreviations.

Text-to-Speech (TTS)Prosody

Transcription#

Speech Translation

Speech translation converts spoken audio in one language into text or speech in another. Cascaded systems chain speech recognition, machine translation and text-to-speech. End-to-end models can translate directly from source-language audio, potentially reducing processing stages and latency but making errors harder to trace to a specific stage.

Multilingual Speech RecognitionCascaded PipelineText-to-Speech (TTS)

Voice agents#

Speech-to-Speech Model

A speech-to-speech model accepts spoken audio and generates spoken audio directly, without requiring a separate text transcript as an intermediate step. This approach can reduce latency and preserve more conversational cues, but direct audio processing can make debugging, evaluation and control more challenging than with separate pipeline stages.

Cascaded PipelineVoice AgentMultimodal AI

Transcription#

Speech-to-Text (STT)

Speech-to-text converts spoken audio into written text and is often used interchangeably with automatic speech recognition (ASR). Speech-to-text systems and APIs differ in accuracy, latency, language coverage, pricing and additional outputs such as speaker labels, timestamps and confidence scores.

Automatic Speech Recognition (ASR)Transcription APIBatch TranscriptionStreaming Transcription

Modulate Transcribe
Deepfake & security#

Spoofing Attack

Presentation attack

A spoofing attack presents fake biometric input to a security system in an attempt to be accepted as genuine. In voice systems, spoofing can involve replayed recordings, synthetic speech or voice-converted audio designed to bypass speaker verification or deceive a person.

Replay AttackInjection AttackAnti-Spoofing

Transcription#

Streaming Transcription

Real-time transcription, Live transcription

Streaming transcription converts audio to text while the speaker is still talking rather than waiting for the recording to finish. Audio is sent in small chunks over a persistent connection, allowing it to return partial and final transcription results with low latency. It powers live captions, voice agents and real-time agent assistance in contact centers.

Batch TranscriptionPartial ResultsEndpointingLatencyWebSocket

Privacy & compliance#

Subprocessor

A subprocessor is a third-party organization that processes customer data on behalf of another company, such as a cloud infrastructure provider or service vendor. Organizations that use subprocessors are typically responsible for ensuring appropriate agreements, security measures and privacy obligations are in place.

GDPRCloud Deployment

Modulate subprocessors
Accuracy#

Substitutions, Deletions and Insertions

These are the three error types counted in WER. A substitution replaces the right word with an incorrect word. A deletion occurs when a word in the reference transcript is missing from the output. An insertion occurs when the recognizer adds a word that is not in the reference transcript. Breaking WER down by type can help reveal whether a model is mishearing, omitting or hallucinating.

Word Error Rate (WER)ASR Hallucination

Voice intelligence#

Summarization

Call summarization generates a concise written account of a conversation, often highlighting key topics, outcomes and next steps. Many systems use a language model to summarize the transcript, while some can process audio more directly. Summaries can reduce after-call documentation and make lengthy conversations easier to review. Their accuracy depends on factors such as transcription quality, available context and the performance of the summarization model.

Large Language Model (LLM)After-Call Work (ACW)Meeting Transcription

AI & ML#

Supervised Learning

Supervised learning trains a model using examples that include both inputs and expected outputs. For speech applications, this may include audio paired with transcripts or calls labeled for fraud, sentiment or policy violations. Model performance depends on the quality, quantity and coverage of labeled training examples.

Self-Supervised LearningData LabelingTraining Data

AI & ML#

Synthetic Data

Synthetic data is artificially generated data used for training, testing or evaluation. In speech AI, it may include text-to-speech audio, simulated conversations or artificially created noise conditions. Synthetic data can expand training coverage but may introduce patterns that differ from real-world data if not carefully balanced.

Training DataText-to-Speech (TTS)

Deepfake & security#

Synthetic Media

Synthetic media is content created or modified using AI or other computational methods, including voice, music, video, images and text. The term does not indicate whether the content is beneficial or harmful. Authentication methods such as provenance tracking and detection models help assess whether synthetic media is trustworthy.

DeepfakeContent ProvenanceMusic Detection

Deepfake & security#

Synthetic Speech

Synthetic speech is audio generated by a machine rather than recorded directly from a human speaker. It includes legitimate applications such as virtual assistants, accessibility tools and navigation systems, as well as malicious uses such as impersonation. Detection systems typically classify speech based on whether it is human-generated or synthetic, regardless of intent.

Text-to-Speech (TTS)Bona Fide SpeechDeepfake Detection

T

21 termsBack to top
Voice intelligence#

Talk-to-Listen Ratio

Talk-to-listen ratio compares the amount of time a speaker talks with the amount of time they listen, often measuring agent and customer talk time in contact centers. Coaching programs use it as one indicator of conversational balance and active listening. The appropriate ratio varies by call type and should be interpreted alongside other measures such as customer outcomes and quality scores.

Speaker DynamicsCall Scoring

Transcription#

Telephony Transcription

Telephony transcription applies speech recognition to phone audio, which is often narrowband audio sampled at 8 kHz and may be affected by codecs, line noise, crosstalk and other channel limitations. Models trained on clean wideband audio may perform less accurately under these conditions, so vendors often offer models or configurations optimized for telephony audio.

Narrowband and Wideband AudioMu-law (G.711)Call Transcription

Accuracy#

Test Set

Held-out data, Evaluation set

A test set is data kept separate from model training and used only to measure performance on unseen examples. If test audio leaks into training or influences model development, scores are artificially inflated. A representative test set should reflect the target use case, including factors such as audio quality, accents, vocabulary and recording conditions.

BenchmarkTraining DataOverfitting

Accuracy#

Text Normalization

Text normalization standardizes reference and predicted transcripts before scoring. It can include lowercasing, removing punctuation and standardizing numbers, abbreviations or other formatting differences. Different normalization rules can produce different WER results on the same transcripts, so benchmark results should specify the normalization method used.

Word Error Rate (WER)Inverse Text Normalization (ITN)

Benchmark methodology
Voice agents#

Text-to-Speech (TTS)

Text-to-speech converts written text into spoken audio. Modern neural TTS systems can generate natural-sounding voices with control over factors such as speaking style and prosody. TTS is the output stage of many voice agents and is also used to create synthetic speech for applications such as accessibility, media and, when misused, voice impersonation.

Voice CloningVocoderSpeech Synthesis Markup Language (SSML)Synthetic Speech

AI & ML#

Throughput

Throughput measures how much work a system can complete in a given amount of time, such as hours of audio processed per hour or the number of requests handled per second. It helps determine system capacity for both batch processing and real-time deployments.

InferenceReal-Time Factor (RTF)Concurrency

Audio#

Timbre

Timbre is the quality of a sound that helps distinguish two voices even when they have the same pitch and loudness. It is shaped by factors such as vocal-tract characteristics, resonance and patterns of voice production. Timbre-related characteristics provide useful information for speaker recognition and are among the features voice-cloning systems attempt to reproduce when synthesizing a particular voice.

FormantSpeaker EmbeddingVoice Cloning

AI & ML#

Tokenization

Tokenization is the process of dividing data into smaller units that a model can process. Text models often use word pieces or subwords, while speech systems may use characters, phonemes or learned audio representations. Token choices affect efficiency, vocabulary coverage and how rare terms are handled.

EmbeddingLarge Language Model (LLM)

Voice intelligence#

Topic Detection

Topic detection identifies the subjects or themes discussed in conversations, either by matching them to a predefined taxonomy or discovering patterns automatically. In contact centers, aggregated topic data can reveal drivers of contact volume, common complaint areas, product issues and opportunities to improve self-service or customer support processes.

Intent DetectionSummarizationSpeech Analytics

Trust & safety#

Toxicity

Toxicity refers to behavior that creates a harmful, hostile or disruptive environment, including harassment, threats, hate speech and targeted abuse. Because the same words can have different meanings depending on tone, relationship and context, effective detection requires more than keyword matching.

Trash Talk vs ToxicityHarassmentHate SpeechVoice Moderation

AI & ML#

Training Data

Training data is the collection of examples used to teach a machine learning model. For speech systems, it may include audio from different speakers, environments, devices, accents and domains, along with labels such as transcripts or classifications. Gaps in training data can create performance limitations.

Data LabelingTest SetSynthetic DataModel Drift

Voice intelligence#

Transcript-Based Analysis

Transcript-based analysis converts speech to text first, then applies language models or other text-based methods to analyze the resulting transcript. It is a common architecture for conversation analytics. Its limitation is that information carried primarily in the audio signal, such as vocal characteristics, speaking style and some synthetic speech indicators, may not be available after transcription alone.

Voice-Native AILarge Language Model (LLM)Speech Analytics

Voice moderation is more than transcription
Transcription#

Transcription

Transcription is the process of turning recorded or live speech into text. It can be done by people, by software or by a combination of the two. Automated transcription is commonly evaluated based on accuracy, processing speed and its performance across conditions such as background noise, accents and specialist vocabulary.

Speech-to-Text (STT)Verbatim TranscriptionWord Error Rate (WER)

Transcription#

Transcription API

A transcription API is a programming interface that accepts audio and returns a transcript, often with additional data such as timestamps, confidence scores and speaker labels. Developers can integrate speech recognition into applications without building and hosting their own speech models. Pricing is typically per audio hour or per minute.

Speech-to-Text (STT)Batch TranscriptionStreaming TranscriptionWebSocket

Transcription APIs explained
Transcription#

Transducer (RNN-T)

A transducer is an end-to-end ASR architecture that combines an audio encoder with a prediction network conditioned on previous outputs. It produces text incrementally, which makes it well-suited for low-latency streaming recognition.

End-to-End ASRConnectionist Temporal Classification (CTC)Streaming Transcription

AI & ML#

Transfer Learning

Transfer learning uses knowledge learned by a model on one task or dataset to improve performance on another task. In speech AI, a general model may be adapted for a specific language, industry, speaker task or detection problem through methods such as fine-tuning.

Fine-TuningFoundation Model

AI & ML#

Transformer

A transformer is a neural network architecture built around attention mechanisms that allow models to process relationships between different parts of an input. Introduced in 2017, transformers power many modern AI systems, including large language models and many speech recognition models.

Attention MechanismLarge Language Model (LLM)End-to-End ASR

Trust & safety#

Trash Talk vs Toxicity

Trash talk is competitive banter that participants understand and accept as part of an interaction. Toxicity involves behavior that becomes harmful, unwanted or disruptive. The distinction depends on factors such as tone, context, relationship between participants and whether the target is a willing participant, making context essential for accurate moderation.

ToxicityFalse Positive and False NegativeVoice Moderation

Trash talk vs toxicity
Trust & safety#

Triage

Triage is the process of sorting flagged incidents so moderation teams can focus attention where it is most needed. Automated systems may rank cases by factors such as severity, confidence and user history, while routing ambiguous or high-risk cases for human review.

Severity ScoringHuman-in-the-LoopModeration Queue

Trust & safety#

Trust and Safety

Trust and safety is the discipline of protecting users and communities from harmful behavior, abuse, fraud, exploitation and other risks on digital platforms. It combines policies, moderation systems, enforcement processes and user reporting tools. Voice chat presents unique challenges because speech is real-time, contextual and often temporary.

Content ModerationVoice ModerationPolicy Enforcement

User safety solutions
Transcription#

Turn Detection

Turn detection determines when control of a conversation passes from one participant to another. In voice agents it helps determine when the system should start responding. It combines acoustic silence, prosody and linguistic context, while accounting for natural pauses, interruptions and backchannels.

EndpointingBarge-InBackchannelVoice Agent

U

2 termsBack to top
Trust & safety#

User Report Correlation

User report correlation compares automated moderation detections with reports submitted by users to evaluate how well systems identify harmful behavior. It can help measure detection quality, identify missed incidents and understand the relationship between reported problems and automated signals.

Reactive ModerationProactive Voice ModerationPrecision and Recall

Correlating ToxMod detections with user reports
Transcription#

Utterance

An utterance is a continuous stretch of speech from one speaker, bounded by silence or a speaker change. Speech recognition systems may use utterance boundaries to segment and process audio. These boundaries can also help determine transcript formatting and when a voice agent should respond.

EndpointingPartial Results

V

18 termsBack to top
Transcription#

Verbatim Transcription

Verbatim transcription preserves speech as accurately as possible, including fillers, false starts, repetitions and relevant, non-speech sounds. Clean verbatim removes fillers, stutters and repeated words while preserving the speaker’s meaning. True verbatim is commonly used when details of how something was said are important, such as in some legal and research contexts, while clean verbatim is often preferred for business transcripts and notes.

TranscriptionSmart FormattingDisfluency

Trust & safety#

Violent Radicalization

Violent radicalization is the process through which individuals adopt beliefs that support or encourage violence, sometimes through online communities or interactions. Safety teams may monitor for signals such as recruitment attempts, threats, praise of violent acts or coordinated targeting while considering context and avoiding reliance on isolated phrases.

Hate SpeechSeverity ScoringEscalation

Protecting players against violent radicalization
Deepfake & security#

Vishing

Vishing (voice phishing) is a phone-based social engineering attack in which a caller impersonates a trusted person or organization to obtain information, access or money. AI voice cloning can make these scams more convincing by allowing attackers to imitate specific individuals.

Voice FraudSocial EngineeringVoice Deepfake

The hidden threat of voice scams
Voice intelligence#

Vocal Stress Indicators

Vocal stress indicators are acoustic features such as changes in pitch, energy, speaking rate, pauses and vocal stability that may be associated with stress or heightened arousal. Voice intelligence systems can use them as one signal among many when analyzing interactions, such as identifying potentially difficult calls or supporting risk analysis. These indicators provide clues, not proof of a person’s emotional state.

ParalinguisticsSpeech Emotion Recognition (SER)Voice Fraud

Voice agents#

Vocoder

A vocoder is a component of many speech synthesis systems that converts an intermediate audio representation, such as a mel spectrogram, into a waveform that can be played as sound. Neural vocoders have significantly improved synthetic speech quality, making automated detection methods more important than relying on human listening alone.

Text-to-Speech (TTS)Mel Spectrogram and MFCCsDeepfake Detection

Transcription#

Voice Activity Detection (VAD)

Voice activity detection classifies segments of audio as speech or non-speech. It can determine what audio is passed to a speech recognizer, trim silence from recordings and help endpointers detect when someone has stopped talking. Good VAD must distinguish speech from background noise across challenging conditions, including noisy environments and compressed telephone audio.

EndpointingTurn DetectionSignal-to-Noise Ratio (SNR)

Voice agents#

Voice Agent

AI voice agent

A voice agent is software that conducts spoken conversations with people in real time to answer questions, complete tasks or route requests. It typically combines speech recognition, a reasoning or dialogue system and speech synthesis. Effective voice agents must manage turn-taking, interruptions, context and latency.

Conversational AICascaded PipelineSpeech-to-Speech ModelTurn DetectionBarge-In

AI voice agents
Deepfake & security#

Voice Biometrics

Voice biometrics identifies or verifies individuals using characteristics of their voice. It can be used in contact centers to authenticate callers without relying solely on passwords or security questions. Because synthetic voices can imitate some biometric characteristics, voice biometric systems are often combined with anti-spoofing and deepfake detection technologies.

Speaker VerificationSpeaker IdentificationVoiceprintDeepfake Detection

Voice agents#

Voice Cloning

Voice cloning creates synthetic speech that imitates the voice characteristics of a specific person using examples of their speech. Some systems can generate a voice from only a short audio sample. Voice cloning has legitimate uses such as accessibility and media localization, but it can also enable impersonation and fraud.

Text-to-Speech (TTS)Voice DeepfakeVoice ConversionZero-Shot Learning

Voice agents#

Voice Conversion

Voice conversion changes the perceived identity or characteristics of a speaker in an existing recording while preserving much of the original speech content, timing and delivery. Unlike text-to-speech voice cloning, it starts with recorded speech rather than text, which can preserve natural conversational cues.

Voice CloningVoice Deepfake

Deepfake & security#

Voice Deepfake

Audio deepfake, Synthetic voice attack

A voice deepfake is AI-generated or AI-modified speech designed to imitate a real person’s voice. It can be created through techniques such as voice cloning or voice conversion and delivered as a recording or generated in real time. Voice deepfakes are used in both legitimate applications and impersonation attacks.

DeepfakeVoice CloningVoice ConversionVishing

How to detect deepfakes
Deepfake & security#

Voice Fraud

Voice-based fraud

Voice fraud is fraudulent activity conducted through spoken communication channels. Examples include impersonating a customer during a support call, posing as a trusted organization or using a cloned voice to authorize actions. Detection systems can analyze behavioral, conversational and audio signals to identify potential risk.

VishingVoice DeepfakeAccount Takeover (ATO)Social Engineering

Fraud has a voice
Voice intelligence#

Voice Intelligence

Voice intelligence analyzes spoken conversations to extract insights such as intent, compliance signals, risk indicators and interaction patterns. It goes beyond transcription by using information from both the words and the audio signal, including tone, speaking style, speaker dynamics and other acoustic cues. Applications include conversation analytics, quality monitoring, fraud prevention and synthetic speech detection.

Audio IntelligenceVoice-Native AISpeech AnalyticsEnsemble Listening Model (ELM)

Velma voice intelligence platform
Trust & safety#

Voice Moderation

Voice moderation uses automated systems and human review processes to identify harmful behavior in spoken interactions, including harassment, hate speech, threats, grooming and other policy violations. Effective voice moderation considers not only words but also factors such as context, conversation patterns and audio signals.

Proactive Voice ModerationToxicityTrash Talk vs ToxicityAudio Event Detection

ToxMod voice moderation
Contact center#

Voice of the Customer (VoC)

Voice of the customer (VoC) programs collect and analyze customer feedback from sources such as surveys, reviews, support interactions and conversations. Analyzing customer interactions at scale can provide additional insight into customer needs, issues and trends beyond survey responses alone.

Customer Satisfaction (CSAT)Topic DetectionSentiment Analysis

Trust & safety#

Voice over IP (VoIP)

Voice over IP (VoIP) transmits voice as digital data over internet networks rather than traditional telephone networks. It powers applications such as game voice chat, video conferencing and many modern communication systems. Factors such as codecs, packet loss, latency and jitter affect the quality of audio received by transcription and moderation systems.

CodecSession Initiation Protocol (SIP)Proximity Voice Chat

VoIP for gaming 101
Voice intelligence#

Voice-Native AI

Audio-native AI

Voice-native AI processes speech or other voice signals directly rather than relying solely on a text transcript before analysis or reasoning. This allows systems to use acoustic and conversational information such as tone, hesitation, overlapping speech and vocal characteristics that may be lost or reduced during transcription. Audio-native AI is a broader term that can also include non-speech sounds.

Transcript-Based AnalysisEnsemble Listening Model (ELM)Paralinguistics

What voice-native means
Deepfake & security#

Voiceprint

A voiceprint is a stored representation of voice characteristics used for biometric comparison, typically in the form of a speaker embedding rather than a recording of a person’s voice. Voiceprints are used for speaker verification and identification. Because they can be linked to an individual, they are treated as biometric data in many privacy frameworks.

Speaker EmbeddingSpeaker VerificationBiometric Privacy Laws

W

8 termsBack to top
Transcription#

Wake Word

A wake word is the trigger phrase that activates a device or voice assistant for further speech processing. It’s often handled by a lightweight, always-on model running locally on the device. After detection, further speech recognition may occur locally or be sent to a larger remote system. Local wake-word detection can reduce computational demands and limit the amount of audio sent for further processing.

Keyword SpottingOn-Device Processing

Audio#

Waveform

A waveform represents how a sound signal’s amplitude changes over time. In digital audio, it is stored as a sequence of sampled amplitude values. Speech models may process the waveform directly or transform it into representations such as spectrograms or other acoustic features before analysis.

Sample RateSpectrogramPulse-Code Modulation (PCM)

AI & ML#

Webhook

A webhook is a method for one system to automatically send information to another system when an event occurs. For example, an AI service may send a webhook when a transcript is complete or a risk alert is generated, allowing connected applications to respond without repeatedly checking for updates.

REST APIReal-Time Processing

Transcription#

WebSocket

A WebSocket is a persistent two-way connection between a client and a server. Streaming speech APIs commonly use it to send audio chunks continuously in one direction and return partial and final transcripts in the other without opening a new HTTP request for each change.

Streaming TranscriptionREST API

Transcription#

Whisper

Whisper is an open-source speech recognition model family released by OpenAI in 2022 and trained on a large multilingual dataset. It is widely used as a benchmark or foundation for other speech recognition systems. Trade-offs include an architecture not designed for native streaming, the potential to hallucinate text unsupported by the audio and no native speaker diarization.

Open-Source Speech RecognitionASR Hallucination

Accuracy#

Word Error Rate (WER)

Word error rate (WER) is the standard measure of transcription accuracy. It adds substitutions, deletions and insertions, then divides by the number of words in the reference transcript. A WER of 10% means roughly one word in ten is wrong. Results depend heavily on the test audio and on how text is normalized before comparison.

Character Error Rate (CER)Substitutions, Deletions and InsertionsText NormalizationBenchmark

Modulate benchmarks
Transcription#

Word Timestamps

Word-level timestamps

Word timestamps indicate when individual words occur in an audio recording, typically with start and end times for each recognized word in a transcript. They let applications jump to the moment a phrase was said, build captions, sync highlights to playback and analyze speech timing. Timestamp accuracy is separate from transcription accuracy, since a word can be recognized correctly but assigned an incorrect time.

Forced AlignmentCaptions and Subtitles

Contact center#

Workforce Management (WFM)

Workforce management (WFM) uses forecasting, scheduling and performance data to ensure contact centers have the right number of agents available at the right times. Conversation analytics can enhance WFM by revealing changes in customer needs, interaction complexity and areas where agents may need additional support.

Occupancy RateAutomatic Call Distributor (ACD)Agent Well-Being

Contact center workforce management

Z

1 termsBack to top
AI & ML#

Zero-Shot Learning

Zero-shot learning allows a model to perform a task or recognize a category without receiving task-specific examples during training. In AI applications, this can include generating a voice from a short sample or handling new categories based on learned representations. It improves flexibility but can create new evaluation and safety challenges.

Voice CloningFoundation Model

No terms matchTry a shorter search or clear the topic filters.
Voice intelligence

Hear what transcripts miss

Modulate's audio-native models read tone, stress, speaker dynamics and acoustic authenticity directly from the audio, on every call, in real time.