Reference · Glossary

TTS Glossary — 30+ Terms

Complete reference of text-to-speech (TTS) terminology. From SSML and neural TTS to voice cloning and mel-spectrograms — every TTS concept defined and explained in plain English.

Text to Speech (TTS)

Technology that converts written text into spoken audio using artificial intelligence. Modern TTS uses neural networks to produce natural, human-like speech.

Neural TTS

TTS powered by deep learning neural networks. Produces human-quality speech far surpassing older concatenative or formant methods.

SSML

Speech Synthesis Markup Language — XML-based markup for controlling TTS pronunciation, speed, pitch, and pauses. Example: <prosody rate="slow">Hello</prosody>

Voice cloning

AI technique that replicates a specific person's voice from a few audio samples. Used by ElevenLabs, Resemble AI, Descript Overdub.

Concatenative synthesis

TTS method that stitches together pre-recorded audio segments (diphones, syllables). Better quality than formant but less flexible than neural.

Formant synthesis

Rule-based TTS that models vocal tract resonances (formants). Sounds robotic but requires minimal storage. Used in early screen readers.

G2P (Grapheme-to-Phoneme)

Conversion process that maps written text (graphemes) to pronunciation symbols (phonemes). Essential for TTS to know how to pronounce words.

Phoneme

Smallest unit of sound in a language. For example, /k/, /æ/, /t/ are the three phonemes in "cat".

Mel-spectrogram

Time-frequency representation of audio used as input/output for neural TTS models. Compresses audio into perceptually-meaningful features.

WaveNet

DeepMind's 2016 neural TTS architecture that generates raw audio waveforms sample-by-sample. First AI to match human speech quality.

Tacotron

Google's neural TTS architecture (2017) that predicts mel-spectrograms from text, then converts to audio via vocoder.

FastSpeech

Non-autoregressive neural TTS (2019). Generates audio in parallel rather than sequentially — much faster than Tacotron.

Prosody

Patterns of stress, intonation, and rhythm in speech. Modern neural TTS models prosody automatically from text context.

Pitch

Perceived frequency of speech. Higher pitch sounds higher in tone. Adjustable in TTS via Hz offset (e.g., +10Hz).

Speech rate

Speed of speech delivery. Measured in words per minute (average human: 150 WPM). TTS adjustable from 0.5× to 2.0×.

Voice font

A specific AI voice model with unique characteristics. Examples: "Aria" (US English female), "Guy" (US English male), "Salma" (Egyptian Arabic female).

Locale

Language-region identifier. Examples: en-US (US English), ar-EG (Egyptian Arabic), zh-CN (Mandarin Chinese).

Dialect

Regional variety of a language with distinct pronunciation, vocabulary, and grammar. Example: Egyptian Arabic vs Saudi Arabic.

Audiobook narration

Long-form TTS for converting books to audio format. Requires consistent voice quality across many hours of audio.

IVR (Interactive Voice Response)

Automated phone menu system using TTS to speak options and STT/DTMF to capture user responses.

Screen reader

Accessibility software that reads screen content aloud via TTS. Examples: NVDA, JAWS, VoiceOver, TalkBack.

Audio description

TTS narration of visual content (movies, TV shows) for visually impaired users. Describes actions, settings, and on-screen text.

STT (Speech to Text)

Opposite of TTS — converts spoken audio into written text. Also called ASR (Automatic Speech Recognition).

ASR (Automatic Speech Recognition)

Same as STT. AI technology that transcribes spoken audio to text. Used in dictation, subtitles, voice search.

Wake word

Phrase that triggers a voice assistant. Examples: "Hey Siri", "OK Google", "Alexa".

TTS API

Programmatic interface for converting text to speech in apps. TTS Now offers a free TTS API at /api/tts.

Emotional TTS

TTS that conveys emotions (happy, sad, angry, whispering, excited). Advanced feature offered by ElevenLabs and Murf AI.

Multi-speaker TTS

TTS system that supports multiple voices in one output. Used for dialogues, interviews, and podcasts. TTS Now supports this via [S1]/[S2] markers.

Latency

Delay between text input and audio output. Lower latency = more responsive TTS. Browser-based TTS has lowest latency.

Streaming TTS

TTS that outputs audio as it generates, reducing perceived latency. User hears audio while generation continues.

Voice banking

Process of recording one's voice for future TTS use, often by people with degenerative diseases like ALS.

WFST decoding

Weighted Finite-State Transducer — algorithm used in older G2P and ASR systems. Replaced by neural approaches in modern systems.

Vocoder

Algorithm that converts mel-spectrograms back into audio waveforms. Examples: Griffin-Lim (older), WaveNet, HiFi-GAN (modern).

Try TTS Now — free

322+ AI voices · 75+ languages · No signup · Free commercial use

Open TTS Now →

Related: What is Text to Speech? · Best Free TTS Tools 2026