TTS Glossary — 30+ Terms
Complete reference of text-to-speech (TTS) terminology. From SSML and neural TTS to voice cloning and mel-spectrograms — every TTS concept defined and explained in plain English.
Text to Speech (TTS)
Technology that converts written text into spoken audio using artificial intelligence. Modern TTS uses neural networks to produce natural, human-like speech.
Neural TTS
TTS powered by deep learning neural networks. Produces human-quality speech far surpassing older concatenative or formant methods.
SSML
Speech Synthesis Markup Language — XML-based markup for controlling TTS pronunciation, speed, pitch, and pauses. Example: <prosody rate="slow">Hello</prosody>
Voice cloning
AI technique that replicates a specific person's voice from a few audio samples. Used by ElevenLabs, Resemble AI, Descript Overdub.
Concatenative synthesis
TTS method that stitches together pre-recorded audio segments (diphones, syllables). Better quality than formant but less flexible than neural.
Formant synthesis
Rule-based TTS that models vocal tract resonances (formants). Sounds robotic but requires minimal storage. Used in early screen readers.
G2P (Grapheme-to-Phoneme)
Conversion process that maps written text (graphemes) to pronunciation symbols (phonemes). Essential for TTS to know how to pronounce words.
Phoneme
Smallest unit of sound in a language. For example, /k/, /æ/, /t/ are the three phonemes in "cat".
Mel-spectrogram
Time-frequency representation of audio used as input/output for neural TTS models. Compresses audio into perceptually-meaningful features.
WaveNet
DeepMind's 2016 neural TTS architecture that generates raw audio waveforms sample-by-sample. First AI to match human speech quality.
Tacotron
Google's neural TTS architecture (2017) that predicts mel-spectrograms from text, then converts to audio via vocoder.
FastSpeech
Non-autoregressive neural TTS (2019). Generates audio in parallel rather than sequentially — much faster than Tacotron.
Prosody
Patterns of stress, intonation, and rhythm in speech. Modern neural TTS models prosody automatically from text context.
Pitch
Perceived frequency of speech. Higher pitch sounds higher in tone. Adjustable in TTS via Hz offset (e.g., +10Hz).
Speech rate
Speed of speech delivery. Measured in words per minute (average human: 150 WPM). TTS adjustable from 0.5× to 2.0×.
Voice font
A specific AI voice model with unique characteristics. Examples: "Aria" (US English female), "Guy" (US English male), "Salma" (Egyptian Arabic female).
Locale
Language-region identifier. Examples: en-US (US English), ar-EG (Egyptian Arabic), zh-CN (Mandarin Chinese).
Dialect
Regional variety of a language with distinct pronunciation, vocabulary, and grammar. Example: Egyptian Arabic vs Saudi Arabic.
Audiobook narration
Long-form TTS for converting books to audio format. Requires consistent voice quality across many hours of audio.
IVR (Interactive Voice Response)
Automated phone menu system using TTS to speak options and STT/DTMF to capture user responses.
Screen reader
Accessibility software that reads screen content aloud via TTS. Examples: NVDA, JAWS, VoiceOver, TalkBack.
Audio description
TTS narration of visual content (movies, TV shows) for visually impaired users. Describes actions, settings, and on-screen text.
STT (Speech to Text)
Opposite of TTS — converts spoken audio into written text. Also called ASR (Automatic Speech Recognition).
ASR (Automatic Speech Recognition)
Same as STT. AI technology that transcribes spoken audio to text. Used in dictation, subtitles, voice search.
Wake word
Phrase that triggers a voice assistant. Examples: "Hey Siri", "OK Google", "Alexa".
TTS API
Programmatic interface for converting text to speech in apps. TTS Now offers a free TTS API at /api/tts.
Emotional TTS
TTS that conveys emotions (happy, sad, angry, whispering, excited). Advanced feature offered by ElevenLabs and Murf AI.
Multi-speaker TTS
TTS system that supports multiple voices in one output. Used for dialogues, interviews, and podcasts. TTS Now supports this via [S1]/[S2] markers.
Latency
Delay between text input and audio output. Lower latency = more responsive TTS. Browser-based TTS has lowest latency.
Streaming TTS
TTS that outputs audio as it generates, reducing perceived latency. User hears audio while generation continues.
Voice banking
Process of recording one's voice for future TTS use, often by people with degenerative diseases like ALS.
WFST decoding
Weighted Finite-State Transducer — algorithm used in older G2P and ASR systems. Replaced by neural approaches in modern systems.
Vocoder
Algorithm that converts mel-spectrograms back into audio waveforms. Examples: Griffin-Lim (older), WaveNet, HiFi-GAN (modern).
Related: What is Text to Speech? · Best Free TTS Tools 2026