Text-to-speech (TTS) gives machines a voice. It reads screens aloud for blind and visually impaired users, delivers health information to people with low literacy, powers voice assistants and navigation, and can present information in local languages over a simple phone call. Neural methods made synthetic speech remarkably natural โ which also created new risks from voice cloning.
The TTS pipeline#
- Text analysis (front-end)
- Text normalisation: expand numbers, dates, abbreviations and symbols into words ("Dr. Rahman, 12/05, 3.5 kg" โ "Doctor Rahman, twelfth of May, three point five kilograms"). Language- and context-dependent โ and a common source of errors.
- Grapheme-to-phoneme (G2P) conversion: map spelling to pronunciation, handling exceptions and homographs ("read" present vs past).
- Prosody prediction: phrasing, stress and intonation (questions rise in pitch).
- Acoustic model: map phonemes/characters to an intermediate acoustic representation, usually a mel spectrogram.
- Vocoder: convert the mel spectrogram into a waveform.
Historical approaches#
- Formant synthesis: rule-based models of the vocal tract โ intelligible but robotic.
- Concatenative synthesis (unit selection): record hours of speech from one speaker, cut into units, and select and join the best-matching units. Natural in-domain, but glitchy at joins and inflexible.
- Statistical parametric synthesis (HMM-based): model acoustic parameters with HMMs โ flexible, but muffled ("buzzy") speech.
Neural TTS#
WaveNet (DeepMind, 2016) generated raw audio sample by sample with dilated causal convolutions, producing strikingly natural speech โ but slowly, one of 16,000โ24,000 samples per second at a time.
Tacotron 2 (2018) combined an attention-based sequence-to-sequence model (characters โ mel spectrogram) with a WaveNet vocoder, reaching naturalness ratings close to recorded human speech on its evaluation.
FastSpeech / FastSpeech 2 (2019โ2020) made acoustic models non-autoregressive: a duration predictor decides how many frames each phoneme lasts, and the whole spectrogram is generated in parallel โ much faster and more robust (no skipped or repeated words from attention failures). FastSpeech 2 also predicts pitch and energy for controllable prosody.
Neural vocoders: WaveRNN, WaveGlow (flow-based), HiFi-GAN (GAN-based, fast and high quality) made real-time high-quality synthesis possible on modest hardware.
End-to-end models such as VITS (2021) combine acoustic model and vocoder into one model trained with variational inference and adversarial losses.
Codec language models and zero-shot voices#
A newer paradigm treats speech as sequences of discrete tokens from a neural audio codec (e.g. EnCodec, SoundStream) and models them with language-model techniques. VALL-E (2023) showed that a model conditioned on text and a three-second recording of an unseen speaker could synthesise speech in that speaker's voice (zero-shot voice cloning). Many modern systems use codec tokens, diffusion or flow matching for highly natural, expressive, multilingual speech.
# A lightweight open multilingual TTS example (Meta's MMS-TTS, VITS-based)
from transformers import VitsModel, AutoTokenizer
import torch, scipy.io.wavfile
model = VitsModel.from_pretrained("facebook/mms-tts-eng")
tok = AutoTokenizer.from_pretrained("facebook/mms-tts-eng")
inputs = tok("The health centre is open from nine in the morning until four.", return_tensors="pt")
with torch.no_grad():
wav = model(**inputs).waveform[0].numpy()
scipy.io.wavfile.write("announcement.wav", model.config.sampling_rate, wav)(Meta's Massively Multilingual Speech project released TTS models for over a thousand languages, including many low-resource ones.)
Evaluation#
- Mean Opinion Score (MOS): listeners rate naturalness from 1 to 5; report confidence intervals.
- Intelligibility: transcribe the synthetic speech with ASR (or humans) and compute WER.
- Speaker similarity for voice cloning (embedding similarity + listener tests).
- Prosody and pronunciation errors: especially names, numbers and loanwords.
- Real-time factor and latency for interactive systems.
Test with native speakers of each target language and dialect; pronunciation errors in local names and places undermine trust quickly.
Ethics of synthetic voices#
Applications for inclusion#
TTS in local languages over basic phones (interactive voice response) can reach people with limited literacy or no smartphone; screen readers depend on high-quality TTS; and educational content can be delivered in learners' mother tongues. Building good voices for under-served languages requires recorded speech from native speakers, careful text normalisation rules and community evaluation.