๐Ÿ’ฌ NLP & Transformers ยท Lecture 28 of 29

Text-to-Speech: From Concatenation to Neural Voices

Text-to-speech systems turn written text into natural-sounding speech. We cover the TTS pipeline, text normalisation and phonemes, acoustic models like Tacotron and FastSpeech, neural vocoders, end-to-end and zero-shot voice models, evaluation, and voice-cloning ethics.

Text-to-speech (TTS) gives machines a voice. It reads screens aloud for blind and visually impaired users, delivers health information to people with low literacy, powers voice assistants and navigation, and can present information in local languages over a simple phone call. Neural methods made synthetic speech remarkably natural โ€” which also created new risks from voice cloning.

The TTS pipeline#

  1. Text analysis (front-end)
    • Text normalisation: expand numbers, dates, abbreviations and symbols into words ("Dr. Rahman, 12/05, 3.5 kg" โ†’ "Doctor Rahman, twelfth of May, three point five kilograms"). Language- and context-dependent โ€” and a common source of errors.
    • Grapheme-to-phoneme (G2P) conversion: map spelling to pronunciation, handling exceptions and homographs ("read" present vs past).
    • Prosody prediction: phrasing, stress and intonation (questions rise in pitch).
  2. Acoustic model: map phonemes/characters to an intermediate acoustic representation, usually a mel spectrogram.
  3. Vocoder: convert the mel spectrogram into a waveform.

Historical approaches#

  • Formant synthesis: rule-based models of the vocal tract โ€” intelligible but robotic.
  • Concatenative synthesis (unit selection): record hours of speech from one speaker, cut into units, and select and join the best-matching units. Natural in-domain, but glitchy at joins and inflexible.
  • Statistical parametric synthesis (HMM-based): model acoustic parameters with HMMs โ€” flexible, but muffled ("buzzy") speech.

Neural TTS#

WaveNet (DeepMind, 2016) generated raw audio sample by sample with dilated causal convolutions, producing strikingly natural speech โ€” but slowly, one of 16,000โ€“24,000 samples per second at a time.

Tacotron 2 (2018) combined an attention-based sequence-to-sequence model (characters โ†’ mel spectrogram) with a WaveNet vocoder, reaching naturalness ratings close to recorded human speech on its evaluation.

FastSpeech / FastSpeech 2 (2019โ€“2020) made acoustic models non-autoregressive: a duration predictor decides how many frames each phoneme lasts, and the whole spectrogram is generated in parallel โ€” much faster and more robust (no skipped or repeated words from attention failures). FastSpeech 2 also predicts pitch and energy for controllable prosody.

Neural vocoders: WaveRNN, WaveGlow (flow-based), HiFi-GAN (GAN-based, fast and high quality) made real-time high-quality synthesis possible on modest hardware.

End-to-end models such as VITS (2021) combine acoustic model and vocoder into one model trained with variational inference and adversarial losses.

Codec language models and zero-shot voices#

A newer paradigm treats speech as sequences of discrete tokens from a neural audio codec (e.g. EnCodec, SoundStream) and models them with language-model techniques. VALL-E (2023) showed that a model conditioned on text and a three-second recording of an unseen speaker could synthesise speech in that speaker's voice (zero-shot voice cloning). Many modern systems use codec tokens, diffusion or flow matching for highly natural, expressive, multilingual speech.

python
# A lightweight open multilingual TTS example (Meta's MMS-TTS, VITS-based)
from transformers import VitsModel, AutoTokenizer
import torch, scipy.io.wavfile

model = VitsModel.from_pretrained("facebook/mms-tts-eng")
tok = AutoTokenizer.from_pretrained("facebook/mms-tts-eng")
inputs = tok("The health centre is open from nine in the morning until four.", return_tensors="pt")
with torch.no_grad():
    wav = model(**inputs).waveform[0].numpy()
scipy.io.wavfile.write("announcement.wav", model.config.sampling_rate, wav)

(Meta's Massively Multilingual Speech project released TTS models for over a thousand languages, including many low-resource ones.)

Evaluation#

  • Mean Opinion Score (MOS): listeners rate naturalness from 1 to 5; report confidence intervals.
  • Intelligibility: transcribe the synthetic speech with ASR (or humans) and compute WER.
  • Speaker similarity for voice cloning (embedding similarity + listener tests).
  • Prosody and pronunciation errors: especially names, numbers and loanwords.
  • Real-time factor and latency for interactive systems.

Test with native speakers of each target language and dialect; pronunciation errors in local names and places undermine trust quickly.

Ethics of synthetic voices#

Applications for inclusion#

TTS in local languages over basic phones (interactive voice response) can reach people with limited literacy or no smartphone; screen readers depend on high-quality TTS; and educational content can be delivered in learners' mother tongues. Building good voices for under-served languages requires recorded speech from native speakers, careful text normalisation rules and community evaluation.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Automatic Speech Recognition: From HMMs to Whisper

Speech recognition converts audio into text. We cover audio features and spectrograms, the classical HMMโ€“GMM pipeline, end-to-end neural models with CTC and attention, self-supervised wav2vec 2.0, Whisper, and evaluation with word error rate.

Intermediateโฑ 5 min#188
๐Ÿ’ฌ NLP & Transformers

Efficient Transformers: Sparse Attention, Linear Attention and FlashAttention

Self-attention's quadratic cost limits context length. We survey sparse and local attention, low-rank and kernel-based linear attention, IO-aware FlashAttention, and alternatives such as state-space models โ€” with their trade-offs.

Advancedโฑ 6 min#190
๐Ÿ’ฌ NLP & Transformers

Dialogue Systems and Chatbots: From Rules to LLM Assistants

We compare rule-based, task-oriented and open-domain dialogue systems; cover intent detection, slot filling and dialogue state tracking; and show how LLM-based assistants with retrieval and tools are designed, evaluated and deployed safely.

Intermediateโฑ 5 min#187