💬 NLP & Transformers · Lecture 27 of 29

Automatic Speech Recognition: From HMMs to Whisper

Speech recognition converts audio into text. We cover audio features and spectrograms, the classical HMM–GMM pipeline, end-to-end neural models with CTC and attention, self-supervised wav2vec 2.0, Whisper, and evaluation with word error rate.

Speech is the most natural human interface, and for people who cannot read or type easily, often the only practical one. Automatic Speech Recognition (ASR) converts spoken audio into text, powering voice assistants, captions for deaf and hard-of-hearing people, transcription of interviews and meetings, and voice interfaces for services in many languages. In the last decade ASR accuracy improved dramatically — but performance still varies widely across languages, accents and recording conditions.

From waveform to features#

Audio is a 1-D signal sampled at, for example, 16,000 samples per second. Speech information lives in frequencies that change over time, so we compute a spectrogram:

  1. Split the signal into short overlapping frames (e.g. 25 ms windows every 10 ms).
  2. Apply a Fourier transform to each frame (Short-Time Fourier Transform).
  3. Map frequencies to the mel scale, which mimics human pitch perception (finer resolution at low frequencies), and take logarithms of the energies → log-mel spectrogram (typically 80 mel bins).

Classical systems further computed MFCCs (mel-frequency cepstral coefficients). Modern neural models consume log-mel spectrograms or even raw waveforms.

python
import torch, torchaudio

waveform, sr = torchaudio.load("speech.wav")                       # (channels, samples)
waveform = torchaudio.functional.resample(waveform, sr, 16000).mean(0, keepdim=True)
mel = torchaudio.transforms.MelSpectrogram(sample_rate=16000, n_fft=400, hop_length=160, n_mels=80)(waveform)
log_mel = torch.log(mel + 1e-6)
print(log_mel.shape)                    # (1, 80, frames) — about 100 frames per second

The classical pipeline (until ~2015)#

$$ \hat{W} = \arg\max_W P(W \mid X) = \arg\max_W \underbrace{P(X \mid W)}_{\text{acoustic model}}\;\underbrace{P(W)}_{\text{language model}} $$
  • Acoustic model: HMMs over phonemes (context-dependent triphones), with Gaussian mixture models (later deep neural networks — the "hybrid" DNN-HMM systems) modelling acoustic features.
  • Pronunciation lexicon: maps words to phoneme sequences.
  • Language model: n-gram model over words.
  • Decoder: searches the combined space with weighted finite-state transducers and beam search.

Powerful but complex, requiring expert-built lexicons and many separately trained components.

End-to-end neural ASR#

Neural models map audio features directly to characters or subword tokens:

  • CTC-based models (e.g. DeepSpeech, Jasper, QuartzNet): an encoder (RNN, CNN or transformer) outputs a distribution per frame; CTC loss (the same algorithm used in OCR) handles the unknown alignment between many audio frames and fewer output tokens.
  • Attention encoder–decoder models (Listen, Attend and Spell): a decoder generates tokens while attending over encoder frames.
  • RNN-Transducer (RNN-T): combines an audio encoder, a prediction network over previous tokens and a joint network; naturally streaming — dominant in on-device assistants.
  • Conformer encoders combine convolution (local acoustic patterns) and self-attention (global context) and are widely used.

External language models can still be fused during decoding to improve domain vocabulary.

Self-supervised speech representations#

wav2vec 2.0 (Baevski et al., 2020) learns from unlabelled audio: a CNN encodes the raw waveform, spans of the latent sequence are masked, and a transformer is trained with a contrastive loss to identify the correct quantised latent for each masked position among distractors. Fine-tuned with CTC on very little labelled speech — in the paper, as little as ten minutes — it produced usable recognisers, and with an hour it achieved strong results on standard benchmarks. HuBERT and XLS-R (multilingual, 128 languages) extended the approach — a breakthrough for low-resource languages.

Whisper: weakly supervised at scale#

OpenAI's Whisper (2022) trained an encoder–decoder transformer on about 680,000 hours of audio paired with (noisy) transcripts from the web, covering many languages, with multitask tokens for transcription, translation into English, language identification and timestamps. It is robust to accents, noise and domains without fine-tuning, and became a popular open model.

python
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="openai/whisper-small", chunk_length_s=30)
print(asr("interview_bn.wav", generate_kwargs={"language": "bengali", "task": "transcribe"})["text"])

Evaluation: word error rate#

$$ \text{WER} = \frac{S + D + I}{N} $$

where $S$, $D$, $I$ are substitutions, deletions and insertions in the minimum edit-distance alignment and $N$ is the number of reference words. Character error rate (CER) is preferred for languages without clear word boundaries or with rich morphology. Normalise text consistently (numbers, punctuation, casing) before scoring.

python
# pip install jiwer
import jiwer
print(jiwer.wer("the clinic opens at nine tomorrow", "the clinic open at nine to morrow"))

Challenges#

  • Accents, dialects and code-switching; studies have found substantially higher error rates for some speaker groups (e.g. Koenecke et al., 2020, found commercial systems made roughly twice as many errors for Black American speakers as for white speakers).
  • Noise, reverberation, far-field microphones, overlapping speakers (diarisation — who spoke when).
  • Low-resource languages and domain-specific vocabulary (names, medical terms).
  • Hallucination in large sequence-to-sequence models — generating text during silence or noise — a known issue for Whisper-style models that matters in medical and legal transcription.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Dialogue Systems and Chatbots: From Rules to LLM Assistants

We compare rule-based, task-oriented and open-domain dialogue systems; cover intent detection, slot filling and dialogue state tracking; and show how LLM-based assistants with retrieval and tools are designed, evaluated and deployed safely.

Intermediate⏱ 5 min#187
💬 NLP & Transformers

Text-to-Speech: From Concatenation to Neural Voices

Text-to-speech systems turn written text into natural-sounding speech. We cover the TTS pipeline, text normalisation and phonemes, acoustic models like Tacotron and FastSpeech, neural vocoders, end-to-end and zero-shot voice models, evaluation, and voice-cloning ethics.

Intermediate⏱ 5 min#189
💬 NLP & Transformers

Multilingual and Low-Resource NLP (with a Focus on Bangla)

Most of the world's languages have little digital data. We examine why this matters, how multilingual models enable cross-lingual transfer, the specific challenges of languages like Bangla, and practical strategies for building NLP in low-resource settings.

Intermediate⏱ 5 min#186