Speech is the most natural human interface, and for people who cannot read or type easily, often the only practical one. Automatic Speech Recognition (ASR) converts spoken audio into text, powering voice assistants, captions for deaf and hard-of-hearing people, transcription of interviews and meetings, and voice interfaces for services in many languages. In the last decade ASR accuracy improved dramatically — but performance still varies widely across languages, accents and recording conditions.
From waveform to features#
Audio is a 1-D signal sampled at, for example, 16,000 samples per second. Speech information lives in frequencies that change over time, so we compute a spectrogram:
- Split the signal into short overlapping frames (e.g. 25 ms windows every 10 ms).
- Apply a Fourier transform to each frame (Short-Time Fourier Transform).
- Map frequencies to the mel scale, which mimics human pitch perception (finer resolution at low frequencies), and take logarithms of the energies → log-mel spectrogram (typically 80 mel bins).
Classical systems further computed MFCCs (mel-frequency cepstral coefficients). Modern neural models consume log-mel spectrograms or even raw waveforms.
import torch, torchaudio
waveform, sr = torchaudio.load("speech.wav") # (channels, samples)
waveform = torchaudio.functional.resample(waveform, sr, 16000).mean(0, keepdim=True)
mel = torchaudio.transforms.MelSpectrogram(sample_rate=16000, n_fft=400, hop_length=160, n_mels=80)(waveform)
log_mel = torch.log(mel + 1e-6)
print(log_mel.shape) # (1, 80, frames) — about 100 frames per secondThe classical pipeline (until ~2015)#
- Acoustic model: HMMs over phonemes (context-dependent triphones), with Gaussian mixture models (later deep neural networks — the "hybrid" DNN-HMM systems) modelling acoustic features.
- Pronunciation lexicon: maps words to phoneme sequences.
- Language model: n-gram model over words.
- Decoder: searches the combined space with weighted finite-state transducers and beam search.
Powerful but complex, requiring expert-built lexicons and many separately trained components.
End-to-end neural ASR#
Neural models map audio features directly to characters or subword tokens:
- CTC-based models (e.g. DeepSpeech, Jasper, QuartzNet): an encoder (RNN, CNN or transformer) outputs a distribution per frame; CTC loss (the same algorithm used in OCR) handles the unknown alignment between many audio frames and fewer output tokens.
- Attention encoder–decoder models (Listen, Attend and Spell): a decoder generates tokens while attending over encoder frames.
- RNN-Transducer (RNN-T): combines an audio encoder, a prediction network over previous tokens and a joint network; naturally streaming — dominant in on-device assistants.
- Conformer encoders combine convolution (local acoustic patterns) and self-attention (global context) and are widely used.
External language models can still be fused during decoding to improve domain vocabulary.
Self-supervised speech representations#
wav2vec 2.0 (Baevski et al., 2020) learns from unlabelled audio: a CNN encodes the raw waveform, spans of the latent sequence are masked, and a transformer is trained with a contrastive loss to identify the correct quantised latent for each masked position among distractors. Fine-tuned with CTC on very little labelled speech — in the paper, as little as ten minutes — it produced usable recognisers, and with an hour it achieved strong results on standard benchmarks. HuBERT and XLS-R (multilingual, 128 languages) extended the approach — a breakthrough for low-resource languages.
Whisper: weakly supervised at scale#
OpenAI's Whisper (2022) trained an encoder–decoder transformer on about 680,000 hours of audio paired with (noisy) transcripts from the web, covering many languages, with multitask tokens for transcription, translation into English, language identification and timestamps. It is robust to accents, noise and domains without fine-tuning, and became a popular open model.
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="openai/whisper-small", chunk_length_s=30)
print(asr("interview_bn.wav", generate_kwargs={"language": "bengali", "task": "transcribe"})["text"])Evaluation: word error rate#
where $S$, $D$, $I$ are substitutions, deletions and insertions in the minimum edit-distance alignment and $N$ is the number of reference words. Character error rate (CER) is preferred for languages without clear word boundaries or with rich morphology. Normalise text consistently (numbers, punctuation, casing) before scoring.
# pip install jiwer
import jiwer
print(jiwer.wer("the clinic opens at nine tomorrow", "the clinic open at nine to morrow"))Challenges#
- Accents, dialects and code-switching; studies have found substantially higher error rates for some speaker groups (e.g. Koenecke et al., 2020, found commercial systems made roughly twice as many errors for Black American speakers as for white speakers).
- Noise, reverberation, far-field microphones, overlapping speakers (diarisation — who spoke when).
- Low-resource languages and domain-specific vocabulary (names, medical terms).
- Hallucination in large sequence-to-sequence models — generating text during silence or noise — a known issue for Whisper-style models that matters in medical and legal transcription.