Between encoder-only BERT and decoder-only GPT sits a third family: encoder–decoder transformers pretrained for sequence-to-sequence tasks. Two landmark models — Google's T5 and Facebook's BART (both 2019) — showed how to pretrain them effectively. They remain strong choices for translation, summarisation and other tasks where an input is transformed into an output.
T5: the text-to-text framework#
Raffel et al.'s "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" framed every task as text-to-text: the input is text with a task prefix, the output is text.
"translate English to German: That is good." → "Das ist gut."
"summarize: <article>" → "<summary>"
"cola sentence: The course is jumping well." → "unacceptable"
"stsb sentence1: ... sentence2: ..." → "3.8"
"question: Who wrote Hamlet? context: ..." → "Shakespeare"Even classification and regression outputs are strings. One model, one loss (cross-entropy on output tokens), one decoding procedure — for everything. This uniformity made large-scale multi-task learning and systematic comparison straightforward.
Span corruption#
T5's pretraining objective replaces random spans of the input (about 15% of tokens, average span length 3) with sentinel tokens, and trains the decoder to output the missing spans:
Input: Thank you <X> me to your party <Y> week.
Target: <X> for inviting <Y> last <Z>Targets are short (only the dropped spans), which makes pretraining efficient.
The C4 dataset and the systematic study#
T5 introduced C4 (Colossal Clean Crawled Corpus), about 750 GB of English web text from Common Crawl, filtered heuristically (removing pages with offensive words, code, very short lines, duplicates). Later audits found that such filtering also removed a disproportionate amount of text from and about some minority groups and dialects — a reminder that data-cleaning choices embed values.
The paper systematically compared architectures, objectives, datasets, transfer strategies and scaling. Key findings:
- An encoder–decoder with a denoising objective performed best in their setting.
- Span corruption was a strong, efficient objective.
- Scale (more parameters, more data, longer training) consistently helped; the largest T5 had 11B parameters.
- Multi-task pretraining followed by fine-tuning worked well.
mT5 extended T5 to 101 languages; Flan-T5 (2022) fine-tuned T5 on more than 1,800 tasks phrased as instructions, markedly improving zero-shot instruction following — an early demonstration of instruction tuning.
BART: a denoising autoencoder#
Lewis et al.'s BART combined a bidirectional encoder (like BERT) with an autoregressive decoder (like GPT). Pretraining corrupts documents and trains the model to reconstruct the original. They compared corruptions:
- token masking, token deletion;
- text infilling (replace spans, including empty spans, with a single mask token — the model must infer how many tokens are missing);
- sentence permutation, document rotation.
Text infilling plus sentence shuffling worked best. BART excelled at abstractive summarisation (e.g. CNN/DailyMail, XSum) and generation tasks, while matching RoBERTa on understanding benchmarks. mBART extended denoising pretraining to many languages for translation.
Using T5 and BART#
from transformers import pipeline, AutoTokenizer, AutoModelForSeq2SeqLM
t5 = pipeline("text2text-generation", model="google/flan-t5-base")
print(t5("Translate English to French: The school opens next Monday.")[0]["generated_text"])
print(t5("Is this review positive or negative? The staff were patient and kind.")[0]["generated_text"])
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
article = ("Heavy monsoon rains caused flooding in several districts. Local authorities opened "
"temporary shelters in schools and distributed clean water. Health workers warned of "
"waterborne diseases and urged families to boil drinking water.")
print(summarizer(article, max_length=40, min_length=10)[0]["summary_text"])Fine-tuning follows the familiar pattern with AutoModelForSeq2SeqLM and Seq2SeqTrainer: tokenise inputs and targets, train with teacher forcing, generate with beam search for evaluation.
When to choose an encoder–decoder#
| Situation | Good choice |
|---|---|
| Input and output are both substantial text (translation, summarisation, rewriting) | Encoder–decoder (T5, BART, mT5, NLLB) |
| Classification, extraction, embeddings | Encoder (BERT family) |
| Open-ended generation, chat, general assistant | Decoder-only LLM |
| Limited compute, specific seq2seq task, need fine-tuning | Small T5/BART fine-tuned — often very cost-effective |
The encoder processes the input bidirectionally once; the decoder attends to it via cross-attention. For tasks with long inputs and short outputs (summarisation, QA), this can be efficient.