"Janin A Apurba, an ICT officer at CNRS-UNHCR, delivered a workshop in Dhaka on 12 March." A Named Entity Recognition (NER) system should identify Janin A Apurba as a PERSON, CNRS-UNHCR as an ORGANISATION, Dhaka as a LOCATION and 12 March as a DATE. NER is the foundation of information extraction: turning unstructured text into structured records for search, knowledge graphs, analytics and redaction of personal data.
Entity types#
Standard schemes (e.g. CoNLL-2003): PER, ORG, LOC, MISC. Richer schemes (OntoNotes) add DATE, TIME, MONEY, PERCENT, GPE (geo-political entity), PRODUCT, EVENT and more. Domains define their own: diseases, drugs and dosages in clinical text; case numbers and document types in administrative records.
Sequence labelling with BIO tags#
NER is framed as token classification with the BIO (or IOB2) scheme:
- B-TYPE: beginning of an entity;
- I-TYPE: inside (continuation) of an entity;
- O: outside any entity.
| Token | Janin | A | Apurba | works | at | CNRS-UNHCR | in | Dhaka |
|---|---|---|---|---|---|---|---|---|
| Tag | B-PER | I-PER | I-PER | O | O | B-ORG | O | B-LOC |
BIO tags let adjacent entities of the same type be separated. BIOES/BILOU adds explicit end and single-token tags and sometimes improves accuracy.
Approaches through history#
- Rules and gazetteers โ lists of names and patterns ("Mr. X", capitalised words after "in"). High precision on known names, poor coverage.
- Feature-based statistical models โ HMMs, then Conditional Random Fields (CRFs) with features such as the word, its shape ("Xxxx", "dd-dd"), prefixes/suffixes, capitalisation, neighbouring words and gazetteer membership.
- BiLSTM-CRF (Lample et al., 2016) โ word and character embeddings, a bidirectional LSTM, and a CRF layer enforcing valid tag transitions (e.g. I-PER cannot follow B-ORG).
- Fine-tuned transformers โ BERT-style encoders with a token-classification head; the current standard.
- LLM-based extraction โ prompting generative models to list entities; flexible for new entity types but needs careful evaluation.
Transformers and subword alignment#
Subword tokenisers split words: "Apurba" might become "Ap", "##ur", "##ba". Labels are per word, so we must align them: typically, give the label to the first subword of each word and ignore (label โ100) the remaining subwords when computing the loss.
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
tok = AutoTokenizer.from_pretrained("bert-base-cased")
words = ["Janin", "A", "Apurba", "works", "at", "CNRS-UNHCR", "in", "Dhaka"]
tags = ["B-PER", "I-PER", "I-PER", "O", "O", "B-ORG", "O", "B-LOC"]
label2id = {l: i for i, l in enumerate(["O", "B-PER", "I-PER", "B-ORG", "I-ORG", "B-LOC", "I-LOC"])}
enc = tok(words, is_split_into_words=True)
aligned, prev = [], None
for wid in enc.word_ids():
if wid is None:
aligned.append(-100) # special tokens
elif wid != prev:
aligned.append(label2id[tags[wid]]) # first subword gets the label
else:
aligned.append(-100) # other subwords ignored in the loss
prev = wid
print(list(zip(tok.convert_ids_to_tokens(enc["input_ids"]), aligned)))
ner = pipeline("ner", model="dslim/bert-base-NER", aggregation_strategy="simple")
print(ner("Janin A Apurba presented the results in Dhaka for UNHCR."))Fine-tune with AutoModelForTokenClassification exactly like sequence classification, using the aligned labels.
Evaluation: entity-level F1#
Token-level accuracy is misleading (most tokens are "O"). The standard is entity-level (exact-match) precision, recall and F1: a predicted entity counts only if both its span and its type exactly match the gold entity (as implemented in seqeval). Partial-match metrics give additional insight for boundary errors.
from seqeval.metrics import classification_report
gold = [["B-PER", "I-PER", "O", "B-LOC"]]
pred = [["B-PER", "O", "O", "B-LOC"]] # truncated person span -> counts as an error
print(classification_report(gold, pred))Challenges#
- Ambiguity: "Jordan" (person, country, river); "Apple" (fruit, company).
- Nested entities: "[[Dhaka] University]" โ a location inside an organisation; flat BIO cannot represent nesting.
- Domain shift: news-trained models perform poorly on medical notes, social media or legal text.
- Low-resource languages and scripts: capitalisation, a key English cue, does not exist in Bangla, Arabic or Hindi.
- Name diversity: models trained mostly on Western names may underperform on names from other cultures โ an important fairness issue, e.g. when NER drives record matching.
Applications#
Knowledge-graph construction, search and question answering, news analytics, clinical information extraction, de-identification/anonymisation (finding and masking names, phone numbers and ID numbers before data sharing), resume parsing and automatically populating case-management forms.