💬 NLP & Transformers · Lecture 25 of 29

Multilingual and Low-Resource NLP (with a Focus on Bangla)

Most of the world's languages have little digital data. We examine why this matters, how multilingual models enable cross-lingual transfer, the specific challenges of languages like Bangla, and practical strategies for building NLP in low-resource settings.

There are roughly 7,000 languages in the world, but most NLP research and data focus on a handful — above all English. Joshi et al. (2020) grouped languages by data availability and found that the vast majority of languages, spoken by well over a billion people in total, have almost no labelled data and little unlabelled text. Bangla, with more than 250 million speakers, is one of the most spoken languages in the world, yet it has historically been under-resourced in NLP. For people who need information and services in their own language — including refugees and migrants — this gap matters.

What makes a language "low-resource"?#

  • Little unlabelled text online.
  • Few labelled datasets for tasks (NER, sentiment, QA).
  • Missing tools: tokenisers, morphological analysers, spell checkers.
  • Few evaluation benchmarks, so progress cannot be measured.
  • Script and encoding issues, non-standard spelling, dialectal variation and code-mixing.

Challenges specific to Bangla (and similar languages)#

  • Script complexity: Bengali script uses conjunct consonants (যুক্তাক্ষর), vowel signs that attach before or after consonants, and multiple Unicode representations of visually identical text — Unicode normalisation is essential.
  • Rich morphology: inflections for case, number, tense and person, plus postpositions attached to words, create many surface forms per lemma.
  • Romanised Bangla ("Banglish") and code-mixing with English on social media: "ami office e jacchi, meeting ta late hobe".
  • Dialects: Sylheti, Chittagonian and others differ substantially from standard Bangla.
  • Tokeniser inefficiency: tokenisers trained mainly on English split Bangla into many more tokens, increasing cost and shortening effective context for multilingual LLMs.

Strategy 1: multilingual pretrained models#

Models pretrained on text in many languages share a single vocabulary and parameters:

  • mBERT (104 languages), XLM-RoBERTa (100 languages, CommonCrawl), mT5, NLLB-200 for translation, multilingual sentence encoders (LaBSE, multilingual E5), and multilingual LLMs.
  • Language-specific models such as BanglaBERT (trained on large Bangla corpora) often outperform multilingual models on Bangla tasks.

Strategy 2: cross-lingual transfer#

Remarkably, a multilingual model fine-tuned on labelled data in one language (e.g. English NER) can perform the task in another language zero-shot. Representations of translations tend to align in the shared space. Transfer works best between related languages and scripts, and degrades for distant, under-represented ones. Benchmarks such as XTREME and XGLUE measure this.

python
from transformers import pipeline
# A multilingual NLI model can classify text in many languages without language-specific training
clf = pipeline("zero-shot-classification", model="joeddav/xlm-roberta-large-xnli")
text = "আমাদের এলাকায় বিশুদ্ধ পানির খুব অভাব, শিশুরা অসুস্থ হয়ে পড়ছে।"   # "Our area lacks clean water; children are falling ill."
print(clf(text, candidate_labels=["water and sanitation", "education", "shelter", "health"], multi_label=True))

Strategy 3: create data efficiently#

  • Translate-train / translate-test: machine-translate English training data into the target language (or test data into English). Cheap, but translation errors and "translationese" limit quality.
  • Annotation projection: project labels (e.g. entities) through word alignments of parallel text.
  • Active learning to prioritise the most informative examples for native-speaker annotators.
  • LLM-assisted annotation: let an LLM pre-label, then have native speakers correct — verify quality carefully, since LLMs are weaker in low-resource languages.
  • Community-driven datasets: initiatives such as Masakhane (African languages) show the power of participatory, community-led NLP with local speakers as researchers, not only annotators.

Strategy 4: adapt the model#

  • Continued pretraining on monolingual target-language text.
  • Vocabulary extension: add target-language tokens to the tokeniser and train their embeddings — reduces fertility and cost.
  • Adapters per language (e.g. MAD-X) to add languages without retraining everything.

Evaluate fairly#

Why it matters#

Language technology determines who can access information, services and opportunities online. Poor support for a language means worse search results, worse translation, weaker content moderation (both over- and under-moderation), and LLM assistants that are less helpful or less safe. Building NLP for low-resource languages is both a fascinating research problem and an act of inclusion. It is also a field where students from those language communities have a unique advantage: native-speaker insight.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

💬 NLP & Transformers

Neural Machine Translation: From Seq2Seq to Multilingual Transformers

Machine translation is one of NLP's oldest and most impactful tasks. We trace its evolution to neural systems, cover training data, subword vocabularies, back-translation, evaluation with BLEU and COMET, and the challenges of low-resource languages.

Intermediate⏱ 5 min#173
💬 NLP & Transformers

Topic Modelling: LDA, NMF and Neural Topic Models

Topic models discover themes in large document collections without labels. We derive Latent Dirichlet Allocation's generative story, compare it with NMF and embedding-based BERTopic, and discuss evaluating and interpreting topics.

Intermediate⏱ 5 min#185
💬 NLP & Transformers

Dialogue Systems and Chatbots: From Rules to LLM Assistants

We compare rule-based, task-oriented and open-domain dialogue systems; cover intent detection, slot filling and dialogue state tracking; and show how LLM-based assistants with retrieval and tools are designed, evaluated and deployed safely.

Intermediate⏱ 5 min#187