๐Ÿ’ฌ NLP & Transformers ยท Lecture 9 of 29

Sentiment Analysis: Opinions, Aspects and Nuance

Sentiment analysis detects opinions and emotions in text. We compare lexicon-based and learned approaches, tackle negation and sarcasm, extend to aspect-based sentiment and emotions, and discuss evaluation and misuse.

"The food was delicious but the service was painfully slow." Is this review positive or negative? Both โ€” about different things. Sentiment analysis (opinion mining) identifies the attitudes expressed in text. It is used to understand customer feedback, monitor public opinion, analyse survey responses and track reactions to services. It is also a perfect case study in how simple-looking NLP tasks hide deep linguistic complexity.

Levels of granularity#

  • Document-level: overall polarity of a review.
  • Sentence-level: polarity of each sentence.
  • Aspect-based: sentiment towards specific aspects ("food: positive; service: negative").
  • Emotion detection: joy, anger, fear, sadness, surprise, disgust โ€” beyond positive/negative.
  • Intensity: how strongly positive or negative (a regression problem).

Lexicon-based approaches#

Use a dictionary of words with sentiment scores and aggregate them, with rules for:

  • Negation: "not good" flips polarity.
  • Intensifiers: "very good" > "good"; "slightly good" < "good".
  • Contrast: "but" shifts emphasis to the clause that follows.
  • Punctuation, capitalisation and emojis: "GREAT!!!" is stronger than "great".

VADER (Hutto & Gilbert, 2014) encodes such rules for social-media English and works without any training data.

python
# pip install vaderSentiment
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
an = SentimentIntensityAnalyzer()
for s in ["The service was good.", "The service was not good.", "The service was VERY good!!!",
          "The food was delicious but the service was painfully slow."]:
    print(f"{an.polarity_scores(s)['compound']:+.3f}  {s}")

Lexicons are transparent and fast, but they miss domain-specific meaning ("unpredictable" is good for a film plot, bad for a car's brakes) and most context.

Machine-learning approaches#

  • TF-IDF n-grams + logistic regression: bigrams like "not good" help with negation; strong baseline.
  • Fine-tuned transformers (BERT, RoBERTa, multilingual XLM-R): the state of the art, capturing context, negation scope and many subtle cues.
  • LLMs zero-/few-shot: good general sentiment judgement without training data; useful for bootstrapping labels.

Hard phenomena#

PhenomenonExampleWhy it is hard
Negation scope"I don't think it's bad at all"Double negation; scope beyond adjacent words
Sarcasm/irony"Great, another two-hour queue."Literal words are positive
Comparatives"Better than the old clinic, worse than expected"Sentiment is relative
Implicit sentiment"The battery lasted three hours."Facts that imply opinion via world knowledge
Domain dependence"long" (battery life vs waiting time)Polarity varies by domain
Code-mixing"Service ta khub bhalo chilo, but slow"Multiple languages in one sentence
Mixed sentimentPros and cons in one textSingle label is inadequate

Aspect-based sentiment analysis (ABSA)#

ABSA extracts (aspect, opinion, polarity) triples:

"The doctor was kind but the waiting room was dirty." โ†’ (doctor, kind, positive), (waiting room, dirty, negative).

Approaches: sequence labelling to find aspect terms, then classification of each aspect's polarity with the aspect as additional input (e.g. "[CLS] sentence [SEP] waiting room [SEP]" to a BERT classifier), or generative models that output the triples directly. For service improvement, ABSA is far more actionable than overall polarity.

python
from transformers import pipeline
clf = pipeline("sentiment-analysis", model="cardiffnlp/twitter-xlm-roberta-base-sentiment")
print(clf(["The registration process was quick and staff were helpful.",
           "Waited 5 hours and nobody explained anything."]))

Evaluation#

Use macro-F1 (neutral is often the hardest and most frequent class), per-aspect metrics for ABSA, and correlation for intensity. Most importantly, evaluate on your domain โ€” a model trained on movie reviews can perform poorly on health-service feedback or on another dialect.

Responsible use#

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ’ฌ NLP & Transformers

Text Classification: From Linear Models to Fine-Tuned Transformers

Text classification is the most widely deployed NLP task. We compare TF-IDF baselines, CNN and RNN classifiers, and fine-tuned transformers, and cover label design, imbalance, multilingual data and evaluation.

Beginnerโฑ 4 min#169
๐Ÿ’ฌ NLP & Transformers

Named Entity Recognition: Finding People, Places and Organisations

NER locates and classifies entity mentions in text. We cover BIO tagging, feature-based and neural approaches, transformer token classification with subword alignment, entity-level evaluation, and domain adaptation.

Intermediateโฑ 5 min#171
๐Ÿ’ฌ NLP & Transformers

Subword Tokenisation: BPE, WordPiece, Unigram and SentencePiece

Modern language models split text into subword units. We derive byte-pair encoding step by step, compare WordPiece and Unigram LM tokenisation, discuss byte-level BPE, and examine how tokenisation affects multilingual fairness and cost.

Intermediateโฑ 5 min#168