"The food was delicious but the service was painfully slow." Is this review positive or negative? Both โ about different things. Sentiment analysis (opinion mining) identifies the attitudes expressed in text. It is used to understand customer feedback, monitor public opinion, analyse survey responses and track reactions to services. It is also a perfect case study in how simple-looking NLP tasks hide deep linguistic complexity.
Levels of granularity#
- Document-level: overall polarity of a review.
- Sentence-level: polarity of each sentence.
- Aspect-based: sentiment towards specific aspects ("food: positive; service: negative").
- Emotion detection: joy, anger, fear, sadness, surprise, disgust โ beyond positive/negative.
- Intensity: how strongly positive or negative (a regression problem).
Lexicon-based approaches#
Use a dictionary of words with sentiment scores and aggregate them, with rules for:
- Negation: "not good" flips polarity.
- Intensifiers: "very good" > "good"; "slightly good" < "good".
- Contrast: "but" shifts emphasis to the clause that follows.
- Punctuation, capitalisation and emojis: "GREAT!!!" is stronger than "great".
VADER (Hutto & Gilbert, 2014) encodes such rules for social-media English and works without any training data.
# pip install vaderSentiment
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer
an = SentimentIntensityAnalyzer()
for s in ["The service was good.", "The service was not good.", "The service was VERY good!!!",
"The food was delicious but the service was painfully slow."]:
print(f"{an.polarity_scores(s)['compound']:+.3f} {s}")Lexicons are transparent and fast, but they miss domain-specific meaning ("unpredictable" is good for a film plot, bad for a car's brakes) and most context.
Machine-learning approaches#
- TF-IDF n-grams + logistic regression: bigrams like "not good" help with negation; strong baseline.
- Fine-tuned transformers (BERT, RoBERTa, multilingual XLM-R): the state of the art, capturing context, negation scope and many subtle cues.
- LLMs zero-/few-shot: good general sentiment judgement without training data; useful for bootstrapping labels.
Hard phenomena#
| Phenomenon | Example | Why it is hard |
|---|---|---|
| Negation scope | "I don't think it's bad at all" | Double negation; scope beyond adjacent words |
| Sarcasm/irony | "Great, another two-hour queue." | Literal words are positive |
| Comparatives | "Better than the old clinic, worse than expected" | Sentiment is relative |
| Implicit sentiment | "The battery lasted three hours." | Facts that imply opinion via world knowledge |
| Domain dependence | "long" (battery life vs waiting time) | Polarity varies by domain |
| Code-mixing | "Service ta khub bhalo chilo, but slow" | Multiple languages in one sentence |
| Mixed sentiment | Pros and cons in one text | Single label is inadequate |
Aspect-based sentiment analysis (ABSA)#
ABSA extracts (aspect, opinion, polarity) triples:
"The doctor was kind but the waiting room was dirty." โ (doctor, kind, positive), (waiting room, dirty, negative).
Approaches: sequence labelling to find aspect terms, then classification of each aspect's polarity with the aspect as additional input (e.g. "[CLS] sentence [SEP] waiting room [SEP]" to a BERT classifier), or generative models that output the triples directly. For service improvement, ABSA is far more actionable than overall polarity.
from transformers import pipeline
clf = pipeline("sentiment-analysis", model="cardiffnlp/twitter-xlm-roberta-base-sentiment")
print(clf(["The registration process was quick and staff were helpful.",
"Waited 5 hours and nobody explained anything."]))Evaluation#
Use macro-F1 (neutral is often the hardest and most frequent class), per-aspect metrics for ABSA, and correlation for intensity. Most importantly, evaluate on your domain โ a model trained on movie reviews can perform poorly on health-service feedback or on another dialect.