CMP422 Data Science and Analytics

Data Science and AnalyticsUnit 811 min read

Text Analytics: NLP, Sentiment, Topic Modeling & Applications

Unit 8 of Data Science and Analytics explores how to extract meaning from unstructured text data using techniques like Natural Language Processing (NLP), sentiment analysis, topic modeling, and text classification, with real-world applications in Nepali and global tech.

TAKEAWAYS:

  • Text analytics transforms raw text (e.g., tweets, reviews, news) into structured data for analysis using NLP techniques like tokenization, stemming, and TF-IDF.
  • Sentiment analysis classifies opinions (positive/negative/neutral) in text, critical for customer feedback (e.g., Daraz reviews) or social media monitoring (e.g., Pathao driver complaints).
  • Topic modeling (LDA) identifies hidden themes in large text corpora, used by NEPSE for stock market news analysis or NTC for customer complaint categorization.
  • Text classification (Naive Bayes, SVM) automates tagging (e.g., spam detection in eSewa messages or news categorization by Kantipur).
  • Preprocessing (cleaning, normalization) is 50% of the work—noise in text (emojis, slang, misspellings) skews results.
  • Ethical challenges (bias, privacy) in text analytics require careful handling of sensitive data (e.g., medical records or political speeches).

1. Introduction to Text Analytics

Text analytics is the process of deriving high-quality information from unstructured text data using computational linguistics, statistics, and machine learning. Unlike structured data (tables, databases), text data (emails, social media, documents) lacks a predefined schema, making it harder to analyze. Key applications include:

  • Customer feedback analysis (e.g., Daraz product reviews).
  • Social media monitoring (e.g., Pathao’s driver complaints on Twitter).
  • News sentiment tracking (e.g., Kantipur’s political headline analysis).
  • Automated content tagging (e.g., YouTube’s video categorization).

Why Text Analytics?

  • 80% of enterprise data is unstructured (Gartner).
  • Nepali challenges: Low-resource language (few pre-trained models), code-mixing (English + Nepali), and informal text (e.g., "k tyo garxa" in WhatsApp).
  • Business impact: Reduces manual work (e.g., Ncell’s customer service chatbots) and uncovers trends (e.g., NEPSE’s market sentiment).

2. Core Techniques in Text Analytics

A. Text Preprocessing

Before analysis, text must be cleaned and normalized. Key steps:

  1. Tokenization: Splitting text into words/tokens.
    • Example: "I love Nepal!" → ["I", "love", "Nepal", "!"]
  2. Stopword Removal: Removing common words (e.g., "the", "is").
  3. Stemming/Lemmatization: Reducing words to root form.
    • Stemming: "running" → "run" (aggressive).
    • Lemmatization: "better" → "good" (context-aware).
  4. Handling Nepali Text:
    • Devanagari script processing (e.g., "नेपाल" → "नेपा" via stemming).
    • Code-mixing: "I’m going to Kathmandu tomorrow" → split into Nepali/English tokens.
graph LR
    A["Raw Text"] --> B["Tokenization"]
    B --> C["Stopword Removal"]
    C --> D["Stemming/Lemmatization"]
    D --> E["Normalization"]
    E --> F["Structured Data"]

B. Feature Extraction

Converts text into numerical features for ML models.

Technique Description Example Use Case
Bag of Words (BoW) Counts word frequencies (ignores grammar). Spam detection in eSewa emails.
TF-IDF Weighs words by importance (high TF-IDF = rare but relevant). News article topic modeling for Kantipur.
Word Embeddings Represents words as vectors (e.g., Word2Vec, GloVe). Sentiment analysis for Daraz reviews.
N-grams Considers word sequences (e.g., bigrams: "machine learning"). Detecting sarcasm in Twitter posts.

Worked Example: TF-IDF for Nepali News Suppose we analyze 3 NEPSE-related articles:

  • Article 1: "नेपालको शेयर बाजार मन्दीमा छ।"
  • Article 2: "बाजारमा वृद्धि भएको छ।"
  • Article 3: "मन्दीले गरीबलाई प्रभावित गरेको छ।"

TF-IDF Calculation:

  • "बाजार" appears in 2/3 articles → lower IDF (common).
  • "मन्दी" appears in 2/3 articles → lower IDF.
  • "गरीब" appears in 1/3 articles → higher IDF (rare but relevant).

3. Sentiment Analysis

Classifies text polarity (positive, negative, neutral). Used by:

  • eSewa: Analyzing customer complaints about failed transactions.
  • Pathao: Monitoring driver reviews for service quality.
  • NTC: Gauging public sentiment about internet shutdowns.

How It Works

  1. Lexicon-Based: Uses predefined word sentiment scores (e.g., "great" = +2, "terrible" = -2).
    • Challenge: Nepali slang ("कुरा छैन" = negative) lacks standardized lexicons.
  2. Machine Learning:
    • Naive Bayes: Fast but assumes word independence.
    • SVM: Better for high-dimensional text (e.g., long reviews).
    • Deep Learning (LSTM): Captures context (e.g., "not bad" → positive).

Worked Example: Daraz Review Sentiment Review: "Product quality is good but delivery took 5 days. Not happy with service."

  • Tokenized: ["product", "quality", "good", "delivery", "took", "5", "days", "not", "happy", "service"]
  • Sentiment Scores (lexicon-based):
    • "good" = +1, "happy" = +1, "not happy" = -1.5, "took" (context: delay) = -0.8
  • Final Score: (-1.5) + (-0.8) + (+1) + (+1) = -0.3 → Negative.
classDiagram
    class Review {
        +text: str
        +sentiment_score: float
    }
    class SentimentAnalyzer {
        +analyze(text) float
        -preprocess(text) list
        -classify(features) str
    }
    Review --> SentimentAnalyzer : "uses"

4. Topic Modeling

Identifies abstract "topics" in a text corpus. Used by:

  • NEPSE: Extracting themes from stock market news (e.g., "inflation", "FDI").
  • NTC: Categorizing customer complaints (e.g., "internet speed", "billing errors").
  • Kantipur: Automating news headline clustering.

Latent Dirichlet Allocation (LDA)

  • Assumption: Each document is a mix of topics; each topic is a mix of words.
  • Example: Analyzing 100 NEPSE news articles → 5 topics:
    1. Economic Growth (words: "GDP", "inflation", "remittance")
    2. Political Instability (words: "election", "protest", "government")
    3. Technology (words: "Fintech", "blockchain", "Khalti")

Worked Example: LDA on Nepali News Input: 50 articles about Nepal’s economy. Output:

  • Topic 1 (Weight: 40%): "आर्थिक वृद्धि", "रोजगारी", "उत्पादन"
  • Topic 2 (Weight: 30%): "महंगाई", "वित्तीय संकट", "ब्याज दर"

5. Text Classification

Assigns predefined labels to text (e.g., spam/ham, topic categories). Used by:

  • eSewa: Filtering spam transaction alerts.
  • Khalti: Categorizing customer support tickets.
  • YouTube: Tagging Nepali music videos vs. tutorials.

Algorithms Compared

Algorithm Pros Cons Best For
Naive Bayes Fast, works well with high dimensions Assumes feature independence Spam detection, sentiment analysis
SVM Effective in high-dimensional spaces Slower training Text classification with clear margins
Random Forest Handles non-linear relationships Prone to overfitting Multi-class classification
Deep Learning (CNN/RNN) Captures context/syntax Needs large data, computationally heavy Complex tasks (e.g., sarcasm detection)

Worked Example: Spam Detection in eSewa Messages

  • Training Data:
    • Spam: "URGENT: Claim your free Khalti cash! Click here."
    • Ham: "Your transaction of Rs. 500 to ABC Store is successful."
  • Features: TF-IDF of words like "free", "click", "urgent", "transaction".
  • Model: Naive Bayes classifier trained on labeled data.
  • Prediction: New message "Win Rs. 10,000! Reply STOP" → Spam (high score for "win", "reply").

6. Challenges in Nepali Text Analytics

Challenge Solution
Low-resource language Use transfer learning (fine-tune English models on Nepali data).
Code-mixing Tokenize at script level (Devanagari vs. Latin).
Informal language Augment datasets with slang (e.g., "k tyo garxa" → "okay").
Lack of labeled data Use weak supervision (e.g., emojis as sentiment labels).

7. Tools and Libraries

Tool/Library Purpose Example Use Case
NLTK Text preprocessing, tokenization Building a Nepali sentiment analyzer
spaCy Industrial-strength NLP pipeline eSewa’s customer feedback system
Gensim Topic modeling (LDA, LSI) NEPSE news topic extraction
Hugging Face Pre-trained transformers (BERT, RoBERTa) Fine-tuning for Nepali language
Python (scikit-learn) ML classification (SVM, Naive Bayes) Daraz review categorization

## In the real world

  1. eSewa’s Fraud Detection

    • Idea Used: Text classification + TF-IDF
    • How: eSewa’s ML model flags suspicious transaction messages (e.g., "URGENT: Verify your account") by analyzing word patterns. TF-IDF helps identify rare but risky phrases like "click this link" in Nepali.
  2. Pathao’s Driver Satisfaction Dashboard

    • Idea Used: Sentiment analysis + N-grams
    • How: Pathao scrapes Twitter/X for mentions like "@PathaoNepal" and uses sentiment analysis to track driver complaints (e.g., "My app crashed 3 times today" → negative). N-grams capture phrases like "bad service" for better accuracy.
  3. NEPSE’s Market Sentiment Index

    • Idea Used: Topic modeling (LDA) + Sentiment Analysis
    • How: NEPSE’s algorithm analyzes news headlines (e.g., "Stocks surge on FDI news") to extract topics ("FDI", "inflation") and sentiment scores. A sudden spike in negative sentiment about "government policies" triggers alerts for traders.

## Exam Tip

  1. Focus on Preprocessing:

    • Exams often test tokenization, stemming, and TF-IDF. Practice: Given a Nepali sentence, show step-by-step preprocessing (e.g., "मेरो खाता ब्लक भएको छ" → tokens → remove stopwords → stem to "खाता ब्लक").
  2. Compare Algorithms:

    • Know when to use Naive Bayes (fast, high-dimensional) vs. SVM (better margins) vs. Deep Learning (context-aware). Example: "Why would you use LSTM for sentiment analysis in Daraz reviews?"
  3. Real-World Applications:

    • Always relate answers to Nepali contexts (e.g., "How would you analyze NTC customer complaints using topic modeling?").
    • Common Exam Questions:
      • "Explain how TF-IDF works with an example from Nepali news."
      • "Compare lexicon-based and ML-based sentiment analysis for Pathao reviews."
      • "What challenges arise when applying BERT to Nepali text? How would you address them?"
  4. Visuals Matter:

    • Draw pipelines (e.g., text preprocessing steps) and tables (e.g., algorithm comparisons). Label every component clearly.
  5. Ethics and Bias:

    • Discuss how sentiment analysis can be biased (e.g., favoring formal Nepali over slang) and how to mitigate it (e.g., balanced datasets).

Final Note: Text analytics is 50% preprocessing, 30% feature engineering, and 20% modeling. Master the basics (tokenization, TF-IDF, Naive Bayes) before diving into deep learning. For Nepali-specific questions, always consider code-mixing, slang, and limited datasets as key challenges.

Based on the PU BE Computer (PU) syllabus for Data Science and Analytics (CMP422), unit 8.

Discussion

Loading…