CMP346 Artificial Intelligence

Artificial IntelligenceUnit 916 min read

NLP: Text Processing, ML Models & Applications

Unit 9 of Artificial Intelligence explores Natural Language Processing (NLP), covering text preprocessing, machine learning models (e.g., word embeddings, transformers), and real-world applications like chatbots, sentiment analysis, and machine translation.

TAKEAWAYS:

  • NLP bridges linguistics and AI to enable machines to understand and generate human language.
  • Text preprocessing (tokenization, stemming, POS tagging) is critical for model input quality.
  • Word embeddings (Word2Vec, GloVe) and transformers (BERT) capture semantic meaning in text.
  • Key NLP tasks include sentiment analysis, machine translation, and question answering.
  • Applications range from chatbots (e.g., WhatsApp Business) to financial NLP (e.g., NEPSE stock analysis).
  • Ethical challenges (bias, privacy) must be addressed in NLP system design.

1. Introduction to NLP

Natural Language Processing (NLP) is a subfield of AI that enables machines to interpret, understand, and generate human language. It combines:

  • Linguistics (grammar, syntax, semantics)
  • Computer Science (algorithms, ML)
  • Domain Knowledge (e.g., finance for NEPSE reports, healthcare for medical records)

Why NLP Matters

NLP powers:

  • Virtual assistants (e.g., Google Assistant, Siri)
  • Customer support (e.g., Daraz’s chatbots)
  • Financial analysis (e.g., NEPSE stock news sentiment)
  • Healthcare (e.g., diagnosing symptoms from patient reports)

Key Challenges in NLP

Challenge Example Solution Approach
Ambiguity "Bank" (financial vs. river) Contextual embeddings (BERT)
Noise (typos, slang) "Thx" → "Thanks" Text normalization (stemming, lemmatization)
Multilingual support Nepali → English translation Multilingual BERT, sequence-to-sequence models
Cultural context Nepali proverbs vs. English idioms Domain-specific fine-tuning

2. Text Preprocessing

Before feeding text into models, we clean and structure it. Key steps:

A. Tokenization

Splitting text into meaningful units (words, subwords, or characters). Example: Input: "I love NLP!" Output: ["I", "love", "NLP", "!"]

flowchart LR
    A["Raw Text: 'I love NLP!'"]
    B["Tokenization"]
    C["Tokens: ['I', 'love', 'NLP', '!']"]
    A --> B --> C

B. Normalization

  • Lowercasing: Convert all text to lowercase ("Hello" → "hello").
  • Removing punctuation: "NLP!" → "NLP".
  • Expanding contractions: "don't" → "do not".

C. Stemming vs. Lemmatization

Technique Example (Input: "running") Output Notes
Stemming Porter Stemmer "run" Aggressive (may produce non-words)
Lemmatization WordNet Lemmatizer "run" Uses vocabulary (slower but accurate)

Worked Example: Input: "The cats are running fast."

  1. Tokenize: ["The", "cats", "are", "running", "fast", "."]
  2. Lowercase: ["the", "cats", "are", "running", "fast", "."]
  3. Remove punctuation: ["the", "cats", "are", "running", "fast"]
  4. Lemmatize: ["the", "cat", "be", "run", "fast"]

3. Word Representations

Machines don’t understand words directly; they need numerical representations.

A. Bag-of-Words (BoW)

Represents text as a vector of word counts (ignores order). Example: Text: "I love NLP" Vocabulary: {"I": 0, "love": 1, "NLP": 2} Vector: [1, 1, 1] (one-hot encoded)

Problem: High dimensionality (sparse vectors for large vocabularies).

B. TF-IDF (Term Frequency-Inverse Document Frequency)

Weighs words by importance:

  • TF: How often a word appears in a document.
  • IDF: How rare the word is across all documents.

Formula:

Example:

Word TF (in doc) IDF (across 100 docs) TF-IDF
"NLP" 2 3.0 6.0
"love" 1 1.0 1.0

C. Word Embeddings (Dense Vectors)

Maps words to low-dimensional dense vectors (e.g., 300D) capturing semantic meaning.

1. Word2Vec

  • Skip-gram: Predicts context words from a target word.
  • CBOW (Continuous Bag of Words): Predicts a target word from context.

Example (Skip-gram): Input: "The cat sat on the mat" Target: "cat" → Predict context: "the", "sat", "on", "the", "mat"

graph LR
    A["Input: 'cat'"] --> B["Context: 'the', 'sat', 'on', 'mat'"]
    B --> C["Word2Vec updates 'cat' vector"]

2. GloVe (Global Vectors)

Uses co-occurrence statistics from a corpus to learn embeddings. Advantage: Captures global word relationships (e.g., "king" - "man" + "woman" ≈ "queen").

3. FastText

Extends Word2Vec by using subword information (helpful for rare words). Example: Word: "unhappiness" → Subwords: "un", "happi", "ness" → Embedding is sum of subword vectors.


D. Contextual Embeddings (Transformers)

Unlike static embeddings (Word2Vec), context matters. Example Models:

  • BERT (Bidirectional Encoder Representations from Transformers): Uses self-attention to weigh word importance dynamically.
  • GPT (Generative Pre-trained Transformer): Predicts next words (used in chatbots).

BERT Example: Input: "The [MASK] sat on the mat." BERT predicts "cat" (not just based on word frequency but context).

graph TD
    A["Input Sentence"] --> B["Tokenized: ['The', 'cat', 'sat', 'on', 'the', 'mat', '.']"]
    B --> C["BERT Self-Attention"]
    C --> D["Contextual Embeddings"]
    D --> E["Predicts 'cat' for [MASK]"]

4. Key NLP Tasks

A. Sentiment Analysis

Classifies text as positive, negative, or neutral. Example (Nepali E-commerce): Input: "Daraz delivery was late, but the product was good." Output: Mixed sentiment (negative for delivery, positive for product).

Approach:

  1. Preprocess text (tokenize, remove stopwords).
  2. Use TF-IDF + Logistic Regression or BERT for fine-tuning.
  3. Train on labeled data (e.g., Nepali product reviews).

B. Machine Translation

Converts text from one language to another (e.g., Nepali → English). Example: Input (Nepali): "मेरो नाम सूर्य हो।" Output (English): "My name is Surya."

Approach:

  • Sequence-to-Sequence (Seq2Seq) Models:
    • Encoder: Converts input sequence to a context vector.
    • Decoder: Generates output sequence word-by-word.
flowchart LR
    A["Nepali Text"] --> B["Encoder (RNN/Transformer)"]
    B --> C["Context Vector"]
    C --> D["Decoder (Generates English)"]
    D --> E["English Text"]

Real-World Use:

  • Google Translate (uses Transformer-based models).
  • Nepali Wikipedia (automated translations for low-resource languages).

C. Named Entity Recognition (NER)

Identifies entities like names, dates, organizations. Example: Text: "Surya Nepal Airlines flew to Kathmandu on 2023-10-01." Output:

  • PERSON: "Surya"
  • ORGANIZATION: "Nepal Airlines"
  • DATE: "2023-10-01"

Approach:

  • Train a CRF (Conditional Random Field) or BERT-NER model.

D. Chatbots & Conversational AI

Simulates human conversation using:

  • Rule-based systems (if-else for FAQs).
  • ML-based (Rasa, Dialogflow for intent recognition).

Example (Khalti Customer Support): User: "My transaction failed." Bot:

  1. Detects intent: "transaction_failed".
  2. Extracts entities: {"amount": "500", "time": "10:30 AM"}.
  3. Responds: "We’ve detected an issue. Please retry or contact support at [number]."

5. Applications of NLP in Nepal

A. Financial NLP (NEPSE, Banks)

  • Sentiment Analysis of Stock News: Input: "NEPSE index drops due to global uncertainty." Output: Negative sentiment → Algorithmic trading signals.
  • Loan Application Analysis: Banks (e.g., Global IME) use NLP to extract key details from handwritten forms.

B. Healthcare (e.g., Kathmandu Model Hospital)

  • Symptom Checker Chatbots: Input: "I have a fever and cough." Output: "You may have flu. Please consult a doctor." (powered by MedQA models).

C. E-Commerce (Daraz, Swoyambu)

  • Product Description Analysis:
    • Auto-tagging products (e.g., "Nike shoes" → tags: sports, footwear).
    • Recommendation systems (users who bought X also bought Y).

D. Traffic & Urban Planning (NTC, Kathmandu Metro)

  • Twitter/News Analysis for Traffic Predictions: Input: "Heavy rain expected in Kathmandu tomorrow." Output: Increased traffic delay risk → NTC adjusts signal timings.

6. Ethical Considerations in NLP

Issue Example Mitigation Strategy
Bias Gender bias in hiring chatbots Audit datasets, use diverse training data
Privacy WhatsApp Business storing chats Anonymize data, GDPR-compliant policies
Misinformation Fake news on Facebook Fact-checking NLP models (e.g., Google Fact Check)
Job Displacement NLP replacing customer service Reskill workers for AI-assisted roles

7. Exam Tip: How to Score Full Marks

A. For Theory Questions

  • Define clearly: Always start with a one-sentence definition (e.g., "Word Embeddings are dense vector representations of words that capture semantic meaning.").
  • Compare models: Use tables (e.g., BoW vs. Word2Vec vs. BERT).
  • Mention real-world use: Link to Nepali examples (e.g., Khalti chatbots, NEPSE sentiment analysis).

B. For Practical Questions

  1. Show preprocessing steps (tokenization → normalization → lemmatization).
  2. Draw diagrams for:
    • Tokenization pipelines.
    • Encoder-decoder architectures (for translation).
    • Attention mechanisms (for transformers).
  3. Use small datasets in examples (e.g., 3-5 sentences for sentiment analysis).

C. Common Pitfalls to Avoid

  • ❌ Ignoring context: Always explain why BERT is better than BoW for tasks like sentiment analysis.
  • ❌ Overcomplicating: For short-answer questions, stick to key points (e.g., "TF-IDF weighs rare words higher").
  • ❌ Assuming perfect data: Mention noise handling (e.g., "Real-world text has typos, so we use spell-checking before tokenization.").

D. Sample Exam Question & Answer

Question: "Explain how BERT improves over Word2Vec for sentiment analysis. Provide a Nepali example."

Model Answer: BERT (Bidirectional Encoder Representations from Transformers) improves over Word2Vec in sentiment analysis through:

  1. Bidirectional Context:

    • Word2Vec processes words left-to-right or right-to-left (limited context).
    • BERT uses self-attention to weigh all words in a sentence dynamically. Example:
    • Word2Vec: "Not good" → "good" gets a positive embedding (ignores "Not").
    • BERT: Detects "Not" negates "good" → Correctly predicts negative sentiment.
  2. Deep Contextual Understanding:

    • Word2Vec: "Bank" (financial vs. river) → Same vector.
    • BERT: Uses surrounding words (e.g., "I deposited money in the bank." → financial context).

Nepali Example (Daraz Review): Input: "Product quality was not good, but delivery was fast."

  • Word2Vec: May misclassify due to "not good" ambiguity.
  • BERT: Correctly identifies mixed sentiment (negative for quality, positive for delivery).

8. Worked Example: Sentiment Analysis Pipeline

Task: Classify Nepali restaurant reviews as positive/negative.

Step 1: Data Collection

Review (Nepali) Label
"खाना स्वादिष्ट थियो।" Positive
"सेवा ढिलो थियो।" Negative
"ठूलो समस्या भएन।" Neutral

Step 2: Preprocessing

  1. Tokenize: "सेवा ढिलो थियो।" → ["सेवा", "ढिलो", "थियो", "."]
  2. Normalize:
    • Remove punctuation: ["सेवा", "ढिलो", "थियो"]
    • Lemmatize: ["सेवा", "ढिलो", "हुनु"]
  3. TF-IDF Vectorization:
    • "सेवा" → High TF-IDF (common in negative reviews).
    • "स्वादिष्ट" → High TF-IDF (positive).

Step 3: Model Training

  • Approach 1: TF-IDF + Logistic Regression.
  • Approach 2: Fine-tuned mBERT (multilingual BERT).

Step 4: Prediction

Input: "भातको स्वाद राम्रो थियो तर दाम महंगो थियो।"

  • BERT predicts: Mixed sentiment (positive for taste, negative for price).

9. Real-World NLP Systems

A. WhatsApp Business API (Nepal)

  • Use Case: Small businesses (e.g., Kathmandu tailors) automate replies.
  • NLP Idea: Intent classification (e.g., "order status" vs. "return request").
  • How It Works:
    1. User: "My order #12345 is delayed."
    2. NLP detects intent: "order_delay".
    3. Bot: "We’re checking. Expected delivery: tomorrow."

B. Google Translate (Nepali-English)

  • Use Case: Tourists, business emails.
  • NLP Idea: Transformer-based Seq2Seq.
  • Example: Input (Nepali): "मेरो नाम सूर्य हो।" Output (English): "My name is Surya."

C. NEPSE Stock News Analysis

  • Use Case: Investors track sentiment in news articles.
  • NLP Idea: Sentiment analysis + keyword extraction.
  • Example: Input: "NEPSE gains 2% as global markets recover." Output:
    • Sentiment: Positive
    • Keywords: NEPSE, gain, global markets
    • Action: Buy signal for algorithms.

10. Summary Table: NLP Techniques

Technique Use Case Pros Cons
BoW Simple text classification Fast, easy to implement Ignores word order/meaning
TF-IDF Document similarity Handles rare words well Still sparse vectors
Word2Vec Word similarity tasks Captures semantic meaning No context awareness
BERT Sentiment, QA, translation State-of-the-art accuracy Computationally expensive
Seq2Seq Machine translation Handles variable-length text Needs large datasets

11. Key Formulas to Remember

  1. TF-IDF:
  2. Softmax (for classification):
  3. Attention Score (BERT):

12. Visualizing NLP Concepts

A. Tokenization Pipeline

flowchart LR
    A["Raw Text: 'I love NLP!'"]
    B["Lowercase: 'i love nlp!'"]
    C["Remove Punctuation: 'i love nlp'"]
    D["Tokenize: ['i', 'love', 'nlp']"]
    E["Lemmatize: ['i', 'love', 'nlp']"]
    A --> B --> C --> D --> E

B. BERT Self-Attention

graph TD
    A["Input: 'The cat sat on the mat'"] --> B["Token Embeddings"]
    B --> C["Positional Encoding"]
    C --> D["Self-Attention Layers"]
    D --> E["Contextualized Output"]
    E --> F["Predicts: 'cat' for masked word"]

C. Seq2Seq for Translation

sequenceDiagram
    participant Encoder
    participant Decoder
    Encoder->>Decoder: Context Vector (Nepali)
    Decoder->>Decoder: Generates English word-by-word
    Decoder-->>User: "My name is Surya."

13. Common Mistakes in Exams

  1. Assuming all NLP tasks need deep learning:
    • BoW/TF-IDF works for simple classification (e.g., spam detection).
  2. Ignoring preprocessing:
    • Always mention tokenization, normalization, and stopword removal.
  3. Overlooking ethical issues:
    • Exams may ask: "How would you handle bias in a Nepali sentiment analysis model?"
  4. Confusing Word2Vec and BERT:
    • Word2Vec = static embeddings; BERT = contextual embeddings.

Tool/Library Purpose Example Use Case
NLTK Text preprocessing Tokenization, stemming
spaCy Industrial-strength NLP Named Entity Recognition
HuggingFace Pre-trained models (BERT, GPT) Fine-tuning for Nepali text
TensorFlow/NLP Deep learning for NLP Building custom Seq2Seq models
Google Colab Free GPU for training models Experimenting with BERT

15. Further Reading

Based on the PU BE Computer (PU) syllabus for Artificial Intelligence (CMP346), unit 9.

Discussion

Loading…