Artificial IntelligenceUnit 917 min read
NLP: Text Processing, Chatbots, and Language Models
Unit 9 of Artificial Intelligence explores Natural Language Processing (NLP), covering text preprocessing, tokenization, parsing, semantic analysis, and real-world applications like chatbots and translation systems. Learn how machines understand human language, from rule-based systems to deep learning models like trans
TAKEAWAYS:
- NLP bridges human language and machine understanding through tokenization, parsing, and semantic analysis, enabling applications like chatbots and sentiment analysis.
- Rule-based systems (e.g., syntax parsers) rely on predefined grammars, while statistical/machine learning models (e.g., transformers) learn patterns from data.
- Text preprocessing (cleaning, tokenization, stemming) is critical before any NLP task, as raw text is noisy and unstructured.
- Semantic analysis (word embeddings, sentiment analysis) captures meaning, while syntax parsing structures sentences grammatically.
- Real-world NLP applications include chatbots (e.g., eSewa’s customer support), translation (Google Translate), and search engines (Google’s ranking algorithms).
- Ethical challenges in NLP include bias in training data, privacy concerns (e.g., voice assistants), and misinformation spread via automated text generation.
1. Introduction to Natural Language Processing (NLP)
NLP is a subfield of AI that enables machines to understand, interpret, and generate human language. It combines:
- Linguistics (grammar, syntax, semantics),
- Computer science (algorithms, machine learning),
- Statistics (probabilistic modeling).
Why NLP Matters
- Human-computer interaction: Voice assistants (e.g., Google Assistant, Siri), chatbots (e.g., eSewa’s automated customer service).
- Information extraction: Summarizing news articles (e.g., Google News), analyzing social media trends (e.g., Facebook’s sentiment analysis).
- Automation: Email filtering (e.g., Gmail’s spam detection), legal document review (e.g., AI tools used by law firms).
Challenges in NLP
- Ambiguity: Words/sentences can have multiple meanings (e.g., "bank" as a financial institution or river edge).
- Context dependence: Meaning changes with context (e.g., "I shot an elephant in my pajamas" – was it a hunting trip or a dream?).
- Noisy data: Typos, slang, sarcasm, and cultural nuances complicate processing.
2. Text Preprocessing: Cleaning and Structuring Data
Before machines can analyze text, it must be cleaned and structured. Key steps:
A. Text Normalization
Convert text to a standardized form:
- Lowercasing: "Hello" → "hello" (to avoid case sensitivity issues).
- Removing punctuation: "Hello!" → "Hello".
- Expanding contractions: "don’t" → "do not".
- Correcting spelling: "teh" → "the" (using tools like
textbloborspacy).
B. Tokenization
Splitting text into words, phrases, or symbols (tokens).
Example:
Input: "I love NLP! It's amazing."
Output: ["I", "love", "NLP", "!", "It", "'s", "amazing", "."]
Types of tokenization:
| Type | Example | Use Case |
|---|---|---|
| Word-level | "I love NLP" → ["I", "love", "NLP"] | Most NLP tasks |
| Subword (BPE) | "running" → ["run", "##ning"] | Handling rare words (e.g., transformers) |
| Character-level | "cat" → ["c", "a", "t"] | Low-resource languages |
C. Stemming and Lemmatization
Reduce words to their base or root form:
- Stemming (aggressive): "running" → "run" (using Porter Stemmer).
- Lemmatization (accurate): "running" → "run" (using WordNet or spaCy).
Worked Example:
| Word | Stemming (Porter) | Lemmatization (spaCy) |
|---|---|---|
| "happiness" | "happi" | "happiness" |
| "better" | "bet" | "good" |
3. Syntax Analysis: Parsing Sentences
Syntax analysis structures sentences grammatically using parse trees. This helps machines understand relationships between words.
A. Part-of-Speech (POS) Tagging
Assign grammatical labels to words:
Example:
Sentence: "The quick brown fox jumps over the lazy dog."
POS tags: ["DET", "ADJ", "ADJ", "NOUN", "VERB", "PREP", "DET", "ADJ", "NOUN", "."]
Tools: spaCy, NLTK, Stanford CoreNLP.
B. Dependency Parsing
Shows how words relate to each other in a sentence.
Example:
Sentence: "Nepal won the match."
Dependency tree:
won
/ \
Nepal match
/
the
- "Nepal" is the subject of "won".
- "match" is the object of "won".
Visualization:
graph TD
won["won"]
nepal["Nepal"] -->|"nsubj"| won
match["match"] -->|"dobj"| won
the["the"] -->|"det"| matchC. Constituency Parsing
Groups words into phrases (noun phrases, verb phrases).
Example:
Sentence: "AI improves healthcare."
Parse tree:
S
/ \
NP VP
| / \
AI V NP
| |
improves healthcare
4. Semantic Analysis: Understanding Meaning
Syntax tells us how words are arranged; semantics tells us what they mean.
A. Word Embeddings
Convert words into vector representations (e.g., 300-dimensional vectors) where similar words are close in space. Example: Word2Vec, GloVe, FastText.
- "king" – "man" + "woman" ≈ "queen" (vector arithmetic).
Visualization of word embeddings (2D projection):
graph TD
king["king"] -->|"similar"| queen["queen"]
man["man"] -->|"similar"| woman["woman"]
paris["Paris"] -->|"similar"| france["France"]B. Sentiment Analysis
Classify text as positive, negative, or neutral.
Example: Analyzing a product review:
Input: "The phone battery is terrible!"
Output: Negative (using VADER, TextBlob, or BERT).
Worked Example (Rule-Based Sentiment):
| Word | Sentiment Score | Polarity |
|---|---|---|
| "terrible" | -2.5 | Negative |
| "battery" | 0.1 | Neutral |
| "phone" | 0.5 | Positive |
| Total | -2.5 + 0.1 + 0.5 = -1.9 | Negative |
C. Named Entity Recognition (NER)
Identify real-world entities in text (people, places, organizations).
Example:
Text: "Pathao is headquartered in Singapore."
Entities:
- Organization: Pathao
- Location: Singapore
Tools: spaCy, Stanford NER, FLAN T5.
5. Machine Learning in NLP
Traditional NLP used rule-based systems, but modern NLP relies on machine learning.
A. Traditional ML Approaches
- Bag-of-Words (BoW): Represents text as word frequency counts (ignores order).
Example:
"I love NLP"→{I:1, love:1, NLP:1}. - TF-IDF: Weighs words by importance (common words like "the" get lower scores).
- Naive Bayes: Classifies text (e.g., spam detection) using probability.
B. Deep Learning in NLP
Modern NLP uses neural networks to learn from large datasets.
1. Recurrent Neural Networks (RNNs)
Process sequential data (e.g., sentences).
- Problem: Vanishing gradients (forgets long-term dependencies).
- Solution: LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit).
2. Transformers
State-of-the-art models (e.g., BERT, GPT-3) use self-attention to weigh word importance dynamically.
Example: BERT predicts missing words in a sentence:
Input: "The capital of Nepal is [MASK]."
Output: "Kathmandu" (with 95% confidence).
Transformer Architecture:
graph TD
Input["Input Tokens"] --> Encoder["Encoder Layers\n(Self-Attention)"]
Encoder --> Pooler["Pooler"]
Pooler --> Output["Output\n(Embeddings)"]
Output --> Decoder["Decoder Layers\n(Masked Self-Attention)"]
Decoder --> Final["Final Prediction"]3. Pretrained Language Models (PLMs)
Models trained on huge datasets (e.g., Wikipedia, books) and fine-tuned for tasks:
- BERT: Bidirectional Encoder Representations from Transformers (Google).
- GPT-3: Generative Pre-trained Transformer (OpenAI).
- T5: Text-to-Text Transfer Transformer (Google).
Example: Fine-tuning BERT for sentiment analysis:
- Train on labeled data (e.g., IMDB reviews).
- Use for new reviews:
"This movie was awful."→ Negative (98%).
6. Applications of NLP in Nepal and Globally
In Nepal
| Application | Example | NLP Technique Used |
|---|---|---|
| Chatbots | eSewa customer support | Rule-based + ML (e.g., Rasa) |
| News Summarization | Kantipur’s AI-generated summaries | Text ranking (e.g., TextRank) |
| Loan Approval | NMB Bank’s automated processing | NER + sentiment analysis (risk assessment) |
| Traffic Management | NTC’s voice-based route queries | Speech-to-text + intent classification |
| Job Matching | Daraz’s resume screening | Keyword extraction + NER |
Global Examples
| Application | Example | NLP Technique Used |
|---|---|---|
| Translation | Google Translate | Neural Machine Translation (NMT) |
| Voice Assistants | Siri, Alexa | Speech recognition + NLP |
| Search Engines | Google Search | Query understanding + ranking (BERT) |
| Automated Journalism | Associated Press sports reports | Template-based + NER |
| Customer Support | WhatsApp Business API | Intent classification + chatbots |
7. Ethical and Societal Impact of NLP
Challenges
- Bias in Data: Models trained on biased data (e.g., gender bias in hiring chatbots).
- Privacy Concerns: Voice assistants (e.g., Google Home) record conversations.
- Misinformation: Deepfakes and AI-generated fake news (e.g., election interference).
- Job Displacement: Automation of customer service roles.
Solutions
- Fairness-aware ML: Audit datasets for bias (e.g., Google’s What-If Tool).
- Differential Privacy: Protect user data (e.g., Apple’s on-device processing).
- Regulation: Governments must set guidelines (e.g., EU’s GDPR).
In the Real World
eSewa’s Chatbot
- Idea Used: Intent classification + rule-based responses.
- How It Works:
- User: "How do I reset my password?"
- NLP model detects intent ("password reset") and triggers a predefined response with a link.
- Real Example: If you type "I forgot my eSewa PIN" in their chatbot, it asks for your email and sends a reset link—no human agent needed.
Pathao’s Driver-User Communication
- Idea Used: Speech-to-text + sentiment analysis.
- How It Works:
- When a user reports a driver issue (e.g., "The driver is late" or "He’s rude!"), Pathao’s NLP system:
- Converts speech to text.
- Uses sentiment analysis to flag complaints.
- Triggers automated penalties (e.g., deactivating the driver temporarily).
- Worked Example:
- Input (audio): "The driver stopped in the middle of the road!"
- Text:
"The driver stopped in the middle of the road!" - Sentiment: Negative (anger detected)
- Action: Driver rated poorly; Pathao sends a warning.
- When a user reports a driver issue (e.g., "The driver is late" or "He’s rude!"), Pathao’s NLP system:
Ncell’s Customer Support Automation
- Idea Used: Named Entity Recognition (NER) + dialogue management.
- How It Works:
- User calls Ncell IVR: "I want to upgrade my plan to 4G."
- NLP extracts:
- Entity: "4G" (product), "upgrade" (action).
- System responds: "Your current plan is 3G. Upgrading to 4G costs NPR 500/month. Proceed? (Yes/No)".
- Real Scenario: If you call and say "My phone has no signal in Kathmandu", NLP identifies:
- Issue: "no signal"
- Location: "Kathmandu"
- Action: Route to technical support or offer a signal booster.
Daraz’s Order Processing
- Idea Used: Text classification + keyword extraction.
- How It Works:
- When a customer messages: "My order #12345 is delayed. What’s the status?"
- NLP:
- Extracts order ID ("12345") via NER.
- Classifies intent ("order status").
- Fetches data from the database and replies: "Your order is out for delivery. ETA: 2 hours."
- Worked Example:
Customer Message NLP Action System Response "Cancel order #67890" Extracts order ID, intent="cancel" "Order #67890 canceled. Refund processed." "Where is my shipment?" Intent="track" "Your shipment is in transit (Kathmandu → Pokhara)."
NEPSE’s Stock Market Summaries
- Idea Used: Text summarization + keyword extraction.
- How It Works:
- NEPSE publishes daily reports like: "The stock market rose by 2% today due to high demand in the banking sector."
- NLP tools (e.g., TextRank) summarize longer reports into: "NEPSE up 2% today. Banking stocks led gains (e.g., NMB +3%, Global IME +2%)."
- Real Use: Investors get quick insights without reading full reports.
Exam Tip
What to Expect in TU/PU Exams
Theory Questions (30-40%)
- Define tokenization vs. stemming vs. lemmatization.
- Explain POS tagging vs. dependency parsing.
- Compare BoW vs. TF-IDF vs. word embeddings.
- Describe how transformers work (self-attention mechanism).
Short Problems (30-40%)
- Text preprocessing: Given a sentence, perform tokenization/stemming. Example: "The quick brown foxes are jumping!" → Tokenize and stem.
- Sentiment analysis: Classify a sentence as positive/negative/neutral. Example: "This laptop is fast and reliable." → Positive.
- NER extraction: Identify entities in a sentence. Example: "Ncell launched 5G in Kathmandu." → Org: Ncell, Location: Kathmandu, Event: 5G launch.
Long Problems (20-30%)
- Design an NLP pipeline for a given task (e.g., chatbot for a bank).
Steps:
- Preprocess text (tokenize, clean).
- Use NER to extract account numbers.
- Classify intent (e.g., "balance check," "transfer").
- Generate response (rule-based or ML).
- Evaluate an NLP model: Given precision/recall, calculate F1-score.
Example:
- Precision = 0.8, Recall = 0.7 → F1 = .
- Design an NLP pipeline for a given task (e.g., chatbot for a bank).
Steps:
Case Studies (10-20%)
- Explain how eSewa’s chatbot uses NLP (intent classification + rule-based responses).
- Discuss bias in NLP models (e.g., gender bias in hiring chatbots).
How to Score Full Marks
- For definitions: Use official terms (e.g., "tokenization splits text into words," not "breaking sentences").
- For examples: Always tie to real-world apps (e.g., "Like Pathao’s complaint system").
- For diagrams: Draw parse trees or transformer architectures clearly.
- For calculations: Show every step (e.g., TF-IDF formula, F1-score derivation).
- For ethical questions: Mention bias, privacy, and regulation (e.g., "EU’s GDPR protects user data").
Practice Questions for Self-Assessment
Tokenization: Given the sentence: "AI is transforming healthcare! #FutureTech"
- Perform word-level tokenization.
- Apply stemming (Porter Stemmer) to the tokens.
POS Tagging: Tag the following sentence: "Nepal won the match against India."
Sentiment Analysis: Classify the following sentences using a rule-based approach (assign +1, 0, or -1):
- "The service was excellent!"
- "I hate waiting in long queues."
- "The product is okay."
NER Extraction: Identify entities in: "Daraz delivered my order to my home in Lalitpur."
Short Answer:
- What is the difference between lemmatization and stemming?
- How does BERT improve over traditional word embeddings like Word2Vec?
Recommended Tools to Try
| Tool | Purpose | Link |
|---|---|---|
| spaCy | POS tagging, NER, dependency parsing | spacy.io |
| NLTK | Text preprocessing, tokenization | nltk.org |
| Hugging Face | Transformers (BERT, GPT) | huggingface.co |
| Google Colab | Run NLP experiments (free GPU) | colab.research.google.com |
| TextBlob | Simple sentiment analysis | textblob.readthedocs.io |
Final Summary
| Concept | Key Idea | Real-World Example |
|---|---|---|
| Tokenization | Splitting text into words/symbols. | eSewa’s chatbot processes user queries word-by-word. |
| POS Tagging | Labeling words as nouns, verbs, etc. | Google Search understands query structure. |
| NER | Extracting entities (names, places). | Pathao identifies driver/user complaints. |
| Sentiment Analysis | Classifying text as positive/negative. | Daraz reviews determine product ratings. |
| Transformers | Self-attention for context-aware language models. | Google Translate uses BERT for accuracy. |
| Ethical NLP | Addressing bias and privacy. | EU’s GDPR protects user data in chatbots. |
Based on the TU BIM syllabus for Artificial Intelligence (IT228), unit 9.
Discussion
Loading…