Artificial IntelligenceUnit 916 min read
NLP: Text Processing, ML Models & Applications
Unit 9 of Artificial Intelligence explores Natural Language Processing (NLP), covering text preprocessing, machine learning models (e.g., word embeddings, transformers), and real-world applications like chatbots, sentiment analysis, and machine translation.
TAKEAWAYS:
- NLP bridges linguistics and AI to enable machines to understand and generate human language.
- Text preprocessing (tokenization, stemming, POS tagging) is critical for model input quality.
- Word embeddings (Word2Vec, GloVe) and transformers (BERT) capture semantic meaning in text.
- Key NLP tasks include sentiment analysis, machine translation, and question answering.
- Applications range from chatbots (e.g., WhatsApp Business) to financial NLP (e.g., NEPSE stock analysis).
- Ethical challenges (bias, privacy) must be addressed in NLP system design.
1. Introduction to NLP
Natural Language Processing (NLP) is a subfield of AI that enables machines to interpret, understand, and generate human language. It combines:
- Linguistics (grammar, syntax, semantics)
- Computer Science (algorithms, ML)
- Domain Knowledge (e.g., finance for NEPSE reports, healthcare for medical records)
Why NLP Matters
NLP powers:
- Virtual assistants (e.g., Google Assistant, Siri)
- Customer support (e.g., Daraz’s chatbots)
- Financial analysis (e.g., NEPSE stock news sentiment)
- Healthcare (e.g., diagnosing symptoms from patient reports)
Key Challenges in NLP
| Challenge | Example | Solution Approach |
|---|---|---|
| Ambiguity | "Bank" (financial vs. river) | Contextual embeddings (BERT) |
| Noise (typos, slang) | "Thx" → "Thanks" | Text normalization (stemming, lemmatization) |
| Multilingual support | Nepali → English translation | Multilingual BERT, sequence-to-sequence models |
| Cultural context | Nepali proverbs vs. English idioms | Domain-specific fine-tuning |
2. Text Preprocessing
Before feeding text into models, we clean and structure it. Key steps:
A. Tokenization
Splitting text into meaningful units (words, subwords, or characters).
Example:
Input: "I love NLP!"
Output: ["I", "love", "NLP", "!"]
flowchart LR
A["Raw Text: 'I love NLP!'"]
B["Tokenization"]
C["Tokens: ['I', 'love', 'NLP', '!']"]
A --> B --> CB. Normalization
- Lowercasing: Convert all text to lowercase (
"Hello"→"hello"). - Removing punctuation:
"NLP!"→"NLP". - Expanding contractions:
"don't"→"do not".
C. Stemming vs. Lemmatization
| Technique | Example (Input: "running") | Output | Notes |
|---|---|---|---|
| Stemming | Porter Stemmer | "run" |
Aggressive (may produce non-words) |
| Lemmatization | WordNet Lemmatizer | "run" |
Uses vocabulary (slower but accurate) |
Worked Example:
Input: "The cats are running fast."
- Tokenize:
["The", "cats", "are", "running", "fast", "."] - Lowercase:
["the", "cats", "are", "running", "fast", "."] - Remove punctuation:
["the", "cats", "are", "running", "fast"] - Lemmatize:
["the", "cat", "be", "run", "fast"]
3. Word Representations
Machines don’t understand words directly; they need numerical representations.
A. Bag-of-Words (BoW)
Represents text as a vector of word counts (ignores order).
Example:
Text: "I love NLP"
Vocabulary: {"I": 0, "love": 1, "NLP": 2}
Vector: [1, 1, 1] (one-hot encoded)
Problem: High dimensionality (sparse vectors for large vocabularies).
B. TF-IDF (Term Frequency-Inverse Document Frequency)
Weighs words by importance:
- TF: How often a word appears in a document.
- IDF: How rare the word is across all documents.
Formula:
Example:
| Word | TF (in doc) | IDF (across 100 docs) | TF-IDF |
|---|---|---|---|
| "NLP" | 2 | 3.0 | 6.0 |
| "love" | 1 | 1.0 | 1.0 |
C. Word Embeddings (Dense Vectors)
Maps words to low-dimensional dense vectors (e.g., 300D) capturing semantic meaning.
1. Word2Vec
- Skip-gram: Predicts context words from a target word.
- CBOW (Continuous Bag of Words): Predicts a target word from context.
Example (Skip-gram):
Input: "The cat sat on the mat"
Target: "cat" → Predict context: "the", "sat", "on", "the", "mat"
graph LR
A["Input: 'cat'"] --> B["Context: 'the', 'sat', 'on', 'mat'"]
B --> C["Word2Vec updates 'cat' vector"]2. GloVe (Global Vectors)
Uses co-occurrence statistics from a corpus to learn embeddings.
Advantage: Captures global word relationships (e.g., "king" - "man" + "woman" ≈ "queen").
3. FastText
Extends Word2Vec by using subword information (helpful for rare words).
Example:
Word: "unhappiness" → Subwords: "un", "happi", "ness" → Embedding is sum of subword vectors.
D. Contextual Embeddings (Transformers)
Unlike static embeddings (Word2Vec), context matters. Example Models:
- BERT (Bidirectional Encoder Representations from Transformers): Uses self-attention to weigh word importance dynamically.
- GPT (Generative Pre-trained Transformer): Predicts next words (used in chatbots).
BERT Example:
Input: "The [MASK] sat on the mat."
BERT predicts "cat" (not just based on word frequency but context).
graph TD
A["Input Sentence"] --> B["Tokenized: ['The', 'cat', 'sat', 'on', 'the', 'mat', '.']"]
B --> C["BERT Self-Attention"]
C --> D["Contextual Embeddings"]
D --> E["Predicts 'cat' for [MASK]"]4. Key NLP Tasks
A. Sentiment Analysis
Classifies text as positive, negative, or neutral.
Example (Nepali E-commerce):
Input: "Daraz delivery was late, but the product was good."
Output: Mixed sentiment (negative for delivery, positive for product).
Approach:
- Preprocess text (tokenize, remove stopwords).
- Use TF-IDF + Logistic Regression or BERT for fine-tuning.
- Train on labeled data (e.g., Nepali product reviews).
B. Machine Translation
Converts text from one language to another (e.g., Nepali → English).
Example:
Input (Nepali): "मेरो नाम सूर्य हो।"
Output (English): "My name is Surya."
Approach:
- Sequence-to-Sequence (Seq2Seq) Models:
- Encoder: Converts input sequence to a context vector.
- Decoder: Generates output sequence word-by-word.
flowchart LR
A["Nepali Text"] --> B["Encoder (RNN/Transformer)"]
B --> C["Context Vector"]
C --> D["Decoder (Generates English)"]
D --> E["English Text"]Real-World Use:
- Google Translate (uses Transformer-based models).
- Nepali Wikipedia (automated translations for low-resource languages).
C. Named Entity Recognition (NER)
Identifies entities like names, dates, organizations.
Example:
Text: "Surya Nepal Airlines flew to Kathmandu on 2023-10-01."
Output:
PERSON: "Surya"ORGANIZATION: "Nepal Airlines"DATE: "2023-10-01"
Approach:
- Train a CRF (Conditional Random Field) or BERT-NER model.
D. Chatbots & Conversational AI
Simulates human conversation using:
- Rule-based systems (if-else for FAQs).
- ML-based (Rasa, Dialogflow for intent recognition).
Example (Khalti Customer Support):
User: "My transaction failed."
Bot:
- Detects intent:
"transaction_failed". - Extracts entities:
{"amount": "500", "time": "10:30 AM"}. - Responds:
"We’ve detected an issue. Please retry or contact support at [number]."
5. Applications of NLP in Nepal
A. Financial NLP (NEPSE, Banks)
- Sentiment Analysis of Stock News:
Input:
"NEPSE index drops due to global uncertainty."Output: Negative sentiment → Algorithmic trading signals. - Loan Application Analysis: Banks (e.g., Global IME) use NLP to extract key details from handwritten forms.
B. Healthcare (e.g., Kathmandu Model Hospital)
- Symptom Checker Chatbots:
Input:
"I have a fever and cough."Output:"You may have flu. Please consult a doctor."(powered by MedQA models).
C. E-Commerce (Daraz, Swoyambu)
- Product Description Analysis:
- Auto-tagging products (e.g.,
"Nike shoes"→ tags:sports, footwear). - Recommendation systems (users who bought
Xalso boughtY).
- Auto-tagging products (e.g.,
D. Traffic & Urban Planning (NTC, Kathmandu Metro)
- Twitter/News Analysis for Traffic Predictions:
Input:
"Heavy rain expected in Kathmandu tomorrow."Output: Increased traffic delay risk → NTC adjusts signal timings.
6. Ethical Considerations in NLP
| Issue | Example | Mitigation Strategy |
|---|---|---|
| Bias | Gender bias in hiring chatbots | Audit datasets, use diverse training data |
| Privacy | WhatsApp Business storing chats | Anonymize data, GDPR-compliant policies |
| Misinformation | Fake news on Facebook | Fact-checking NLP models (e.g., Google Fact Check) |
| Job Displacement | NLP replacing customer service | Reskill workers for AI-assisted roles |
7. Exam Tip: How to Score Full Marks
A. For Theory Questions
- Define clearly: Always start with a one-sentence definition (e.g., "Word Embeddings are dense vector representations of words that capture semantic meaning.").
- Compare models: Use tables (e.g., BoW vs. Word2Vec vs. BERT).
- Mention real-world use: Link to Nepali examples (e.g., Khalti chatbots, NEPSE sentiment analysis).
B. For Practical Questions
- Show preprocessing steps (tokenization → normalization → lemmatization).
- Draw diagrams for:
- Tokenization pipelines.
- Encoder-decoder architectures (for translation).
- Attention mechanisms (for transformers).
- Use small datasets in examples (e.g., 3-5 sentences for sentiment analysis).
C. Common Pitfalls to Avoid
- ❌ Ignoring context: Always explain why BERT is better than BoW for tasks like sentiment analysis.
- ❌ Overcomplicating: For short-answer questions, stick to key points (e.g., "TF-IDF weighs rare words higher").
- ❌ Assuming perfect data: Mention noise handling (e.g., "Real-world text has typos, so we use spell-checking before tokenization.").
D. Sample Exam Question & Answer
Question: "Explain how BERT improves over Word2Vec for sentiment analysis. Provide a Nepali example."
Model Answer: BERT (Bidirectional Encoder Representations from Transformers) improves over Word2Vec in sentiment analysis through:
Bidirectional Context:
- Word2Vec processes words left-to-right or right-to-left (limited context).
- BERT uses self-attention to weigh all words in a sentence dynamically. Example:
- Word2Vec:
"Not good"→"good"gets a positive embedding (ignores "Not"). - BERT: Detects "Not" negates
"good"→ Correctly predicts negative sentiment.
Deep Contextual Understanding:
- Word2Vec:
"Bank"(financial vs. river) → Same vector. - BERT: Uses surrounding words (e.g.,
"I deposited money in the bank."→ financial context).
- Word2Vec:
Nepali Example (Daraz Review):
Input: "Product quality was not good, but delivery was fast."
- Word2Vec: May misclassify due to
"not good"ambiguity. - BERT: Correctly identifies mixed sentiment (negative for quality, positive for delivery).
8. Worked Example: Sentiment Analysis Pipeline
Task: Classify Nepali restaurant reviews as positive/negative.
Step 1: Data Collection
| Review (Nepali) | Label |
|---|---|
"खाना स्वादिष्ट थियो।" |
Positive |
"सेवा ढिलो थियो।" |
Negative |
"ठूलो समस्या भएन।" |
Neutral |
Step 2: Preprocessing
- Tokenize:
"सेवा ढिलो थियो।"→["सेवा", "ढिलो", "थियो", "."] - Normalize:
- Remove punctuation:
["सेवा", "ढिलो", "थियो"] - Lemmatize:
["सेवा", "ढिलो", "हुनु"]
- Remove punctuation:
- TF-IDF Vectorization:
"सेवा"→ High TF-IDF (common in negative reviews)."स्वादिष्ट"→ High TF-IDF (positive).
Step 3: Model Training
- Approach 1: TF-IDF + Logistic Regression.
- Approach 2: Fine-tuned mBERT (multilingual BERT).
Step 4: Prediction
Input: "भातको स्वाद राम्रो थियो तर दाम महंगो थियो।"
- BERT predicts: Mixed sentiment (positive for taste, negative for price).
9. Real-World NLP Systems
A. WhatsApp Business API (Nepal)
- Use Case: Small businesses (e.g., Kathmandu tailors) automate replies.
- NLP Idea: Intent classification (e.g.,
"order status"vs."return request"). - How It Works:
- User:
"My order #12345 is delayed." - NLP detects intent:
"order_delay". - Bot:
"We’re checking. Expected delivery: tomorrow."
- User:
B. Google Translate (Nepali-English)
- Use Case: Tourists, business emails.
- NLP Idea: Transformer-based Seq2Seq.
- Example:
Input (Nepali):
"मेरो नाम सूर्य हो।"Output (English):"My name is Surya."
C. NEPSE Stock News Analysis
- Use Case: Investors track sentiment in news articles.
- NLP Idea: Sentiment analysis + keyword extraction.
- Example:
Input:
"NEPSE gains 2% as global markets recover."Output:- Sentiment: Positive
- Keywords:
NEPSE, gain, global markets - Action: Buy signal for algorithms.
10. Summary Table: NLP Techniques
| Technique | Use Case | Pros | Cons |
|---|---|---|---|
| BoW | Simple text classification | Fast, easy to implement | Ignores word order/meaning |
| TF-IDF | Document similarity | Handles rare words well | Still sparse vectors |
| Word2Vec | Word similarity tasks | Captures semantic meaning | No context awareness |
| BERT | Sentiment, QA, translation | State-of-the-art accuracy | Computationally expensive |
| Seq2Seq | Machine translation | Handles variable-length text | Needs large datasets |
11. Key Formulas to Remember
- TF-IDF:
- Softmax (for classification):
- Attention Score (BERT):
12. Visualizing NLP Concepts
A. Tokenization Pipeline
flowchart LR
A["Raw Text: 'I love NLP!'"]
B["Lowercase: 'i love nlp!'"]
C["Remove Punctuation: 'i love nlp'"]
D["Tokenize: ['i', 'love', 'nlp']"]
E["Lemmatize: ['i', 'love', 'nlp']"]
A --> B --> C --> D --> EB. BERT Self-Attention
graph TD
A["Input: 'The cat sat on the mat'"] --> B["Token Embeddings"]
B --> C["Positional Encoding"]
C --> D["Self-Attention Layers"]
D --> E["Contextualized Output"]
E --> F["Predicts: 'cat' for masked word"]C. Seq2Seq for Translation
sequenceDiagram
participant Encoder
participant Decoder
Encoder->>Decoder: Context Vector (Nepali)
Decoder->>Decoder: Generates English word-by-word
Decoder-->>User: "My name is Surya."13. Common Mistakes in Exams
- Assuming all NLP tasks need deep learning:
- BoW/TF-IDF works for simple classification (e.g., spam detection).
- Ignoring preprocessing:
- Always mention tokenization, normalization, and stopword removal.
- Overlooking ethical issues:
- Exams may ask: "How would you handle bias in a Nepali sentiment analysis model?"
- Confusing Word2Vec and BERT:
- Word2Vec = static embeddings; BERT = contextual embeddings.
14. Recommended Tools & Libraries
| Tool/Library | Purpose | Example Use Case |
|---|---|---|
| NLTK | Text preprocessing | Tokenization, stemming |
| spaCy | Industrial-strength NLP | Named Entity Recognition |
| HuggingFace | Pre-trained models (BERT, GPT) | Fine-tuning for Nepali text |
| TensorFlow/NLP | Deep learning for NLP | Building custom Seq2Seq models |
| Google Colab | Free GPU for training models | Experimenting with BERT |
15. Further Reading
- Books:
- Speech and Language Processing (Jurafsky & Martin)
- Natural Language Processing with Python (Bird et al.)
- Courses:
- Fast.ai Practical Deep Learning for Coders (NLP section)
- HuggingFace Course
- Nepali Resources:
Based on the PU BE Computer (PU) syllabus for Artificial Intelligence (CMP346), unit 9.
Discussion
Loading…