Artificial IntelligenceUnit 97 min read
Natural Language Processing – Key Concepts and Applications
Unit 9 of Artificial Intelligence: covers the fundamentals of NLP, from text preprocessing to advanced language models, and demonstrates how these techniques power modern applications.
Key points
- NLP transforms raw text into structured data through tokenization, POS tagging, and parsing.
- Statistical and neural methods enable robust language understanding and generation.
- Word embeddings capture semantic similarity, enabling tasks like analogy solving and machine translation.
- Evaluation metrics such as BLEU, ROUGE, and F1-score quantify system performance.
- Real‑world systems (eSewa chatbot, Pathao driver matching, NEPSE sentiment analysis) rely on NLP for user interaction and decision support.
1. Introduction to NLP
Natural Language Processing (NLP) is the subfield of AI that enables machines to understand, interpret, and generate human language. It combines linguistics, computer science, and machine learning to process text and speech.
| Domain | Typical Input | Typical Output |
|---|---|---|
| Text mining | News articles | Topic clusters |
| Dialogue systems | User utterances | System responses |
| Machine translation | Source sentence | Target sentence |
| Speech recognition | Audio waveform | Transcribed text |
2. Text Pre‑processing
Pre‑processing cleans raw text and converts it into a format suitable for modeling.
2.1 Tokenization
Splitting text into tokens (words, sub‑words, or characters).
flowchart LR A["Raw sentence"] --> B["Tokenizer"] B --> C["Token list"]
2.2 Normalization
Lowercasing, removing punctuation, and handling accents.
2.3 Stop‑word Removal
Filtering high‑frequency, low‑informative words (e.g., “the”, “and”).
2.4 Stemming & Lemmatization
Reducing words to their base form.
- Stemming: Porter stemmer example: “running” → “run”.
- Lemmatization: Uses POS tags; “better” → “good”.
Worked Example – Stemming
Input sentence: “The children were running happily.”
- Tokenize → [The, children, were, running, happily]
- Stem each token → [the, child, were, run, happily]
3. Part‑of‑Speech (POS) Tagging
Assigning grammatical categories to tokens.
| POS Tag | Example | Description |
|---|---|---|
| NN | dog | Noun, singular |
| VB | run | Verb, base form |
| JJ | quick | Adjective |
3.1 Rule‑Based vs Statistical vs Neural
- Rule‑Based: Hand‑crafted grammar rules.
- Statistical: Hidden Markov Models (HMMs).
- Neural: Bi‑LSTM‑CRF, BERT.
Worked Example – POS Tagging (Rule‑Based)
Sentence: “The quick brown fox jumps over the lazy dog.”
| Token | POS |
|---|---|
| The | DT |
| quick | JJ |
| brown | JJ |
| fox | NN |
| jumps | VBZ |
| over | IN |
| the | DT |
| lazy | JJ |
| dog | NN |
4. Named Entity Recognition (NER)
Identifying entities such as names, locations, and organizations.
| Entity | Example |
|---|---|
| PERSON | “Rahul Sharma” |
| LOCATION | “Kathmandu” |
| ORGANIZATION | “Ncell” |
4.1 NER Models
- CRF: Conditional Random Fields.
- Bi‑LSTM‑CRF: Captures context.
- BERT‑NER: Fine‑tuned transformer.
5. Parsing
Understanding syntactic structure.
5.1 Constituency Parsing
Tree representation of phrases.
5.2 Dependency Parsing
Shows grammatical relations between words.
6. Semantic Analysis
Mapping words to meaning.
6.1 Word Sense Disambiguation (WSD)
Choosing correct sense of a polysemous word.
6.2 Semantic Role Labeling (SRL)
Identifying predicate‑argument structures.
7. Word Embeddings
Dense vector representations capturing semantics.
7.1 Word2Vec (CBOW & Skip‑gram)
Trains embeddings by predicting context words.
7.2 GloVe
Global co‑occurrence statistics.
7.3 FastText
Sub‑word information for rare words.
Worked Example – Analogy
Compute vector(king) – vector(man) + vector(woman).
Result ≈ vector(queen).
8. Language Models
Predicting next word or sequence probability.
8.1 N‑gram Models
8.2 Neural Language Models
- RNN: Captures long‑term dependencies.
- Transformer: Self‑attention, state‑of‑the‑art.
8.3 Evaluation – Perplexity
9. Machine Translation (MT)
Translating text from source to target language.
9.1 Phrase‑Based MT
Aligns phrases, uses statistical models.
9.2 Neural MT (NMT)
Encoder‑decoder architecture with attention.
Worked Example – Phrase‑Based MT
Source: “I love Nepal.”
- Phrase alignment: “I” → “म”, “love” → “प्रेम गर्छु”, “Nepal” → “नेपाल”
- Re‑ordering: “म प्रेम गर्छु नेपाल” → “म नेपाल प्रेम गर्छु” (corrected by re‑ordering model).
10. Sentiment Analysis
Determining polarity of text.
10.1 Lexicon‑Based
Using sentiment dictionaries.
10.2 Machine Learning
Feature extraction + classifier (SVM, Naïve Bayes).
10.3 Deep Learning
CNN or LSTM on word embeddings.
Worked Example – Sentiment Scoring
Sentence: “The product quality is excellent.”
- Lexicon score: +2 (excellent) +0 (product) +0 (quality) +0 (is) +0 (is) → Total +2 → Positive.
11. Speech Recognition
Converting audio to text.
11.1 Acoustic Model
Maps audio frames to phonemes.
11.2 Language Model
Guides phoneme sequence to words.
11.3 End‑to‑End Models
DeepSpeech, wav2vec 2.0.
Worked Example – Phoneme Mapping
Audio: /k/ /æ/ /t/ → “cat”.
12. NLP Pipeline – End‑to‑End Flow
flowchart LR A["Raw Input"] --> B["Pre‑processing"] B --> C["Feature Extraction"] C --> D["Model Inference"] D --> E["Post‑processing"] E --> F["Output"]
13. Evaluation Metrics
| Task | Metric | Formula |
|---|---|---|
| Classification | F1‑score | |
| Machine Translation | BLEU | |
| Summarization | ROUGE |
14. Comparison of NLP Approaches
| Approach | Strengths | Weaknesses | Typical Use |
|---|---|---|---|
| Rule‑Based | Transparent, low data | Poor scalability | Simple chatbots |
| Statistical | Handles noise | Requires labeled data | POS tagging |
| Neural | State‑of‑the‑art performance | Data hungry, opaque | Machine translation |
15. In the real world
| Product | NLP Idea Used | How It Works |
|---|---|---|
| eSewa Chatbot | Intent Classification & NER | Detects user intent (“transfer money”) and extracts account numbers. |
| Pathao Driver Matching | Named Entity Recognition & Sentiment | Parses rider requests, extracts pickup/drop locations, and gauges urgency via sentiment. |
| NEPSE News Sentiment | Sentiment Analysis & Time‑Series | Scores daily news articles to predict stock movement. |
| Google Translate | Neural Machine Translation | Encoder‑decoder transformer predicts target language sentence. |
| WhatsApp Translation | Language Detection & MT | Detects language of a message and offers real‑time translation. |
| YouTube Auto‑Captioning | Speech Recognition | Converts spoken video content into captions using end‑to‑end models. |
Worked Real‑World Example – NEPSE Sentiment
- Collect 1000 news headlines.
- Pre‑process → tokenize, remove stop‑words.
- Apply pre‑trained BERT sentiment classifier → scores.
- Aggregate daily average sentiment → +0.45.
- Correlate with NSE index → positive correlation (r = 0.6).
16. Exam tip
- Understand the pipeline: Know each preprocessing step and its purpose.
- Know evaluation metrics: Be able to compute F1, BLEU, ROUGE manually.
- Compare methods: Be ready to discuss rule‑based vs statistical vs neural for a given task.
- Trace examples: Practice POS tagging, NER, and word embedding arithmetic.
- Real‑world mapping: Relate concepts to applications like eSewa or NEPSE.
flowchart LR A["Raw Text"] --> B["Tokenization"] B --> C["POS Tagging"] C --> D["NER"] D --> E["Parsing"] E --> F["Semantic Analysis"] F --> G["Word Embeddings"] G --> H["Language Model"] H --> I["Output"]
Based on the TU BITM syllabus for Artificial Intelligence (IT228), unit 9.
Discussion
Loading…