Artificial IntelligenceUnit 97 min read

Natural Language Processing – Key Concepts and Applications

Unit 9 of Artificial Intelligence: covers the fundamentals of NLP, from text preprocessing to advanced language models, and demonstrates how these techniques power modern applications.

Key points

  • NLP transforms raw text into structured data through tokenization, POS tagging, and parsing.
  • Statistical and neural methods enable robust language understanding and generation.
  • Word embeddings capture semantic similarity, enabling tasks like analogy solving and machine translation.
  • Evaluation metrics such as BLEU, ROUGE, and F1-score quantify system performance.
  • Real‑world systems (eSewa chatbot, Pathao driver matching, NEPSE sentiment analysis) rely on NLP for user interaction and decision support.

1. Introduction to NLP

Natural Language Processing (NLP) is the subfield of AI that enables machines to understand, interpret, and generate human language. It combines linguistics, computer science, and machine learning to process text and speech.

Domain Typical Input Typical Output
Text mining News articles Topic clusters
Dialogue systems User utterances System responses
Machine translation Source sentence Target sentence
Speech recognition Audio waveform Transcribed text

2. Text Pre‑processing

Pre‑processing cleans raw text and converts it into a format suitable for modeling.

2.1 Tokenization

Splitting text into tokens (words, sub‑words, or characters).

flowchart LR
  A["Raw sentence"] --> B["Tokenizer"]
  B --> C["Token list"]

2.2 Normalization

Lowercasing, removing punctuation, and handling accents.

2.3 Stop‑word Removal

Filtering high‑frequency, low‑informative words (e.g., “the”, “and”).

2.4 Stemming & Lemmatization

Reducing words to their base form.

  • Stemming: Porter stemmer example: “running” → “run”.
  • Lemmatization: Uses POS tags; “better” → “good”.

Worked Example – Stemming

Input sentence: “The children were running happily.”

  1. Tokenize → [The, children, were, running, happily]
  2. Stem each token → [the, child, were, run, happily]

3. Part‑of‑Speech (POS) Tagging

Assigning grammatical categories to tokens.

POS Tag Example Description
NN dog Noun, singular
VB run Verb, base form
JJ quick Adjective

3.1 Rule‑Based vs Statistical vs Neural

  • Rule‑Based: Hand‑crafted grammar rules.
  • Statistical: Hidden Markov Models (HMMs).
  • Neural: Bi‑LSTM‑CRF, BERT.

Worked Example – POS Tagging (Rule‑Based)

Sentence: “The quick brown fox jumps over the lazy dog.”

Token POS
The DT
quick JJ
brown JJ
fox NN
jumps VBZ
over IN
the DT
lazy JJ
dog NN

4. Named Entity Recognition (NER)

Identifying entities such as names, locations, and organizations.

Entity Example
PERSON “Rahul Sharma”
LOCATION “Kathmandu”
ORGANIZATION “Ncell”

4.1 NER Models

  • CRF: Conditional Random Fields.
  • Bi‑LSTM‑CRF: Captures context.
  • BERT‑NER: Fine‑tuned transformer.

5. Parsing

Understanding syntactic structure.

5.1 Constituency Parsing

Tree representation of phrases.

5.2 Dependency Parsing

Shows grammatical relations between words.

6. Semantic Analysis

Mapping words to meaning.

6.1 Word Sense Disambiguation (WSD)

Choosing correct sense of a polysemous word.

6.2 Semantic Role Labeling (SRL)

Identifying predicate‑argument structures.

7. Word Embeddings

Dense vector representations capturing semantics.

7.1 Word2Vec (CBOW & Skip‑gram)

Trains embeddings by predicting context words.

7.2 GloVe

Global co‑occurrence statistics.

7.3 FastText

Sub‑word information for rare words.

Worked Example – Analogy

Compute vector(king) – vector(man) + vector(woman).
Result ≈ vector(queen).

8. Language Models

Predicting next word or sequence probability.

8.1 N‑gram Models

8.2 Neural Language Models

  • RNN: Captures long‑term dependencies.
  • Transformer: Self‑attention, state‑of‑the‑art.

8.3 Evaluation – Perplexity

9. Machine Translation (MT)

Translating text from source to target language.

9.1 Phrase‑Based MT

Aligns phrases, uses statistical models.

9.2 Neural MT (NMT)

Encoder‑decoder architecture with attention.

Worked Example – Phrase‑Based MT

Source: “I love Nepal.”

  1. Phrase alignment: “I” → “म”, “love” → “प्रेम गर्छु”, “Nepal” → “नेपाल”
  2. Re‑ordering: “म प्रेम गर्छु नेपाल” → “म नेपाल प्रेम गर्छु” (corrected by re‑ordering model).

10. Sentiment Analysis

Determining polarity of text.

10.1 Lexicon‑Based

Using sentiment dictionaries.

10.2 Machine Learning

Feature extraction + classifier (SVM, Naïve Bayes).

10.3 Deep Learning

CNN or LSTM on word embeddings.

Worked Example – Sentiment Scoring

Sentence: “The product quality is excellent.”

  • Lexicon score: +2 (excellent) +0 (product) +0 (quality) +0 (is) +0 (is) → Total +2 → Positive.

11. Speech Recognition

Converting audio to text.

11.1 Acoustic Model

Maps audio frames to phonemes.

11.2 Language Model

Guides phoneme sequence to words.

11.3 End‑to‑End Models

DeepSpeech, wav2vec 2.0.

Worked Example – Phoneme Mapping

Audio: /k/ /æ/ /t/ → “cat”.

12. NLP Pipeline – End‑to‑End Flow

flowchart LR
  A["Raw Input"] --> B["Pre‑processing"]
  B --> C["Feature Extraction"]
  C --> D["Model Inference"]
  D --> E["Post‑processing"]
  E --> F["Output"]

13. Evaluation Metrics

Task Metric Formula
Classification F1‑score
Machine Translation BLEU
Summarization ROUGE

14. Comparison of NLP Approaches

Approach Strengths Weaknesses Typical Use
Rule‑Based Transparent, low data Poor scalability Simple chatbots
Statistical Handles noise Requires labeled data POS tagging
Neural State‑of‑the‑art performance Data hungry, opaque Machine translation

15. In the real world

Product NLP Idea Used How It Works
eSewa Chatbot Intent Classification & NER Detects user intent (“transfer money”) and extracts account numbers.
Pathao Driver Matching Named Entity Recognition & Sentiment Parses rider requests, extracts pickup/drop locations, and gauges urgency via sentiment.
NEPSE News Sentiment Sentiment Analysis & Time‑Series Scores daily news articles to predict stock movement.
Google Translate Neural Machine Translation Encoder‑decoder transformer predicts target language sentence.
WhatsApp Translation Language Detection & MT Detects language of a message and offers real‑time translation.
YouTube Auto‑Captioning Speech Recognition Converts spoken video content into captions using end‑to‑end models.

Worked Real‑World Example – NEPSE Sentiment

  1. Collect 1000 news headlines.
  2. Pre‑process → tokenize, remove stop‑words.
  3. Apply pre‑trained BERT sentiment classifier → scores.
  4. Aggregate daily average sentiment → +0.45.
  5. Correlate with NSE index → positive correlation (r = 0.6).

16. Exam tip

  • Understand the pipeline: Know each preprocessing step and its purpose.
  • Know evaluation metrics: Be able to compute F1, BLEU, ROUGE manually.
  • Compare methods: Be ready to discuss rule‑based vs statistical vs neural for a given task.
  • Trace examples: Practice POS tagging, NER, and word embedding arithmetic.
  • Real‑world mapping: Relate concepts to applications like eSewa or NEPSE.

flowchart LR
  A["Raw Text"] --> B["Tokenization"]
  B --> C["POS Tagging"]
  C --> D["NER"]
  D --> E["Parsing"]
  E --> F["Semantic Analysis"]
  F --> G["Word Embeddings"]
  G --> H["Language Model"]
  H --> I["Output"]

Based on the TU BITM syllabus for Artificial Intelligence (IT228), unit 9.

Discussion

Loading…