Knowledge EngineeringUnit 59 min read
NLP Techniques: Parsing, POS Tagging, NER & Semantic Analysis
Unit 5 of Knowledge Engineering explores core NLP techniques—parsing, POS tagging, named entity recognition (NER), and semantic analysis—with real-world applications in Nepali/English processing, chatbots, and information extraction. Learn how morphology, syntax, and semantics interact through visual traces, worked exa
TAKEAWAYS:
- Parsing breaks sentences into syntactic trees (e.g., constituency vs. dependency parsing) to enable machine understanding of grammar.
- POS tagging assigns grammatical labels (noun, verb, adjective) to words, critical for disambiguation in Nepali/English (e.g., "run" as noun vs. verb).
- Named Entity Recognition (NER) extracts real-world entities (people, dates, locations) from text, powering search and recommendation systems like Daraz or Ncell.
- Semantic analysis resolves ambiguity (e.g., "bank" as financial vs. river) using ontologies (OWL) and word embeddings (Word2Vec, GloVe).
- Morphology (word structure) and lexicon (vocabulary) are the building blocks for accurate NLP pipelines in Nepali (e.g., handling sandhi rules).
- Real-world tie-ins: WhatsApp’s message parsing, eSewa’s NER for transaction validation, and Google’s semantic search rely on these techniques.
Core NLP Techniques: How Machines Understand Language
1. Morphology and Lexicon: The Building Blocks
NLP starts with morphology (studying word forms) and lexicon (vocabulary). In Nepali, morphology is complex due to:
- Sandhi rules: Word combinations change forms (e.g., "राम + लाई" → "रामलाई").
- Inflection: Verbs change based on tense/person (e.g., "खानु" → "खाइस", "खान्छ").
Worked Example: Nepali Verb Conjugation Sentence: "मेरो पुस्तक खान्छ" (My book eats) → Incorrect. Correct: "मेरो पुस्तक खाइन्छ" (My book is eaten). Trace:
- Identify root: "खानु" (to eat).
- Apply passive rule: "खाइन्छ" (is eaten).
- Assign POS: "खाइन्छ" → verb (passive, present).
graph LR
A["खानु (root)"] --> B["खान्छ (active)"] --> C["खाइन्छ (passive)"]
A --> D["खाइस (past)"] --> E["खाइएको (past participle)"]2. Part-of-Speech (POS) Tagging: Labeling Words
POS tagging assigns grammatical labels to words. Example for English:
Sentence: "The quick brown fox jumps over the lazy dog."
Tags: DET ADJ ADJ NOUN VERB PREP DET ADJ NOUN .
Nepali Example:
Sentence: "रामले पुस्तक पढ्यो।"
Tags: PROPER VERB NOUN VERB .
(रामले = subject, पुस्तक = object, पढ्यो = verb).
Why It Matters:
- Disambiguates homographs (e.g., "run" as noun vs. verb).
- Enables syntax parsing (next step).
Worked Example: Nepali POS Tagging Sentence: "सुरुमा मलाई थाहा थिएन।"
- Tokenize: ["सुरुमा", "मलाई", "थाहा", "थिएन", "।"]
- Tag:
- सुरुमा →
ADV(adverb) - मलाई →
PRON(pronoun) - थाहा →
NOUN(noun) - थिएन →
VERB(past negative) - । →
PUNCT
- सुरुमा →
graph TD
A["सुरुमा\nADV"] --> B["मलाई\nPRON"]
B --> C["थाहा\nNOUN"]
C --> D["थिएन\nVERB"]
D --> E["।\nPUNCT"]3. Parsing: Building Syntax Trees
Parsing converts sentences into structured trees to understand relationships. Two main types:
- Constituency Parsing: Hierarchical phrase structure (like a family tree).
Example: "रामले पुस्तक पढ्यो" →
S ├── NP: रामले └── VP ├── V: पढ्यो └── NP: पुस्तक - Dependency Parsing: Word-to-word relationships (who does what to whom).
Example:
पढ्यो (root) ├── subject: रामले └── object: पुस्तक
Worked Example: Nepali Dependency Parse Sentence: "मेरो गाडी खराब भएको छ।"
- Tokenize: ["मेरो", "गाडी", "खराब", "भएको", "छ", "।"]
- Dependencies:
- भएको (root)
- subject: गाडी
- modifier: मेरो
- copula: छ
- adjective: खराब
- भएको (root)
graph TD
A["भएको\n(root)"] --> B["गाडी\n(subject)"]
A --> C["छ\n(copula)"]
B --> D["मेरो\n(modifier)"]
A --> E["खराब\n(adjective)"]4. Named Entity Recognition (NER): Extracting Real-World Entities
NER identifies and classifies entities like:
- Person: राम, Barack Obama
- Location: काठमाडौं, New York
- Organization: NTC, Google
- Date: २०८० चैत १, 2023-10-05
- Money: ₹५००, $100
Real-World Use in Nepal:
- eSewa: Extracts user names, transaction IDs, and amounts from chat inputs.
- Ncell: Identifies phone numbers in customer queries for routing.
- Daraz: Pulls product names, prices, and seller info from reviews.
Worked Example: Nepali NER Sentence: "रामले २०८० चैत १ मा काठमाडौंमा ५००० रुपैयाँ जमा गर्यो।" Entities:
| Entity Type | Text | Confidence |
|---|---|---|
| Person | राम | 0.98 |
| Date | २०८० चैत १ | 0.95 |
| Location | काठमाडौंमा | 0.99 |
| Money | ५००० रुपैयाँ | 0.97 |
5. Semantic Analysis: Beyond Words to Meaning
Semantics resolves ambiguity and links words to real-world knowledge. Key techniques:
- Word Sense Disambiguation (WSD):
- "Bank" → financial (NTC) vs. river (Ganga).
- Tools: WordNet, BabelNet.
- Ontologies (OWL):
- Links "लेखक" (author) to "व्यक्ति" (person) in Nepali ontologies.
- Word Embeddings:
- Word2Vec, GloVe: Represent words as vectors (e.g., "काठमाडौं" close to "नेपाल").
Worked Example: Nepali WSD Sentence: "मेरो बैंकमा पैसा छ।"
- Possible senses for "बैंक":
- Financial institution (NTC).
- River bank (Ganga).
- Context clues: "पैसा" (money) → NTC (financial).
Visualizing Word Embeddings:
graph TD
A["नेपाल"] --> B["काठमाडौं"]
A --> C["पोखरा"]
B --> D["ललितपुर"]
C --> DClose vectors indicate semantic similarity.6. NLP Pipelines: Putting It All Together
A typical NLP pipeline for Nepali text processing:
flowchart LR
A["Text Input"] --> B["Tokenization"]
B --> C["POS Tagging"]
C --> D["Parsing"]
D --> E["NER"]
E --> F["Semantic Analysis"]
F --> G["Output: Structured Data"]Real-World Pipeline: WhatsApp Chat Analysis
- Input: "रामले २०८० चैत १ मा ५००० जमा गर्यो।"
- Tokenization: Split into words.
- POS Tagging: Identify verbs, dates.
- NER: Extract राम, २०८० चैत १, ५०००.
- Semantic Analysis: Link "जमा गर्यो" to financial transaction.
- Output: Structured data for eSewa/Khalti validation.
In the Real World
eSewa/Khalti:
- NER + Semantic Analysis: Extracts user names, transaction amounts, and service codes (e.g., "बिजुली बिल") from chat inputs to auto-fill forms.
- Example: Input "रामले २००० बिजुली बिल भुक्तानी गर्नुहोस्" → System identifies:
- User: राम
- Amount: २०००
- Service: बिजुली बिल
Daraz Customer Support:
- Parsing + POS Tagging: Routes queries like "मेरो order ID DZ123456 भएको छ, कहाँ छ?" by extracting:
- Order ID: DZ123456 (via NER).
- Intent: status inquiry (via POS + parsing).
- Parsing + POS Tagging: Routes queries like "मेरो order ID DZ123456 भएको छ, कहाँ छ?" by extracting:
Ncell Chatbot:
- Morphology Handling: Corrects Nepali grammar errors (e.g., "मेरो नम्बर कस्तो छ" → "मेरो नम्बर के हो?") before processing.
- Semantic Search: Matches "तपाईंको नम्बर के हो" to "What’s my number?" using word embeddings.
Nepali Wikipedia/BabelNet:
- Ontology Integration: Links Nepali terms (e.g., "लिच्छवी") to English ("Licchavi") and historical context using OWL-based ontologies.
Comparisons: NLP Techniques vs. Semantic Web
| Feature | NLP Techniques | Semantic Web (OWL, RDF) |
|---|---|---|
| Purpose | Process natural language text. | Represent knowledge machine-readable. |
| Example | POS tagging, NER. | Ontology of Nepali historical figures. |
| Data Format | Text, trees, vectors. | Triples (subject-predicate-object). |
| Use Case | Chatbots, search. | Linked Data, knowledge graphs. |
| Nepali Example | Khalti’s transaction parsing. | Nepali Wikipedia’s structured data. |
Exam Tip
For Short Questions (1+4 marks):
- Parsing: Always draw a dependency tree for the given sentence. Label nodes clearly (e.g., "subject", "object").
- POS Tagging: Show tokenization + tags in a table. Example:
Word POS रामले PROPN पुस्तक NOUN पढ्यो VERB
For Long Questions (5+5 marks):
- Structure: Use the pipeline flowchart (tokenization → POS → parsing → NER → semantics).
- Real-World Tie-Ins: Pick one Nepali app (eSewa, Daraz) and explain how two techniques (e.g., NER + semantics) work together.
- Nepali Examples: Always use Nepali sentences for parsing/POS tagging. Examiners favor localized examples.
Common Pitfalls:
- Forgetting morphology for Nepali (e.g., sandhi rules in POS tagging).
- Confusing constituency vs. dependency parsing. Always ask: "Does it show phrases or word relationships?"
- Overlooking semantics in NER (e.g., "bank" can be both financial and geographical).
Visuals in Exams:
- If asked to "explain parsing," draw a dependency tree (even if rough).
- For NER, use a table with columns: Entity Type | Text | Confidence.
- For semantics, sketch a simple ontology graph (e.g., "लेखक" → "व्यक्ति").
Pro Tip: Memorize one Nepali sentence for each technique (e.g., POS tagging, parsing) and practice tracing it step-by-step. Examiners often reuse similar examples!
Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 5.
Discussion
Loading…