CACS458 Knowledge Engineering

Knowledge EngineeringUnit 59 min read

NLP Techniques: Parsing, POS Tagging, NER & Semantic Analysis

Unit 5 of Knowledge Engineering explores core NLP techniques—parsing, POS tagging, named entity recognition (NER), and semantic analysis—with real-world applications in Nepali/English processing, chatbots, and information extraction. Learn how morphology, syntax, and semantics interact through visual traces, worked exa

TAKEAWAYS:

  • Parsing breaks sentences into syntactic trees (e.g., constituency vs. dependency parsing) to enable machine understanding of grammar.
  • POS tagging assigns grammatical labels (noun, verb, adjective) to words, critical for disambiguation in Nepali/English (e.g., "run" as noun vs. verb).
  • Named Entity Recognition (NER) extracts real-world entities (people, dates, locations) from text, powering search and recommendation systems like Daraz or Ncell.
  • Semantic analysis resolves ambiguity (e.g., "bank" as financial vs. river) using ontologies (OWL) and word embeddings (Word2Vec, GloVe).
  • Morphology (word structure) and lexicon (vocabulary) are the building blocks for accurate NLP pipelines in Nepali (e.g., handling sandhi rules).
  • Real-world tie-ins: WhatsApp’s message parsing, eSewa’s NER for transaction validation, and Google’s semantic search rely on these techniques.

Core NLP Techniques: How Machines Understand Language

1. Morphology and Lexicon: The Building Blocks

NLP starts with morphology (studying word forms) and lexicon (vocabulary). In Nepali, morphology is complex due to:

  • Sandhi rules: Word combinations change forms (e.g., "राम + लाई" → "रामलाई").
  • Inflection: Verbs change based on tense/person (e.g., "खानु" → "खाइस", "खान्छ").

Worked Example: Nepali Verb Conjugation Sentence: "मेरो पुस्तक खान्छ" (My book eats) → Incorrect. Correct: "मेरो पुस्तक खाइन्छ" (My book is eaten). Trace:

  1. Identify root: "खानु" (to eat).
  2. Apply passive rule: "खाइन्छ" (is eaten).
  3. Assign POS: "खाइन्छ" → verb (passive, present).
graph LR
    A["खानु (root)"] --> B["खान्छ (active)"] --> C["खाइन्छ (passive)"]
    A --> D["खाइस (past)"] --> E["खाइएको (past participle)"]

2. Part-of-Speech (POS) Tagging: Labeling Words

POS tagging assigns grammatical labels to words. Example for English: Sentence: "The quick brown fox jumps over the lazy dog." Tags: DET ADJ ADJ NOUN VERB PREP DET ADJ NOUN .

Nepali Example: Sentence: "रामले पुस्तक पढ्यो।" Tags: PROPER VERB NOUN VERB . (रामले = subject, पुस्तक = object, पढ्यो = verb).

Why It Matters:

  • Disambiguates homographs (e.g., "run" as noun vs. verb).
  • Enables syntax parsing (next step).

Worked Example: Nepali POS Tagging Sentence: "सुरुमा मलाई थाहा थिएन।"

  1. Tokenize: ["सुरुमा", "मलाई", "थाहा", "थिएन", "।"]
  2. Tag:
    • सुरुमा → ADV (adverb)
    • मलाई → PRON (pronoun)
    • थाहा → NOUN (noun)
    • थिएन → VERB (past negative)
    • । → PUNCT
graph TD
    A["सुरुमा\nADV"] --> B["मलाई\nPRON"]
    B --> C["थाहा\nNOUN"]
    C --> D["थिएन\nVERB"]
    D --> E["।\nPUNCT"]

3. Parsing: Building Syntax Trees

Parsing converts sentences into structured trees to understand relationships. Two main types:

  1. Constituency Parsing: Hierarchical phrase structure (like a family tree). Example: "रामले पुस्तक पढ्यो" →
    S
    ├── NP: रामले
    └── VP
        ├── V: पढ्यो
        └── NP: पुस्तक
    
  2. Dependency Parsing: Word-to-word relationships (who does what to whom). Example:
    पढ्यो (root)
    ├── subject: रामले
    └── object: पुस्तक
    

Worked Example: Nepali Dependency Parse Sentence: "मेरो गाडी खराब भएको छ।"

  1. Tokenize: ["मेरो", "गाडी", "खराब", "भएको", "छ", "।"]
  2. Dependencies:
    • भएको (root)
      • subject: गाडी
      • modifier: मेरो
      • copula: छ
      • adjective: खराब
graph TD
    A["भएको\n(root)"] --> B["गाडी\n(subject)"]
    A --> C["छ\n(copula)"]
    B --> D["मेरो\n(modifier)"]
    A --> E["खराब\n(adjective)"]

4. Named Entity Recognition (NER): Extracting Real-World Entities

NER identifies and classifies entities like:

  • Person: राम, Barack Obama
  • Location: काठमाडौं, New York
  • Organization: NTC, Google
  • Date: २०८० चैत १, 2023-10-05
  • Money: ₹५००, $100

Real-World Use in Nepal:

  • eSewa: Extracts user names, transaction IDs, and amounts from chat inputs.
  • Ncell: Identifies phone numbers in customer queries for routing.
  • Daraz: Pulls product names, prices, and seller info from reviews.

Worked Example: Nepali NER Sentence: "रामले २०८० चैत १ मा काठमाडौंमा ५००० रुपैयाँ जमा गर्यो।" Entities:

Entity Type Text Confidence
Person राम 0.98
Date २०८० चैत १ 0.95
Location काठमाडौंमा 0.99
Money ५००० रुपैयाँ 0.97

5. Semantic Analysis: Beyond Words to Meaning

Semantics resolves ambiguity and links words to real-world knowledge. Key techniques:

  1. Word Sense Disambiguation (WSD):
    • "Bank" → financial (NTC) vs. river (Ganga).
    • Tools: WordNet, BabelNet.
  2. Ontologies (OWL):
    • Links "लेखक" (author) to "व्यक्ति" (person) in Nepali ontologies.
  3. Word Embeddings:
    • Word2Vec, GloVe: Represent words as vectors (e.g., "काठमाडौं" close to "नेपाल").

Worked Example: Nepali WSD Sentence: "मेरो बैंकमा पैसा छ।"

  • Possible senses for "बैंक":
    1. Financial institution (NTC).
    2. River bank (Ganga).
  • Context clues: "पैसा" (money) → NTC (financial).

Visualizing Word Embeddings:

graph TD
    A["नेपाल"] --> B["काठमाडौं"]
    A --> C["पोखरा"]
    B --> D["ललितपुर"]
    C --> D
Close vectors indicate semantic similarity.

6. NLP Pipelines: Putting It All Together

A typical NLP pipeline for Nepali text processing:

flowchart LR
    A["Text Input"] --> B["Tokenization"]
    B --> C["POS Tagging"]
    C --> D["Parsing"]
    D --> E["NER"]
    E --> F["Semantic Analysis"]
    F --> G["Output: Structured Data"]

Real-World Pipeline: WhatsApp Chat Analysis

  1. Input: "रामले २०८० चैत १ मा ५००० जमा गर्यो।"
  2. Tokenization: Split into words.
  3. POS Tagging: Identify verbs, dates.
  4. NER: Extract राम, २०८० चैत १, ५०००.
  5. Semantic Analysis: Link "जमा गर्यो" to financial transaction.
  6. Output: Structured data for eSewa/Khalti validation.

In the Real World

  1. eSewa/Khalti:

    • NER + Semantic Analysis: Extracts user names, transaction amounts, and service codes (e.g., "बिजुली बिल") from chat inputs to auto-fill forms.
    • Example: Input "रामले २००० बिजुली बिल भुक्तानी गर्नुहोस्" → System identifies:
      • User: राम
      • Amount: २०००
      • Service: बिजुली बिल
  2. Daraz Customer Support:

    • Parsing + POS Tagging: Routes queries like "मेरो order ID DZ123456 भएको छ, कहाँ छ?" by extracting:
      • Order ID: DZ123456 (via NER).
      • Intent: status inquiry (via POS + parsing).
  3. Ncell Chatbot:

    • Morphology Handling: Corrects Nepali grammar errors (e.g., "मेरो नम्बर कस्तो छ" → "मेरो नम्बर के हो?") before processing.
    • Semantic Search: Matches "तपाईंको नम्बर के हो" to "What’s my number?" using word embeddings.
  4. Nepali Wikipedia/BabelNet:

    • Ontology Integration: Links Nepali terms (e.g., "लिच्छवी") to English ("Licchavi") and historical context using OWL-based ontologies.

Comparisons: NLP Techniques vs. Semantic Web

Feature NLP Techniques Semantic Web (OWL, RDF)
Purpose Process natural language text. Represent knowledge machine-readable.
Example POS tagging, NER. Ontology of Nepali historical figures.
Data Format Text, trees, vectors. Triples (subject-predicate-object).
Use Case Chatbots, search. Linked Data, knowledge graphs.
Nepali Example Khalti’s transaction parsing. Nepali Wikipedia’s structured data.

Exam Tip

  1. For Short Questions (1+4 marks):

    • Parsing: Always draw a dependency tree for the given sentence. Label nodes clearly (e.g., "subject", "object").
    • POS Tagging: Show tokenization + tags in a table. Example:
      Word POS
      रामले PROPN
      पुस्तक NOUN
      पढ्यो VERB
  2. For Long Questions (5+5 marks):

    • Structure: Use the pipeline flowchart (tokenization → POS → parsing → NER → semantics).
    • Real-World Tie-Ins: Pick one Nepali app (eSewa, Daraz) and explain how two techniques (e.g., NER + semantics) work together.
    • Nepali Examples: Always use Nepali sentences for parsing/POS tagging. Examiners favor localized examples.
  3. Common Pitfalls:

    • Forgetting morphology for Nepali (e.g., sandhi rules in POS tagging).
    • Confusing constituency vs. dependency parsing. Always ask: "Does it show phrases or word relationships?"
    • Overlooking semantics in NER (e.g., "bank" can be both financial and geographical).
  4. Visuals in Exams:

    • If asked to "explain parsing," draw a dependency tree (even if rough).
    • For NER, use a table with columns: Entity Type | Text | Confidence.
    • For semantics, sketch a simple ontology graph (e.g., "लेखक" → "व्यक्ति").

Pro Tip: Memorize one Nepali sentence for each technique (e.g., POS tagging, parsing) and practice tracing it step-by-step. Examiners often reuse similar examples!

Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 5.

Discussion

Loading…