Artificial IntelligenceUnit 59 min read
NLP Core: PEAS, Parsing, Semantics, Pragmatics & Challenges
Unit 5 of Artificial Intelligence covers Natural Language Processing (NLP) fundamentals: PEAS framework for intelligent agents, NLP pipeline stages (tokenization, morphological/syntactic/semantic/pragmatic analysis), real-world applications in Nepali/English systems, and key challenges like ambiguity and context depend
What is NLP?
Natural Language Processing (NLP) is the field of AI that enables computers to understand, generate, and manipulate human language. It bridges linguistics and computer science to build systems that can:
- Interpret user queries (e.g., "Book a flight to Kathmandu")
- Generate human-like responses (e.g., chatbots)
- Translate languages (e.g., Google Translate)
- Extract structured data from text (e.g., eSewa parsing bills)
Why Machines Need NLP
Humans communicate primarily through language (90%+ of data is unstructured text). Machines need NLP to:
- Interact naturally (e.g., voice assistants like Google Assistant).
- Automate tasks (e.g., eSewa processing payments from text).
- Extract insights (e.g., Ncell analyzing customer complaints).
PEAS Framework for Intelligent Agents
The Performance, Environment, Actuators, Sensors framework defines an agent’s operational context. Let’s analyze two agents:
Example 1: Internet Shopping Assistant (e.g., Daraz)
- Performance: Must recommend relevant products (90%+ accuracy) and process orders in <2s.
- Environment: Dynamic (prices change, trends shift).
- Actuators: Sends recommendations, confirms orders.
- Sensors: User search queries, product catalog, reviews.
Example 2: English Language Tutor (e.g., Duolingo)
- Performance: Corrects 95% of grammar errors, teaches 1000+ words/year.
- Environment: Personalized (adapts to user mistakes).
- Actuators: Provides hints, quizzes.
- Sensors: Records user responses, tracks mistakes.
In the Real World
- eSewa: Uses NLP to parse payment instructions from SMS/text (e.g., "Pay Rs. 500 to ABC for electricity"). The semantic analysis step extracts the amount, recipient, and service type.
- Khalti: When you say, "Pay Rs. 1000 to my brother," the pragmatic analysis resolves "brother" to a contact in your phone’s address book.
- Pathao Driver App: The syntactic parser breaks down "Take me to Thamel" into subject (user), verb (go), and object (Thamel), then maps it to GPS coordinates.
NLP Pipeline: From Text to Meaning
The NLP pipeline converts raw text into structured data. Here’s how it works for the query: "Book a flight to Kathmandu for Rs. 15,000 on 15th June."
Step 1: Tokenization
Split text into words/tokens:
["Book", "a", "flight", "to", "Kathmandu", "for", "Rs.", "15,000", "on", "15th", "June"]
Step 2: Morphological Analysis
Break words into stems and identify parts of speech (POS):
| Token | Stem | POS Tag |
|---|---|---|
| Book | book | Verb |
| flight | flight | Noun |
| Kathmandu | Kathmandu | Proper Noun |
| Rs. | Rs. | Currency |
| 15,000 | 15000 | Number |
Step 3: Syntactic Analysis (Parsing)
Build a parse tree to understand sentence structure:
graph TD
S["Sentence"] --> NP1["NP: Book"]
S --> VP["VP: flight to Kathmandu"]
VP --> V["V: book"]
VP --> PP1["PP: to Kathmandu"]
PP1 --> P["P: to"]
PP1 --> NP2["NP: Kathmandu"]
VP --> PP2["PP: for Rs. 15,000 on 15th June"]
PP2 --> P2["P: for"]
PP2 --> NP3["NP: Rs. 15,000"]
PP2 --> PP3["PP: on 15th June"]Key Insight: The verb "book" governs the flight (direct object) and the time/money (prepositional phrases).
Step 4: Semantic Analysis
Extract meaning using a knowledge base (e.g., flight database):
- Entities:
- Action:
book_flight - Destination:
Kathmandu - Price:
15000 NPR - Date:
2024-06-15
- Action:
- Relations:
book_flight(destination=Kathmandu, price=15000, date=2024-06-15)
Step 5: Pragmatic Analysis
Resolve ambiguity and context:
- If you’re logged into Yeti Airlines, the system assumes you want a Yeti flight.
- If you say "Book a flight to Kathmandu for my mom," the system may ask: "Which ticket: adult or child?"
Challenges in NLP
| Challenge | Example | Solution Approach |
|---|---|---|
| Ambiguity | "I saw the man with the telescope." (Who has the telescope?) | Use coreference resolution (e.g., statistical models). |
| Context Dependency | "She ate an apple. It was red." (What does "it" refer to?) | Discourse analysis (track conversation history). |
| Sarcasm/Irony | "Great, another power cut." (Negative tone) | Sentiment analysis + pragmatic rules. |
| Domain-Specific Terms | "Pull the trigger" (gun vs. camera) | Domain adaptation (train on relevant data). |
| Low-Resource Languages | Nepali NLP lacks large datasets | Transfer learning (e.g., fine-tune on Hindi data). |
Real Picture: NLP in Action
A screenshot of Google Translate converting "मेरो नाम सूर्य हो" to "My name is Surya." Caption: Semantic and syntactic analysis at work: the system maps Nepali words to English equivalents while preserving grammar.
Worked Example: Parsing a Nepali Sentence
Input: "मेरो बाइकले पेट्रोल भर्नु पर्छ।" (My bike needs petrol.)
Step 1: Tokenization
["मेरो", "बाइकले", "पेट्रोल", "भर्नु", "पर्छ", "।"]
Step 2: POS Tagging
| Token | POS Tag (Nepali) |
|---|---|
| मेरो | Possessive Pronoun |
| बाइकले | Noun + Postposition |
| पेट्रोल | Noun |
| भर्नु | Verb (Infinitive) |
| पर्छ | Auxiliary Verb |
Step 3: Dependency Parse Tree
Translation: need(petrol, my_bike)
Syntactic vs. Semantic Analysis
| Feature | Syntactic Analysis | Semantic Analysis |
|---|---|---|
| Focus | Grammar and sentence structure | Meaning and knowledge |
| Example | "She [subject] runs [verb] quickly [adverb]." | "She" refers to a female runner; "quickly" implies speed. |
| Tools | Parse trees, CFG (Context-Free Grammar) | Ontologies, word embeddings (e.g., Word2Vec) |
| Challenge | Ambiguous grammars (e.g., "Time flies like an arrow.") | Polysemy (e.g., "bank" as river or institution) |
Natural Language Generation (NLG) vs. Understanding
| Aspect | NLG (Generating Text) | NLU (Understanding Text) |
|---|---|---|
| Goal | Create human-like text | Extract meaning from text |
| Example | Chatbots replying "Your order #123 is confirmed." | Parsing "Cancel order #123" to extract cancel(order_id=123). |
| Steps | 1. Plan content <br> 2. Sentence planning <br> 3. Realization | 1. Tokenization <br> 2. Parsing <br> 3. Semantic analysis |
| Challenge | Fluency and coherence | Disambiguation and context |
Morphological Analysis in Nepali
Nepali is highly inflected, meaning words change based on context. Example:
- "बाटो" (bāṭo) can mean:
- Noun: "road" (बाटो = road)
- Verb: "to go" (बाटो = from, postposition)
- Adjective: "from which" (बाटो = postposition)
Worked Example: Analyzing "बाटोमा"
- Token: "बाटोमा"
- Breakdown:
- Stem: "बाटो" (root: "बाट" = path)
- Suffix: "मा" (locative case marker)
- Meaning: "on/at the road"
Visual:
Exam Tip
- PEAS Framework: Always define Performance (metrics), Environment (context), Actuators (outputs), and Sensors (inputs). Use real examples (e.g., eSewa for actuators = payment confirmation).
- NLP Pipeline: Memorize the 5 steps (tokenization → morphology → syntax → semantics → pragmatics) and one example per step. For syntax, draw a parse tree (even if simple).
- Challenges: Link challenges to real-world failures:
- Ambiguity → "I shot an elephant in my pajamas" (was the elephant in pajamas?).
- Context → "The stock market crashed. He died." (Who is "he"?).
- Comparisons: For NLG vs. NLU, use a table and one sentence per cell.
- Nepali Focus: Highlight morphological complexity (e.g., case markers, verb conjugations) and low-resource challenges (e.g., lack of labeled data).
Summary Checklist
- Can define PEAS for two agents (e.g., Daraz + tutor).
- Can draw a parse tree for a simple sentence.
- Knows 3 NLP challenges + solutions.
- Can tokenize and POS-tag a Nepali sentence.
- Understands NLG vs. NLU with examples.
Based on the TU BSc CSIT syllabus for Artificial Intelligence (CSC266), unit 5.
Discussion
Loading…