Knowledge EngineeringUnit 1018 min read
Applications of Knowledge Engineering: Systems, AI, and Real-World Impact
Unit 10 of Knowledge Engineering explores how ontologies, NLP, semantic web, probabilistic reasoning, and AI-driven systems (like chatbots, recommendation engines, and healthcare diagnostics) are implemented in modern industries—with Nepalese and global case studies, technical workflows, and exam-focused comparisons.
TAKEAWAYS:
- Ontologies act as structured "vocabularies" for machines (e.g., eSewa’s service categories) and enable cross-system knowledge sharing via OWL/RDF.
- NLP pipelines (tokenization → POS tagging → semantic parsing) power apps like Khalti’s chat support and Daraz’s product search—real examples show how morphology/lexicon/syntax work together.
- Knowledge-Based Systems (KBS) in healthcare (e.g., NTC’s fault prediction) use rule engines and case-based reasoning to handle uncertainty, but struggle with dynamic data.
- Semantic Web (RDF/SPARQL) lets NEPSE and global platforms (Google Knowledge Graph) link disparate datasets—visualize how triples connect entities like "Company → IPO → StockPrice."
- Probabilistic reasoning (Bayes’ theorem, Dempster-Shafer) models uncertainty in real systems: e.g., Pathao’s route optimization or bank loan approvals with incomplete data.
- AI + Knowledge Engineering fusion drives modern apps: WhatsApp’s intent classification (NLP), YouTube’s recommendation (collaborative filtering + ontologies), and Ncell’s predictive maintenance (SVM + time-series data).
1. Ontologies in Action: Structuring Knowledge for Machines
Ontologies are formal, machine-readable specifications of concepts, relationships, and rules in a domain. They act as the "DNA" of knowledge-sharing systems, enabling interoperability between disparate databases.
How Ontologies Work
- Classes & Hierarchies: Define categories (e.g.,
Vehicle → {Car, Bike, Truck}). - Properties & Relationships: Link entities (e.g.,
Car → hasEngine → Engine). - Axioms/Rules: Enforce constraints (e.g.,
if Vehicle has "electric" then must have "batteryCapacity").
Visual: Ontology Hierarchy for eSewa Services
Real Example: eSewa’s Ontology eSewa uses an ontology to classify services (bill payments, remittances, top-ups). When you select "Electricity Bill," the system:
- Checks the
ElectricityBillclass for required slots (customerID,dueDate). - Validates against the
ServiceProvider(NTC) rules (e.g., "dueDate must be ≤ today"). - Generates an RDF triple:
:Bill_123 a eSewa:ElectricityBill ; eSewa:customerID "KATHMANDU-456" ; eSewa:dueDate "2024-05-15" ; eSewa:provider :NTC .
Worked Example: Daraz’s Product Ontology Daraz’s search engine uses an ontology to map user queries to products. For the query "Nike running shoes size 9":
- Tokenization: Split into
["Nike", "running", "shoes", "size", "9"]. - Semantic Matching:
- "Nike" →
Brandclass. - "running shoes" →
ProductTypesubclass ofFootwear. - "size 9" →
hasSizeproperty.
- "Nike" →
- SPARQL Query:
SELECT ?product WHERE { ?product a daraz:Product ; daraz:brand "Nike" ; daraz:productType daraz:RunningShoes ; daraz:size "9" . }
2. Natural Language Processing (NLP): From Text to Knowledge
NLP bridges human language and machine understanding. In Knowledge Engineering, it’s used for:
- Information extraction (e.g., parsing Ncell customer complaints).
- Chatbots (e.g., Khalti’s FAQ system).
- Sentiment analysis (e.g., Daraz reviews).
NLP Pipeline with Real Numbers
Input: "My Daraz order #DZ12345 was delivered late. The package was damaged."
| Step | Process | Output |
|---|---|---|
| Tokenization | Split into words/punctuation | ["My", "Daraz", "order", "#DZ12345", "was", "delivered", "late", ...] |
| POS Tagging | Label parts of speech | ["PRON", "PROPN", "NOUN", "HASHTAG", "VERB", "ADJ", ...] |
| Named Entity Recognition (NER) | Identify entities | OrderID: "DZ12345", Issue: "delayed_delivery", Issue: "damaged" |
| Dependency Parsing | Relationships between words | delivered ← order (subject-verb), damaged ← package |
| Semantic Role Labeling | Extract key actions/issues | Event: "delivery", Role: "recipient" → "customer", Role: "issue" → "damage" |
Visual: NLP Pipeline for Customer Complaint
flowchart LR
A["Input Text: 'Order DZ12345 was delivered late.'"] --> B["Tokenization"]
B --> C["POS Tagging: NOUN, VERB, ADJ"]
C --> D["NER: OrderID=DZ12345, Issue=delay"]
D --> E["Dependency Parse: order ← delivered"]
E --> F["Semantic Role: Event=delivery, Role=delay"]
F --> G["Knowledge Base Update: Add complaint record"]Real Example: Khalti’s Chatbot Khalti’s customer support bot uses:
- Intent Classification: Detects user intent (e.g., "refund," "transaction status").
- Slot Filling: Extracts entities like
transactionID,amount. - Rule Engine: Matches intent to predefined responses or escalates to human agent.
- Example: If user says "I want a refund for transaction ID KH56789", the bot:
- Checks the
Transactionontology forstatus = "pending". - Generates a refund request in the
Refundclass.
- Checks the
- Example: If user says "I want a refund for transaction ID KH56789", the bot:
3. Semantic Web: Linking Data Like the Web of Knowledge
The Semantic Web extends the traditional web by adding meaning to data via:
- RDF (Resource Description Framework): Triples of
(subject, predicate, object). - OWL (Web Ontology Language): Defines classes, properties, and constraints.
- SPARQL: Query language for RDF data.
Visual: RDF Triples for NEPSE Stock Data
graph TD
A["NEPSE:Company"] -->|"hasSymbol"| B["NEPSE:Symbol"]
A -->|"hasName"| C["NEPSE:Name"]
A -->|"hasPrice"| D["NEPSE:Price"]
B -->|"value"| E["NEPSE:NPAL50"]
C -->|"value"| F["NEPSE:Nepal Bank Ltd"]
D -->|"value"| G["NEPSE:1250.50"]Real Example: Google Knowledge Graph When you search "Sagarmatha National Park":
- Google’s semantic engine queries its RDF knowledge base for triples like:
:SagarmathaPark a schema:NationalPark ; schema:name "Sagarmatha National Park" ; schema:location :Nepal ; schema:area "1148" ; schema:protectedAreaDesignation schema:NationalPark . - It also links to related entities:
:SagarmathaPark schema:contains :Everest ; schema:managedBy :DepartmentOfNationalParksAndWildlifeConservation .
Worked Example: NTC’s Fault Prediction with RDF NTC uses RDF to model power grid faults:
- Triple 1:
:Transformer_456 a ntc:Transformer ; ntc:location "Kathmandu" ; ntc:status "overheating" . - Triple 2:
:Transformer_456 ntc:connectedTo :Substation_789 . - SPARQL Query to find all overheating transformers:
SELECT ?transformer WHERE { ?transformer a ntc:Transformer ; ntc:status "overheating" . }
4. Knowledge-Based Systems (KBS) in Healthcare and Beyond
KBS use rules, cases, and heuristics to solve problems. Key components:
- Rule Engine:
IF condition THEN action(e.g., "IF blood pressure > 140 THEN alert doctor"). - Case-Based Reasoning (CBR): Reuses past solutions (e.g., "Patient X had symptoms A,B,C → treatment Y").
- Expert Systems: Mimic human expertise (e.g., IBM Watson for Oncology).
Advantages/Disadvantages Table
| Advantage | Disadvantage | Example |
|---|---|---|
| Handles complex, rule-based domains | Struggles with dynamic/unstructured data | NTC’s fault diagnosis system |
| Explains decisions (transparency) | Requires manual rule maintenance | Bank loan approval systems |
| Works well with incomplete data | Poor at learning from new data | Pathao’s route optimization (early versions) |
Real Example: NTC’s Fault Prediction KBS
- Knowledge Base:
- Rules:
IF transformer_temperature > 90°C THEN classify_as "high_risk". - Cases: Past faults stored as
(symptoms → cause → solution).
- Rules:
- Inference Engine:
- Input: Sensor data shows
Transformer_456at 92°C. - Output: Triggers alert → dispatch technician.
- Input: Sensor data shows
- CBR Step:
- Matches to past case: "Transformer_123 at 91°C → cause: loose connection → solution: tighten bolts."
- Suggests same solution for
Transformer_456.
Visual: KBS for Healthcare Diagnosis
5. Probabilistic Reasoning: Handling Uncertainty
Real-world data is noisy. Probabilistic methods (Bayes’ theorem, Dempster-Shafer) quantify uncertainty.
Key Concepts
- Bayesian Networks: Graphical models of dependencies (e.g., "Rain → Wet Ground").
- Dempster-Shafer Theory: Handles incomplete evidence (e.g., "Is this email spam?" with 70% confidence).
Worked Example: Pathao’s Route Optimization Pathao uses probabilistic reasoning to predict delays:
- Variables:
TrafficJam(probability = 0.6),Accident(0.2),Construction(0.1).
- Bayesian Network:
- Calculation:
- If
TrafficJamis observed, updateP(Delay|TrafficJam)using Bayes’ theorem: - Suppose:
- (prior probability).
- .
- .
- Then:
- Pathao adjusts ETA: "Your ride may be delayed by 15 minutes (67% confidence)."
- If
Real Example: Bank Loan Approval A bank uses probabilistic reasoning to approve loans:
- Features: Income, credit score, loan amount.
- Decision Rule:
- If , reject.
- Bayesian Update:
- Prior: .
- Likelihood: .
- Posterior:
6. AI + Knowledge Engineering: The Modern Fusion
Modern systems combine AI (machine learning) with knowledge engineering for robustness.
Examples
| Application | Knowledge Engineering Component | AI Component | Example |
|---|---|---|---|
| YouTube Recommendations | Ontology of user preferences | Collaborative filtering | "Users who liked X also liked Y" |
| WhatsApp Business Chatbots | Intent classification rules | NLP (transformers) | "Hi, I want to book a ticket." → Intent: "Booking" |
| Ncell Predictive Maintenance | Fault ontology (symptoms → causes) | SVM on sensor data | Predicts engine failure 3 days early |
| Google Search | Knowledge Graph (RDF triples) | RankBrain (deep learning) | Answers questions using structured data |
Visual: AI + KBS Pipeline for Ncell Maintenance
flowchart LR
A["Sensor Data"] --> B["Preprocess: Clean & Normalize"]
B --> C["KBS: Match to Fault Ontology"]
C --> D["AI: SVM Classifier"]
D --> E["Probability: 85% Failure Risk"]
E --> F["Action: Schedule Maintenance"]In the Real World
eSewa’s Ontology-Driven Services
- Idea Used: Ontology engineering (OWL/RDF) to classify services like
BillPayment,Remittance, andTopUp. - How: When you select "Electricity Bill," eSewa’s backend queries its ontology to validate required fields (e.g.,
customerID,dueDate) before processing. This reduces errors and enables seamless integration with NTC’s systems. - Visual: The ontology hierarchy (shown above) ensures that only valid combinations (e.g.,
ElectricityBillcannot have arecipientAddress—it’s forRemittance) are allowed.
- Idea Used: Ontology engineering (OWL/RDF) to classify services like
Khalti’s NLP-Powered Chatbot
- Idea Used: NLP pipeline (tokenization → intent classification → slot filling).
- How: When a user types "I forgot my Khalti PIN," the bot:
- Tokenizes:
["I", "forgot", "my", "Khalti", "PIN"]. - Classifies Intent: "Forgot PIN" → triggers
resetPINworkflow. - Fills Slots: Extracts
accountIDfrom user profile. - Executes Rule: Checks if
lastLogin < 24h(security check), then sends OTP.
- Tokenizes:
- Real Impact: Handles 60% of customer queries without human intervention, reducing wait times.
NTC’s Fault Prediction with KBS
- Idea Used: Rule-based reasoning + case-based reasoning (CBR).
- How: During the 2023 Kathmandu blackout:
- Sensors detected
Transformer_456at 95°C (rule:IF temperature > 90 THEN alert). - CBR matched to past case: "Transformer_123 at 92°C → cause: loose connection → solution: tighten bolts."
- Technicians acted within 2 hours, restoring power faster than manual diagnosis.
- Sensors detected
- Probabilistic Twist: NTC now uses Bayesian networks to predict which transformers are most likely to fail next, prioritizing maintenance.
Pathao’s Probabilistic Route Planning
- Idea Used: Bayesian inference for uncertainty in traffic data.
- How: When you book a ride from Thapathali to Lakankhel:
- Pathao’s system checks real-time traffic (e.g.,
P(TrafficJam) = 0.7on Ring Road). - Uses Bayes’ theorem to update
P(Delay|TrafficJam)and adjusts ETA dynamically. - If
P(Delay) > 0.8, it suggests an alternative route via Swoyambhu.
- Pathao’s system checks real-time traffic (e.g.,
- Result: Reduces average delay by 20% compared to deterministic routing.
NEPSE’s Semantic Web Integration
- Idea Used: RDF/SPARQL for linking stock data.
- How: Investors query NEPSE’s semantic database to find:
- "Show me all companies in the Banking sector with P/E ratio > 20."
- Translated to SPARQL:
SELECT ?company WHERE { ?company a nepse:Company ; nepse:sector "Banking" ; nepse:peRatio ?ratio . FILTER (?ratio > 20) }
- Output: Returns
Nepal Bank Ltd,Global IME Bank, etc., with linked data on market cap, dividend yield.
Exam Tip
This unit is conceptual but applied—expect:
Compare/Contrast Questions:
- "How does ontology differ from a traditional database schema?"
- Answer: Ontologies use classes/properties/axioms (OWL) for semantic relationships, while databases use tables/rows with rigid schemas. Example: An ontology can say
Car → hasEngine → Engine, but a database would need aJOINtable.
- Answer: Ontologies use classes/properties/axioms (OWL) for semantic relationships, while databases use tables/rows with rigid schemas. Example: An ontology can say
- "NLP vs. Rule-Based Systems in chatbots."
- Table:
Feature NLP (e.g., Khalti) Rule-Based (e.g., IVR) Handles Ambiguity Yes (e.g., "book ticket" → intent) No (requires exact phrases) Scalability High (learns from data) Low (rules must be manually added) Example "Reset my Khalti PIN" "Press 1 for balance, 2 for mini-statement"
- Table:
- "How does ontology differ from a traditional database schema?"
Application-Based Questions:
- "Design a KBS for Daraz’s customer complaint system."
- Structure:
- Knowledge Base:
- Rules:
IF complaint_type = "damaged" AND delivery_status = "delivered" THEN trigger refund. - Cases: Past complaints stored as
(symptoms → resolution).
- Rules:
- Inference Engine:
- Input: User uploads photo of damaged shoes + order ID.
- Output: "Refund initiated. Estimated time: 3–5 days."
- Visual: Draw a KBS flowchart (like the healthcare example above) with Daraz-specific nodes.
- Knowledge Base:
- Structure:
- "Design a KBS for Daraz’s customer complaint system."
Probability/Uncertainty Questions:
- "Calculate the probability that a loan applicant defaults given their income and credit score."
- Steps:
- Define events: = Default, = Income < 50k, = Score < 650.
- Use Bayes’ theorem:
- Plug in numbers (from exam data or assume values like , ).
- Steps:
- "Calculate the probability that a loan applicant defaults given their income and credit score."
Semantic Web Questions:
- "Write RDF triples for a university course ontology."
- Example:
:CACS458 a cacs:Course ; cacs:courseCode "CACS458" ; cacs:title "Knowledge Engineering" ; cacs:credits "3" ; cacs:prerequisite :CACS357 .
- Example:
- "How does SPARQL differ from SQL?"
- Table:
Feature SPARQL SQL Data Model Triples (subject-predicate-object) Tables (rows/columns) Query Focus Semantic relationships Tabular data Example SELECT ?student WHERE { ?student rdf:type ex:Student }SELECT * FROM Students
- Table:
- "Write RDF triples for a university course ontology."
Short-Answer Tricks:
- For "What is ontology?", define it as:
"A formal, machine-readable specification of a domain’s concepts, relationships, and constraints, enabling knowledge sharing across systems (e.g., eSewa’s service ontology)."
- For "Probabilistic reasoning," give a real-world analogy:
"Like a doctor diagnosing a patient: If a patient has a fever (70% chance of flu) and cough (60% chance of flu), the doctor combines these probabilities to conclude ‘80% chance of flu’ rather than assuming certainty."
- For "What is ontology?", define it as:
Final Advice:
- Draw diagrams for ontologies (class hierarchies), NLP pipelines (flowcharts), and KBS (rule engines).
- Use Nepalese examples (eSewa, NTC, Khalti) to stand out—they’re less common in model answers.
- For probability questions, always show the Bayes’ theorem formula and plug in numbers, even if assumed.
- Memorize RDF/OWL basics: Know that RDF uses triples and OWL adds logic (e.g.,
subClassOf,equivalentClass).
Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 10.
Discussion
Loading…