CACS458 Knowledge Engineering

Knowledge EngineeringUnit 1018 min read

Applications of Knowledge Engineering: Systems, AI, and Real-World Impact

Unit 10 of Knowledge Engineering explores how ontologies, NLP, semantic web, probabilistic reasoning, and AI-driven systems (like chatbots, recommendation engines, and healthcare diagnostics) are implemented in modern industries—with Nepalese and global case studies, technical workflows, and exam-focused comparisons.

TAKEAWAYS:

  • Ontologies act as structured "vocabularies" for machines (e.g., eSewa’s service categories) and enable cross-system knowledge sharing via OWL/RDF.
  • NLP pipelines (tokenization → POS tagging → semantic parsing) power apps like Khalti’s chat support and Daraz’s product search—real examples show how morphology/lexicon/syntax work together.
  • Knowledge-Based Systems (KBS) in healthcare (e.g., NTC’s fault prediction) use rule engines and case-based reasoning to handle uncertainty, but struggle with dynamic data.
  • Semantic Web (RDF/SPARQL) lets NEPSE and global platforms (Google Knowledge Graph) link disparate datasets—visualize how triples connect entities like "Company → IPO → StockPrice."
  • Probabilistic reasoning (Bayes’ theorem, Dempster-Shafer) models uncertainty in real systems: e.g., Pathao’s route optimization or bank loan approvals with incomplete data.
  • AI + Knowledge Engineering fusion drives modern apps: WhatsApp’s intent classification (NLP), YouTube’s recommendation (collaborative filtering + ontologies), and Ncell’s predictive maintenance (SVM + time-series data).

1. Ontologies in Action: Structuring Knowledge for Machines

Ontologies are formal, machine-readable specifications of concepts, relationships, and rules in a domain. They act as the "DNA" of knowledge-sharing systems, enabling interoperability between disparate databases.

How Ontologies Work

  • Classes & Hierarchies: Define categories (e.g., Vehicle → {Car, Bike, Truck}).
  • Properties & Relationships: Link entities (e.g., Car → hasEngine → Engine).
  • Axioms/Rules: Enforce constraints (e.g., if Vehicle has "electric" then must have "batteryCapacity").

Visual: Ontology Hierarchy for eSewa Services

ElectricityBillBillPaymentService
Hierarchy: eSewa Service Ontology (abstract Service → BillPayment → ElectricityBill)

Real Example: eSewa’s Ontology eSewa uses an ontology to classify services (bill payments, remittances, top-ups). When you select "Electricity Bill," the system:

  1. Checks the ElectricityBill class for required slots (customerID, dueDate).
  2. Validates against the ServiceProvider (NTC) rules (e.g., "dueDate must be ≤ today").
  3. Generates an RDF triple:
    :Bill_123 a eSewa:ElectricityBill ;
              eSewa:customerID "KATHMANDU-456" ;
              eSewa:dueDate "2024-05-15" ;
              eSewa:provider :NTC .
    

Worked Example: Daraz’s Product Ontology Daraz’s search engine uses an ontology to map user queries to products. For the query "Nike running shoes size 9":

  1. Tokenization: Split into ["Nike", "running", "shoes", "size", "9"].
  2. Semantic Matching:
    • "Nike" → Brand class.
    • "running shoes" → ProductType subclass of Footwear.
    • "size 9" → hasSize property.
  3. SPARQL Query:
    SELECT ?product WHERE {
      ?product a daraz:Product ;
               daraz:brand "Nike" ;
               daraz:productType daraz:RunningShoes ;
               daraz:size "9" .
    }
    

2. Natural Language Processing (NLP): From Text to Knowledge

NLP bridges human language and machine understanding. In Knowledge Engineering, it’s used for:

  • Information extraction (e.g., parsing Ncell customer complaints).
  • Chatbots (e.g., Khalti’s FAQ system).
  • Sentiment analysis (e.g., Daraz reviews).

NLP Pipeline with Real Numbers

Input: "My Daraz order #DZ12345 was delivered late. The package was damaged."

Step Process Output
Tokenization Split into words/punctuation ["My", "Daraz", "order", "#DZ12345", "was", "delivered", "late", ...]
POS Tagging Label parts of speech ["PRON", "PROPN", "NOUN", "HASHTAG", "VERB", "ADJ", ...]
Named Entity Recognition (NER) Identify entities OrderID: "DZ12345", Issue: "delayed_delivery", Issue: "damaged"
Dependency Parsing Relationships between words delivered ← order (subject-verb), damaged ← package
Semantic Role Labeling Extract key actions/issues Event: "delivery", Role: "recipient" → "customer", Role: "issue" → "damage"

Visual: NLP Pipeline for Customer Complaint

flowchart LR
    A["Input Text: 'Order DZ12345 was delivered late.'"] --> B["Tokenization"]
    B --> C["POS Tagging: NOUN, VERB, ADJ"]
    C --> D["NER: OrderID=DZ12345, Issue=delay"]
    D --> E["Dependency Parse: order ← delivered"]
    E --> F["Semantic Role: Event=delivery, Role=delay"]
    F --> G["Knowledge Base Update: Add complaint record"]

Real Example: Khalti’s Chatbot Khalti’s customer support bot uses:

  1. Intent Classification: Detects user intent (e.g., "refund," "transaction status").
  2. Slot Filling: Extracts entities like transactionID, amount.
  3. Rule Engine: Matches intent to predefined responses or escalates to human agent.
    • Example: If user says "I want a refund for transaction ID KH56789", the bot:
      • Checks the Transaction ontology for status = "pending".
      • Generates a refund request in the Refund class.

3. Semantic Web: Linking Data Like the Web of Knowledge

The Semantic Web extends the traditional web by adding meaning to data via:

  • RDF (Resource Description Framework): Triples of (subject, predicate, object).
  • OWL (Web Ontology Language): Defines classes, properties, and constraints.
  • SPARQL: Query language for RDF data.

Visual: RDF Triples for NEPSE Stock Data

graph TD
    A["NEPSE:Company"] -->|"hasSymbol"| B["NEPSE:Symbol"]
    A -->|"hasName"| C["NEPSE:Name"]
    A -->|"hasPrice"| D["NEPSE:Price"]
    B -->|"value"| E["NEPSE:NPAL50"]
    C -->|"value"| F["NEPSE:Nepal Bank Ltd"]
    D -->|"value"| G["NEPSE:1250.50"]

Real Example: Google Knowledge Graph When you search "Sagarmatha National Park":

  1. Google’s semantic engine queries its RDF knowledge base for triples like:
    :SagarmathaPark a schema:NationalPark ;
                    schema:name "Sagarmatha National Park" ;
                    schema:location :Nepal ;
                    schema:area "1148" ;
                    schema:protectedAreaDesignation schema:NationalPark .
    
  2. It also links to related entities:
    :SagarmathaPark schema:contains :Everest ;
                    schema:managedBy :DepartmentOfNationalParksAndWildlifeConservation .
    

Worked Example: NTC’s Fault Prediction with RDF NTC uses RDF to model power grid faults:

  • Triple 1: :Transformer_456 a ntc:Transformer ; ntc:location "Kathmandu" ; ntc:status "overheating" .
  • Triple 2: :Transformer_456 ntc:connectedTo :Substation_789 .
  • SPARQL Query to find all overheating transformers:
    SELECT ?transformer WHERE {
      ?transformer a ntc:Transformer ;
                   ntc:status "overheating" .
    }
    

4. Knowledge-Based Systems (KBS) in Healthcare and Beyond

KBS use rules, cases, and heuristics to solve problems. Key components:

  • Rule Engine: IF condition THEN action (e.g., "IF blood pressure > 140 THEN alert doctor").
  • Case-Based Reasoning (CBR): Reuses past solutions (e.g., "Patient X had symptoms A,B,C → treatment Y").
  • Expert Systems: Mimic human expertise (e.g., IBM Watson for Oncology).

Advantages/Disadvantages Table

Advantage Disadvantage Example
Handles complex, rule-based domains Struggles with dynamic/unstructured data NTC’s fault diagnosis system
Explains decisions (transparency) Requires manual rule maintenance Bank loan approval systems
Works well with incomplete data Poor at learning from new data Pathao’s route optimization (early versions)

Real Example: NTC’s Fault Prediction KBS

  1. Knowledge Base:
    • Rules: IF transformer_temperature > 90°C THEN classify_as "high_risk".
    • Cases: Past faults stored as (symptoms → cause → solution).
  2. Inference Engine:
    • Input: Sensor data shows Transformer_456 at 92°C.
    • Output: Triggers alert → dispatch technician.
  3. CBR Step:
    • Matches to past case: "Transformer_123 at 91°C → cause: loose connection → solution: tighten bolts."
    • Suggests same solution for Transformer_456.

Visual: KBS for Healthcare Diagnosis

IF fever AND cough THENIF chest pain AND shortness THENPossiblePossibleCheckEmergency ProtocolSymptomsRule EngineFluHeart AttackLab ResultsEmergency
KBS for Healthcare Diagnosis (Flu vs. Heart Attack Pathways)

5. Probabilistic Reasoning: Handling Uncertainty

Real-world data is noisy. Probabilistic methods (Bayes’ theorem, Dempster-Shafer) quantify uncertainty.

Key Concepts

  • Bayesian Networks: Graphical models of dependencies (e.g., "Rain → Wet Ground").
  • Dempster-Shafer Theory: Handles incomplete evidence (e.g., "Is this email spam?" with 70% confidence).

Worked Example: Pathao’s Route Optimization Pathao uses probabilistic reasoning to predict delays:

  1. Variables:
    • TrafficJam (probability = 0.6), Accident (0.2), Construction (0.1).
  2. Bayesian Network:
-5-4-3-2-1123450.10.20.30.40.50.6yTrafficJamAccidentConstruction
Pathao Delay Probabilities (Bayesian Network Weights)
  1. Calculation:
    • If TrafficJam is observed, update P(Delay|TrafficJam) using Bayes’ theorem:
    • Suppose:
      • (prior probability).
      • .
      • .
    • Then:
    • Pathao adjusts ETA: "Your ride may be delayed by 15 minutes (67% confidence)."

Real Example: Bank Loan Approval A bank uses probabilistic reasoning to approve loans:

  • Features: Income, credit score, loan amount.
  • Decision Rule:
    • If , reject.
  • Bayesian Update:
    • Prior: .
    • Likelihood: .
    • Posterior:

6. AI + Knowledge Engineering: The Modern Fusion

Modern systems combine AI (machine learning) with knowledge engineering for robustness.

startTemperature > 80°CVibration > ThresholdCooling AppliedNormalWarningFailure
Ncell Predictive Maintenance State Machine (Fault Ontology Trigger)

Examples

Application Knowledge Engineering Component AI Component Example
YouTube Recommendations Ontology of user preferences Collaborative filtering "Users who liked X also liked Y"
WhatsApp Business Chatbots Intent classification rules NLP (transformers) "Hi, I want to book a ticket." → Intent: "Booking"
Ncell Predictive Maintenance Fault ontology (symptoms → causes) SVM on sensor data Predicts engine failure 3 days early
Google Search Knowledge Graph (RDF triples) RankBrain (deep learning) Answers questions using structured data

Visual: AI + KBS Pipeline for Ncell Maintenance

flowchart LR
    A["Sensor Data"] --> B["Preprocess: Clean & Normalize"]
    B --> C["KBS: Match to Fault Ontology"]
    C --> D["AI: SVM Classifier"]
    D --> E["Probability: 85% Failure Risk"]
    E --> F["Action: Schedule Maintenance"]

In the Real World

  1. eSewa’s Ontology-Driven Services

    • Idea Used: Ontology engineering (OWL/RDF) to classify services like BillPayment, Remittance, and TopUp.
    • How: When you select "Electricity Bill," eSewa’s backend queries its ontology to validate required fields (e.g., customerID, dueDate) before processing. This reduces errors and enables seamless integration with NTC’s systems.
    • Visual: The ontology hierarchy (shown above) ensures that only valid combinations (e.g., ElectricityBill cannot have a recipientAddress—it’s for Remittance) are allowed.
  2. Khalti’s NLP-Powered Chatbot

    • Idea Used: NLP pipeline (tokenization → intent classification → slot filling).
    • How: When a user types "I forgot my Khalti PIN," the bot:
      • Tokenizes: ["I", "forgot", "my", "Khalti", "PIN"].
      • Classifies Intent: "Forgot PIN" → triggers resetPIN workflow.
      • Fills Slots: Extracts accountID from user profile.
      • Executes Rule: Checks if lastLogin < 24h (security check), then sends OTP.
    • Real Impact: Handles 60% of customer queries without human intervention, reducing wait times.
  3. NTC’s Fault Prediction with KBS

    • Idea Used: Rule-based reasoning + case-based reasoning (CBR).
    • How: During the 2023 Kathmandu blackout:
      • Sensors detected Transformer_456 at 95°C (rule: IF temperature > 90 THEN alert).
      • CBR matched to past case: "Transformer_123 at 92°C → cause: loose connection → solution: tighten bolts."
      • Technicians acted within 2 hours, restoring power faster than manual diagnosis.
    • Probabilistic Twist: NTC now uses Bayesian networks to predict which transformers are most likely to fail next, prioritizing maintenance.
  4. Pathao’s Probabilistic Route Planning

    • Idea Used: Bayesian inference for uncertainty in traffic data.
    • How: When you book a ride from Thapathali to Lakankhel:
      • Pathao’s system checks real-time traffic (e.g., P(TrafficJam) = 0.7 on Ring Road).
      • Uses Bayes’ theorem to update P(Delay|TrafficJam) and adjusts ETA dynamically.
      • If P(Delay) > 0.8, it suggests an alternative route via Swoyambhu.
    • Result: Reduces average delay by 20% compared to deterministic routing.
  5. NEPSE’s Semantic Web Integration

    • Idea Used: RDF/SPARQL for linking stock data.
    • How: Investors query NEPSE’s semantic database to find:
      • "Show me all companies in the Banking sector with P/E ratio > 20."
      • Translated to SPARQL:
        SELECT ?company WHERE {
          ?company a nepse:Company ;
                   nepse:sector "Banking" ;
                   nepse:peRatio ?ratio .
          FILTER (?ratio > 20)
        }
        
    • Output: Returns Nepal Bank Ltd, Global IME Bank, etc., with linked data on market cap, dividend yield.

Exam Tip

This unit is conceptual but applied—expect:

  1. Compare/Contrast Questions:

    • "How does ontology differ from a traditional database schema?"
      • Answer: Ontologies use classes/properties/axioms (OWL) for semantic relationships, while databases use tables/rows with rigid schemas. Example: An ontology can say Car → hasEngine → Engine, but a database would need a JOIN table.
    • "NLP vs. Rule-Based Systems in chatbots."
      • Table:
        Feature NLP (e.g., Khalti) Rule-Based (e.g., IVR)
        Handles Ambiguity Yes (e.g., "book ticket" → intent) No (requires exact phrases)
        Scalability High (learns from data) Low (rules must be manually added)
        Example "Reset my Khalti PIN" "Press 1 for balance, 2 for mini-statement"
  2. Application-Based Questions:

    • "Design a KBS for Daraz’s customer complaint system."
      • Structure:
        1. Knowledge Base:
          • Rules: IF complaint_type = "damaged" AND delivery_status = "delivered" THEN trigger refund.
          • Cases: Past complaints stored as (symptoms → resolution).
        2. Inference Engine:
          • Input: User uploads photo of damaged shoes + order ID.
          • Output: "Refund initiated. Estimated time: 3–5 days."
        • Visual: Draw a KBS flowchart (like the healthcare example above) with Daraz-specific nodes.
  3. Probability/Uncertainty Questions:

    • "Calculate the probability that a loan applicant defaults given their income and credit score."
      • Steps:
        1. Define events: = Default, = Income < 50k, = Score < 650.
        2. Use Bayes’ theorem:
        3. Plug in numbers (from exam data or assume values like , ).
  4. Semantic Web Questions:

    • "Write RDF triples for a university course ontology."
      • Example:
        :CACS458 a cacs:Course ;
                cacs:courseCode "CACS458" ;
                cacs:title "Knowledge Engineering" ;
                cacs:credits "3" ;
                cacs:prerequisite :CACS357 .
        
    • "How does SPARQL differ from SQL?"
      • Table:
        Feature SPARQL SQL
        Data Model Triples (subject-predicate-object) Tables (rows/columns)
        Query Focus Semantic relationships Tabular data
        Example SELECT ?student WHERE { ?student rdf:type ex:Student } SELECT * FROM Students
  5. Short-Answer Tricks:

    • For "What is ontology?", define it as:

      "A formal, machine-readable specification of a domain’s concepts, relationships, and constraints, enabling knowledge sharing across systems (e.g., eSewa’s service ontology)."

    • For "Probabilistic reasoning," give a real-world analogy:

      "Like a doctor diagnosing a patient: If a patient has a fever (70% chance of flu) and cough (60% chance of flu), the doctor combines these probabilities to conclude ‘80% chance of flu’ rather than assuming certainty."


Final Advice:

  • Draw diagrams for ontologies (class hierarchies), NLP pipelines (flowcharts), and KBS (rule engines).
  • Use Nepalese examples (eSewa, NTC, Khalti) to stand out—they’re less common in model answers.
  • For probability questions, always show the Bayes’ theorem formula and plug in numbers, even if assumed.
  • Memorize RDF/OWL basics: Know that RDF uses triples and OWL adds logic (e.g., subClassOf, equivalentClass).

Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 10.

Discussion

Loading…