CACS458 Knowledge Engineering

Knowledge EngineeringUnit 814 min read

Knowledge-Based Systems & Case-Based Reasoning: Models, Workflows & Applications

Unit 8 of Knowledge Engineering explores how knowledge-based systems (KBS) encode domain expertise into rule-based or case-based reasoning engines, and how case-based reasoning (CBR) solves new problems by adapting past solutions. Covers architectures, CBR cycles, real-world deployments (e.g., medical diagnosis, legal

TAKEAWAYS:

  • Knowledge-Based Systems (KBS) combine domain knowledge (rules, facts) with inference engines to solve problems like medical diagnosis or tax calculation—unlike ML, they explain their reasoning.
  • Case-Based Reasoning (CBR) solves new problems by retrieving, adapting, and reusing past cases (e.g., a bank loan officer approving a new applicant by comparing to similar past loans).
  • The CBR cycle (Retrieve → Reuse → Revise → Retain) turns experience into a feedback loop, improving over time (visualized as a circular flowchart).
  • Advantages of KBS: Transparency (rules are inspectable), works with incomplete data, and handles uncertainty via probabilistic reasoning.
  • Limitations: Brittleness (rules must cover all cases), high maintenance cost, and struggles with unstructured data (unlike deep learning).
  • Real-world tie: eSewa uses CBR to flag fraudulent transactions by matching new payments to past fraud patterns; Daraz’s customer service adapts solutions from resolved complaints.

Core Concepts: Knowledge-Based Systems (KBS)

Order(O123)Customer(C456)FactsIF Damaged THEN RefundIF Premium THEN Offer_ReplacementRulesKnowledge Base
Hierarchical structure of a KBS Knowledge Base (Facts + Rules)

What is a Knowledge-Based System?

A Knowledge-Based System (KBS) is an AI system that uses domain-specific knowledge (facts, rules, heuristics) to solve problems in a way humans would. It consists of:

  1. Knowledge Base (KB): Stores facts (e.g., "All birds can fly") and rules (e.g., "IF wings THEN can_fly").
  2. Inference Engine: Applies logical rules to derive new knowledge (e.g., "Penguin has wings → Penguin can fly" unless it’s an exception).
  3. Working Memory: Holds current data and intermediate results.

Why KBS?

  • Explainability: Unlike black-box ML, KBS shows how it reached a decision (critical for healthcare/legal domains).
  • Efficiency: Solves problems without retraining (e.g., a tax calculator uses fixed rules).
  • Uncertainty Handling: Uses probabilistic logic (e.g., "80% chance of rain → carry umbrella").

How KBS Works: A Worked Example

Scenario: A Daraz customer support agent uses a KBS to resolve order disputes. The KB contains:

  • Facts:
    • Order(O123, "Laptop", "Delivered", "Damaged")
    • Customer(C456, "Premium")
    • Policy(P1, "Premium customers get free replacement")
  • Rules:
    1. IF Order.Status = "Delivered" AND Order.Condition = "Damaged" THEN Flag_for_Refund()
    2. IF Customer.Tier = "Premium" AND Flagged_for_Refund THEN Offer_Replacement()

Trace:

  1. New dispute: Order(O123, ...) is loaded into working memory.
  2. Rule 1 fires → Flag_for_Refund(O123).
  3. Rule 2 fires (since Customer(C456) is Premium) → Offer_Replacement(O123).
  4. Output: "Replace your laptop for free under Policy P1."

Visual: KBS Architecture

flowchart LR
    A["User Input\n(e.g., 'My order is damaged')"] --> B["Working Memory\n(Facts: Order(O123), Customer(C456))"]
    B --> C["Inference Engine\n(Rules: IF Damaged THEN Refund)"]
    C --> D["Knowledge Base\n(Facts + Rules)"]
    C --> E["Output\n('Replace laptop for free')"]
    D -->|"Update"| B

Types of KBS

Type Description Example Limitations
Rule-Based (RB) Uses IF-THEN rules (e.g., expert systems). MYCIN (medical diagnosis) Brittle; rules must cover all cases.
Frame-Based Organizes knowledge as objects with slots/values (e.g., Patient(age=30, symptoms=[fever])). Medical record systems Hard to update dynamically.
Case-Based (CBR) Solves new problems by adapting past cases (see next section). eSewa fraud detection Needs large, labeled case database.
Hybrid Combines rules + ML (e.g., rules for high-stakes decisions, ML for data prep). Bank loan approval systems Complex to maintain.

Case-Based Reasoning (CBR): The 4R Cycle

What is CBR?

CBR solves new problems by reusing solutions to similar past problems. Unlike rule-based systems, it learns from experience. Key idea:

"If a problem is similar to one you’ve solved before, use that solution as a starting point."

Example in Nepal:

  • NTC’s network outage resolution: When a new outage occurs in Kathmandu, technicians retrieve past outages in the same area, adapt the fix (e.g., "Last time, it was a broken pole—check the same pole today"), and update the case database for future use.

The CBR Cycle: Retrieve → Reuse → Revise → Retain

[object Object][object Object][object Object][object Object][object Object]New ProblemRetrieveReuseReviseRetain
The CBR Cycle: Retrieve → Reuse → Revise → Retain (with feedback loop)

Worked Example: eSewa Fraud Detection

Problem: A user reports a suspicious transaction of ₹50,000 to "Nepal Telecom" (but NTC’s real ID is NTCLTD). CBR Steps:

  1. Retrieve:

    • Past fraud cases with:
      • Amount: ₹40,000–₹60,000
      • Recipient: Non-NTC IDs (e.g., NTCLTD123 vs. NTCLTD)
      • Time: Last 24 hours (peak fraud hours: 2–4 AM).
    • Top match: A ₹55,000 scam to NTCLTD999 (similarity score: 0.85).
  2. Reuse:

    • Apply solution from past case: "Flag as fraud, block transaction, notify user."
  3. Revise:

    • Adjust: "Since this is ₹50K (below ₹60K threshold), only block and warn—don’t auto-refund."
  4. Retain:

    • Add new case to database: Case(50000, "NTCLTD123", "Blocked", "User warned").

Real Output:



Similarity Measures in CBR

To find the "most similar" past case, CBR uses metrics like:

  • Numeric attributes: Euclidean distance (e.g., age difference).
  • Categorical attributes: Hamming distance (e.g., "NTC" vs. "NTCLTD" → 1 mismatch).
  • Weighted scores: Combine metrics (e.g., 60% weight to amount, 40% to time).

Example Calculation:

Attribute New Case Past Case Distance Weight Score
Amount (₹) 50,000 55,000 50,000–55,000 = 5,000
Time (hours) 3 AM 2 AM 3–2 = 1 hour
Total 3.4
(Lower score = more similar. Threshold: 4.0 → reject if score > 4.0.)

KBS vs. CBR vs. Machine Learning

Feature Knowledge-Based Systems (KBS) Case-Based Reasoning (CBR) Machine Learning (ML)
Knowledge Source Hand-coded rules/expert input Past cases (experience) Data (labeled/unlabeled)
Learning No learning; static rules Learns by updating case base Learns from data (training)
Explainability High (rules are transparent) Medium (can trace case similarities) Low (black box)
Data Requirements Rules + small datasets Large case database Large labeled datasets
Adaptability Low (rules must be updated manually) High (adapts to new cases) High (generalizes to new data)
Example in Nepal Ncell’s IVR menu ("Press 1 for balance") Daraz’s customer support chatbot NEPSE stock prediction models

Applications of KBS and CBR

1. Healthcare

  • KBS Example: MYCIN (1970s) diagnosed bacterial infections by applying medical rules (e.g., "IF organism is Streptococcus AND patient is allergic to penicillin THEN prescribe erythromycin").
  • CBR Example: Radiology case matching – New X-ray images are compared to past cases with similar symptoms to suggest diagnoses.
flowchart LR
  A["New X-ray
(Pneumonia suspected)"] --> B["Retrieve
Past cases with:
- Cough + fever
- Lung opacity >50%"]
  B --> C["Reuse
Diagnosis: 'Pneumonia (85% match)'"]
  C --> D["Revise
Adjust for:
- Patient age (child vs. adult)
- Allergies (e.g., penicillin)"]
  D --> E["Retain
Add case: 'Child, pneumonia, treated with Amoxicillin'"]
  E -->|"Feedback"| A

2. Finance

  • KBS: Credit scoring (e.g., "IF income > ₹50K AND credit_score > 700 THEN approve loan").
  • CBR: Fraud detection (e.g., eSewa matching new transactions to past scams).
  • CBR: Legal case retrieval – Lawyers input new cases to find similar precedents (e.g., "This contract breach resembles Case X from 2018").

4. Customer Support

  • CBR: Daraz/Pathao chatbots resolve complaints by matching to past resolved issues (e.g., "Delivery delayed in Lalitpur → offer ₹200 coupon").

Advantages and Limitations

Advantages of KBS/CBR

  • Transparency: Rules/cases are human-readable (critical for healthcare/legal domains).
  • No Training Data Needed: KBS works with expert rules; CBR learns from cases.
  • Handles Uncertainty: Uses probabilistic logic (e.g., "70% chance of fraud").
  • Fast for Niche Domains: Outperforms ML when rules/cases are well-defined (e.g., tax calculation).

Limitations

Limitation Impact Mitigation
Brittleness Fails on edge cases not covered by rules/cases. Hybrid systems (rules + ML).
High Maintenance Rules/cases must be updated manually. Automated case retrieval tools.
Scalability Struggles with high-dimensional data (e.g., images). Use ML for feature extraction.
No Generalization CBR only works within the case database. Combine with symbolic reasoning.

Probabilistic Reasoning in KBS

Many KBS handle uncertainty using probabilistic logic (e.g., Bayesian networks). Example:

0.10.20.30.40.50.60.70.80.910.20.40.60.81xyApproval Probability (Credit Score)Base Probability (Employment)Final P=0.9Stable Employment
Probability adjustment for loan approval (Credit Score × Employment Stability)

Scenario: A bank loan officer uses a KBS to approve loans. The KB includes:

  • Facts:
    • Loan_Amount(₹500,000)
    • Credit_Score(720)
    • Employment_Status("Stable")
  • Probabilistic Rules:
    1. IF Credit_Score > 700 THEN Approval_Probability = 0.9
    2. IF Employment_Status = "Unstable" THEN Approval_Probability *= 0.5

Calculation:

  • Base probability: 0.9 (from credit score).
  • Adjusted for stable employment: 0.9 * 1.0 = 0.9 → 90% approval chance.

Visual: Probability Tree


Exam Tip: How to Score Full Marks

  1. Define Clearly: Always start with precise definitions (e.g., "CBR is a problem-solving paradigm that reuses past cases...").
  2. Use Diagrams: Draw the CBR cycle or KBS architecture in exams—it’s worth 3–5 marks.
  3. Compare Tables: For questions like "KBS vs. ML," use a side-by-side table (as above).
  4. Real-World Tie: Link examples to Nepali contexts (e.g., eSewa, Daraz, NTC) for +2 marks.
  5. Step-by-Step Traces: For CBR/KBS questions, show each phase (Retrieve → Reuse → Revise) with concrete data.
  6. Limitations: Always mention 1–2 limitations (e.g., "KBS struggles with unstructured data like images").
  7. Probabilistic Logic: If asked about uncertainty, show a probability tree or calculation.

Example Answer Structure:

Question: Explain the CBR cycle with an example from Daraz. Answer: The CBR cycle consists of four phases:

  1. Retrieve: For a complaint about a delayed order in Kathmandu, the system finds past cases with:
    • Order value: ₹3,000–₹5,000
    • Delay: 2–4 days
    • Location: Kathmandu/Lalitpur. Visual: [Draw a similarity score table as above.]
  2. Reuse: The top match (similarity score: 0.88) suggests offering a ₹100 coupon.
  3. Revise: Since the customer is a "Premium" user, the solution is revised to a ₹200 coupon.
  4. Retain: The new case is stored for future use. Limitation: CBR may fail if no similar past cases exist (e.g., first-ever complaint type).

Key Formulas and Shortcuts

  1. Similarity Score (CBR): (Lower score = more similar.)

  2. Probabilistic Rule Application:

  3. Rule-Based Inference:

    • Forward Chaining: Starts with facts → applies rules to derive conclusions.
    • Backward Chaining: Starts with a hypothesis → checks if facts support it.

Common Pitfalls in Exams

  • Confusing KBS and ML: KBS uses rules; ML learns patterns from data. Avoid saying "KBS learns like ML."
  • Skipping the Cycle: For CBR, always mention all 4 phases (Retrieve/Reuse/Revise/Retain).
  • Ignoring Real-World Context: Nepal-specific examples (e.g., eSewa, Daraz) fetch extra marks.
  • Overcomplicating Probability: Use simple percentages (e.g., "70% chance") unless asked for Bayes’ theorem.

Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 8.

Discussion

Loading…