Knowledge EngineeringUnit 814 min read
Knowledge-Based Systems & Case-Based Reasoning: Models, Workflows & Applications
Unit 8 of Knowledge Engineering explores how knowledge-based systems (KBS) encode domain expertise into rule-based or case-based reasoning engines, and how case-based reasoning (CBR) solves new problems by adapting past solutions. Covers architectures, CBR cycles, real-world deployments (e.g., medical diagnosis, legal
TAKEAWAYS:
- Knowledge-Based Systems (KBS) combine domain knowledge (rules, facts) with inference engines to solve problems like medical diagnosis or tax calculation—unlike ML, they explain their reasoning.
- Case-Based Reasoning (CBR) solves new problems by retrieving, adapting, and reusing past cases (e.g., a bank loan officer approving a new applicant by comparing to similar past loans).
- The CBR cycle (Retrieve → Reuse → Revise → Retain) turns experience into a feedback loop, improving over time (visualized as a circular flowchart).
- Advantages of KBS: Transparency (rules are inspectable), works with incomplete data, and handles uncertainty via probabilistic reasoning.
- Limitations: Brittleness (rules must cover all cases), high maintenance cost, and struggles with unstructured data (unlike deep learning).
- Real-world tie: eSewa uses CBR to flag fraudulent transactions by matching new payments to past fraud patterns; Daraz’s customer service adapts solutions from resolved complaints.
Core Concepts: Knowledge-Based Systems (KBS)
What is a Knowledge-Based System?
A Knowledge-Based System (KBS) is an AI system that uses domain-specific knowledge (facts, rules, heuristics) to solve problems in a way humans would. It consists of:
- Knowledge Base (KB): Stores facts (e.g., "All birds can fly") and rules (e.g., "IF wings THEN can_fly").
- Inference Engine: Applies logical rules to derive new knowledge (e.g., "Penguin has wings → Penguin can fly" unless it’s an exception).
- Working Memory: Holds current data and intermediate results.
Why KBS?
- Explainability: Unlike black-box ML, KBS shows how it reached a decision (critical for healthcare/legal domains).
- Efficiency: Solves problems without retraining (e.g., a tax calculator uses fixed rules).
- Uncertainty Handling: Uses probabilistic logic (e.g., "80% chance of rain → carry umbrella").
How KBS Works: A Worked Example
Scenario: A Daraz customer support agent uses a KBS to resolve order disputes. The KB contains:
- Facts:
Order(O123, "Laptop", "Delivered", "Damaged")Customer(C456, "Premium")Policy(P1, "Premium customers get free replacement")
- Rules:
IF Order.Status = "Delivered" AND Order.Condition = "Damaged" THEN Flag_for_Refund()IF Customer.Tier = "Premium" AND Flagged_for_Refund THEN Offer_Replacement()
Trace:
- New dispute:
Order(O123, ...)is loaded into working memory. - Rule 1 fires →
Flag_for_Refund(O123). - Rule 2 fires (since
Customer(C456)is Premium) →Offer_Replacement(O123). - Output: "Replace your laptop for free under Policy P1."
Visual: KBS Architecture
flowchart LR
A["User Input\n(e.g., 'My order is damaged')"] --> B["Working Memory\n(Facts: Order(O123), Customer(C456))"]
B --> C["Inference Engine\n(Rules: IF Damaged THEN Refund)"]
C --> D["Knowledge Base\n(Facts + Rules)"]
C --> E["Output\n('Replace laptop for free')"]
D -->|"Update"| BTypes of KBS
| Type | Description | Example | Limitations |
|---|---|---|---|
| Rule-Based (RB) | Uses IF-THEN rules (e.g., expert systems). |
MYCIN (medical diagnosis) | Brittle; rules must cover all cases. |
| Frame-Based | Organizes knowledge as objects with slots/values (e.g., Patient(age=30, symptoms=[fever])). |
Medical record systems | Hard to update dynamically. |
| Case-Based (CBR) | Solves new problems by adapting past cases (see next section). | eSewa fraud detection | Needs large, labeled case database. |
| Hybrid | Combines rules + ML (e.g., rules for high-stakes decisions, ML for data prep). | Bank loan approval systems | Complex to maintain. |
Case-Based Reasoning (CBR): The 4R Cycle
What is CBR?
CBR solves new problems by reusing solutions to similar past problems. Unlike rule-based systems, it learns from experience. Key idea:
"If a problem is similar to one you’ve solved before, use that solution as a starting point."
Example in Nepal:
- NTC’s network outage resolution: When a new outage occurs in Kathmandu, technicians retrieve past outages in the same area, adapt the fix (e.g., "Last time, it was a broken pole—check the same pole today"), and update the case database for future use.
The CBR Cycle: Retrieve → Reuse → Revise → Retain
Worked Example: eSewa Fraud Detection
Problem: A user reports a suspicious transaction of ₹50,000 to "Nepal Telecom" (but NTC’s real ID is NTCLTD).
CBR Steps:
Retrieve:
- Past fraud cases with:
- Amount: ₹40,000–₹60,000
- Recipient: Non-NTC IDs (e.g.,
NTCLTD123vs.NTCLTD) - Time: Last 24 hours (peak fraud hours: 2–4 AM).
- Top match: A ₹55,000 scam to
NTCLTD999(similarity score: 0.85).
- Past fraud cases with:
Reuse:
- Apply solution from past case: "Flag as fraud, block transaction, notify user."
Revise:
- Adjust: "Since this is ₹50K (below ₹60K threshold), only block and warn—don’t auto-refund."
Retain:
- Add new case to database:
Case(50000, "NTCLTD123", "Blocked", "User warned").
- Add new case to database:
Real Output:
Similarity Measures in CBR
To find the "most similar" past case, CBR uses metrics like:
- Numeric attributes: Euclidean distance (e.g., age difference).
- Categorical attributes: Hamming distance (e.g., "NTC" vs. "NTCLTD" → 1 mismatch).
- Weighted scores: Combine metrics (e.g., 60% weight to amount, 40% to time).
Example Calculation:
| Attribute | New Case | Past Case | Distance | Weight | Score |
|---|---|---|---|---|---|
| Amount (₹) | 50,000 | 55,000 | 50,000–55,000 | = 5,000 | |
| Time (hours) | 3 AM | 2 AM | 3–2 | = 1 hour | |
| Total | 3.4 |
KBS vs. CBR vs. Machine Learning
| Feature | Knowledge-Based Systems (KBS) | Case-Based Reasoning (CBR) | Machine Learning (ML) |
|---|---|---|---|
| Knowledge Source | Hand-coded rules/expert input | Past cases (experience) | Data (labeled/unlabeled) |
| Learning | No learning; static rules | Learns by updating case base | Learns from data (training) |
| Explainability | High (rules are transparent) | Medium (can trace case similarities) | Low (black box) |
| Data Requirements | Rules + small datasets | Large case database | Large labeled datasets |
| Adaptability | Low (rules must be updated manually) | High (adapts to new cases) | High (generalizes to new data) |
| Example in Nepal | Ncell’s IVR menu ("Press 1 for balance") | Daraz’s customer support chatbot | NEPSE stock prediction models |
Applications of KBS and CBR
1. Healthcare
- KBS Example: MYCIN (1970s) diagnosed bacterial infections by applying medical rules (e.g., "IF organism is Streptococcus AND patient is allergic to penicillin THEN prescribe erythromycin").
- CBR Example: Radiology case matching – New X-ray images are compared to past cases with similar symptoms to suggest diagnoses.
flowchart LR A["New X-ray (Pneumonia suspected)"] --> B["Retrieve Past cases with: - Cough + fever - Lung opacity >50%"] B --> C["Reuse Diagnosis: 'Pneumonia (85% match)'"] C --> D["Revise Adjust for: - Patient age (child vs. adult) - Allergies (e.g., penicillin)"] D --> E["Retain Add case: 'Child, pneumonia, treated with Amoxicillin'"] E -->|"Feedback"| A
2. Finance
- KBS: Credit scoring (e.g., "IF income > ₹50K AND credit_score > 700 THEN approve loan").
- CBR: Fraud detection (e.g., eSewa matching new transactions to past scams).
3. Legal Systems
- CBR: Legal case retrieval – Lawyers input new cases to find similar precedents (e.g., "This contract breach resembles Case X from 2018").
4. Customer Support
- CBR: Daraz/Pathao chatbots resolve complaints by matching to past resolved issues (e.g., "Delivery delayed in Lalitpur → offer ₹200 coupon").
Advantages and Limitations
Advantages of KBS/CBR
- Transparency: Rules/cases are human-readable (critical for healthcare/legal domains).
- No Training Data Needed: KBS works with expert rules; CBR learns from cases.
- Handles Uncertainty: Uses probabilistic logic (e.g., "70% chance of fraud").
- Fast for Niche Domains: Outperforms ML when rules/cases are well-defined (e.g., tax calculation).
Limitations
| Limitation | Impact | Mitigation |
|---|---|---|
| Brittleness | Fails on edge cases not covered by rules/cases. | Hybrid systems (rules + ML). |
| High Maintenance | Rules/cases must be updated manually. | Automated case retrieval tools. |
| Scalability | Struggles with high-dimensional data (e.g., images). | Use ML for feature extraction. |
| No Generalization | CBR only works within the case database. | Combine with symbolic reasoning. |
Probabilistic Reasoning in KBS
Many KBS handle uncertainty using probabilistic logic (e.g., Bayesian networks). Example:
Scenario: A bank loan officer uses a KBS to approve loans. The KB includes:
- Facts:
Loan_Amount(₹500,000)Credit_Score(720)Employment_Status("Stable")
- Probabilistic Rules:
IF Credit_Score > 700 THEN Approval_Probability = 0.9IF Employment_Status = "Unstable" THEN Approval_Probability *= 0.5
Calculation:
- Base probability: 0.9 (from credit score).
- Adjusted for stable employment:
0.9 * 1.0 = 0.9→ 90% approval chance.
Visual: Probability Tree
Exam Tip: How to Score Full Marks
- Define Clearly: Always start with precise definitions (e.g., "CBR is a problem-solving paradigm that reuses past cases...").
- Use Diagrams: Draw the CBR cycle or KBS architecture in exams—it’s worth 3–5 marks.
- Compare Tables: For questions like "KBS vs. ML," use a side-by-side table (as above).
- Real-World Tie: Link examples to Nepali contexts (e.g., eSewa, Daraz, NTC) for +2 marks.
- Step-by-Step Traces: For CBR/KBS questions, show each phase (Retrieve → Reuse → Revise) with concrete data.
- Limitations: Always mention 1–2 limitations (e.g., "KBS struggles with unstructured data like images").
- Probabilistic Logic: If asked about uncertainty, show a probability tree or calculation.
Example Answer Structure:
Question: Explain the CBR cycle with an example from Daraz. Answer: The CBR cycle consists of four phases:
- Retrieve: For a complaint about a delayed order in Kathmandu, the system finds past cases with:
- Order value: ₹3,000–₹5,000
- Delay: 2–4 days
- Location: Kathmandu/Lalitpur. Visual: [Draw a similarity score table as above.]
- Reuse: The top match (similarity score: 0.88) suggests offering a ₹100 coupon.
- Revise: Since the customer is a "Premium" user, the solution is revised to a ₹200 coupon.
- Retain: The new case is stored for future use. Limitation: CBR may fail if no similar past cases exist (e.g., first-ever complaint type).
Key Formulas and Shortcuts
Similarity Score (CBR): (Lower score = more similar.)
Probabilistic Rule Application:
Rule-Based Inference:
- Forward Chaining: Starts with facts → applies rules to derive conclusions.
- Backward Chaining: Starts with a hypothesis → checks if facts support it.
Common Pitfalls in Exams
- Confusing KBS and ML: KBS uses rules; ML learns patterns from data. Avoid saying "KBS learns like ML."
- Skipping the Cycle: For CBR, always mention all 4 phases (Retrieve/Reuse/Revise/Retain).
- Ignoring Real-World Context: Nepal-specific examples (e.g., eSewa, Daraz) fetch extra marks.
- Overcomplicating Probability: Use simple percentages (e.g., "70% chance") unless asked for Bayes’ theorem.
Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 8.
Discussion
Loading…