Decision Support System and Expert SystemUnit 109 min read
Case-Based Reasoning & Data Mining in DSS/ES: Techniques, Tools & Applications
Unit 10 of Decision Support System and Expert System explores Case-Based Reasoning (CBR)—how systems reuse past solutions—and Data Mining—extracting patterns from DSS/ES data—to solve unstructured problems. Covers techniques (k-NN, association rules), tools (Weka, RapidMiner), and real-world applications in finance, he
Key Concepts: Case-Based Reasoning (CBR) in DSS/ES
CBR solves new problems by adapting solutions from similar past cases (stored in a case base). Unlike rule-based systems, it avoids explicit knowledge engineering by leveraging memory and analogy.
1. CBR Cycle: Retrieve → Reuse → Revise → Retain
How it works:
- Retrieve: Find k most similar cases using metrics like Euclidean distance or cosine similarity.
- Reuse: Apply the solution from the top-k cases (e.g., majority vote for classification).
- Revise: Adjust the solution if it fails (e.g., tweak loan approval rules).
- Retain: Store the new case for future use.
2. Worked Example: Daraz Customer Complaint Resolution
Scenario: Daraz receives a complaint about a delayed order. The CBR system matches it to past cases:
- Case 1: Delayed by 3 days → Refunded 50%.
- Case 2: Delayed by 5 days → Full refund + discount coupon.
- Case 3: Delayed by 2 days → No refund, but expedited shipping.
Steps:
- Feature Vector:
[delay_days=4, order_value=12000, customer_rating=4.2]. - Similarity Metric: Euclidean distance to stored cases.
- Retrieve Top-3: Cases 1, 2, and 3 (distances: 0.8, 1.2, 1.5).
- Reuse: Majority vote suggests partial refund + discount (Case 1 + Case 2).
- Revise: Agent overrides to full refund (customer is premium user).
- Retain: New case added:
[delay_days=4, action=full_refund, customer_type=premium].
3. Data Mining in DSS/ES: Extracting Patterns
Data mining discovers hidden patterns in DSS/ES data to support decisions. Key techniques:
| Technique | Purpose | Example in Nepal | Tools |
|---|---|---|---|
| Classification | Predict categories (e.g., loan approval) | Ncell churn prediction: Identify customers likely to switch. | Weka, Orange |
| Clustering | Group similar data (e.g., market segments) | NEPSE stock clustering: Group stocks by volatility. | K-Means, DBSCAN |
| Association Rules | Find co-occurring items (e.g., "buy X, buy Y") | Khalti transactions: "Users who book movie tickets also order food." | Apriori, FP-Growth |
| Regression | Predict continuous values (e.g., sales) | Pathao driver earnings: Predict monthly income based on trips. | Linear Regression, XGBoost |
4. Worked Example: NEPSE Stock Trend Prediction (Regression)
Problem: Predict tomorrow’s stock price for Nepal Bank Limited (NBL) using past 30 days of data.
Steps:
- Data Collection: Features =
[opening_price, volume, moving_avg_7day]; Target =next_day_price. - Model: Linear Regression (simplified):
- Training: Fit model on 25 days of data.
- Prediction: For Day 30:
- Volume = 500,000; MovingAvg = 1,200 →
- DSS Use: Trader uses this to decide whether to buy/sell/hold.
In the Real World
eSewa’s Fraud Detection
- CBR: Matches new transactions to past fraud cases (e.g., "user in Kathmandu paying for a service in Pokhara").
- Data Mining: Uses clustering to flag unusual spending patterns (e.g., sudden high-value transfers).
Pathao’s Driver Routing
- Association Rules: "Drivers in Lalitpur with >3 trips/day earn 20% more."
- Regression: Predicts demand spikes during Dashain (e.g., +40% trips on Day 1).
NTC’s Network Outage Prediction
- Classification: Trained on past outages (e.g., "rain + old cables → 80% chance of outage").
- CBR: New outage in Chitwan? Matches to similar outages in Makwanpur.
5. Data Mining Tools for DSS/ES
| Tool | Type | Use Case | Nepal Example |
|---|---|---|---|
| Weka | Open-source ML | Classify customer segments for Daraz. | Segment users by purchase frequency. |
| RapidMiner | Drag-and-drop mining | Predict loan defaults for banks. | NMB Bank: Flag high-risk borrowers. |
| Orange | Visual workflows | Explore NEPSE stock correlations. | "Does oil price affect stock markets?" |
| KNIME | Enterprise mining | Fraud detection for Khalti. | Build a pipeline for real-time alerts. |
6. Case-Based Reasoning vs. Data Mining: Key Differences
| Feature | Case-Based Reasoning (CBR) | Data Mining |
|---|---|---|
| Approach | Reuses past solutions. | Discovers patterns from data. |
| Data Requirement | Needs a case base (structured past cases). | Needs large datasets (structured/unstructured). |
| Output | Specific solution for a new problem. | General patterns/rules (e.g., "IF-THEN"). |
| Example in Nepal | eSewa resolving complaints by matching to past cases. | NTC using clustering to predict outages. |
| Strength | Works well with small, expert-curated data. | Scales to big data (e.g., Daraz transactions). |
7. Challenges and Limitations
CBR Pitfalls
- Cold Start Problem: No cases for new scenarios (e.g., first COVID-19 lockdown).
- Similarity Metric Bias: Euclidean distance may fail for high-dimensional data.
- Maintenance Overhead: Case base must be updated regularly.
Data Mining Challenges
- Garbage In, Garbage Out (GIGO): Poor data quality → useless patterns.
- Overfitting: Model works on training data but fails in real-world (e.g., NEPSE model trained only on bull markets).
- Interpretability: Black-box models (e.g., deep learning) are hard to explain to stakeholders.
Exam Tip
For CBR:
- Always explain the 4-step cycle (Retrieve-Reuse-Revise-Retain) with a real example (e.g., Daraz, eSewa).
- Compare CBR with rule-based systems (CBR uses past cases; rule-based uses IF-THEN).
- Graph the similarity metric (e.g., Euclidean distance formula for 2D/3D cases).
For Data Mining:
- Define the technique (classification/clustering/association) and give a Nepalese example.
- Show the math for simple cases (e.g., linear regression equation with real numbers).
- Tool names matter: Weka, RapidMiner, and Orange are high-yield for short-answer questions.
Common Exam Traps:
- ❌ Saying "data mining is the same as big data."
- ❌ Forgetting to link DSS/ES (e.g., "This clustering helps the bank approve loans faster").
- ❌ Using vague examples (e.g., "Amazon uses data mining" → specify association rules for recommendations).
Visual Summary:
mindmap
root((Case-Based Reasoning & Data Mining))
CBR
Retrieve["Find Similar Cases\n(Euclidean/Cosine Similarity)"]
Reuse["Apply Past Solutions"]
Revise["Adjust for New Context"]
Retain["Update Case Base"]
Data Mining
Classification["Predict Categories\n(e.g., Loan Approval)"]
Clustering["Group Data\n(e.g., NEPSE Stock Segments)"]
Association["Find Patterns\n(e.g., Khalti Transaction Rules)"]
Regression["Predict Values\n(e.g., Pathao Earnings)"]
Tools
Weka["Open-Source ML"]
RapidMiner["Drag-and-Drop"]
Orange["Visual Workflows"]
Real World
eSewa["Fraud Detection via CBR"]
Daraz["Customer Complaints via CBR"]
NTC["Outage Prediction via Clustering"]Based on the TU BSc CSIT syllabus for Decision Support System and Expert System (CSC469), unit 10.
Discussion
Loading…