CSC469 Decision Support System and Expert System

Decision Support System and Expert SystemUnit 109 min read

Case-Based Reasoning & Data Mining in DSS/ES: Techniques, Tools & Applications

Unit 10 of Decision Support System and Expert System explores Case-Based Reasoning (CBR)—how systems reuse past solutions—and Data Mining—extracting patterns from DSS/ES data—to solve unstructured problems. Covers techniques (k-NN, association rules), tools (Weka, RapidMiner), and real-world applications in finance, he

Key Concepts: Case-Based Reasoning (CBR) in DSS/ES

CBR solves new problems by adapting solutions from similar past cases (stored in a case base). Unlike rule-based systems, it avoids explicit knowledge engineering by leveraging memory and analogy.

1. CBR Cycle: Retrieve → Reuse → Revise → Retain

[object Object]New ProblemRetrieve Similar CasesReuse Adapt SolutionRevise Validate SolutionRetain Store New Case
CBR Cycle: Retrieve → Reuse → Revise → Retain (with feedback loop)

How it works:

  • Retrieve: Find k most similar cases using metrics like Euclidean distance or cosine similarity.
  • Reuse: Apply the solution from the top-k cases (e.g., majority vote for classification).
  • Revise: Adjust the solution if it fails (e.g., tweak loan approval rules).
  • Retain: Store the new case for future use.

2. Worked Example: Daraz Customer Complaint Resolution

Scenario: Daraz receives a complaint about a delayed order. The CBR system matches it to past cases:

  • Case 1: Delayed by 3 days → Refunded 50%.
  • Case 2: Delayed by 5 days → Full refund + discount coupon.
  • Case 3: Delayed by 2 days → No refund, but expedited shipping.
Complaint 1: Delayed Delivery0Complaint 2: Damaged Product1Complaint 3: Refund Issue2
Example case base (highlighted: retrieved similar complaint for reuse)

Steps:

  1. Feature Vector: [delay_days=4, order_value=12000, customer_rating=4.2].
  2. Similarity Metric: Euclidean distance to stored cases.
  3. Retrieve Top-3: Cases 1, 2, and 3 (distances: 0.8, 1.2, 1.5).
  4. Reuse: Majority vote suggests partial refund + discount (Case 1 + Case 2).
  5. Revise: Agent overrides to full refund (customer is premium user).
  6. Retain: New case added: [delay_days=4, action=full_refund, customer_type=premium].

3. Data Mining in DSS/ES: Extracting Patterns

Data mining discovers hidden patterns in DSS/ES data to support decisions. Key techniques:

Technique Purpose Example in Nepal Tools
Classification Predict categories (e.g., loan approval) Ncell churn prediction: Identify customers likely to switch. Weka, Orange
Clustering Group similar data (e.g., market segments) NEPSE stock clustering: Group stocks by volatility. K-Means, DBSCAN
Association Rules Find co-occurring items (e.g., "buy X, buy Y") Khalti transactions: "Users who book movie tickets also order food." Apriori, FP-Growth
Regression Predict continuous values (e.g., sales) Pathao driver earnings: Predict monthly income based on trips. Linear Regression, XGBoost

4. Worked Example: NEPSE Stock Trend Prediction (Regression)

Problem: Predict tomorrow’s stock price for Nepal Bank Limited (NBL) using past 30 days of data.

2468101214161820101214161820yPredicted Trend LineActual Data (Sample)Observed: 13.5Observed: 14.8
Regression model fitting NEPSE stock data (simplified)

Steps:

  1. Data Collection: Features = [opening_price, volume, moving_avg_7day]; Target = next_day_price.
  2. Model: Linear Regression (simplified):
  3. Training: Fit model on 25 days of data.
  4. Prediction: For Day 30:
    • Volume = 500,000; MovingAvg = 1,200 →
  5. DSS Use: Trader uses this to decide whether to buy/sell/hold.

In the Real World

  1. eSewa’s Fraud Detection

    • CBR: Matches new transactions to past fraud cases (e.g., "user in Kathmandu paying for a service in Pokhara").
    • Data Mining: Uses clustering to flag unusual spending patterns (e.g., sudden high-value transfers).
  2. Pathao’s Driver Routing

    • Association Rules: "Drivers in Lalitpur with >3 trips/day earn 20% more."
    • Regression: Predicts demand spikes during Dashain (e.g., +40% trips on Day 1).
  3. NTC’s Network Outage Prediction

    • Classification: Trained on past outages (e.g., "rain + old cables → 80% chance of outage").
    • CBR: New outage in Chitwan? Matches to similar outages in Makwanpur.

5. Data Mining Tools for DSS/ES

Tool Type Use Case Nepal Example
Weka Open-source ML Classify customer segments for Daraz. Segment users by purchase frequency.
RapidMiner Drag-and-drop mining Predict loan defaults for banks. NMB Bank: Flag high-risk borrowers.
Orange Visual workflows Explore NEPSE stock correlations. "Does oil price affect stock markets?"
KNIME Enterprise mining Fraud detection for Khalti. Build a pipeline for real-time alerts.

6. Case-Based Reasoning vs. Data Mining: Key Differences

Feature Case-Based Reasoning (CBR) Data Mining
Approach Reuses past solutions. Discovers patterns from data.
Data Requirement Needs a case base (structured past cases). Needs large datasets (structured/unstructured).
Output Specific solution for a new problem. General patterns/rules (e.g., "IF-THEN").
Example in Nepal eSewa resolving complaints by matching to past cases. NTC using clustering to predict outages.
Strength Works well with small, expert-curated data. Scales to big data (e.g., Daraz transactions).

7. Challenges and Limitations

CBR Pitfalls

  • Cold Start Problem: No cases for new scenarios (e.g., first COVID-19 lockdown).
  • Similarity Metric Bias: Euclidean distance may fail for high-dimensional data.
  • Maintenance Overhead: Case base must be updated regularly.

Data Mining Challenges

  • Garbage In, Garbage Out (GIGO): Poor data quality → useless patterns.
  • Overfitting: Model works on training data but fails in real-world (e.g., NEPSE model trained only on bull markets).
  • Interpretability: Black-box models (e.g., deep learning) are hard to explain to stakeholders.

Exam Tip

  1. For CBR:

    • Always explain the 4-step cycle (Retrieve-Reuse-Revise-Retain) with a real example (e.g., Daraz, eSewa).
    • Compare CBR with rule-based systems (CBR uses past cases; rule-based uses IF-THEN).
    • Graph the similarity metric (e.g., Euclidean distance formula for 2D/3D cases).
  2. For Data Mining:

    • Define the technique (classification/clustering/association) and give a Nepalese example.
    • Show the math for simple cases (e.g., linear regression equation with real numbers).
    • Tool names matter: Weka, RapidMiner, and Orange are high-yield for short-answer questions.
  3. Common Exam Traps:

    • ❌ Saying "data mining is the same as big data."
    • ❌ Forgetting to link DSS/ES (e.g., "This clustering helps the bank approve loans faster").
    • ❌ Using vague examples (e.g., "Amazon uses data mining" → specify association rules for recommendations).

Visual Summary:

mindmap
  root((Case-Based Reasoning & Data Mining))
    CBR
      Retrieve["Find Similar Cases\n(Euclidean/Cosine Similarity)"]
      Reuse["Apply Past Solutions"]
      Revise["Adjust for New Context"]
      Retain["Update Case Base"]
    Data Mining
      Classification["Predict Categories\n(e.g., Loan Approval)"]
      Clustering["Group Data\n(e.g., NEPSE Stock Segments)"]
      Association["Find Patterns\n(e.g., Khalti Transaction Rules)"]
      Regression["Predict Values\n(e.g., Pathao Earnings)"]
    Tools
      Weka["Open-Source ML"]
      RapidMiner["Drag-and-Drop"]
      Orange["Visual Workflows"]
    Real World
      eSewa["Fraud Detection via CBR"]
      Daraz["Customer Complaints via CBR"]
      NTC["Outage Prediction via Clustering"]

Based on the TU BSc CSIT syllabus for Decision Support System and Expert System (CSC469), unit 10.

Discussion

Loading…