CSC410 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 1112 min read

Evaluation & Practical Applications of Data Mining

Unit 11 of Data Warehousing and Data Mining explores how to evaluate data mining results, measure performance, and apply techniques to real-world problems like fraud detection, recommendation systems, and business intelligence—with case studies from Nepalese and global companies.

TAKEAWAYS:

  • Evaluation metrics (accuracy, precision, recall, F1-score) quantify how well a model performs, with trade-offs between false positives and false negatives.
  • Cross-validation (k-fold) ensures robust model evaluation by testing on unseen data without wasting training samples.
  • Business impact matters more than technical metrics: a 1% accuracy gain may not justify a 50% increase in processing time.
  • Practical applications span fraud detection (e.g., Ncell’s SIM fraud alerts), personalized recommendations (e.g., Daraz’s "Customers who bought this also bought"), and public policy (e.g., NTC’s traffic pattern analysis).
  • Ethical considerations include bias in training data (e.g., loan approval models favoring urban areas) and privacy risks (e.g., social network analysis exposing sensitive relationships).
  • End-to-end pipelines (data → preprocessing → modeling → evaluation → deployment) must balance speed, cost, and accuracy for real-world use.

1. Why Evaluate Data Mining Results?

Data mining produces models, rules, or patterns—but how do you know they’re good? Evaluation answers:

  • Is the model accurate? (Does it predict correctly?)
  • Is it useful? (Does it solve the business problem?)
  • Can it be deployed? (Is it fast enough, scalable, and interpretable?)

Without evaluation, you risk:

  • Overfitting: A model that works perfectly on training data but fails on real data (like memorizing exam answers instead of understanding concepts).
  • Bias: A model that reflects historical discrimination (e.g., a hiring tool trained on data from male-dominated fields).
  • Wasted resources: Deploying a slow or expensive model when a simpler one works just as well.

2. Key Evaluation Metrics

Metrics depend on the type of problem:

  • Classification (e.g., spam detection, loan approval): Confusion matrix → Accuracy, Precision, Recall, F1-score.
  • Regression (e.g., house price prediction): Mean Absolute Error (MAE), Root Mean Squared Error (RMSE).
  • Clustering (e.g., customer segmentation): Silhouette Score, Davies-Bouldin Index.
  • Association Rules (e.g., market basket analysis): Support, Confidence, Lift.
-3-2-1123-3-2-1123xyLogistic Sigmoid (Decision Boundary)Linear Decision BoundaryClass 0Class 1
Precision-Recall trade-off visualized: Sigmoid vs. linear thresholds

Worked Example: Evaluating a Loan Approval Model

Scenario: A Nepalese bank (e.g., NMB) uses a decision tree to approve loans. The model predicts "Approved" or "Rejected" based on income, credit score, and employment history.

Actual \ Predicted Approved (P) Rejected (N)
Approved (P) True Positive (TP) = 80 False Negative (FN) = 20
Rejected (N) False Positive (FP) = 10 True Negative (TN) = 90

Calculations:

  • Accuracy = (TP + TN) / (TP + TN + FP + FN) = (80 + 90) / 200 = 85%
  • Precision = TP / (TP + FP) = 80 / (80 + 10) = 88.9% (Of all predicted "Approved," 88.9% were correct.)
  • Recall (Sensitivity) = TP / (TP + FN) = 80 / (80 + 20) = 80% (Of all actual "Approved" loans, 80% were caught.)
  • F1-Score = 2 × (Precision × Recall) / (Precision + Recall) = 84.4%

Business Interpretation:

  • High Precision (88.9%): The bank rarely approves bad loans (few FP).
  • Lower Recall (80%): Some good loan applicants are rejected (FN = 20). Trade-off: Increasing recall might raise FP (approving riskier loans).

graph TD
    A["Confusion Matrix"] --> B["TP: Correct Approvals"]
    A --> C["FN: Missed Approvals"]
    A --> D["FP: Wrong Approvals"]
    A --> E["TN: Correct Rejections"]
    B --> F["Precision = TP/(TP+FP)"]
    C --> G["Recall = TP/(TP+FN)"]
    F & G --> H["F1-Score = Harmonic Mean"]

3. Real-World Evaluation Challenges

Example 1: Ncell’s Fraud Detection

  • Problem: Detecting fake SIM registrations (e.g., stolen identities).
  • Model: Uses clustering to flag unusual registration patterns (e.g., same address, multiple phones).
  • Evaluation:
    • Precision: Must be high (false alarms annoy customers).
    • Recall: Must catch most fraud (low recall = lost revenue).
  • Real Trade-off: Ncell might tolerate a few false positives (e.g., blocking a legitimate user’s second SIM) to stop fraudsters.

Example 2: Daraz’s Recommendation System

  • Problem: "Customers who bought this also bought" suggestions.
  • Model: Association rule mining (e.g., "Diapers → Beer" effect).
  • Evaluation:
    • Lift: How much more likely a user buys Beer if they bought Diapers? (Lift > 1 = useful rule.)
    • Support: 10% of orders include both → not too rare.
  • Business Impact: A 1% increase in cross-sell conversions = millions in revenue.

Example 3: NTC’s Traffic Pattern Analysis

  • Problem: Predicting congestion hotspots in Kathmandu.
  • Model: Time-series clustering of GPS data.
  • Evaluation:
    • Silhouette Score: How distinct are clusters? (Score near 1 = well-separated traffic patterns.)
    • Deployment Cost: A high-accuracy model that runs on a supercomputer is useless if NTC’s servers can’t handle it.

4. Cross-Validation: Testing Without Wasting Data

Problem: If you train on 90% of data and test on 10%, you might get lucky (or unlucky) with that split. Solution: k-Fold Cross-Validation

  1. Split data into k equal folds (e.g., k=5).
  2. Train on k-1 folds, test on the remaining fold.
  3. Repeat k times; average the results.

Worked Example: Evaluating a House Price Predictor Data: 1000 houses in Lalitpur with features: area, bedrooms, age, location. Steps:

  1. Split into 5 folds (200 houses each).
  2. Fold 1: Train on folds 2–5 (800 houses), test on fold 1.
    • RMSE = 25,000 NPR.
  3. Fold 2: Train on folds 1,3–5, test on fold 2.
    • RMSE = 22,000 NPR.
  4. Repeat for all folds → Average RMSE = 23,500 NPR.

Why This Matters:

  • More reliable than a single train-test split.
  • Uses all data for training (unlike holding out a test set forever).

flowchart LR
    A["Data"] --> B["Split into 5 folds"]
    B --> C["Train on 4 folds\nTest on 1st fold"]
    B --> D["Train on 4 folds\nTest on 2nd fold"]
    D --> E["..."]
    E --> F["Average RMSE\n= 23,500 NPR"]

5. Practical Applications and Case Studies

Application Technique Used Nepalese/Global Example Evaluation Metric
Fraud detection Anomaly detection, clustering Ncell (SIM fraud), Khalti (transaction fraud) Precision, Recall, F1-score
Customer segmentation K-means, hierarchical clustering Daraz (VIP vs. budget customers) Silhouette Score, Davies-Bouldin
Recommendation systems Association rules, collaborative filtering Daraz, Amazon ("Frequently bought together") Lift, Support, Confidence
Loan approval Decision trees, logistic regression NMB, Global IME (credit scoring) AUC-ROC, Precision-Recall
Traffic prediction Time-series forecasting NTC (Kathmandu traffic lights) MAE, RMSE
Social network analysis Graph mining Facebook (friend recommendations) Clustering coefficient, Modularity

6. Ethical and Practical Considerations

A. Bias in Data

  • Example: A loan approval model trained on historical data might reject more women if past lenders discriminated.
  • Fix: Audit training data for bias; use fairness metrics (e.g., demographic parity).
021.2542.563.7585Male Applicants85Female Applicants15Approval Rate (%)
Example of biased loan approval data (hypothetical Ncell dataset)

B. Privacy Risks

  • Example: Graph mining on Facebook data can expose sensitive relationships (e.g., secret affairs).
  • Fix: Anonymize data; use differential privacy (add "noise" to protect individuals).

C. Model Interpretability

  • Example: A deep learning model predicting stock prices (like NEPSE’s) may be a "black box."
  • Fix: Use decision trees or LIME for explainable AI.

D. Deployment Constraints

  • Example: A slow Python model for Pathao’s ride-matching might cause delays.
  • Fix: Optimize with C++ or edge computing; trade off accuracy for speed.

7. End-to-End Data Mining Pipeline

Every project follows this cycle:

Worked Example: Building a Khalti Fraud Alert System

  1. Data: 10,000 transactions with labels (fraud/legit).
  2. Preprocessing: Remove duplicates, handle missing values (e.g., "amount" = 0 → flag).
  3. Model: Isolation Forest (anomaly detection).
  4. Evaluation: 92% precision (only 8% false alarms), 85% recall (catches 85% of fraud).
  5. Deployment: API that flags transactions > threshold.
  6. Monitoring: Retrain monthly as fraud patterns evolve.

8. Common Pitfalls and How to Avoid Them

Pitfall Cause Solution
Overfitting Model memorizes noise in training data Use cross-validation, regularization
Ignoring class imbalance 99% "not fraud," 1% "fraud" Use F1-score, SMOTE (oversampling minority)
Wrong evaluation metric Choosing accuracy for imbalanced data Use precision-recall curve, AUC-ROC
Not considering deployment costs High-accuracy model too slow Profile runtime; optimize for edge devices
Ethical blind spots Unchecked bias in training data Audit data; consult stakeholders

In the Real World

  1. Ncell’s SIM Fraud Detection

    • Idea: Uses graph mining to detect fake SIM registrations by analyzing call/SMS patterns.
    • How: Nodes = phone numbers; edges = calls/SMS. Clusters with high connectivity but no real-world ties (e.g., same SIM card used in multiple cities) are flagged.
    • Evaluation: False positive rate < 5% (critical for customer trust).
  2. Daraz’s "Frequently Bought Together"

    • Idea: Association rule mining (Apriori algorithm).
    • How: Finds items often bought together (e.g., phone + charger). Rules like {phone} → {charger} with support > 10% and confidence > 60% are displayed.
    • Evaluation: Lift > 1.5 means the rule is actionable (e.g., placing items near each other increases sales).
  3. NTC’s Traffic Light Optimization

    • Idea: Time-series clustering of GPS data.
    • How: Groups similar traffic patterns (e.g., rush hour vs. night). Adjusts light timings dynamically.
    • Evaluation: Reduces average wait time by 15% (measured via probe vehicles).

Exam Tip

What Examiners Want to See:

  1. Definitions with Examples:
    • Don’t just say "Precision = TP/(TP+FP)." Show it with a confusion matrix and a real scenario (e.g., Ncell fraud).
  2. Trade-offs:
    • Always discuss precision vs. recall or accuracy vs. speed. Example: "A bank might prefer 90% precision (fewer bad loans) over 95% recall (missing some good applicants)."
  3. Practical Applications:
    • Link techniques to Nepalese companies (e.g., Khalti for fraud, Daraz for recommendations). Use small numbers for calculations (e.g., 2×2 confusion matrix).
  4. Evaluation Metrics:
    • For classification: Confusion matrix → metrics → interpretation.
    • For clustering: Silhouette Score → how it measures cluster separation.
    • For association rules: Support, Confidence, Lift → why Lift > 1 matters.
  5. Diagrams:
    • Draw a cross-validation pipeline or a confusion matrix for classification questions.
    • For graph mining, sketch a small network (3–5 nodes) with labeled edges (e.g., "friendship" or "transaction").

Avoid:

  • Memorizing formulas without context (e.g., just writing "F1 = 2PR/(P+R)" without explaining P/R trade-offs).
  • Ignoring real-world constraints (e.g., "The model is 99% accurate but takes 1 hour per prediction—useless for Pathao!").
  • Vague answers like "Data mining is used everywhere." Specify (e.g., "Khalti uses anomaly detection for fraud; Daraz uses association rules for cross-selling.").

Final Checklist for Full Marks: ✅ Define the concept clearly (e.g., "Cross-validation splits data into k folds to avoid overfitting"). ✅ Show a visual (confusion matrix, pipeline, or small example graph). ✅ Tie to a real Nepalese/global example (Ncell, Daraz, NTC). ✅ Discuss trade-offs (speed vs. accuracy, precision vs. recall). ✅ End with a business/practical implication (e.g., "This means Ncell can block 85% of fraud with only 5% false alarms.").

Based on the TU BSc CSIT syllabus for Data Warehousing and Data Mining (CSC410), unit 11.

Discussion

Loading…