Data Warehousing and Data MiningUnit 1112 min read
Evaluation & Practical Applications of Data Mining
Unit 11 of Data Warehousing and Data Mining explores how to evaluate data mining results, measure performance, and apply techniques to real-world problems like fraud detection, recommendation systems, and business intelligence—with case studies from Nepalese and global companies.
TAKEAWAYS:
- Evaluation metrics (accuracy, precision, recall, F1-score) quantify how well a model performs, with trade-offs between false positives and false negatives.
- Cross-validation (k-fold) ensures robust model evaluation by testing on unseen data without wasting training samples.
- Business impact matters more than technical metrics: a 1% accuracy gain may not justify a 50% increase in processing time.
- Practical applications span fraud detection (e.g., Ncell’s SIM fraud alerts), personalized recommendations (e.g., Daraz’s "Customers who bought this also bought"), and public policy (e.g., NTC’s traffic pattern analysis).
- Ethical considerations include bias in training data (e.g., loan approval models favoring urban areas) and privacy risks (e.g., social network analysis exposing sensitive relationships).
- End-to-end pipelines (data → preprocessing → modeling → evaluation → deployment) must balance speed, cost, and accuracy for real-world use.
1. Why Evaluate Data Mining Results?
Data mining produces models, rules, or patterns—but how do you know they’re good? Evaluation answers:
- Is the model accurate? (Does it predict correctly?)
- Is it useful? (Does it solve the business problem?)
- Can it be deployed? (Is it fast enough, scalable, and interpretable?)
Without evaluation, you risk:
- Overfitting: A model that works perfectly on training data but fails on real data (like memorizing exam answers instead of understanding concepts).
- Bias: A model that reflects historical discrimination (e.g., a hiring tool trained on data from male-dominated fields).
- Wasted resources: Deploying a slow or expensive model when a simpler one works just as well.
2. Key Evaluation Metrics
Metrics depend on the type of problem:
- Classification (e.g., spam detection, loan approval): Confusion matrix → Accuracy, Precision, Recall, F1-score.
- Regression (e.g., house price prediction): Mean Absolute Error (MAE), Root Mean Squared Error (RMSE).
- Clustering (e.g., customer segmentation): Silhouette Score, Davies-Bouldin Index.
- Association Rules (e.g., market basket analysis): Support, Confidence, Lift.
Worked Example: Evaluating a Loan Approval Model
Scenario: A Nepalese bank (e.g., NMB) uses a decision tree to approve loans. The model predicts "Approved" or "Rejected" based on income, credit score, and employment history.
| Actual \ Predicted | Approved (P) | Rejected (N) |
|---|---|---|
| Approved (P) | True Positive (TP) = 80 | False Negative (FN) = 20 |
| Rejected (N) | False Positive (FP) = 10 | True Negative (TN) = 90 |
Calculations:
- Accuracy = (TP + TN) / (TP + TN + FP + FN) = (80 + 90) / 200 = 85%
- Precision = TP / (TP + FP) = 80 / (80 + 10) = 88.9% (Of all predicted "Approved," 88.9% were correct.)
- Recall (Sensitivity) = TP / (TP + FN) = 80 / (80 + 20) = 80% (Of all actual "Approved" loans, 80% were caught.)
- F1-Score = 2 × (Precision × Recall) / (Precision + Recall) = 84.4%
Business Interpretation:
- High Precision (88.9%): The bank rarely approves bad loans (few FP).
- Lower Recall (80%): Some good loan applicants are rejected (FN = 20). Trade-off: Increasing recall might raise FP (approving riskier loans).
graph TD
A["Confusion Matrix"] --> B["TP: Correct Approvals"]
A --> C["FN: Missed Approvals"]
A --> D["FP: Wrong Approvals"]
A --> E["TN: Correct Rejections"]
B --> F["Precision = TP/(TP+FP)"]
C --> G["Recall = TP/(TP+FN)"]
F & G --> H["F1-Score = Harmonic Mean"]3. Real-World Evaluation Challenges
Example 1: Ncell’s Fraud Detection
- Problem: Detecting fake SIM registrations (e.g., stolen identities).
- Model: Uses clustering to flag unusual registration patterns (e.g., same address, multiple phones).
- Evaluation:
- Precision: Must be high (false alarms annoy customers).
- Recall: Must catch most fraud (low recall = lost revenue).
- Real Trade-off: Ncell might tolerate a few false positives (e.g., blocking a legitimate user’s second SIM) to stop fraudsters.
Example 2: Daraz’s Recommendation System
- Problem: "Customers who bought this also bought" suggestions.
- Model: Association rule mining (e.g., "Diapers → Beer" effect).
- Evaluation:
- Lift: How much more likely a user buys Beer if they bought Diapers? (Lift > 1 = useful rule.)
- Support: 10% of orders include both → not too rare.
- Business Impact: A 1% increase in cross-sell conversions = millions in revenue.
Example 3: NTC’s Traffic Pattern Analysis
- Problem: Predicting congestion hotspots in Kathmandu.
- Model: Time-series clustering of GPS data.
- Evaluation:
- Silhouette Score: How distinct are clusters? (Score near 1 = well-separated traffic patterns.)
- Deployment Cost: A high-accuracy model that runs on a supercomputer is useless if NTC’s servers can’t handle it.
4. Cross-Validation: Testing Without Wasting Data
Problem: If you train on 90% of data and test on 10%, you might get lucky (or unlucky) with that split. Solution: k-Fold Cross-Validation
- Split data into k equal folds (e.g., k=5).
- Train on k-1 folds, test on the remaining fold.
- Repeat k times; average the results.
Worked Example: Evaluating a House Price Predictor Data: 1000 houses in Lalitpur with features: area, bedrooms, age, location. Steps:
- Split into 5 folds (200 houses each).
- Fold 1: Train on folds 2–5 (800 houses), test on fold 1.
- RMSE = 25,000 NPR.
- Fold 2: Train on folds 1,3–5, test on fold 2.
- RMSE = 22,000 NPR.
- Repeat for all folds → Average RMSE = 23,500 NPR.
Why This Matters:
- More reliable than a single train-test split.
- Uses all data for training (unlike holding out a test set forever).
flowchart LR
A["Data"] --> B["Split into 5 folds"]
B --> C["Train on 4 folds\nTest on 1st fold"]
B --> D["Train on 4 folds\nTest on 2nd fold"]
D --> E["..."]
E --> F["Average RMSE\n= 23,500 NPR"]5. Practical Applications and Case Studies
| Application | Technique Used | Nepalese/Global Example | Evaluation Metric |
|---|---|---|---|
| Fraud detection | Anomaly detection, clustering | Ncell (SIM fraud), Khalti (transaction fraud) | Precision, Recall, F1-score |
| Customer segmentation | K-means, hierarchical clustering | Daraz (VIP vs. budget customers) | Silhouette Score, Davies-Bouldin |
| Recommendation systems | Association rules, collaborative filtering | Daraz, Amazon ("Frequently bought together") | Lift, Support, Confidence |
| Loan approval | Decision trees, logistic regression | NMB, Global IME (credit scoring) | AUC-ROC, Precision-Recall |
| Traffic prediction | Time-series forecasting | NTC (Kathmandu traffic lights) | MAE, RMSE |
| Social network analysis | Graph mining | Facebook (friend recommendations) | Clustering coefficient, Modularity |
6. Ethical and Practical Considerations
A. Bias in Data
- Example: A loan approval model trained on historical data might reject more women if past lenders discriminated.
- Fix: Audit training data for bias; use fairness metrics (e.g., demographic parity).
B. Privacy Risks
- Example: Graph mining on Facebook data can expose sensitive relationships (e.g., secret affairs).
- Fix: Anonymize data; use differential privacy (add "noise" to protect individuals).
C. Model Interpretability
- Example: A deep learning model predicting stock prices (like NEPSE’s) may be a "black box."
- Fix: Use decision trees or LIME for explainable AI.
D. Deployment Constraints
- Example: A slow Python model for Pathao’s ride-matching might cause delays.
- Fix: Optimize with C++ or edge computing; trade off accuracy for speed.
7. End-to-End Data Mining Pipeline
Every project follows this cycle:
Worked Example: Building a Khalti Fraud Alert System
- Data: 10,000 transactions with labels (fraud/legit).
- Preprocessing: Remove duplicates, handle missing values (e.g., "amount" = 0 → flag).
- Model: Isolation Forest (anomaly detection).
- Evaluation: 92% precision (only 8% false alarms), 85% recall (catches 85% of fraud).
- Deployment: API that flags transactions > threshold.
- Monitoring: Retrain monthly as fraud patterns evolve.
8. Common Pitfalls and How to Avoid Them
| Pitfall | Cause | Solution |
|---|---|---|
| Overfitting | Model memorizes noise in training data | Use cross-validation, regularization |
| Ignoring class imbalance | 99% "not fraud," 1% "fraud" | Use F1-score, SMOTE (oversampling minority) |
| Wrong evaluation metric | Choosing accuracy for imbalanced data | Use precision-recall curve, AUC-ROC |
| Not considering deployment costs | High-accuracy model too slow | Profile runtime; optimize for edge devices |
| Ethical blind spots | Unchecked bias in training data | Audit data; consult stakeholders |
In the Real World
Ncell’s SIM Fraud Detection
- Idea: Uses graph mining to detect fake SIM registrations by analyzing call/SMS patterns.
- How: Nodes = phone numbers; edges = calls/SMS. Clusters with high connectivity but no real-world ties (e.g., same SIM card used in multiple cities) are flagged.
- Evaluation: False positive rate < 5% (critical for customer trust).
Daraz’s "Frequently Bought Together"
- Idea: Association rule mining (Apriori algorithm).
- How: Finds items often bought together (e.g., phone + charger). Rules like
{phone} → {charger}with support > 10% and confidence > 60% are displayed. - Evaluation: Lift > 1.5 means the rule is actionable (e.g., placing items near each other increases sales).
NTC’s Traffic Light Optimization
- Idea: Time-series clustering of GPS data.
- How: Groups similar traffic patterns (e.g., rush hour vs. night). Adjusts light timings dynamically.
- Evaluation: Reduces average wait time by 15% (measured via probe vehicles).
Exam Tip
What Examiners Want to See:
- Definitions with Examples:
- Don’t just say "Precision = TP/(TP+FP)." Show it with a confusion matrix and a real scenario (e.g., Ncell fraud).
- Trade-offs:
- Always discuss precision vs. recall or accuracy vs. speed. Example: "A bank might prefer 90% precision (fewer bad loans) over 95% recall (missing some good applicants)."
- Practical Applications:
- Link techniques to Nepalese companies (e.g., Khalti for fraud, Daraz for recommendations). Use small numbers for calculations (e.g., 2×2 confusion matrix).
- Evaluation Metrics:
- For classification: Confusion matrix → metrics → interpretation.
- For clustering: Silhouette Score → how it measures cluster separation.
- For association rules: Support, Confidence, Lift → why Lift > 1 matters.
- Diagrams:
- Draw a cross-validation pipeline or a confusion matrix for classification questions.
- For graph mining, sketch a small network (3–5 nodes) with labeled edges (e.g., "friendship" or "transaction").
Avoid:
- Memorizing formulas without context (e.g., just writing "F1 = 2PR/(P+R)" without explaining P/R trade-offs).
- Ignoring real-world constraints (e.g., "The model is 99% accurate but takes 1 hour per prediction—useless for Pathao!").
- Vague answers like "Data mining is used everywhere." Specify (e.g., "Khalti uses anomaly detection for fraud; Daraz uses association rules for cross-selling.").
Final Checklist for Full Marks: ✅ Define the concept clearly (e.g., "Cross-validation splits data into k folds to avoid overfitting"). ✅ Show a visual (confusion matrix, pipeline, or small example graph). ✅ Tie to a real Nepalese/global example (Ncell, Daraz, NTC). ✅ Discuss trade-offs (speed vs. accuracy, precision vs. recall). ✅ End with a business/practical implication (e.g., "This means Ncell can block 85% of fraud with only 5% false alarms.").
Based on the TU BSc CSIT syllabus for Data Warehousing and Data Mining (CSC410), unit 11.
Discussion
Loading…