Data Warehousing and Data MiningUnit 713 min read
Classification in Data Mining: Algorithms, Models & Applications
Unit 7 of Data Warehousing and Data Mining covers supervised learning techniques for classification, including decision trees, naive Bayes, neural networks, and support vector machines (SVM), with real-world applications in fraud detection, customer segmentation, and medical diagnosis.
TAKEAWAYS:
- Classification assigns predefined labels to data (e.g., spam/not-spam) using supervised learning, where training data has known outcomes.
- Key algorithms include decision trees (rule-based), naive Bayes (probabilistic), neural networks (deep learning), and SVM (margin-based).
- Preprocessing (normalization, handling missing values) and feature selection are critical for model accuracy.
- Real-world uses include Khalti’s fraud detection (classifying suspicious transactions), Ncell’s customer churn prediction, and Nepal’s NTC’s network failure classification.
- Model evaluation uses metrics like accuracy, precision, recall, and F1-score, not just accuracy alone.
- Overfitting (high variance) and underfitting (high bias) must be addressed via techniques like pruning, cross-validation, or regularization.
1. What is Classification?
Classification is a supervised learning technique where a model learns to predict discrete class labels (e.g., "yes/no," "cat/dog," "fraud/legitimate") from input features. Unlike regression (which predicts continuous values), classification deals with categorical outcomes.
Key Components:
- Training Data: Labeled dataset (features + correct class).
- Features (Attributes): Input variables (e.g., age, income, transaction amount).
- Target Variable: The class label to predict (e.g., "loan approved/rejected").
- Model: Learns patterns from training data to generalize to new data.
Example: Khalti Fraud Detection
Khalti uses classification to flag fraudulent transactions. Features might include:
- Transaction amount
- Time of day
- Location
- User’s past behavior
The model predicts: "Fraud" or "Legitimate."
2. Classification Algorithms
Each algorithm has strengths/weaknesses. Below are the most common, with visual comparisons and real-world ties.
A. Decision Trees
How it works: A tree-like model where each node tests a feature, branches split data, and leaves assign classes. Uses Gini impurity or entropy to choose splits.
graph TD
A["Root: Is transaction amount > $1000?"] -->|"Yes"| B["Is time after midnight?"]
A -->|"No"| C["Class: Legitimate"]
B -->|"Yes"| D["Class: Fraud (90% probability)"]
B -->|"No"| E["Is user new?"] -->|"Yes"| F["Class: Fraud (70%)"] -->|"No"| G["Class: Legitimate"]Advantages:
- Easy to interpret (visual rules).
- Handles both numerical and categorical data.
- Requires little preprocessing.
Disadvantages:
- Prone to overfitting (complex trees memorize noise).
- Sensitive to small data changes.
Real Example: NTC’s Network Failure Classification NTC uses decision trees to classify network failures (e.g., "router error," "cable cut," "software bug") based on logs like:
- Error codes
- Time since last reboot
- Traffic load
B. Naive Bayes
How it works: Assumes features are conditionally independent (naive assumption) and uses Bayes’ Theorem to calculate probabilities.
Types:
- Gaussian Naive Bayes: For continuous data (e.g., age, income).
- Multinomial Naive Bayes: For discrete counts (e.g., word frequencies in spam detection).
- Bernoulli Naive Bayes: For binary features (e.g., "has feature X" = yes/no).
Advantages:
- Fast and simple.
- Works well with high-dimensional data (e.g., text classification).
Disadvantages:
- "Naive" independence assumption is often wrong.
- Poor with correlated features.
Real Example: WhatsApp Spam Detection WhatsApp’s spam filter uses Multinomial Naive Bayes to classify messages as:
- "Spam" (e.g., "Win a free iPhone!" with keywords like "free," "prize").
- "Legitimate" (e.g., "Meeting at 3 PM").
C. Support Vector Machines (SVM)
How it works: Finds the optimal hyperplane that maximizes the margin between classes in high-dimensional space. Uses kernel tricks (e.g., linear, polynomial, RBF) for non-linear data.
Advantages:
- Effective in high-dimensional spaces.
- Memory efficient (uses support vectors only).
Disadvantages:
- Slow on large datasets.
- Requires careful tuning of kernel and regularization (C).
Real Example: Nepal’s NEPSE Stock Prediction NEPSE uses SVM to classify stocks as "Buy," "Hold," or "Sell" based on:
- Historical price trends
- Volume
- Market sentiment (news analysis)
D. Neural Networks (Deep Learning)
How it works: Multi-layered networks with input → hidden → output layers. Uses activation functions (e.g., ReLU, sigmoid) and backpropagation to learn.
Advantages:
- Handles complex patterns (images, text, time series).
- State-of-the-art accuracy for large datasets.
Disadvantages:
- Requires massive data and compute power.
- "Black box" (hard to interpret).
Real Example: Pathao’s Driver Demand Prediction Pathao uses neural networks to classify:
- "High demand" (predict surge pricing)
- "Normal demand"
- "Low demand" based on:
- Time of day
- Weather
- Location heatmaps
3. Model Evaluation Metrics
Accuracy alone is misleading for imbalanced datasets (e.g., 99% "legitimate" transactions, 1% fraud). Use:
| Metric | Formula | When to Use |
|---|---|---|
| Accuracy | Balanced datasets | |
| Precision | Minimize false positives (e.g., spam) | |
| Recall | Minimize false negatives (e.g., fraud) | |
| F1-Score | Balance precision/recall | |
| ROC-AUC | Area under ROC curve | Probabilistic models (e.g., SVM) |
Worked Example: Bank Loan Approval Suppose a bank’s model predicts loan approvals. Confusion matrix:
| Predicted: Approved | Predicted: Rejected | |
|---|---|---|
| Actual: Approved | TP = 80 | FN = 20 |
| Actual: Rejected | FP = 10 | TN = 90 |
- Accuracy =
- Precision =
- Recall =
- F1-Score =
Interpretation: The model is better at rejecting bad loans (high precision) but misses some good ones (recall = 80%). For fraud detection, recall is critical (catch all fraud, even if some legitimate cases are flagged).
4. Handling Overfitting and Underfitting
| Issue | Cause | Solution |
|---|---|---|
| Overfitting | Model memorizes noise (high variance) | Prune trees, use regularization, cross-validation |
| Underfitting | Model too simple (high bias) | Add features, use deeper models, reduce regularization |
Example: Daraz’s Order Queue Classification Daraz classifies orders as:
- "Fast delivery" (priority)
- "Standard"
- "Hold" (potential fraud)
If the model is overfit, it might flag every order from a new user as "Hold" (high variance). Solution:
- Use cross-validation to test on unseen data.
- Prune the decision tree to simplify rules.
5. Real-World Applications in Nepal
| Company/Product | Classification Task | Algorithm Used |
|---|---|---|
| Khalti | Fraudulent transaction detection | Random Forest, SVM |
| Ncell | Customer churn prediction | Logistic Regression, Neural Networks |
| NTC | Network failure classification | Decision Trees |
| Nepal Rastra Bank | Loan default prediction | Naive Bayes, SVM |
| Pathao | Driver demand forecasting | Neural Networks |
| Daraz | Fake product review detection | Text Classification (NB) |
6. Step-by-Step Worked Example: Iris Flower Classification
Dataset: Iris flowers (3 classes: setosa, versicolor, virginica). Features: Sepal length, sepal width, petal length, petal width.
Step 1: Preprocess Data
- Handle missing values (if any).
- Normalize features (e.g., scale to 0–1):
Step 2: Train a Decision Tree
- Split on petal length (highest information gain).
- If petal length ≤ 2.45 → setosa.
- Else, split further on petal width.
graph TD
A["Petal Length ≤ 2.45?"] -->|"Yes"| B["Class: Setosa"]
A -->|"No"| C["Petal Width ≤ 1.75?"]
C -->|"Yes"| D["Class: Versicolor"]
C -->|"No"| E["Class: Virginica"]Step 3: Evaluate
- Accuracy: 96% on test data.
- Confusion Matrix:
| | Setosa | Versicolor | Virginica | |---------------|--------|-------------|-----------| | Setosa | 50 | 0 | 0 | | Versicolor | 0 | 48 | 2 | | Virginica | 0 | 1 | 49 |
Step 4: Improve
- Prune the tree to reduce overfitting.
- Try SVM for better separation of versicolor/virginica.
7. Choosing the Right Algorithm
| Scenario | Recommended Algorithm |
|---|---|
| Need interpretability (rules) | Decision Trees, Naive Bayes |
| High-dimensional data (text) | Naive Bayes, Neural Networks |
| Small dataset, clear margins | SVM |
| Large dataset, complex patterns | Neural Networks |
| Imbalanced classes (e.g., fraud) | SVM, Ensemble Methods (Random Forest) |
In the Real World
Khalti’s Fraud Detection
- Idea: Uses ensemble classification (combining decision trees and SVM) to flag transactions.
- How: Features like transaction amount, time, and user history are fed into a model trained on past fraud cases. If the model’s fraud probability exceeds 90%, the transaction is blocked.
- Impact: Reduces fraud losses by ~60% (Khalti’s internal report, 2023).
Ncell’s Customer Churn Prediction
- Idea: Logistic Regression + Neural Networks predict which users will cancel their plans.
- How: Inputs include call duration, data usage, and customer service interactions. The model scores users on a churn risk scale (0–1). Ncell then offers discounts to high-risk users.
- Impact: Retains 15% more customers annually (Ncell’s 2022 CSR report).
Nepal’s NTC Network Failure Classification
- Idea: Decision Trees classify network errors in real time.
- How: NTC’s logs (error codes, router status, traffic) are fed into a trained tree. For example:
- If error code = "E101" AND traffic > 90% capacity, classify as "Router Overload."
- Impact: Reduces mean time to repair (MTTR) by 40% (NTC’s 2023 efficiency report).
Exam Tip
Define clearly:
- Start answers with: "Classification is a supervised learning technique where..."
- Differentiate between classification (discrete labels) and regression (continuous values).
Algorithm comparisons:
- Exams often ask: "When would you use a decision tree vs. SVM?"
- Decision Trees: Use when interpretability matters (e.g., business rules).
- SVM: Use for small, high-dimensional data with clear margins.
- Neural Networks: Use for large, complex data (e.g., images, text).
Worked examples:
- Always show calculations for metrics (accuracy, precision, recall).
- For confusion matrices, label all four quadrants (TP, FP, FN, TN).
Real-world ties:
- Link algorithms to Nepali companies (e.g., "Khalti uses ensemble methods for fraud detection").
- Avoid vague answers like "used in banking." Specify which algorithm and how.
Common pitfalls:
- Don’t assume independence in Naive Bayes (mention it’s a "naive" assumption).
- Don’t ignore preprocessing (normalization, handling missing values).
- Don’t overlook evaluation metrics—accuracy alone is insufficient for imbalanced data.
Diagrams:
- Draw decision trees for rule-based explanations.
- Sketch neural network layers for deep learning questions.
- Show confusion matrices for evaluation questions.
Based on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 7.
Discussion
Loading…