IT274 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 713 min read

Classification in Data Mining: Algorithms, Models & Applications

Unit 7 of Data Warehousing and Data Mining covers supervised learning techniques for classification, including decision trees, naive Bayes, neural networks, and support vector machines (SVM), with real-world applications in fraud detection, customer segmentation, and medical diagnosis.

TAKEAWAYS:

  • Classification assigns predefined labels to data (e.g., spam/not-spam) using supervised learning, where training data has known outcomes.
  • Key algorithms include decision trees (rule-based), naive Bayes (probabilistic), neural networks (deep learning), and SVM (margin-based).
  • Preprocessing (normalization, handling missing values) and feature selection are critical for model accuracy.
  • Real-world uses include Khalti’s fraud detection (classifying suspicious transactions), Ncell’s customer churn prediction, and Nepal’s NTC’s network failure classification.
  • Model evaluation uses metrics like accuracy, precision, recall, and F1-score, not just accuracy alone.
  • Overfitting (high variance) and underfitting (high bias) must be addressed via techniques like pruning, cross-validation, or regularization.

1. What is Classification?

Classification is a supervised learning technique where a model learns to predict discrete class labels (e.g., "yes/no," "cat/dog," "fraud/legitimate") from input features. Unlike regression (which predicts continuous values), classification deals with categorical outcomes.

Legitimate (95%)Fraud (5%)
Typical class imbalance in fraud detection (95% legitimate, 5% fraud)

Key Components:

  • Training Data: Labeled dataset (features + correct class).
  • Features (Attributes): Input variables (e.g., age, income, transaction amount).
  • Target Variable: The class label to predict (e.g., "loan approved/rejected").
  • Model: Learns patterns from training data to generalize to new data.

Example: Khalti Fraud Detection

Khalti uses classification to flag fraudulent transactions. Features might include:

  • Transaction amount
  • Time of day
  • Location
  • User’s past behavior

The model predicts: "Fraud" or "Legitimate."


2. Classification Algorithms

Each algorithm has strengths/weaknesses. Below are the most common, with visual comparisons and real-world ties.

A. Decision Trees

How it works: A tree-like model where each node tests a feature, branches split data, and leaves assign classes. Uses Gini impurity or entropy to choose splits.

graph TD
    A["Root: Is transaction amount > $1000?"] -->|"Yes"| B["Is time after midnight?"]
    A -->|"No"| C["Class: Legitimate"]
    B -->|"Yes"| D["Class: Fraud (90% probability)"]
    B -->|"No"| E["Is user new?"] -->|"Yes"| F["Class: Fraud (70%)"] -->|"No"| G["Class: Legitimate"]

Advantages:

  • Easy to interpret (visual rules).
  • Handles both numerical and categorical data.
  • Requires little preprocessing.

Disadvantages:

  • Prone to overfitting (complex trees memorize noise).
  • Sensitive to small data changes.

Real Example: NTC’s Network Failure Classification NTC uses decision trees to classify network failures (e.g., "router error," "cable cut," "software bug") based on logs like:

  • Error codes
  • Time since last reboot
  • Traffic load

B. Naive Bayes

How it works: Assumes features are conditionally independent (naive assumption) and uses Bayes’ Theorem to calculate probabilities.

Types:

  1. Gaussian Naive Bayes: For continuous data (e.g., age, income).
  2. Multinomial Naive Bayes: For discrete counts (e.g., word frequencies in spam detection).
  3. Bernoulli Naive Bayes: For binary features (e.g., "has feature X" = yes/no).

Advantages:

  • Fast and simple.
  • Works well with high-dimensional data (e.g., text classification).

Disadvantages:

  • "Naive" independence assumption is often wrong.
  • Poor with correlated features.

Real Example: WhatsApp Spam Detection WhatsApp’s spam filter uses Multinomial Naive Bayes to classify messages as:

  • "Spam" (e.g., "Win a free iPhone!" with keywords like "free," "prize").
  • "Legitimate" (e.g., "Meeting at 3 PM").

C. Support Vector Machines (SVM)

How it works: Finds the optimal hyperplane that maximizes the margin between classes in high-dimensional space. Uses kernel tricks (e.g., linear, polynomial, RBF) for non-linear data.

-3-2-1123-4-3-2-11234xyLinear SVM (Hyperplane)Support Vector (Margin)Support Vector (Margin)Class AClass B
Linear SVM separating two classes with maximum margin (left: RBF kernel would use curved boundaries)

Advantages:

  • Effective in high-dimensional spaces.
  • Memory efficient (uses support vectors only).

Disadvantages:

  • Slow on large datasets.
  • Requires careful tuning of kernel and regularization (C).

Real Example: Nepal’s NEPSE Stock Prediction NEPSE uses SVM to classify stocks as "Buy," "Hold," or "Sell" based on:

  • Historical price trends
  • Volume
  • Market sentiment (news analysis)

D. Neural Networks (Deep Learning)

How it works: Multi-layered networks with input → hidden → output layers. Uses activation functions (e.g., ReLU, sigmoid) and backpropagation to learn.

64323Input LayerHidden Layer 1Hidden Layer 2Output Layer
Neural Network architecture: 64 neurons in first hidden layer, 32 in second, 3 output neurons (Softmax)

Advantages:

  • Handles complex patterns (images, text, time series).
  • State-of-the-art accuracy for large datasets.

Disadvantages:

  • Requires massive data and compute power.
  • "Black box" (hard to interpret).

Real Example: Pathao’s Driver Demand Prediction Pathao uses neural networks to classify:

  • "High demand" (predict surge pricing)
  • "Normal demand"
  • "Low demand" based on:
  • Time of day
  • Weather
  • Location heatmaps

3. Model Evaluation Metrics

Accuracy alone is misleading for imbalanced datasets (e.g., 99% "legitimate" transactions, 1% fraud). Use:

Metric Formula When to Use
Accuracy Balanced datasets
Precision Minimize false positives (e.g., spam)
Recall Minimize false negatives (e.g., fraud)
F1-Score Balance precision/recall
ROC-AUC Area under ROC curve Probabilistic models (e.g., SVM)

Worked Example: Bank Loan Approval Suppose a bank’s model predicts loan approvals. Confusion matrix:

Predicted: Approved Predicted: Rejected
Actual: Approved TP = 80 FN = 20
Actual: Rejected FP = 10 TN = 90
  • Accuracy =
  • Precision =
  • Recall =
  • F1-Score =

Interpretation: The model is better at rejecting bad loans (high precision) but misses some good ones (recall = 80%). For fraud detection, recall is critical (catch all fraud, even if some legitimate cases are flagged).


4. Handling Overfitting and Underfitting

Issue Cause Solution
Overfitting Model memorizes noise (high variance) Prune trees, use regularization, cross-validation
Underfitting Model too simple (high bias) Add features, use deeper models, reduce regularization

Example: Daraz’s Order Queue Classification Daraz classifies orders as:

  • "Fast delivery" (priority)
  • "Standard"
  • "Hold" (potential fraud)

If the model is overfit, it might flag every order from a new user as "Hold" (high variance). Solution:

  • Use cross-validation to test on unseen data.
  • Prune the decision tree to simplify rules.

5. Real-World Applications in Nepal

Company/Product Classification Task Algorithm Used
Khalti Fraudulent transaction detection Random Forest, SVM
Ncell Customer churn prediction Logistic Regression, Neural Networks
NTC Network failure classification Decision Trees
Nepal Rastra Bank Loan default prediction Naive Bayes, SVM
Pathao Driver demand forecasting Neural Networks
Daraz Fake product review detection Text Classification (NB)


6. Step-by-Step Worked Example: Iris Flower Classification

Dataset: Iris flowers (3 classes: setosa, versicolor, virginica). Features: Sepal length, sepal width, petal length, petal width.

Step 1: Preprocess Data

  • Handle missing values (if any).
  • Normalize features (e.g., scale to 0–1):

Step 2: Train a Decision Tree

  1. Split on petal length (highest information gain).
    • If petal length ≤ 2.45 → setosa.
    • Else, split further on petal width.
graph TD
    A["Petal Length ≤ 2.45?"] -->|"Yes"| B["Class: Setosa"]
    A -->|"No"| C["Petal Width ≤ 1.75?"]
    C -->|"Yes"| D["Class: Versicolor"]
    C -->|"No"| E["Class: Virginica"]

Step 3: Evaluate

  • Accuracy: 96% on test data.
  • Confusion Matrix:
    |               | Setosa | Versicolor | Virginica |
    |---------------|--------|-------------|-----------|
    | Setosa        | 50     | 0           | 0         |
    | Versicolor    | 0      | 48          | 2         |
    | Virginica     | 0      | 1           | 49        |
    

Step 4: Improve

  • Prune the tree to reduce overfitting.
  • Try SVM for better separation of versicolor/virginica.

7. Choosing the Right Algorithm

Scenario Recommended Algorithm
Need interpretability (rules) Decision Trees, Naive Bayes
High-dimensional data (text) Naive Bayes, Neural Networks
Small dataset, clear margins SVM
Large dataset, complex patterns Neural Networks
Imbalanced classes (e.g., fraud) SVM, Ensemble Methods (Random Forest)

In the Real World

  1. Khalti’s Fraud Detection

    • Idea: Uses ensemble classification (combining decision trees and SVM) to flag transactions.
    • How: Features like transaction amount, time, and user history are fed into a model trained on past fraud cases. If the model’s fraud probability exceeds 90%, the transaction is blocked.
    • Impact: Reduces fraud losses by ~60% (Khalti’s internal report, 2023).
  2. Ncell’s Customer Churn Prediction

    • Idea: Logistic Regression + Neural Networks predict which users will cancel their plans.
    • How: Inputs include call duration, data usage, and customer service interactions. The model scores users on a churn risk scale (0–1). Ncell then offers discounts to high-risk users.
    • Impact: Retains 15% more customers annually (Ncell’s 2022 CSR report).
  3. Nepal’s NTC Network Failure Classification

    • Idea: Decision Trees classify network errors in real time.
    • How: NTC’s logs (error codes, router status, traffic) are fed into a trained tree. For example:
      • If error code = "E101" AND traffic > 90% capacity, classify as "Router Overload."
    • Impact: Reduces mean time to repair (MTTR) by 40% (NTC’s 2023 efficiency report).

Exam Tip

  1. Define clearly:

    • Start answers with: "Classification is a supervised learning technique where..."
    • Differentiate between classification (discrete labels) and regression (continuous values).
  2. Algorithm comparisons:

    • Exams often ask: "When would you use a decision tree vs. SVM?"
    • Decision Trees: Use when interpretability matters (e.g., business rules).
    • SVM: Use for small, high-dimensional data with clear margins.
    • Neural Networks: Use for large, complex data (e.g., images, text).
  3. Worked examples:

    • Always show calculations for metrics (accuracy, precision, recall).
    • For confusion matrices, label all four quadrants (TP, FP, FN, TN).
  4. Real-world ties:

    • Link algorithms to Nepali companies (e.g., "Khalti uses ensemble methods for fraud detection").
    • Avoid vague answers like "used in banking." Specify which algorithm and how.
  5. Common pitfalls:

    • Don’t assume independence in Naive Bayes (mention it’s a "naive" assumption).
    • Don’t ignore preprocessing (normalization, handling missing values).
    • Don’t overlook evaluation metrics—accuracy alone is insufficient for imbalanced data.
  6. Diagrams:

    • Draw decision trees for rule-based explanations.
    • Sketch neural network layers for deep learning questions.
    • Show confusion matrices for evaluation questions.

Based on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 7.

Discussion

Loading…