IT274 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 79 min read

Classification: Models, Algorithms & Real-World Applications

Unit 7 of Data Warehousing and Data Mining covers supervised learning techniques for classification, including decision trees, naive Bayes, neural networks, and support vector machines, with step-by-step examples, algorithm comparisons, and real-world applications in Nepalese and global industries.

Classification: Predicting Categories from Data

What is Classification?

Classification is a supervised learning technique used to predict categorical labels (classes) from input data. Unlike regression (which predicts continuous values), classification assigns predefined classes (e.g., "spam" or "not spam," "loan approved" or "rejected").

Key Idea: Given a dataset with labeled examples (features + class), the model learns patterns to classify new, unseen data.


1. Decision Trees: Splitting Data Like a Detective

Decision trees classify data by recursively splitting it based on feature values. Each internal node represents a test (e.g., "Is age > 30?"), branches represent outcomes, and leaves represent class labels.

Loan Approved: NoLoan Approved: NoLoan Approved: YesCredit Score ≥ 700?Income ≤ 40,000?Loan Approved: YesLoan Approved: NoCredit Score ≥ 650?Age ≤ 30?
Decision tree with two splits (age and income) for loan approval
Loan Approved: NoLoan Approved: NoLoan Approved: YesCredit Score ≥ 700?Income ≤ 40,000?
Decision tree for loan approval (explored path marked)

How It Works

  1. Select the best feature to split using metrics like Gini impurity or entropy.
  2. Split the data into subsets.
  3. Repeat until all leaves are pure (or a stopping criterion is met).

Example: Loan Approval Prediction

Dataset:

Age Income (₹) Credit Score Loan Approved?
25 30,000 650 No
35 50,000 720 Yes
40 45,000 680 Yes

Decision Tree Steps:

  1. First Split: Choose "Income" (highest information gain).
    • If Income ≤ 40,000 → No
    • If Income > 40,000 → Check next feature (e.g., Credit Score).
  2. Second Split: If Credit Score ≥ 700 → Yes, else No.

Advantages:

  • Easy to understand and visualize.
  • Requires little data preprocessing.

Disadvantages:

  • Prone to overfitting (use pruning).
  • Sensitive to small data changes.

2. Naive Bayes: Probability in Action

Naive Bayes uses Bayes’ Theorem with an assumption of feature independence (hence "naive"). It calculates the probability of a class given features.

0.10.20.30.40.50.60.70.80.910.10.20.30.40.5yP(Spam|Free, Offer)P(Not Spam|Free, Offer)P(Spam) = 0.4P(Not Spam) = 0.6
Probability curves for 'Free Offer' spam classification (x = P(Free|Spam))

Formula: Where:

  • = Prior probability of the class.
  • = Likelihood (probability of features given the class).

Example: Spam Detection (Khalti Transactions)

Dataset:

Word Spam (P) Not Spam (P)
"Free" 0.9 0.1
"Offer" 0.8 0.2

Question: Is "Free Offer" spam? Assume:

  • , .
  • (naive assumption).

Calculation: Since , classify as Spam.

Advantages:

  • Fast and works well with high-dimensional data.
  • Performs well even with limited training data.

Disadvantages:

  • Assumes feature independence (often unrealistic).

3. Support Vector Machines (SVM): Finding the Best Boundary

SVM finds the optimal hyperplane that separates classes with the maximum margin. Useful for both linear and non-linear data (with kernels).

111PremiumStandardNon-Premium
SVM decision boundary for Daraz customer segmentation (margin maximized)

How It Works

  1. Linear SVM: Maximize the margin between classes.
  2. Non-linear SVM: Use kernels (e.g., RBF) to transform data into higher dimensions.

Example: Customer Segmentation (Daraz Orders)

Dataset:

Age Purchase Frequency Segment
25 High Premium
40 Low Standard

SVM Decision Boundary:

  • Find the line that best separates "Premium" and "Standard" customers.
  • Use a kernel (e.g., polynomial) if data isn’t linearly separable.

Advantages:

  • Effective in high-dimensional spaces.
  • Memory efficient.

Disadvantages:

  • Computationally expensive for large datasets.
  • Requires careful kernel and parameter tuning.

4. Neural Networks for Classification

Neural networks (a type of multilayer perceptron) use layers of interconnected nodes to learn complex patterns.

Input LayerHidden Layer 1Hidden Layer 2Output Layer
Neural network architecture for handwritten digit recognition (Nepali bank cheques)

Structure:

Input Layer → Hidden Layers → Output Layer (Softmax for classification)

Example: Handwritten Digit Recognition (Nepali Bank Cheques)

  • Input: Pixel values of a digit image.
  • Hidden Layers: Learn features (edges, curves).
  • Output: Probability distribution over digits (0–9).

Advantages:

  • Can model highly non-linear relationships.
  • State-of-the-art performance for complex tasks.

Disadvantages:

  • Requires large datasets and computational power.
  • Black-box nature (hard to interpret).

In the Real World

  1. eSewa (Nepal):

    • Uses decision trees to classify transaction risks (e.g., fraud detection).
    • Example: If a user logs in from a new device + high transaction amount → flag as suspicious.
  2. Pathao (Ride-Hailing):

    • Naive Bayes predicts rider destinations based on historical data (e.g., "If time > 8 PM and location = Thapathali, likely going to home").
    • Neural networks analyze driver behavior to classify "safe" vs. "risky" drivers.
  3. Nepal Stock Exchange (NEPSE):

    • SVM classifies stock trends (e.g., "Buy," "Hold," "Sell") based on technical indicators.
    • Example: If moving average > 50-day MA → classify as "Buy."

Comparison of Classification Algorithms

Algorithm Best For Pros Cons
Decision Trees Interpretability, small datasets Easy to visualize Overfitting
Naive Bayes Text classification, speed Fast, works with limited data Feature independence assumption
SVM High-dimensional data Effective in complex spaces Slow for large datasets
Neural Networks Complex patterns (images, speech) High accuracy Needs big data, black-box

Exam Tip

  1. Understand the Math:

    • Know Gini impurity, entropy, and Bayes’ Theorem formulas.
    • Example: For a decision tree, calculate information gain for a feature.
  2. Practical Scenarios:

    • Relate algorithms to Nepali examples (e.g., Khalti fraud, Daraz recommendations).
    • Always justify why an algorithm is chosen (e.g., "Naive Bayes for spam because it’s fast").
  3. Diagrams Are Key:

    • Draw decision trees, SVM margins, and neural network layers in exams.
    • Label axes, nodes, and probabilities clearly.
  4. Common Pitfalls:

    • Avoid overfitting in decision trees (mention pruning).
    • Don’t assume feature independence in real-world Naive Bayes (but explain why it’s still used).

Final Note: Classification is everywhere—from loan approvals (banks) to disease diagnosis (hospitals). Master the trade-offs between accuracy, speed, and interpretability to ace your exam!

In the real world

  • Khalti (Nepal): Uses Naive Bayes to flag suspicious transactions by analyzing word patterns in user messages (e.g., 'free offer' triggers spam alerts).
  • Pathao (Ride-Hailing): Employs Decision Trees to classify driver risk scores based on trip history (e.g., late-night rides in high-crime zones).
  • Nepal Stock Exchange (NEPSE): Applies SVM to classify stock trends (e.g., 'Buy' vs. 'Sell') using technical indicators like moving averages and volume spikes.

Based on the TU BIM syllabus for Data Warehousing and Data Mining (IT274), unit 7.

Discussion

Loading…