Data Warehousing and Data MiningUnit 79 min read
Classification: Models, Algorithms & Real-World Applications
Unit 7 of Data Warehousing and Data Mining covers supervised learning techniques for classification, including decision trees, naive Bayes, neural networks, and support vector machines, with step-by-step examples, algorithm comparisons, and real-world applications in Nepalese and global industries.
Classification: Predicting Categories from Data
What is Classification?
Classification is a supervised learning technique used to predict categorical labels (classes) from input data. Unlike regression (which predicts continuous values), classification assigns predefined classes (e.g., "spam" or "not spam," "loan approved" or "rejected").
Key Idea: Given a dataset with labeled examples (features + class), the model learns patterns to classify new, unseen data.
1. Decision Trees: Splitting Data Like a Detective
Decision trees classify data by recursively splitting it based on feature values. Each internal node represents a test (e.g., "Is age > 30?"), branches represent outcomes, and leaves represent class labels.
How It Works
- Select the best feature to split using metrics like Gini impurity or entropy.
- Split the data into subsets.
- Repeat until all leaves are pure (or a stopping criterion is met).
Example: Loan Approval Prediction
Dataset:
| Age | Income (₹) | Credit Score | Loan Approved? |
|---|---|---|---|
| 25 | 30,000 | 650 | No |
| 35 | 50,000 | 720 | Yes |
| 40 | 45,000 | 680 | Yes |
Decision Tree Steps:
- First Split: Choose "Income" (highest information gain).
- If Income ≤ 40,000 → No
- If Income > 40,000 → Check next feature (e.g., Credit Score).
- Second Split: If Credit Score ≥ 700 → Yes, else No.
Advantages:
- Easy to understand and visualize.
- Requires little data preprocessing.
Disadvantages:
- Prone to overfitting (use pruning).
- Sensitive to small data changes.
2. Naive Bayes: Probability in Action
Naive Bayes uses Bayes’ Theorem with an assumption of feature independence (hence "naive"). It calculates the probability of a class given features.
Formula: Where:
- = Prior probability of the class.
- = Likelihood (probability of features given the class).
Example: Spam Detection (Khalti Transactions)
Dataset:
| Word | Spam (P) | Not Spam (P) |
|---|---|---|
| "Free" | 0.9 | 0.1 |
| "Offer" | 0.8 | 0.2 |
Question: Is "Free Offer" spam? Assume:
- , .
- (naive assumption).
Calculation: Since , classify as Spam.
Advantages:
- Fast and works well with high-dimensional data.
- Performs well even with limited training data.
Disadvantages:
- Assumes feature independence (often unrealistic).
3. Support Vector Machines (SVM): Finding the Best Boundary
SVM finds the optimal hyperplane that separates classes with the maximum margin. Useful for both linear and non-linear data (with kernels).
How It Works
- Linear SVM: Maximize the margin between classes.
- Non-linear SVM: Use kernels (e.g., RBF) to transform data into higher dimensions.
Example: Customer Segmentation (Daraz Orders)
Dataset:
| Age | Purchase Frequency | Segment |
|---|---|---|
| 25 | High | Premium |
| 40 | Low | Standard |
SVM Decision Boundary:
- Find the line that best separates "Premium" and "Standard" customers.
- Use a kernel (e.g., polynomial) if data isn’t linearly separable.
Advantages:
- Effective in high-dimensional spaces.
- Memory efficient.
Disadvantages:
- Computationally expensive for large datasets.
- Requires careful kernel and parameter tuning.
4. Neural Networks for Classification
Neural networks (a type of multilayer perceptron) use layers of interconnected nodes to learn complex patterns.
Structure:
Input Layer → Hidden Layers → Output Layer (Softmax for classification)
Example: Handwritten Digit Recognition (Nepali Bank Cheques)
- Input: Pixel values of a digit image.
- Hidden Layers: Learn features (edges, curves).
- Output: Probability distribution over digits (0–9).
Advantages:
- Can model highly non-linear relationships.
- State-of-the-art performance for complex tasks.
Disadvantages:
- Requires large datasets and computational power.
- Black-box nature (hard to interpret).
In the Real World
eSewa (Nepal):
- Uses decision trees to classify transaction risks (e.g., fraud detection).
- Example: If a user logs in from a new device + high transaction amount → flag as suspicious.
Pathao (Ride-Hailing):
- Naive Bayes predicts rider destinations based on historical data (e.g., "If time > 8 PM and location = Thapathali, likely going to home").
- Neural networks analyze driver behavior to classify "safe" vs. "risky" drivers.
Nepal Stock Exchange (NEPSE):
- SVM classifies stock trends (e.g., "Buy," "Hold," "Sell") based on technical indicators.
- Example: If moving average > 50-day MA → classify as "Buy."
Comparison of Classification Algorithms
| Algorithm | Best For | Pros | Cons |
|---|---|---|---|
| Decision Trees | Interpretability, small datasets | Easy to visualize | Overfitting |
| Naive Bayes | Text classification, speed | Fast, works with limited data | Feature independence assumption |
| SVM | High-dimensional data | Effective in complex spaces | Slow for large datasets |
| Neural Networks | Complex patterns (images, speech) | High accuracy | Needs big data, black-box |
Exam Tip
Understand the Math:
- Know Gini impurity, entropy, and Bayes’ Theorem formulas.
- Example: For a decision tree, calculate information gain for a feature.
Practical Scenarios:
- Relate algorithms to Nepali examples (e.g., Khalti fraud, Daraz recommendations).
- Always justify why an algorithm is chosen (e.g., "Naive Bayes for spam because it’s fast").
Diagrams Are Key:
- Draw decision trees, SVM margins, and neural network layers in exams.
- Label axes, nodes, and probabilities clearly.
Common Pitfalls:
- Avoid overfitting in decision trees (mention pruning).
- Don’t assume feature independence in real-world Naive Bayes (but explain why it’s still used).
Final Note: Classification is everywhere—from loan approvals (banks) to disease diagnosis (hospitals). Master the trade-offs between accuracy, speed, and interpretability to ace your exam!
In the real world
- Khalti (Nepal): Uses Naive Bayes to flag suspicious transactions by analyzing word patterns in user messages (e.g., 'free offer' triggers spam alerts).
- Pathao (Ride-Hailing): Employs Decision Trees to classify driver risk scores based on trip history (e.g., late-night rides in high-crime zones).
- Nepal Stock Exchange (NEPSE): Applies SVM to classify stock trends (e.g., 'Buy' vs. 'Sell') using technical indicators like moving averages and volume spikes.
Based on the TU BIM syllabus for Data Warehousing and Data Mining (IT274), unit 7.
Discussion
Loading…