Data Warehousing and Data MiningUnit 610 min read
Classification Techniques: Algorithms, Models & Evaluation
Unit 6 of Data Warehousing and Data Mining covers supervised learning classification techniques—decision trees, rule-based systems, neural networks, and evaluation metrics—with real-world applications in fraud detection, recommendation systems, and medical diagnosis.
TAKEAWAYS:
- Classification assigns predefined labels to data (e.g., spam/ham, loan approved/rejected) using algorithms like decision trees, SVM, or neural networks.
- Decision trees split data recursively based on feature thresholds (e.g., "Age > 30?"), visualized as hierarchical trees with pruning to avoid overfitting.
- Neural networks model complex patterns via layered perceptrons (input → hidden → output), trained via backpropagation to minimize loss (e.g., cross-entropy).
- Evaluation metrics (accuracy, precision, recall, F1-score) use confusion matrices to compare true vs. predicted labels, critical for imbalanced datasets.
- Real-world use: eSewa uses classification to flag fraudulent transactions; banks predict loan defaults via decision trees; YouTube recommends videos using collaborative filtering (a classification variant).
- Trade-offs: Decision trees are interpretable but prone to overfitting; neural networks excel at high-dimensional data but require massive training data.
1. What Is Classification?
Classification is a supervised learning technique that predicts discrete class labels from input features. Unlike regression (which predicts continuous values), classification outputs categories (e.g., "Pass/Fail," "Spam/Not Spam").
Key Concepts
- Features (X): Input variables (e.g.,
Confident,Studiedin the exam dataset). - Labels (Y): Target classes (e.g.,
Pass/Fail). - Training Data: Labeled examples used to learn patterns.
- Test Data: Unseen data to evaluate model performance.
Example: Student Performance Prediction
Given the dataset from past exams:
Confident | Studied | Sick | Result
----------|---------|------|-------
Yes | No | No | Fail
Yes | No | No | Pass
No | Yes | Yes | Fail
...
Question: Predict the result for Confident=Yes, Studied=No, Sick=Yes.
Answer: Requires building a classifier (e.g., decision tree) to generalize from the data.
2. Classification Algorithms
A. Decision Trees
Definition: Hierarchical models that split data based on feature thresholds (e.g., "If Age > 30, then Risk=High").
How It Works:
- Root Node: Best feature to split (e.g.,
Confident). - Recursive Splitting: Use metrics like Gini impurity or entropy to choose splits.
- Leaf Nodes: Final class labels (e.g.,
Pass).
Visual: Decision Tree for Student Data
graph TD
A["Confident?"] -->|"Yes"| B["Studied?"]
A -->|"No"| C["Result: Fail"]
B -->|"Yes"| D["Result: Pass"]
B -->|"No"| E["Sick?"]
E -->|"Yes"| F["Result: Fail"]
E -->|"No"| G["Result: Pass"]Worked Example: Predict Result for [Yes, No, Yes].
- Path:
Yes→No→Yes→Fail.
Advantages:
- Interpretable (visual rules).
- Handles non-linear relationships.
Disadvantages:
- Overfitting (solved via pruning).
- Sensitive to small data changes.
B. Rule-Based Classification (e.g., RIPPER)
Definition: Learns "IF-THEN" rules from data (e.g., "IF Age=Mid AND Competition=Yes THEN Type=HW").
Example: From the dataset:
IF Confident=Yes AND Studied=No THEN Result=Fail (support=2)
Algorithm Steps:
- Grow rules by adding conditions that maximize accuracy.
- Prune rules to remove redundant conditions.
C. Neural Networks for Classification
Definition: Multi-layer perceptrons (MLPs) with backpropagation for complex patterns.
Key Components:
- Input Layer: Features (e.g.,
Confident,Studied). - Hidden Layers: Learn non-linear transformations.
- Output Layer: Softmax activation for probabilities (e.g.,
P(Pass)=0.8).
Backpropagation Algorithm (for a 2-class problem):
- Forward Pass: Compute output .
- Loss: Cross-entropy .
- Backward Pass: Update weights .
Visual: Neural Network for Student Data
graph TD
A["Confident"] --> B["Input Layer"]
C["Studied"] --> B
D["Sick"] --> B
B --> E["Hidden Layer (3 neurons)"]
E --> F["Output Layer: Pass/Fail"]Worked Example: Train on [Yes, No, No] → Pass with learning rate .
- Initialize weights , bias .
- Forward pass: .
- Loss: .
- Update using gradient descent.
3. Evaluation Metrics
Confusion Matrix
| Predicted Pass | Predicted Fail | |
|---|---|---|
| Actual Pass | True Positive (TP) | False Negative (FN) |
| Actual Fail | False Positive (FP) | True Negative (TN) |
Metrics:
- Accuracy: .
- Precision: (e.g., "Of predicted passes, how many are correct?").
- Recall: (e.g., "Of actual passes, how many did we catch?").
- F1-Score: Harmonic mean of precision/recall.
Example: For a loan approval model:
- TP = 50 (correctly approved loans),
- FP = 10 (wrongly approved),
- FN = 5 (wrongly rejected).
- Precision = .
4. Classification vs. Regression
| Aspect | Classification | Regression |
|---|---|---|
| Output | Discrete labels (e.g., "Pass/Fail") | Continuous values (e.g., "Salary") |
| Algorithms | Decision trees, SVM, Neural Networks | Linear regression, Polynomial regression |
| Evaluation | Confusion matrix, Accuracy | MSE, R² |
| Example | Spam detection | House price prediction |
5. When to Use Neural Networks?
Neural networks outperform other methods when:
- Data is high-dimensional (e.g., images, text).
- Relationships are non-linear (e.g., stock prices).
- Example: YouTube’s recommendation system uses deep learning to classify user preferences.
Visual: Neural Network Layers
graph TD
A["Input: User Features"] --> B["Hidden Layer 1: 64 neurons"]
B --> C["Hidden Layer 2: 32 neurons"]
C --> D["Output: Video Categories"]In the Real World
eSewa Fraud Detection:
- Uses decision trees to classify transactions as fraudulent/legitimate based on features like
Amount,Time,Location. - Example: A transaction with
Amount > Rs. 50,000andTime=3 AMtriggers a "Fraud" flag.
- Uses decision trees to classify transactions as fraudulent/legitimate based on features like
Khalti Loan Approval:
- Employs neural networks to predict loan defaults by analyzing
Income,Credit Score,Loan History. - Worked Example: A user with
Income=Rs. 80k,Credit Score=650gets a 70% probability of default.
- Employs neural networks to predict loan defaults by analyzing
Pathao Driver Recommendation:
- Uses collaborative filtering (a classification variant) to match drivers to rides based on
Location,Driver Rating,Trip History. - Visual: Pathao’s algorithm is like a decision tree where each split is "Is driver in Zone A?" → "Is rating > 4.5?".
- Uses collaborative filtering (a classification variant) to match drivers to rides based on
Exam Tip
For theoretical questions:
- Define terms precisely (e.g., "Support vector machines maximize the margin between classes").
- Draw decision trees or neural network diagrams to illustrate answers.
- Example: For "Define support vector," write:
"A support vector is a data point closest to the hyperplane that separates classes in SVM. It defines the margin of separation."
For numerical problems:
- Show every step of calculations (e.g., entropy, Gini impurity, weight updates).
- Use small datasets (like the student example) to practice ID3 or RIPPER.
For evaluation:
- Always construct a confusion matrix before calculating metrics.
- Example: Given TP=30, FP=5, FN=10, TN=20, compute accuracy and recall.
Common pitfalls:
- Overfitting: Mention pruning for decision trees or regularization for neural networks.
- Class imbalance: Discuss precision/recall over accuracy (e.g., fraud detection has few positives).
6. Practical Example: Bank Loan Classification
Scenario: A bank uses a decision tree to classify loan applications as Approved/Rejected based on:
Income(Low/Medium/High),Credit Score(300–850),Employment Status(Stable/Unstable).
Dataset:
| Income | Credit Score | Employment | Loan Status |
|---|---|---|---|
| Medium | 700 | Stable | Approved |
| Low | 500 | Unstable | Rejected |
Decision Tree:
graph TD
A["Credit Score > 600?"] -->|"Yes"| B["Loan Status: Approved"]
A -->|"No"| C["Employment=Stable?"]
C -->|"Yes"| D["Loan Status: Approved"]
C -->|"No"| E["Loan Status: Rejected"]Prediction: For Income=Medium, Credit Score=650, Employment=Unstable → Approved (path: Yes → Approved).
7. Real Picture: Neural Network Output
(Note: The site will replace this with an actual image of a neural network prediction interface.)
8. Summary Table: Classification Algorithms
| Algorithm | Best For | Pros | Cons |
|---|---|---|---|
| Decision Trees | Interpretable rules | Fast, easy to visualize | Prone to overfitting |
| Neural Networks | High-dimensional data (images, text) | High accuracy | Needs large data, black-box |
| Rule-Based (RIPPER) | Simple IF-THEN rules | Human-readable | Less flexible than trees |
| SVM | Linearly separable data | Effective in high dimensions | Slow on large datasets |
Final Checklist for Exams
- Can you draw a decision tree from a dataset?
- Do you know the backpropagation steps for a neural network?
- Can you compute precision/recall from a confusion matrix?
- Can you compare classification vs. regression in a table?
- Do you recognize real-world applications (eSewa, Khalti, YouTube)?
Pro Tip: For past exam questions like "Train ID3 classifier for [dataset]", always:
- Calculate entropy/Gini for each feature.
- Choose the best split (lowest entropy).
- Prune the tree if overfitting is suspected.
Based on the TU BSc CSIT syllabus for Data Warehousing and Data Mining (CSC410), unit 6.
Discussion
Loading…