CSC410 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 610 min read

Classification Techniques: Algorithms, Models & Evaluation

Unit 6 of Data Warehousing and Data Mining covers supervised learning classification techniques—decision trees, rule-based systems, neural networks, and evaluation metrics—with real-world applications in fraud detection, recommendation systems, and medical diagnosis.

TAKEAWAYS:

  • Classification assigns predefined labels to data (e.g., spam/ham, loan approved/rejected) using algorithms like decision trees, SVM, or neural networks.
  • Decision trees split data recursively based on feature thresholds (e.g., "Age > 30?"), visualized as hierarchical trees with pruning to avoid overfitting.
  • Neural networks model complex patterns via layered perceptrons (input → hidden → output), trained via backpropagation to minimize loss (e.g., cross-entropy).
  • Evaluation metrics (accuracy, precision, recall, F1-score) use confusion matrices to compare true vs. predicted labels, critical for imbalanced datasets.
  • Real-world use: eSewa uses classification to flag fraudulent transactions; banks predict loan defaults via decision trees; YouTube recommends videos using collaborative filtering (a classification variant).
  • Trade-offs: Decision trees are interpretable but prone to overfitting; neural networks excel at high-dimensional data but require massive training data.

1. What Is Classification?

Classification is a supervised learning technique that predicts discrete class labels from input features. Unlike regression (which predicts continuous values), classification outputs categories (e.g., "Pass/Fail," "Spam/Not Spam").

Key Concepts

  • Features (X): Input variables (e.g., Confident, Studied in the exam dataset).
  • Labels (Y): Target classes (e.g., Pass/Fail).
  • Training Data: Labeled examples used to learn patterns.
  • Test Data: Unseen data to evaluate model performance.

Example: Student Performance Prediction

Given the dataset from past exams:

Confident | Studied | Sick | Result
----------|---------|------|-------
Yes       | No      | No   | Fail
Yes       | No      | No   | Pass
No        | Yes     | Yes  | Fail
...

Question: Predict the result for Confident=Yes, Studied=No, Sick=Yes. Answer: Requires building a classifier (e.g., decision tree) to generalize from the data.


2. Classification Algorithms

A. Decision Trees

Definition: Hierarchical models that split data based on feature thresholds (e.g., "If Age > 30, then Risk=High").

How It Works:

  1. Root Node: Best feature to split (e.g., Confident).
  2. Recursive Splitting: Use metrics like Gini impurity or entropy to choose splits.
  3. Leaf Nodes: Final class labels (e.g., Pass).

Visual: Decision Tree for Student Data

graph TD
    A["Confident?"] -->|"Yes"| B["Studied?"]
    A -->|"No"| C["Result: Fail"]
    B -->|"Yes"| D["Result: Pass"]
    B -->|"No"| E["Sick?"]
    E -->|"Yes"| F["Result: Fail"]
    E -->|"No"| G["Result: Pass"]

Worked Example: Predict Result for [Yes, No, Yes].

  • Path: Yes → No → Yes → Fail.

Advantages:

  • Interpretable (visual rules).
  • Handles non-linear relationships.

Disadvantages:

  • Overfitting (solved via pruning).
  • Sensitive to small data changes.

B. Rule-Based Classification (e.g., RIPPER)

Definition: Learns "IF-THEN" rules from data (e.g., "IF Age=Mid AND Competition=Yes THEN Type=HW").

Example: From the dataset:

IF Confident=Yes AND Studied=No THEN Result=Fail (support=2)

Algorithm Steps:

  1. Grow rules by adding conditions that maximize accuracy.
  2. Prune rules to remove redundant conditions.

C. Neural Networks for Classification

Definition: Multi-layer perceptrons (MLPs) with backpropagation for complex patterns.

Key Components:

  • Input Layer: Features (e.g., Confident, Studied).
  • Hidden Layers: Learn non-linear transformations.
  • Output Layer: Softmax activation for probabilities (e.g., P(Pass)=0.8).

Backpropagation Algorithm (for a 2-class problem):

  1. Forward Pass: Compute output .
  2. Loss: Cross-entropy .
  3. Backward Pass: Update weights .

Visual: Neural Network for Student Data

graph TD
    A["Confident"] --> B["Input Layer"]
    C["Studied"] --> B
    D["Sick"] --> B
    B --> E["Hidden Layer (3 neurons)"]
    E --> F["Output Layer: Pass/Fail"]

Worked Example: Train on [Yes, No, No] → Pass with learning rate .

  • Initialize weights , bias .
  • Forward pass: .
  • Loss: .
  • Update using gradient descent.

3. Evaluation Metrics

Confusion Matrix

Predicted Pass Predicted Fail
Actual Pass True Positive (TP) False Negative (FN)
Actual Fail False Positive (FP) True Negative (TN)

Metrics:

  • Accuracy: .
  • Precision: (e.g., "Of predicted passes, how many are correct?").
  • Recall: (e.g., "Of actual passes, how many did we catch?").
  • F1-Score: Harmonic mean of precision/recall.

Example: For a loan approval model:

  • TP = 50 (correctly approved loans),
  • FP = 10 (wrongly approved),
  • FN = 5 (wrongly rejected).
  • Precision = .

4. Classification vs. Regression

Aspect Classification Regression
Output Discrete labels (e.g., "Pass/Fail") Continuous values (e.g., "Salary")
Algorithms Decision trees, SVM, Neural Networks Linear regression, Polynomial regression
Evaluation Confusion matrix, Accuracy MSE, R²
Example Spam detection House price prediction

5. When to Use Neural Networks?

Neural networks outperform other methods when:

  • Data is high-dimensional (e.g., images, text).
  • Relationships are non-linear (e.g., stock prices).
  • Example: YouTube’s recommendation system uses deep learning to classify user preferences.

Visual: Neural Network Layers

graph TD
    A["Input: User Features"] --> B["Hidden Layer 1: 64 neurons"]
    B --> C["Hidden Layer 2: 32 neurons"]
    C --> D["Output: Video Categories"]

In the Real World

  1. eSewa Fraud Detection:

    • Uses decision trees to classify transactions as fraudulent/legitimate based on features like Amount, Time, Location.
    • Example: A transaction with Amount > Rs. 50,000 and Time=3 AM triggers a "Fraud" flag.
  2. Khalti Loan Approval:

    • Employs neural networks to predict loan defaults by analyzing Income, Credit Score, Loan History.
    • Worked Example: A user with Income=Rs. 80k, Credit Score=650 gets a 70% probability of default.
  3. Pathao Driver Recommendation:

    • Uses collaborative filtering (a classification variant) to match drivers to rides based on Location, Driver Rating, Trip History.
    • Visual: Pathao’s algorithm is like a decision tree where each split is "Is driver in Zone A?" → "Is rating > 4.5?".

Exam Tip

  1. For theoretical questions:

    • Define terms precisely (e.g., "Support vector machines maximize the margin between classes").
    • Draw decision trees or neural network diagrams to illustrate answers.
    • Example: For "Define support vector," write:

      "A support vector is a data point closest to the hyperplane that separates classes in SVM. It defines the margin of separation."

  2. For numerical problems:

    • Show every step of calculations (e.g., entropy, Gini impurity, weight updates).
    • Use small datasets (like the student example) to practice ID3 or RIPPER.
  3. For evaluation:

    • Always construct a confusion matrix before calculating metrics.
    • Example: Given TP=30, FP=5, FN=10, TN=20, compute accuracy and recall.
  4. Common pitfalls:

    • Overfitting: Mention pruning for decision trees or regularization for neural networks.
    • Class imbalance: Discuss precision/recall over accuracy (e.g., fraud detection has few positives).

6. Practical Example: Bank Loan Classification

Scenario: A bank uses a decision tree to classify loan applications as Approved/Rejected based on:

  • Income (Low/Medium/High),
  • Credit Score (300–850),
  • Employment Status (Stable/Unstable).

Dataset:

Income Credit Score Employment Loan Status
Medium 700 Stable Approved
Low 500 Unstable Rejected

Decision Tree:

graph TD
    A["Credit Score > 600?"] -->|"Yes"| B["Loan Status: Approved"]
    A -->|"No"| C["Employment=Stable?"]
    C -->|"Yes"| D["Loan Status: Approved"]
    C -->|"No"| E["Loan Status: Rejected"]

Prediction: For Income=Medium, Credit Score=650, Employment=Unstable → Approved (path: Yes → Approved).


7. Real Picture: Neural Network Output

(Note: The site will replace this with an actual image of a neural network prediction interface.)


8. Summary Table: Classification Algorithms

Algorithm Best For Pros Cons
Decision Trees Interpretable rules Fast, easy to visualize Prone to overfitting
Neural Networks High-dimensional data (images, text) High accuracy Needs large data, black-box
Rule-Based (RIPPER) Simple IF-THEN rules Human-readable Less flexible than trees
SVM Linearly separable data Effective in high dimensions Slow on large datasets

Final Checklist for Exams

  1. Can you draw a decision tree from a dataset?
  2. Do you know the backpropagation steps for a neural network?
  3. Can you compute precision/recall from a confusion matrix?
  4. Can you compare classification vs. regression in a table?
  5. Do you recognize real-world applications (eSewa, Khalti, YouTube)?

Pro Tip: For past exam questions like "Train ID3 classifier for [dataset]", always:

  1. Calculate entropy/Gini for each feature.
  2. Choose the best split (lowest entropy).
  3. Prune the tree if overfitting is suspected.

Based on the TU BSc CSIT syllabus for Data Warehousing and Data Mining (CSC410), unit 6.

Discussion

Loading…