CACS486 Machine Learning

Machine LearningUnit 28 min read

Supervised Learning: Regression & Classification

Unit 2 of Machine Learning: Covers regression (predicting continuous values) and classification (categorizing data), including linear regression, logistic regression, decision boundaries, evaluation metrics, and bias-variance trade-offs with real-world examples and step-by-step math.

TAKEAWAYS:

  • Supervised learning predicts outputs from labeled data: regression for numbers (e.g., house prices), classification for categories (e.g., spam/ham).
  • Linear regression fits a straight line to minimize error using least squares or gradient descent; logistic regression uses sigmoid functions for binary classification.
  • Evaluation metrics like accuracy, precision, recall, and the confusion matrix quantify model performance on unseen data.
  • The bias-variance trade-off balances underfitting (high bias) and overfitting (high variance) by tuning model complexity.
  • Decision boundaries (e.g., SVM’s hyperplane) separate classes in feature space; kernel tricks extend this to non-linear data.
  • Real-world tools like Daraz’s price prediction (regression) and eSewa’s fraud detection (classification) use these techniques daily.

1. Supervised Learning: Core Concepts

Supervised learning trains models on labeled datasets to predict outputs. It splits into two tasks:

  • Regression: Predicts continuous values (e.g., salary, temperature).
  • Classification: Predicts discrete labels (e.g., cat/dog, spam/not spam).

Key Difference

flowchart TD
    A["Supervised Learning"] --> B["Regression<br/>(Continuous Output)"]
    A --> C["Classification<br/>(Discrete Output)"]
    B --> D["Example: House Price Prediction"]
    C --> E["Example: Email Spam Detection"]

2. Regression: Predicting Continuous Values

Linear Regression

Fits a straight line to minimize error between predicted and actual values.

-4-224681020406080100xyRegression Line (y = 0.5x + 2)Actual Data (Simulated)(0, 2)(2, 3)(4, 4)(6, 5)(8, 6)
Example of a linear regression line fitted to simulated house price data (y = 0.5x + 2).

Formula: where:

  • : intercept,
  • : slope,
  • : error term.

Worked Example: Daraz Price Prediction Daraz uses regression to estimate product prices based on features like brand, ratings, and discounts. Given data:

X (Brand Rating) Y (Price in Rs.)
4.5 1200
3.8 950
4.2 1100

Step 1: Calculate means

Step 2: Compute slope () and intercept () Best-fit line: Prediction for :

Visualization:

Least Squares Method

Minimizes the sum of squared errors (SSE): Gradient Descent for Linear Regression Updates weights iteratively: where is the learning rate.


3. Classification: Categorizing Data

Logistic Regression

Predicts probabilities for binary classification using the sigmoid function: Output: Probability . Threshold at 0.5 to classify.

Example: eSewa Fraud Detection eSewa uses logistic regression to flag suspicious transactions. Given features:

Transaction Amount (X) Is Fraud (Y)
500 No
2000 Yes
1000 No

Worked Trace:

  1. Train model on historical data to learn .
  2. For a new transaction :
  3. If , flag as fraud.

Decision Boundary:


4. Evaluation Metrics

Confusion Matrix

For binary classification:

          | Predicted: Yes | Predicted: No
----------|----------------|--------------
Actual: Yes | True Positive (TP) | False Negative (FN)
Actual: No  | False Positive (FP) | True Negative (TN)

Metrics:

  • Accuracy:
  • Precision: (How many predicted positives are correct?)
  • Recall (Sensitivity): (How many actual positives are caught?)
  • F1-Score: Harmonic mean of precision and recall.

Example: Given:

Predicted Yes Predicted No
Actual Yes 80 20
Actual No 10 90

Calculations:

  • Accuracy: (85%)
  • Precision: (89%)
  • Recall: (80%)

5. Bias-Variance Trade-off

  • High Bias (Underfitting): Model is too simple (e.g., linear regression for non-linear data).
  • High Variance (Overfitting): Model fits noise (e.g., polynomial regression with degree 20).

Visualization:

Worked Example: Loan Approval A bank uses a linear model to predict loan defaults.

  • Low complexity: Ignores risk factors → high bias (rejects many valid loans).
  • High complexity: Captures noise → high variance (approves risky loans).

6. Advanced Classification: Decision Trees

Splits data based on feature thresholds to classify. Example: Pathao Ride Classification Features: distance, time, weather.

  • If distance > 5 km → "Long Ride"
  • Else if weather = "rainy" → "Short Rainy Ride"
  • Else → "Short Normal Ride"
503070204080
Decision tree for Pathao ride classification (highlighted node: split on 'weather = rainy').

Visualization:

flowchart TD
    A["Distance > 5 km?"] -->|"Yes"| B["Long Ride"]
    A -->|"No"| C["Weather Rainy?"]
    C -->|"Yes"| D["Short Rainy Ride"]
    C -->|"No"| E["Short Normal Ride"]
    A -->|"Else"| F["Short Normal Ride"]

In the Real World

  1. Daraz (Regression)

    • Idea: Linear regression predicts product prices based on seller ratings, discounts, and brand.
    • How: Uses historical sales data to adjust prices dynamically, improving profit margins.
    • Worked Example: If a seller’s rating drops from 4.5 to 3.8, Daraz’s model predicts a 15% price increase (as seen in the Daraz Price Prediction example above).
  2. eSewa (Classification)

    • Idea: Logistic regression classifies transactions as fraudulent or legitimate.
    • How: Analyzes transaction amount, time, and user history to flag anomalies (e.g., sudden large transfers).
    • Worked Example: A transaction of Rs. 2000 at 3 AM with a new device triggers a 0.85 probability of fraud, prompting a manual review.
  3. Ncell (Customer Churn Prediction)

    • Idea: Decision trees predict whether a subscriber will cancel their plan.
    • How: Uses features like call duration, data usage, and customer service interactions to split data into "churn" or "retain" nodes.
    • Worked Example: Subscribers with <100 call minutes/month and no data usage are 80% likely to churn, prompting targeted retention offers.

Exam Tip

  • Focus on:
    • Deriving the linear regression equation from scratch (means, slope, intercept).
    • Confusion matrix metrics (precision, recall, F1-score) with numerical examples.
    • Bias-variance trade-off with real-world analogies (e.g., "a simple rule vs. a complex rule").
    • Decision trees: how splits are made and visualized.
  • Common Pitfalls:
    • Mixing up precision and recall (precision = correct positives / predicted positives; recall = correct positives / actual positives).
    • Forgetting to normalize data before logistic regression or gradient descent.
    • Overcomplicating the bias-variance trade-off—stick to high/low bias and variance examples.
  • Time Management:
    • Spend 30% of time on definitions (e.g., "What is logistic regression?").
    • Allocate 40% to worked examples (e.g., derive the regression line or build a confusion matrix).
    • Save 30% for comparisons (e.g., "Why is SVM better than logistic regression for non-linear data?").

Visual Recap:

mindmap
  root((Supervised Learning))
    Regression
      Linear Regression
        Least Squares
        Gradient Descent
      Example: Daraz Price Prediction
    Classification
      Logistic Regression
        Sigmoid Function
        Decision Boundary
      Decision Trees
        Splits & Nodes
        Example: Pathao Ride Classification
    Evaluation
      Confusion Matrix
        Precision, Recall, F1-Score
      Bias-Variance Trade-off
        Underfitting vs. Overfitting
    Real-World
      eSewa Fraud Detection
      Ncell Customer Churn

Based on the TU BCA syllabus for Machine Learning (CACS486), unit 2.

Discussion

Loading…