Machine LearningUnit 28 min read
Supervised Learning: Regression & Classification
Unit 2 of Machine Learning: Covers regression (predicting continuous values) and classification (categorizing data), including linear regression, logistic regression, decision boundaries, evaluation metrics, and bias-variance trade-offs with real-world examples and step-by-step math.
TAKEAWAYS:
- Supervised learning predicts outputs from labeled data: regression for numbers (e.g., house prices), classification for categories (e.g., spam/ham).
- Linear regression fits a straight line to minimize error using least squares or gradient descent; logistic regression uses sigmoid functions for binary classification.
- Evaluation metrics like accuracy, precision, recall, and the confusion matrix quantify model performance on unseen data.
- The bias-variance trade-off balances underfitting (high bias) and overfitting (high variance) by tuning model complexity.
- Decision boundaries (e.g., SVM’s hyperplane) separate classes in feature space; kernel tricks extend this to non-linear data.
- Real-world tools like Daraz’s price prediction (regression) and eSewa’s fraud detection (classification) use these techniques daily.
1. Supervised Learning: Core Concepts
Supervised learning trains models on labeled datasets to predict outputs. It splits into two tasks:
- Regression: Predicts continuous values (e.g., salary, temperature).
- Classification: Predicts discrete labels (e.g., cat/dog, spam/not spam).
Key Difference
flowchart TD
A["Supervised Learning"] --> B["Regression<br/>(Continuous Output)"]
A --> C["Classification<br/>(Discrete Output)"]
B --> D["Example: House Price Prediction"]
C --> E["Example: Email Spam Detection"]2. Regression: Predicting Continuous Values
Linear Regression
Fits a straight line to minimize error between predicted and actual values.
Formula: where:
- : intercept,
- : slope,
- : error term.
Worked Example: Daraz Price Prediction Daraz uses regression to estimate product prices based on features like brand, ratings, and discounts. Given data:
| X (Brand Rating) | Y (Price in Rs.) |
|---|---|
| 4.5 | 1200 |
| 3.8 | 950 |
| 4.2 | 1100 |
Step 1: Calculate means
Step 2: Compute slope () and intercept () Best-fit line: Prediction for :
Visualization:
Least Squares Method
Minimizes the sum of squared errors (SSE): Gradient Descent for Linear Regression Updates weights iteratively: where is the learning rate.
3. Classification: Categorizing Data
Logistic Regression
Predicts probabilities for binary classification using the sigmoid function: Output: Probability . Threshold at 0.5 to classify.
Example: eSewa Fraud Detection eSewa uses logistic regression to flag suspicious transactions. Given features:
| Transaction Amount (X) | Is Fraud (Y) |
|---|---|
| 500 | No |
| 2000 | Yes |
| 1000 | No |
Worked Trace:
- Train model on historical data to learn .
- For a new transaction :
- If , flag as fraud.
Decision Boundary:
4. Evaluation Metrics
Confusion Matrix
For binary classification:
| Predicted: Yes | Predicted: No
----------|----------------|--------------
Actual: Yes | True Positive (TP) | False Negative (FN)
Actual: No | False Positive (FP) | True Negative (TN)
Metrics:
- Accuracy:
- Precision: (How many predicted positives are correct?)
- Recall (Sensitivity): (How many actual positives are caught?)
- F1-Score: Harmonic mean of precision and recall.
Example: Given:
| Predicted Yes | Predicted No | |
|---|---|---|
| Actual Yes | 80 | 20 |
| Actual No | 10 | 90 |
Calculations:
- Accuracy: (85%)
- Precision: (89%)
- Recall: (80%)
5. Bias-Variance Trade-off
- High Bias (Underfitting): Model is too simple (e.g., linear regression for non-linear data).
- High Variance (Overfitting): Model fits noise (e.g., polynomial regression with degree 20).
Visualization:
Worked Example: Loan Approval A bank uses a linear model to predict loan defaults.
- Low complexity: Ignores risk factors → high bias (rejects many valid loans).
- High complexity: Captures noise → high variance (approves risky loans).
6. Advanced Classification: Decision Trees
Splits data based on feature thresholds to classify. Example: Pathao Ride Classification Features: distance, time, weather.
- If distance > 5 km → "Long Ride"
- Else if weather = "rainy" → "Short Rainy Ride"
- Else → "Short Normal Ride"
Visualization:
flowchart TD
A["Distance > 5 km?"] -->|"Yes"| B["Long Ride"]
A -->|"No"| C["Weather Rainy?"]
C -->|"Yes"| D["Short Rainy Ride"]
C -->|"No"| E["Short Normal Ride"]
A -->|"Else"| F["Short Normal Ride"]In the Real World
Daraz (Regression)
- Idea: Linear regression predicts product prices based on seller ratings, discounts, and brand.
- How: Uses historical sales data to adjust prices dynamically, improving profit margins.
- Worked Example: If a seller’s rating drops from 4.5 to 3.8, Daraz’s model predicts a 15% price increase (as seen in the Daraz Price Prediction example above).
eSewa (Classification)
- Idea: Logistic regression classifies transactions as fraudulent or legitimate.
- How: Analyzes transaction amount, time, and user history to flag anomalies (e.g., sudden large transfers).
- Worked Example: A transaction of Rs. 2000 at 3 AM with a new device triggers a 0.85 probability of fraud, prompting a manual review.
Ncell (Customer Churn Prediction)
- Idea: Decision trees predict whether a subscriber will cancel their plan.
- How: Uses features like call duration, data usage, and customer service interactions to split data into "churn" or "retain" nodes.
- Worked Example: Subscribers with <100 call minutes/month and no data usage are 80% likely to churn, prompting targeted retention offers.
Exam Tip
- Focus on:
- Deriving the linear regression equation from scratch (means, slope, intercept).
- Confusion matrix metrics (precision, recall, F1-score) with numerical examples.
- Bias-variance trade-off with real-world analogies (e.g., "a simple rule vs. a complex rule").
- Decision trees: how splits are made and visualized.
- Common Pitfalls:
- Mixing up precision and recall (precision = correct positives / predicted positives; recall = correct positives / actual positives).
- Forgetting to normalize data before logistic regression or gradient descent.
- Overcomplicating the bias-variance trade-off—stick to high/low bias and variance examples.
- Time Management:
- Spend 30% of time on definitions (e.g., "What is logistic regression?").
- Allocate 40% to worked examples (e.g., derive the regression line or build a confusion matrix).
- Save 30% for comparisons (e.g., "Why is SVM better than logistic regression for non-linear data?").
Visual Recap:
mindmap
root((Supervised Learning))
Regression
Linear Regression
Least Squares
Gradient Descent
Example: Daraz Price Prediction
Classification
Logistic Regression
Sigmoid Function
Decision Boundary
Decision Trees
Splits & Nodes
Example: Pathao Ride Classification
Evaluation
Confusion Matrix
Precision, Recall, F1-Score
Bias-Variance Trade-off
Underfitting vs. Overfitting
Real-World
eSewa Fraud Detection
Ncell Customer ChurnBased on the TU BCA syllabus for Machine Learning (CACS486), unit 2.
Discussion
Loading…