Machine LearningUnit 312 min read

Linear & Logistic Regression: Models, Math & Applications

Unit 3 of Machine Learning covers linear regression (predicting continuous outputs) and logistic regression (binary classification), including their mathematical foundations, assumptions, real-world applications, and implementation steps with Python-style pseudocode.

TAKEAWAYS:

  • Linear regression minimizes Mean Squared Error (MSE) to fit a straight-line model to data, while logistic regression uses the sigmoid function for probability outputs.
  • Gradient Descent iteratively updates weights via , where is the learning rate and is the loss function.
  • Logistic regression outputs probabilities (e.g., 0.8 for "spam") and uses thresholding (e.g., 0.5) to classify.
  • Overfitting occurs when the model fits noise; regularization (L1/L2) or cross-validation helps mitigate it.
  • Real-world uses: Khalti (fraud detection via logistic regression), Ncell (customer churn prediction), and Daraz (demand forecasting with linear regression).
  • Exam focus: Derive loss functions, explain gradient descent steps, and compare linear vs. logistic regression.

1. Linear Regression: Predicting Continuous Outcomes

Linear regression models the relationship between a dependent variable (e.g., house price) and independent variables (e.g., area) as: where:

  • : bias (intercept),
  • : weight (slope),
  • : error term (noise).

How It Works: Least Squares Method

The goal is to minimize the sum of squared errors (SSE) between predicted () and actual () values: where = number of training examples.

-2-1.5-1-0.50.511.52-2-11234xyŷ = w₀ + w₁x (Predicted)True y (with noise)Data Point 1Data Point 2
Least squares minimizes vertical distances (residuals) between predicted (line) and true values

Gradient Descent Update Rules: where:

Worked Example: Predicting Daraz Order Delays

Suppose Daraz tracks order delays () based on distance () from the warehouse. Given data:

Distance (km) Delay (hours)
5 2
10 4
15 5

Step 1: Initialize , , learning rate . Step 2: Compute predictions . Step 3: Update weights for 100 iterations (pseudocode):

for epoch in range(100):
    for (x, y) in data:
        prediction = w0 + w1 * x
        error = prediction - y
        w0 -= alpha * error
        w1 -= alpha * error * x

Final weights: , . Prediction: For km, hours.

Assumptions and Limitations

Assumption Violation Solution
Linear relationship Non-linear data (e.g., exponential) Use polynomial features or kernels
Homoscedasticity (constant variance) Heteroscedasticity (uneven spread) Transform (e.g., log) or use robust regression
No multicollinearity Highly correlated features Remove features or use PCA

2. Logistic Regression: Binary Classification

Logistic regression predicts probabilities for binary outcomes (e.g., "spam" or "not spam") using the sigmoid function: where .

Loss Function: Log Loss (Cross-Entropy)

The loss for a single example is: Gradient Descent Updates:

0.10.20.30.40.50.60.70.80.910.511.522.5xyp=0.1, y=1p=0.9, y=1
Log loss penalizes wrong predictions more heavily when confidence is high (e.g., p=0.1 vs p=0.9 for y=1)

Worked Example: Khalti Fraud Detection

Khalti flags transactions as fraudulent () or legitimate () based on transaction amount (). Given:

Amount (₹) Fraud ()
500 0
1000 1
2000 1

Step 1: Initialize , , . Step 2: Predict probability . Step 3: Update weights (after 50 iterations):

for epoch in range(50):
    for (x, y) in data:
        z = w0 + w1 * x
        prediction = 1 / (1 + exp(-z))
        error = prediction - y
        w0 -= alpha * error
        w1 -= alpha * error * x

Final weights: , . Prediction: For , . → Classify as fraud (threshold = 0.5).

Decision Boundary and Thresholding

  • The decision boundary is where , i.e., .
  • Threshold tuning: Adjust the threshold (e.g., 0.3 for imbalanced data) to prioritize precision/recall.

3. Comparing Linear and Logistic Regression

Feature Linear Regression Logistic Regression
Output Continuous (e.g., price, temperature) Probability (0 to 1)
Loss Function Mean Squared Error (MSE) Log Loss (Cross-Entropy)
Use Case Prediction (regression) Classification (binary/multi-class)
Gradient Descent Updates weights for Updates weights for
Interpretability Coefficients show feature impact (e.g., +$100 per m²) Odds ratios show feature impact (e.g., 2x risk per ₹1000)
Example Predicting house prices Detecting spam emails or fraudulent transactions

4. Real-World Applications

In Nepal

  1. Khalti (Fraud Detection)

    • Idea: Logistic regression classifies transactions as fraudulent or legitimate.
    • How: Uses features like transaction amount, time, and user history to compute . Transactions with are flagged.
    • Impact: Reduces false positives by 30% compared to rule-based systems.
  2. Ncell (Customer Churn Prediction)

    • Idea: Logistic regression predicts whether a customer will leave (churn) based on call duration, data usage, and complaints.
    • How: Models . Customers with are targeted for retention offers.
    • Example: A customer with 500 minutes/month, 2GB data, and 3 complaints has .
  3. Daraz (Demand Forecasting)

    • Idea: Linear regression predicts product demand based on seasonality, promotions, and historical sales.
    • How: Fits . Helps optimize inventory.
    • Example: During Dashain, Daraz predicts a 40% increase in diya sales using linear regression on past 5 years of data.

Globally

  1. Google (Ad Click Prediction)

    • Idea: Logistic regression predicts whether a user will click an ad based on features like location, time, and ad relevance.
    • How: Computes . Ads with are shown.
  2. Netflix (Movie Recommendation)

    • Idea: Linear regression (with regularization) predicts user ratings for movies based on genre preferences and past ratings.
    • How: Models .

5. Overfitting and Regularization

Problem: The model fits training data too closely, performing poorly on unseen data. Solutions:

  1. L1 Regularization (Lasso): Adds to the loss function; can shrink some weights to zero (feature selection).
  2. L2 Regularization (Ridge): Adds ; shrinks weights but keeps all features.
  3. Cross-Validation: Split data into training/validation sets to tune hyperparameters (e.g., , learning rate).

Example: Regularized Linear Regression for NEPSE Stock Prediction Suppose we predict NEPSE index movement () using past 5 days' movements ( to ) and a news sentiment score ().

  • Without regularization: Overfits to noise in (e.g., a single outlier day).
  • With L2 (): Weights for shrink to near-zero, improving generalization.

6. Extensions: Multivariate and Polynomial Regression

Multivariate Linear Regression

Extends to multiple features: Example: Predicting Kathmandu traffic time () based on:

  • : Time of day (hours),
  • : Rainfall (mm),
  • : Number of accidents yesterday.

Polynomial Regression

Adds non-linear terms to capture curves: Example: Modeling Ncell data usage growth over time (quadratic term captures acceleration).


Figure 1: Machine Learning Pipeline (Unit 3 fits between preprocessing and evaluation).


123456789101020304050xyNcell Data Usage (Quadratic Fit)Linear Baseline202020222024
Quadratic term (x²) captures accelerating data usage growth over time (hypothetical Ncell data)

Figure 2: Gradient Descent Workflow for Linear/Logistic Regression.


7. Exam Tip: What to Expect

  1. Derivations: Be ready to derive the gradient descent update for both linear and logistic regression from scratch. For example:

    • For linear regression, show how .
    • For logistic regression, recall that .
  2. Interpretation: Explain the meaning of coefficients in context. For example:

    • In a linear regression for Daraz sales, means "each additional km from the warehouse increases delivery time by 0.5 hours."
  3. Comparison Questions: Expect questions like:

    • "When would you use linear regression over logistic regression?" Answer: Use linear regression for continuous outputs (e.g., predicting temperature) and logistic for binary outcomes (e.g., spam detection).
  4. Practical Scenarios: Problems may involve:

    • Threshold tuning: Given a logistic regression model for Khalti, how would you adjust the threshold to reduce false positives? Answer: Increase the threshold (e.g., from 0.5 to 0.7) to require higher confidence before flagging a transaction as fraudulent.
    • Feature scaling: Why is it important in gradient descent? Answer: Features on larger scales (e.g., house area in m² vs. price in ₹) can dominate updates. Normalize to or standardize to zero mean.
  5. Code Snippets: Be familiar with Python-like pseudocode for:

    • Computing predictions,
    • Updating weights,
    • Evaluating loss.

8. Common Pitfalls

  • Ignoring feature scaling: Gradient descent converges slowly if features have vastly different scales.
  • Misinterpreting logistic regression outputs: means 30% probability, not "30% chance of error."
  • Overfitting: Always validate on a hold-out set or use regularization.
  • Assuming linearity: If the data shows a curve, try polynomial features or switch to decision trees.

9. Summary Checklist

Before the exam, ensure you can:

  1. Write the equations for linear and logistic regression.
  2. Derive the gradient descent updates for both.
  3. Explain the sigmoid function and its role in classification.
  4. Describe how regularization (L1/L2) helps prevent overfitting.
  5. Give two real-world examples (one from Nepal, one global) for each model.
  6. Interpret coefficients in context (e.g., "a ₹1000 increase in transaction amount increases fraud probability by X%").

Based on the PU BE Computer (PU) syllabus for Machine Learning (CMP364), unit 3.

Discussion

Loading…