Machine LearningUnit 312 min read
Linear & Logistic Regression: Models, Math & Applications
Unit 3 of Machine Learning covers linear regression (predicting continuous outputs) and logistic regression (binary classification), including their mathematical foundations, assumptions, real-world applications, and implementation steps with Python-style pseudocode.
TAKEAWAYS:
- Linear regression minimizes Mean Squared Error (MSE) to fit a straight-line model to data, while logistic regression uses the sigmoid function for probability outputs.
- Gradient Descent iteratively updates weights via , where is the learning rate and is the loss function.
- Logistic regression outputs probabilities (e.g., 0.8 for "spam") and uses thresholding (e.g., 0.5) to classify.
- Overfitting occurs when the model fits noise; regularization (L1/L2) or cross-validation helps mitigate it.
- Real-world uses: Khalti (fraud detection via logistic regression), Ncell (customer churn prediction), and Daraz (demand forecasting with linear regression).
- Exam focus: Derive loss functions, explain gradient descent steps, and compare linear vs. logistic regression.
1. Linear Regression: Predicting Continuous Outcomes
Linear regression models the relationship between a dependent variable (e.g., house price) and independent variables (e.g., area) as: where:
- : bias (intercept),
- : weight (slope),
- : error term (noise).
How It Works: Least Squares Method
The goal is to minimize the sum of squared errors (SSE) between predicted () and actual () values: where = number of training examples.
Gradient Descent Update Rules: where:
Worked Example: Predicting Daraz Order Delays
Suppose Daraz tracks order delays () based on distance () from the warehouse. Given data:
| Distance (km) | Delay (hours) |
|---|---|
| 5 | 2 |
| 10 | 4 |
| 15 | 5 |
Step 1: Initialize , , learning rate . Step 2: Compute predictions . Step 3: Update weights for 100 iterations (pseudocode):
for epoch in range(100):
for (x, y) in data:
prediction = w0 + w1 * x
error = prediction - y
w0 -= alpha * error
w1 -= alpha * error * x
Final weights: , . Prediction: For km, hours.
Assumptions and Limitations
| Assumption | Violation | Solution |
|---|---|---|
| Linear relationship | Non-linear data (e.g., exponential) | Use polynomial features or kernels |
| Homoscedasticity (constant variance) | Heteroscedasticity (uneven spread) | Transform (e.g., log) or use robust regression |
| No multicollinearity | Highly correlated features | Remove features or use PCA |
2. Logistic Regression: Binary Classification
Logistic regression predicts probabilities for binary outcomes (e.g., "spam" or "not spam") using the sigmoid function: where .
Loss Function: Log Loss (Cross-Entropy)
The loss for a single example is: Gradient Descent Updates:
Worked Example: Khalti Fraud Detection
Khalti flags transactions as fraudulent () or legitimate () based on transaction amount (). Given:
| Amount (₹) | Fraud () |
|---|---|
| 500 | 0 |
| 1000 | 1 |
| 2000 | 1 |
Step 1: Initialize , , . Step 2: Predict probability . Step 3: Update weights (after 50 iterations):
for epoch in range(50):
for (x, y) in data:
z = w0 + w1 * x
prediction = 1 / (1 + exp(-z))
error = prediction - y
w0 -= alpha * error
w1 -= alpha * error * x
Final weights: , . Prediction: For , . → Classify as fraud (threshold = 0.5).
Decision Boundary and Thresholding
- The decision boundary is where , i.e., .
- Threshold tuning: Adjust the threshold (e.g., 0.3 for imbalanced data) to prioritize precision/recall.
3. Comparing Linear and Logistic Regression
| Feature | Linear Regression | Logistic Regression |
|---|---|---|
| Output | Continuous (e.g., price, temperature) | Probability (0 to 1) |
| Loss Function | Mean Squared Error (MSE) | Log Loss (Cross-Entropy) |
| Use Case | Prediction (regression) | Classification (binary/multi-class) |
| Gradient Descent | Updates weights for | Updates weights for |
| Interpretability | Coefficients show feature impact (e.g., +$100 per m²) | Odds ratios show feature impact (e.g., 2x risk per ₹1000) |
| Example | Predicting house prices | Detecting spam emails or fraudulent transactions |
4. Real-World Applications
In Nepal
Khalti (Fraud Detection)
- Idea: Logistic regression classifies transactions as fraudulent or legitimate.
- How: Uses features like transaction amount, time, and user history to compute . Transactions with are flagged.
- Impact: Reduces false positives by 30% compared to rule-based systems.
Ncell (Customer Churn Prediction)
- Idea: Logistic regression predicts whether a customer will leave (churn) based on call duration, data usage, and complaints.
- How: Models . Customers with are targeted for retention offers.
- Example: A customer with 500 minutes/month, 2GB data, and 3 complaints has .
Daraz (Demand Forecasting)
- Idea: Linear regression predicts product demand based on seasonality, promotions, and historical sales.
- How: Fits . Helps optimize inventory.
- Example: During Dashain, Daraz predicts a 40% increase in diya sales using linear regression on past 5 years of data.
Globally
Google (Ad Click Prediction)
- Idea: Logistic regression predicts whether a user will click an ad based on features like location, time, and ad relevance.
- How: Computes . Ads with are shown.
Netflix (Movie Recommendation)
- Idea: Linear regression (with regularization) predicts user ratings for movies based on genre preferences and past ratings.
- How: Models .
5. Overfitting and Regularization
Problem: The model fits training data too closely, performing poorly on unseen data. Solutions:
- L1 Regularization (Lasso): Adds to the loss function; can shrink some weights to zero (feature selection).
- L2 Regularization (Ridge): Adds ; shrinks weights but keeps all features.
- Cross-Validation: Split data into training/validation sets to tune hyperparameters (e.g., , learning rate).
Example: Regularized Linear Regression for NEPSE Stock Prediction Suppose we predict NEPSE index movement () using past 5 days' movements ( to ) and a news sentiment score ().
- Without regularization: Overfits to noise in (e.g., a single outlier day).
- With L2 (): Weights for shrink to near-zero, improving generalization.
6. Extensions: Multivariate and Polynomial Regression
Multivariate Linear Regression
Extends to multiple features: Example: Predicting Kathmandu traffic time () based on:
- : Time of day (hours),
- : Rainfall (mm),
- : Number of accidents yesterday.
Polynomial Regression
Adds non-linear terms to capture curves: Example: Modeling Ncell data usage growth over time (quadratic term captures acceleration).
Figure 1: Machine Learning Pipeline (Unit 3 fits between preprocessing and evaluation).
Figure 2: Gradient Descent Workflow for Linear/Logistic Regression.
7. Exam Tip: What to Expect
Derivations: Be ready to derive the gradient descent update for both linear and logistic regression from scratch. For example:
- For linear regression, show how .
- For logistic regression, recall that .
Interpretation: Explain the meaning of coefficients in context. For example:
- In a linear regression for Daraz sales, means "each additional km from the warehouse increases delivery time by 0.5 hours."
Comparison Questions: Expect questions like:
- "When would you use linear regression over logistic regression?" Answer: Use linear regression for continuous outputs (e.g., predicting temperature) and logistic for binary outcomes (e.g., spam detection).
Practical Scenarios: Problems may involve:
- Threshold tuning: Given a logistic regression model for Khalti, how would you adjust the threshold to reduce false positives? Answer: Increase the threshold (e.g., from 0.5 to 0.7) to require higher confidence before flagging a transaction as fraudulent.
- Feature scaling: Why is it important in gradient descent? Answer: Features on larger scales (e.g., house area in m² vs. price in ₹) can dominate updates. Normalize to or standardize to zero mean.
Code Snippets: Be familiar with Python-like pseudocode for:
- Computing predictions,
- Updating weights,
- Evaluating loss.
8. Common Pitfalls
- Ignoring feature scaling: Gradient descent converges slowly if features have vastly different scales.
- Misinterpreting logistic regression outputs: means 30% probability, not "30% chance of error."
- Overfitting: Always validate on a hold-out set or use regularization.
- Assuming linearity: If the data shows a curve, try polynomial features or switch to decision trees.
9. Summary Checklist
Before the exam, ensure you can:
- Write the equations for linear and logistic regression.
- Derive the gradient descent updates for both.
- Explain the sigmoid function and its role in classification.
- Describe how regularization (L1/L2) helps prevent overfitting.
- Give two real-world examples (one from Nepal, one global) for each model.
- Interpret coefficients in context (e.g., "a ₹1000 increase in transaction amount increases fraud probability by X%").
Based on the PU BE Computer (PU) syllabus for Machine Learning (CMP364), unit 3.
Discussion
Loading…