Neural NetworksUnit 39 min read
Regression in Neural Networks: Linear, Logistic & Model Building
Unit 3 of Neural Networks explores regression techniques—linear, logistic, and polynomial—as foundational tools for model building in neural networks, covering mathematical formulations, real-world applications, and implementation steps with visual traces.
TAKEAWAYS:
- Regression in neural networks maps input data to continuous or discrete outputs using linear/logistic functions, forming the basis for supervised learning.
- Linear regression minimizes squared error via gradient descent, while logistic regression uses sigmoid activation for binary classification.
- Polynomial regression extends linear models by adding non-linear terms (e.g., , ) to fit complex patterns.
- Model evaluation relies on metrics like Mean Squared Error (MSE), Mean Absolute Error (MAE), and R² score.
- Real-world applications include Khalti’s fraud detection (logistic regression), NTC’s network traffic prediction (linear regression), and Daraz’s demand forecasting (polynomial regression).
- Overfitting is mitigated via regularization (L1/L2) and cross-validation.
1. Introduction to Regression in Neural Networks
Regression is a supervised learning technique that predicts continuous (regression) or discrete (classification) outputs from input data. In neural networks, regression serves as:
- A single-neuron model (perceptron with linear activation).
- A building block for multi-layer perceptrons (MLPs).
- A baseline for comparing non-linear models (e.g., kernel methods).
Key Types of Regression
| Type | Output | Activation Function | Use Case |
|---|---|---|---|
| Linear Regression | Continuous () | Linear () | Predicting house prices, stock trends. |
| Logistic Regression | Binary (0 or 1) | Sigmoid () | Spam detection, medical diagnosis. |
| Polynomial Regression | Continuous | Linear (with terms) | Non-linear trend fitting (e.g., temperature vs. ice cream sales). |
2. Linear Regression: The Foundation
Linear regression models the relationship between input and output as: where:
- = weight (slope),
- = bias (intercept).
How It Works: Gradient Descent
The goal is to minimize the Mean Squared Error (MSE): Steps:
- Initialize and randomly.
- Compute gradients:
- Update weights: where = learning rate.
Worked Example: Predicting NTC’s Network Traffic
Scenario: NTC wants to predict daily internet usage () based on temperature () in Kathmandu. Data:
| Temperature (°C) | Usage (GB) |
|---|---|
| 10 | 500 |
| 20 | 1200 |
| 30 | 2000 |
Steps:
- Plot data (linear trend observed):
Temperature vs. Internet Usage (NTC Data) (Image: Sewaqu, Public domain, via Wikimedia Commons)
*(Imagine a scatter plot with a best-fit line.)*
2. **Initialize:** , , .
3. **Iteration 1:**
- Predict (vs. actual 500).
- Compute gradients and update , .
4. **After 100 iterations:** , .
- Final model: .
**Interpretation:**
- For every 1°C increase, usage rises by **60 GB**.
- **MSE = 12,000** (high error → need more data or polynomial terms).
---
### **3. Logistic Regression: Binary Classification**
Logistic regression predicts **probabilities** (0 to 1) using the **sigmoid function**:
**Output Interpretation:**
- → Class 0.
- → Class 1.
#### **Loss Function: Cross-Entropy**
**Gradient Descent Update:**
#### **Worked Example: Khalti’s Fraud Detection**
**Scenario:** Khalti flags transactions as fraudulent (1) or legitimate (0) based on amount () and time ().
**Data:**
| Amount (₹) | Time (hrs) | Fraud (1/0) |
|------------|------------|-------------|
| 500 | 14 | 1 |
| 2000 | 9 | 0 |
| 1000 | 16 | 1 |
**Steps:**
1. **Feature Engineering:** Combine and into .
2. **Sigmoid Output:** .
3. **Train Model:**
- After 50 iterations: , , .
- For ₹500 at 14 hrs: → → **Flag as fraud (1)**.
**Visualization:**
Logistic Regression Decision Boundary (Image: Eric.LEWIN, CC BY 4.0, via Wikimedia Commons)
*(Shows how maps to probability and class labels.)*
---
### **4. Polynomial Regression: Capturing Non-Linearity**
Linear regression fails when data follows **curves**. Polynomial regression adds higher-order terms:
#### **Worked Example: Daraz’s Demand Forecasting**
**Scenario:** Daraz observes that product demand () peaks at moderate prices () but drops at extremes.
**Data:**
| Price (₹) | Demand (units) |
|-----------|----------------|
| 500 | 100 |
| 1000 | 500 |
| 1500 | 300 |
| 2000 | 50 |
**Steps:**
1. **Add term:**
2. **Train Model:**
- After optimization: .
3. **Plot Fit:**
*(Shows a parabola peaking at ₹1200.)*
**Interpretation:**
- Demand **increases** up to ₹1200, then **decreases** (non-linear effect).
---
### **5. Model Evaluation Metrics**
| **Metric** | **Formula** | **Interpretation** |
|--------------------------|--------------------------------------|---------------------------------------------|
| **MSE** | | Average squared error (lower = better). |
| **MAE** | | Average absolute error. |
| **R² Score** | | % variance explained (1 = perfect fit). |
**Example:**
For NTC’s traffic model ():
- **MSE = 12,000**, **R² = 0.85** → 85% variance explained.
---
### **6. Overfitting and Regularization**
**Problem:** High-degree polynomials fit training data but fail on test data.
**Solutions:**
1. **L1 Regularization (Lasso):**
(Encourages sparsity; some .)
2. **L2 Regularization (Ridge):**
(Shrinks weights smoothly.)
**Example:**
For Daraz’s demand model, add **L2 penalty ()** to prevent overfitting.
---
### **7. Real-World Applications in Nepal**
| **Company/App** | **Regression Type** | **Use Case** |
|------------------|---------------------------|-----------------------------------------------|
| **Khalti** | Logistic Regression | Fraud detection (binary: fraud/legit). |
| **NTC** | Linear Regression | Predicting network traffic spikes. |
| **Daraz** | Polynomial Regression | Optimizing product pricing. |
| **NEPSE** | Linear Regression | Stock price trend analysis. |
| **Pathao** | Ridge Regression | Demand forecasting for driver allocation. |
---
### **8. Exam Tip: What to Focus On**
1. **Derive the update rules** for linear/logistic regression from scratch (show gradients).
2. **Compare MSE vs. MAE vs. R²** in terms of sensitivity to outliers.
3. **Explain overfitting** with a **polynomial regression plot** (high-degree vs. low-degree).
4. **Link to neural networks:**
- A single neuron with linear activation = linear regression.
- A neuron with sigmoid activation = logistic regression.
5. **Practical question:** Given data, **choose** between linear/logistic/polynomial regression and justify.
---
```mermaid
flowchart LR
A[Input Data] --> B[Feature Engineering]
B --> C[Choose Regression Type]
C -->|Linear| D[Linear Regression\n(MSE Loss)]
C -->|Binary| E[Logistic Regression\n(Cross-Entropy Loss)]
C -->|Non-linear| F[Polynomial Regression\n(Add x², x³)]
D & E & F --> G[Train Model\n(Gradient Descent)]
G --> H[Evaluate\n(MSE, R², MAE)]
H -->|Overfit?| I[Apply L1/L2 Regularization]
I --> J[Final Model]
Based on the TU BSc CSIT syllabus for Neural Networks, unit 3.
Discussion
Loading…