Neural NetworksUnit 39 min read

Regression in Neural Networks: Linear, Logistic & Model Building

Unit 3 of Neural Networks explores regression techniques—linear, logistic, and polynomial—as foundational tools for model building in neural networks, covering mathematical formulations, real-world applications, and implementation steps with visual traces.

TAKEAWAYS:

  • Regression in neural networks maps input data to continuous or discrete outputs using linear/logistic functions, forming the basis for supervised learning.
  • Linear regression minimizes squared error via gradient descent, while logistic regression uses sigmoid activation for binary classification.
  • Polynomial regression extends linear models by adding non-linear terms (e.g., , ) to fit complex patterns.
  • Model evaluation relies on metrics like Mean Squared Error (MSE), Mean Absolute Error (MAE), and R² score.
  • Real-world applications include Khalti’s fraud detection (logistic regression), NTC’s network traffic prediction (linear regression), and Daraz’s demand forecasting (polynomial regression).
  • Overfitting is mitigated via regularization (L1/L2) and cross-validation.

1. Introduction to Regression in Neural Networks

Regression is a supervised learning technique that predicts continuous (regression) or discrete (classification) outputs from input data. In neural networks, regression serves as:

  • A single-neuron model (perceptron with linear activation).
  • A building block for multi-layer perceptrons (MLPs).
  • A baseline for comparing non-linear models (e.g., kernel methods).

Key Types of Regression

Type Output Activation Function Use Case
Linear Regression Continuous () Linear () Predicting house prices, stock trends.
Logistic Regression Binary (0 or 1) Sigmoid () Spam detection, medical diagnosis.
Polynomial Regression Continuous Linear (with terms) Non-linear trend fitting (e.g., temperature vs. ice cream sales).

2. Linear Regression: The Foundation

Linear regression models the relationship between input and output as: where:

  • = weight (slope),
  • = bias (intercept).

How It Works: Gradient Descent

The goal is to minimize the Mean Squared Error (MSE): Steps:

  1. Initialize and randomly.
  2. Compute gradients:
  3. Update weights: where = learning rate.

Worked Example: Predicting NTC’s Network Traffic

Scenario: NTC wants to predict daily internet usage () based on temperature () in Kathmandu. Data:

Temperature (°C) Usage (GB)
10 500
20 1200
30 2000

Steps:

  1. Plot data (linear trend observed):
    
    

scatter plot linear regressionTemperature vs. Internet Usage (NTC Data) (Image: Sewaqu, Public domain, via Wikimedia Commons)

*(Imagine a scatter plot with a best-fit line.)*

2. **Initialize:** , , .
3. **Iteration 1:**
- Predict  (vs. actual 500).
- Compute gradients and update , .
4. **After 100 iterations:** , .
- Final model: .

**Interpretation:**
- For every 1°C increase, usage rises by **60 GB**.
- **MSE = 12,000** (high error → need more data or polynomial terms).

---

### **3. Logistic Regression: Binary Classification**
Logistic regression predicts **probabilities** (0 to 1) using the **sigmoid function**:

**Output Interpretation:**
-  → Class 0.
-  → Class 1.

#### **Loss Function: Cross-Entropy**

**Gradient Descent Update:**


#### **Worked Example: Khalti’s Fraud Detection**
**Scenario:** Khalti flags transactions as fraudulent (1) or legitimate (0) based on amount () and time ().
**Data:**
| Amount (₹) | Time (hrs) | Fraud (1/0) |
|------------|------------|-------------|
| 500        | 14         | 1           |
| 2000       | 9          | 0           |
| 1000       | 16         | 1           |

**Steps:**
1. **Feature Engineering:** Combine  and  into .
2. **Sigmoid Output:** .
3. **Train Model:**
- After 50 iterations: , , .
- For ₹500 at 14 hrs:  →  → **Flag as fraud (1)**.

**Visualization:**

sigmoid function graphLogistic Regression Decision Boundary (Image: Eric.LEWIN, CC BY 4.0, via Wikimedia Commons)

*(Shows how  maps to probability and class labels.)*

---

### **4. Polynomial Regression: Capturing Non-Linearity**
Linear regression fails when data follows **curves**. Polynomial regression adds higher-order terms:


#### **Worked Example: Daraz’s Demand Forecasting**
**Scenario:** Daraz observes that product demand () peaks at moderate prices () but drops at extremes.
**Data:**
| Price (₹) | Demand (units) |
|-----------|----------------|
| 500       | 100            |
| 1000      | 500            |
| 1500      | 300            |
| 2000      | 50             |

**Steps:**
1. **Add  term:**
   
2. **Train Model:**
   - After optimization: .
3. **Plot Fit:**
*(Shows a parabola peaking at ₹1200.)*

**Interpretation:**
- Demand **increases** up to ₹1200, then **decreases** (non-linear effect).

---

### **5. Model Evaluation Metrics**
| **Metric**               | **Formula**                          | **Interpretation**                          |
|--------------------------|--------------------------------------|---------------------------------------------|
| **MSE**                  |  | Average squared error (lower = better).    |
| **MAE**                  |  | Average absolute error.                     |
| **R² Score**             |  | % variance explained (1 = perfect fit).   |

**Example:**
For NTC’s traffic model ():
- **MSE = 12,000**, **R² = 0.85** → 85% variance explained.

---

### **6. Overfitting and Regularization**
**Problem:** High-degree polynomials fit training data but fail on test data.
**Solutions:**
1. **L1 Regularization (Lasso):**

(Encourages sparsity; some .)
2. **L2 Regularization (Ridge):**

(Shrinks weights smoothly.)

**Example:**
For Daraz’s demand model, add **L2 penalty ()** to prevent overfitting.

---

### **7. Real-World Applications in Nepal**
| **Company/App**  | **Regression Type**       | **Use Case**                                  |
|------------------|---------------------------|-----------------------------------------------|
| **Khalti**       | Logistic Regression      | Fraud detection (binary: fraud/legit).        |
| **NTC**          | Linear Regression        | Predicting network traffic spikes.            |
| **Daraz**        | Polynomial Regression    | Optimizing product pricing.                   |
| **NEPSE**        | Linear Regression        | Stock price trend analysis.                   |
| **Pathao**       | Ridge Regression         | Demand forecasting for driver allocation.    |

---

### **8. Exam Tip: What to Focus On**
1. **Derive the update rules** for linear/logistic regression from scratch (show gradients).
2. **Compare MSE vs. MAE vs. R²** in terms of sensitivity to outliers.
3. **Explain overfitting** with a **polynomial regression plot** (high-degree vs. low-degree).
4. **Link to neural networks:**
- A single neuron with linear activation = linear regression.
- A neuron with sigmoid activation = logistic regression.
5. **Practical question:** Given data, **choose** between linear/logistic/polynomial regression and justify.

---
```mermaid
flowchart LR
 A[Input Data] --> B[Feature Engineering]
 B --> C[Choose Regression Type]
 C -->|Linear| D[Linear Regression\n(MSE Loss)]
 C -->|Binary| E[Logistic Regression\n(Cross-Entropy Loss)]
 C -->|Non-linear| F[Polynomial Regression\n(Add x², x³)]
 D & E & F --> G[Train Model\n(Gradient Descent)]
 G --> H[Evaluate\n(MSE, R², MAE)]
 H -->|Overfit?| I[Apply L1/L2 Regularization]
 I --> J[Final Model]

Based on the TU BSc CSIT syllabus for Neural Networks, unit 3.

Discussion

Loading…