Neural NetworksUnit 59 min read
Multilayer Perceptron: Architecture, Training & Applications
Unit 5 of Neural Networks explores the Multilayer Perceptron (MLP), its architecture (input, hidden, output layers), forward/backward propagation, activation functions, and real-world applications like pattern recognition and regression. Includes comparisons with single-layer perceptrons, training algorithms (e.g., bac
1. What is a Multilayer Perceptron (MLP)?
A Multilayer Perceptron (MLP) is a feedforward artificial neural network with at least one hidden layer between the input and output layers. Unlike the single-layer perceptron (Rosenblatt’s model), MLPs can learn non-linear decision boundaries by combining multiple layers of neurons.
Key Components of an MLP
Why multiple layers?
- Single-layer perceptrons can only model linear decisions (e.g.,
y = w₁x₁ + w₂x₂ + b). - MLPs use non-linear activation functions (e.g., sigmoid, ReLU) to approximate complex functions (e.g., XOR, spiral data).
2. Architecture of an MLP
An MLP consists of:
- Input Layer: Raw features (e.g., pixel values in an image, loan applicant data).
- Hidden Layers (1 or more): Perform feature transformation (e.g., edge detection in images).
- Output Layer: Final prediction (e.g., class label, regression value).
Example: Loan Approval System (Nepalese Bank)
Input Layer: Applicant data (income, credit score, employment status). Hidden Layer 1: Detects patterns (e.g., "high income + stable job → likely approval"). Output Layer: Binary decision (approve/reject).
3. Forward Propagation
Forward propagation computes the output by passing input through each layer sequentially.
Step-by-Step Calculation
For a 2-layer MLP (1 hidden layer):
- Input to Hidden Layer:
- Hidden to Output Layer:
Worked Example: XOR Problem
Input: (0,0), (0,1), (1,0), (1,1)
Output: (0), (1), (1), (0)
| Layer | Weights (W) | Biases (b) | Activation | Output |
|---|---|---|---|---|
| Hidden | [[1,1],[-1,-1]] |
[0, 0] |
ReLU | max(0, z) |
| Output | [[1], [1]] |
[0] |
Sigmoid |
Result:
- For
(1,1): Hidden layer computesz = [0, 0]→a = [0, 0]→ Outputσ(0) = 0.5(but we want0). Fix: Adjust weights via backpropagation (next section).
4. Backpropagation: Training the MLP
Backpropagation optimizes weights by minimizing the loss function (e.g., Mean Squared Error for regression, Cross-Entropy for classification).
Steps
- Compute Loss: Compare predicted () vs. true () output.
- Propagate Error Backward:
- Compute gradients of loss w.r.t. weights ().
- Update weights using gradient descent: where = learning rate (e.g., 0.01).
Example: Loan Approval Loss
Suppose:
- Predicted approval probability:
- True label: (approved)
- Loss (Cross-Entropy):
- Gradient: is computed via chain rule and used to update weights.
5. Activation Functions
Activation functions introduce non-linearity, enabling MLPs to model complex patterns.
| Function | Formula | Output Range | Use Case |
|---|---|---|---|
| Sigmoid | (0, 1) | Binary classification | |
| ReLU | [0, ∞) | Hidden layers (avoids vanishing gradients) | |
| Tanh | (-1, 1) | Hidden layers (zero-centered) | |
| Softmax | (0, 1), sums to 1 | Multi-class classification |
6. Advantages and Limitations of MLPs
Advantages
- Universal Approximation: Can approximate any continuous function (given enough layers/neurons).
- Flexibility: Works for classification, regression, and pattern recognition.
- Scalability: Can be extended to deep learning (e.g., CNNs, RNNs).
Limitations
- Vanishing Gradients: Deep MLPs may struggle with long training times.
- Black Box: Hard to interpret (unlike linear models).
- Hyperparameter Sensitivity: Requires tuning (learning rate, layers, neurons).
7. Applications in Nepal and Globally
In the Real World
eSewa (Nepal):
- Use: Fraud detection in online payments.
- How: MLP analyzes transaction patterns (amount, time, location) to flag anomalies.
- Example: Detects unusual login from Kathmandu → blocks payment.
Ncell (Nepal):
- Use: Customer churn prediction.
- How: MLP processes call logs, data usage, and billing history to predict if a user will switch providers.
Google’s RankBrain:
- Use: Search result ranking.
- How: Deep MLP interprets ambiguous queries (e.g., "best restaurants near me") to reorder results dynamically.
Pathao (Nepal):
- Use: Driver demand forecasting.
- How: MLP predicts peak hours in Lalitpur vs. Bhaktapur to optimize driver dispatch.
8. Comparison: MLP vs. Single-Layer Perceptron
| Feature | Single-Layer Perceptron (SLP) | Multilayer Perceptron (MLP) |
|---|---|---|
| Layers | 1 (input + output) | ≥1 hidden layer |
| Decision Boundary | Linear | Non-linear |
| XOR Problem | Cannot solve | Can solve |
| Training | Perceptron learning rule | Backpropagation |
| Use Case | Simple linear tasks | Complex pattern recognition |
9. Worked Example: House Price Prediction (Regression)
Task: Predict house prices in Kathmandu based on:
- Features:
size (sq.ft),bedrooms,location (0=city, 1=suburb) - Target:
price (in lakhs)
Step 1: Data Preparation
| Size | Bedrooms | Location | Price (Lakhs) |
|---|---|---|---|
| 1000 | 2 | 0 | 30 |
| 1500 | 3 | 1 | 45 |
| 1200 | 2 | 0 | 35 |
Step 2: MLP Architecture
- Input Layer: 3 neurons (size, bedrooms, location).
- Hidden Layer: 4 neurons (ReLU activation).
- Output Layer: 1 neuron (linear activation for regression).
Step 3: Training (Simplified)
- Initialize weights randomly (e.g.,
W = [[0.1, -0.2, 0.3], ...]). - Forward Pass:
- Compute hidden layer output: .
- Apply ReLU: .
- Output: .
- Backpropagation: Adjust weights to minimize MSE loss.
Step 4: Prediction
After training, for a new house:
- Input:
(1300, 2, 0) - Predicted price: ₹38 lakhs (vs. actual ₹40 lakhs).
Exam Tip
What to Focus On
- Architecture: Draw and label an MLP with input/hidden/output layers.
- Forward/Backward Pass: Show calculations for a 2-layer MLP (e.g., XOR).
- Activation Functions: Know when to use sigmoid vs. ReLU vs. softmax.
- Applications: Link MLPs to real-world systems (e.g., fraud detection, recommendation systems).
- Limitations: Mention vanishing gradients and how deep networks mitigate it.
Common Pitfalls
- Confusing SLP and MLP: SLPs cannot solve XOR; MLPs can.
- Activation Functions: Using sigmoid in hidden layers leads to vanishing gradients.
- Loss Functions: Use MSE for regression, cross-entropy for classification.
Sample Exam Questions
- Short Answer:
- "Explain backpropagation in 3 steps."
- "Why is ReLU preferred over sigmoid in hidden layers?"
- Long Answer:
- "Design an MLP to classify handwritten digits (MNIST). Include layers, activations, and loss function."
- Numerical:
- "Given a 2-layer MLP, compute the output for input
(1, 0)with provided weights."
- "Given a 2-layer MLP, compute the output for input
Based on the TU BSc CSIT syllabus for Neural Networks, unit 5.
Discussion
Loading…