Neural NetworksUnit 59 min read

Multilayer Perceptron: Architecture, Training & Applications

Unit 5 of Neural Networks explores the Multilayer Perceptron (MLP), its architecture (input, hidden, output layers), forward/backward propagation, activation functions, and real-world applications like pattern recognition and regression. Includes comparisons with single-layer perceptrons, training algorithms (e.g., bac

1. What is a Multilayer Perceptron (MLP)?

A Multilayer Perceptron (MLP) is a feedforward artificial neural network with at least one hidden layer between the input and output layers. Unlike the single-layer perceptron (Rosenblatt’s model), MLPs can learn non-linear decision boundaries by combining multiple layers of neurons.

Key Components of an MLP

Forward passForward passForward passBackpropagation (error gradient)BackpropagationBackpropagationInputLayerHiddenLayer1HiddenLayer2OutputLayer
MLP architecture showing forward and backward passes (simplified to 2 hidden layers for clarity)

Why multiple layers?

  • Single-layer perceptrons can only model linear decisions (e.g., y = w₁x₁ + w₂x₂ + b).
  • MLPs use non-linear activation functions (e.g., sigmoid, ReLU) to approximate complex functions (e.g., XOR, spiral data).

2. Architecture of an MLP

An MLP consists of:

  1. Input Layer: Raw features (e.g., pixel values in an image, loan applicant data).
  2. Hidden Layers (1 or more): Perform feature transformation (e.g., edge detection in images).
  3. Output Layer: Final prediction (e.g., class label, regression value).

Example: Loan Approval System (Nepalese Bank)

Input Layer: Applicant data (income, credit score, employment status). Hidden Layer 1: Detects patterns (e.g., "high income + stable job → likely approval"). Output Layer: Binary decision (approve/reject).

[object Object][object Object][object Object][object Object]
Loan approval MLP example with annotated layers (weights/biases omitted for simplicity)

3. Forward Propagation

Forward propagation computes the output by passing input through each layer sequentially.

0.500.810.32
Example weight vector W = [0.5, 0.8, 0.3] for a single neuron's input combination

Step-by-Step Calculation

For a 2-layer MLP (1 hidden layer):

  1. Input to Hidden Layer:
  2. Hidden to Output Layer:

Worked Example: XOR Problem

Input: (0,0), (0,1), (1,0), (1,1) Output: (0), (1), (1), (0)

Layer Weights (W) Biases (b) Activation Output
Hidden [[1,1],[-1,-1]] [0, 0] ReLU max(0, z)
Output [[1], [1]] [0] Sigmoid

Result:

  • For (1,1): Hidden layer computes z = [0, 0] → a = [0, 0] → Output σ(0) = 0.5 (but we want 0). Fix: Adjust weights via backpropagation (next section).

4. Backpropagation: Training the MLP

Backpropagation optimizes weights by minimizing the loss function (e.g., Mean Squared Error for regression, Cross-Entropy for classification).

-1-0.8-0.6-0.4-0.20.20.40.60.810.20.40.60.811.2xyLoss function (MSE)Gradient descent updateMinimum
Loss landscape and gradient descent path (simplified 1D view)

Steps

  1. Compute Loss: Compare predicted () vs. true () output.
  2. Propagate Error Backward:
    • Compute gradients of loss w.r.t. weights ().
    • Update weights using gradient descent: where = learning rate (e.g., 0.01).

Example: Loan Approval Loss

Suppose:

  • Predicted approval probability:
  • True label: (approved)
  • Loss (Cross-Entropy):
  • Gradient: is computed via chain rule and used to update weights.

5. Activation Functions

Activation functions introduce non-linearity, enabling MLPs to model complex patterns.

Function Formula Output Range Use Case
Sigmoid (0, 1) Binary classification
ReLU [0, ∞) Hidden layers (avoids vanishing gradients)
Tanh (-1, 1) Hidden layers (zero-centered)
Softmax (0, 1), sums to 1 Multi-class classification

6. Advantages and Limitations of MLPs

Advantages

  • Universal Approximation: Can approximate any continuous function (given enough layers/neurons).
  • Flexibility: Works for classification, regression, and pattern recognition.
  • Scalability: Can be extended to deep learning (e.g., CNNs, RNNs).

Limitations

  • Vanishing Gradients: Deep MLPs may struggle with long training times.
  • Black Box: Hard to interpret (unlike linear models).
  • Hyperparameter Sensitivity: Requires tuning (learning rate, layers, neurons).

7. Applications in Nepal and Globally

In the Real World

  1. eSewa (Nepal):

    • Use: Fraud detection in online payments.
    • How: MLP analyzes transaction patterns (amount, time, location) to flag anomalies.
    • Example: Detects unusual login from Kathmandu → blocks payment.
  2. Ncell (Nepal):

    • Use: Customer churn prediction.
    • How: MLP processes call logs, data usage, and billing history to predict if a user will switch providers.
  3. Google’s RankBrain:

    • Use: Search result ranking.
    • How: Deep MLP interprets ambiguous queries (e.g., "best restaurants near me") to reorder results dynamically.
  4. Pathao (Nepal):

    • Use: Driver demand forecasting.
    • How: MLP predicts peak hours in Lalitpur vs. Bhaktapur to optimize driver dispatch.

8. Comparison: MLP vs. Single-Layer Perceptron

Feature Single-Layer Perceptron (SLP) Multilayer Perceptron (MLP)
Layers 1 (input + output) ≥1 hidden layer
Decision Boundary Linear Non-linear
XOR Problem Cannot solve Can solve
Training Perceptron learning rule Backpropagation
Use Case Simple linear tasks Complex pattern recognition

9. Worked Example: House Price Prediction (Regression)

Task: Predict house prices in Kathmandu based on:

  • Features: size (sq.ft), bedrooms, location (0=city, 1=suburb)
  • Target: price (in lakhs)

Step 1: Data Preparation

Size Bedrooms Location Price (Lakhs)
1000 2 0 30
1500 3 1 45
1200 2 0 35

Step 2: MLP Architecture

  • Input Layer: 3 neurons (size, bedrooms, location).
  • Hidden Layer: 4 neurons (ReLU activation).
  • Output Layer: 1 neuron (linear activation for regression).

Step 3: Training (Simplified)

  1. Initialize weights randomly (e.g., W = [[0.1, -0.2, 0.3], ...]).
  2. Forward Pass:
    • Compute hidden layer output: .
    • Apply ReLU: .
  3. Output: .
  4. Backpropagation: Adjust weights to minimize MSE loss.

Step 4: Prediction

After training, for a new house:

  • Input: (1300, 2, 0)
  • Predicted price: ₹38 lakhs (vs. actual ₹40 lakhs).

Exam Tip

What to Focus On

  1. Architecture: Draw and label an MLP with input/hidden/output layers.
  2. Forward/Backward Pass: Show calculations for a 2-layer MLP (e.g., XOR).
  3. Activation Functions: Know when to use sigmoid vs. ReLU vs. softmax.
  4. Applications: Link MLPs to real-world systems (e.g., fraud detection, recommendation systems).
  5. Limitations: Mention vanishing gradients and how deep networks mitigate it.

Common Pitfalls

  • Confusing SLP and MLP: SLPs cannot solve XOR; MLPs can.
  • Activation Functions: Using sigmoid in hidden layers leads to vanishing gradients.
  • Loss Functions: Use MSE for regression, cross-entropy for classification.

Sample Exam Questions

  1. Short Answer:
    • "Explain backpropagation in 3 steps."
    • "Why is ReLU preferred over sigmoid in hidden layers?"
  2. Long Answer:
    • "Design an MLP to classify handwritten digits (MNIST). Include layers, activations, and loss function."
  3. Numerical:
    • "Given a 2-layer MLP, compute the output for input (1, 0) with provided weights."

Based on the TU BSc CSIT syllabus for Neural Networks, unit 5.

Discussion

Loading…