CMP346 Artificial Intelligence

Artificial IntelligenceUnit 813 min read

Neural Networks: Architectures, Learning & Applications

Unit 8 of Artificial Intelligence explores neural networks—how they mimic human brain neurons, process data through layers, and learn via backpropagation. Covers architectures (feedforward, CNN, RNN), activation functions, loss metrics, and real-world deployments like recommendation systems and autonomous vehicles.

TAKEAWAYS:

  • Neural networks are computational models inspired by biological neurons, organized into layers (input, hidden, output) that transform data through weighted connections.
  • Forward propagation computes outputs, while backpropagation adjusts weights using gradient descent to minimize loss (e.g., mean squared error or cross-entropy).
  • Key architectures include feedforward networks (static data), convolutional neural networks (CNNs) (images), and recurrent neural networks (RNNs/LSTMs) (sequences like text or time-series).
  • Activation functions (ReLU, sigmoid, tanh) introduce non-linearity; loss functions measure prediction error to guide learning.
  • Real-world applications span Nepali apps (e.g., Khalti’s fraud detection via anomaly detection networks) to global tech (Google’s image recognition, Ncell’s predictive maintenance).
  • Challenges include vanishing gradients, overfitting, and computational cost, addressed via techniques like dropout, batch normalization, and hardware acceleration (GPUs/TPUs).

1. Biological Inspiration and Artificial Neurons

Neural networks emulate the human brain’s neurons, which communicate via electrical/chemical signals. An artificial neuron (perceptron) is the building block:

  • Inputs: Received signals multiplied by weights .
  • Bias: A trainable offset (like a neuron’s resting potential).
  • Activation: A non-linear function (e.g., ) determines output.
graph LR
    A["Inputs \(x_1, x_2\)"] -->|"Weights w_1, w_2"| B["Summation: \(z = w_1x_1 + w_2x_2 + b\)"]
    B --> C["Activation: \(f(z)\)"]
    C --> D["Output"]

Why non-linearity? Linear functions (e.g., ) cannot model complex patterns. Activation functions enable hierarchical feature learning.



2. Neural Network Architectures

A. Feedforward Neural Networks (FNNs)

  • Structure: Layers connected unidirectionally (input → hidden → output).
  • Use case: Tabular data (e.g., predicting house prices from features like size, location).
  • Example: A 3-layer FNN for classifying handwritten digits (MNIST dataset).
Input Layer (784 neurons)Hidden Layer (128 neurons, ReLU)Output Layer (10 neurons, Softmax)
3-layer FNN architecture for MNIST handwritten digit classification (784 input features → 128 hidden neurons → 10 output classes)

Worked Example: Predicting Daraz Order Delays

  • Input: Order ID, distance (km), weather (rainy: 1, sunny: 0), time of day (hour).
  • Hidden Layer: 8 neurons with ReLU activation.
  • Output: Predicted delay (hours) via linear activation.
  • Training Data:
    Order ID Distance (km) Rainy Hour Delay (hours)
    1001 5 1 18 2.1
    1002 3 0 10 0.5

Forward Pass Calculation: For Order 1001:

  1. Input vector: .
  2. Hidden layer weights (random init), bias . . Suppose (1×4), then: .
  3. Apply ReLU: .
  4. Repeat for all 8 hidden neurons → output vector .
  5. Output layer weights (8×1) compute delay: .

B. Convolutional Neural Networks (CNNs)

  • Key Idea: Exploit local connectivity and parameter sharing for spatial data (images).
  • Components:
    • Convolutional Layer: Applies filters (kernels) to detect features (edges, textures).
    • Pooling Layer: Reduces spatial dimensions (e.g., max pooling).
    • Fully Connected Layer: Classifies features (e.g., "cat" vs. "dog").
Input Image (32x32x3)Conv Layer (5x5 Filter)ReLU ActivationMax Pooling (2x2)Conv Layer (3x3 Filter)FlattenFC Layer (Softmax)
CNN architecture for image classification (e.g., CIFAR-10). Shows convolution → activation → pooling → flattening → classification.

Real-World Example: Khalti’s Receipt Scanner

  • Problem: Extract text from handwritten receipts.
  • Solution: CNN to detect characters → OCR (Optical Character Recognition) to convert to digital text.
  • Why CNN?
    • Filters learn to detect strokes (e.g., "₹" symbol) regardless of position/size.
    • Pooling reduces noise (e.g., smudged ink).


C. Recurrent Neural Networks (RNNs)

  • Key Idea: Maintain hidden state to model sequences (time-series or text).
  • Types:
    • Vanilla RNN: Suffer from vanishing gradients.
    • LSTM/GRU: Use gating mechanisms to retain long-term memory.
startx_thidden stateoutputh_{t-1}LSTM Cellh_ty_t
LSTM cell structure: input x_t, hidden state h_{t-1}, and output y_t with updated hidden state h_t. Arrows show data flow and memory retention.

Worked Example: Ncell’s Charging Station Demand Prediction

  • Input: Hourly data for 7 days (temperature, holidays, past demand).
  • RNN Task: Predict demand at hour.
  • LSTM Advantage: Captures patterns like "demand spikes on Fridays at 6 PM."

3. Activation Functions

Activation functions introduce non-linearity. Common choices:

-5-4-3-2-112345-150-100-5050100150xySigmoidTanh
Common activation functions: ReLU (linear for x > 0), Sigmoid (S-shaped, bounded [0,1]), Tanh (S-shaped, bounded [-1,1]).
Function Formula Output Range Use Case
Sigmoid (0, 1) Binary classification (output layer)
Tanh (-1, 1) Hidden layers (centered outputs)
ReLU Hidden layers (fast training)
Leaky ReLU if ; else Avoids "dead neurons"

Graph of Activation Functions: Why ReLU Dominates:

  • Computationally efficient (no exponentials).
  • Mitigates vanishing gradient problem (vs. sigmoid/tanh).

4. Loss Functions and Optimization

-5-4-3-2-112345510152025xyMean Squared Error (MSE)Binary Cross-Entropy
Loss landscapes: MSE (convex, for regression) vs. Binary Cross-Entropy (for binary classification).

A. Loss Functions

Measure error between predictions () and true values ():

Loss Function Formula Use Case
Mean Squared Error (MSE) Regression (e.g., house prices)
Cross-Entropy Classification (e.g., MNIST)
Binary Cross-Entropy Binary classification (e.g., spam detection)

Example: NEPSE Stock Prediction

  • Task: Predict tomorrow’s stock price (regression).
  • Loss: MSE between predicted and actual price.
  • Goal: Minimize .

B. Optimization: Gradient Descent

Adjust weights to minimize loss using calculus:

  1. Compute gradient of loss w.r.t. weights ().
  2. Update weights: , where = learning rate.

Variants:

  • Batch Gradient Descent: Use entire dataset (slow but stable).
  • Stochastic GD: Update per sample (noisy but fast).
  • Mini-Batch GD: Compromise (e.g., batches of 32 samples).

Worked Example: Training a Simple Network

  • Network: 1 input, 1 hidden neuron (ReLU), 1 output (linear).
  • Data: , (linear relationship ).
  • Initial Weights: , , , .
  • Forward Pass for :
    • Hidden: , .
    • Output: .
    • Loss (MSE): .
  • Backpropagation:
    • .
    • Update : .

5. Training Neural Networks

A. Backpropagation Algorithm

  1. Forward Pass: Compute predictions.
  2. Compute Loss: Compare predictions to true values.
  3. Backward Pass: Propagate error backward to compute gradients.
  4. Update Weights: Adjust weights using gradients and learning rate.

Mermaid Diagram:

B. Challenges and Solutions

Challenge Solution
Vanishing Gradients Use ReLU/LSTM, residual connections
Overfitting Dropout, L2 regularization
Slow Training Batch normalization, GPUs/TPUs
Local Minima Momentum, adaptive optimizers (Adam)

Example: Pathao’s Traffic Route Optimization

  • Problem: Overfitting to training data (e.g., memorizing exact routes in Kathmandu).
  • Solution: Dropout (randomly deactivate 20% of neurons during training) to generalize better.

6. Real-World Applications in Nepal and Globally

A. Nepali Examples

  1. Khalti’s Fraud Detection

    • Idea: Anomaly detection using autoencoders (a type of neural network).
    • How: Train on normal transactions; flag transactions with high reconstruction error as fraudulent.
    • Output: Alerts like "Transaction ID 12345: 95% fraud probability."
  2. Ncell’s Predictive Maintenance

    • Idea: LSTM networks analyze tower temperature/vibration data.
    • How: Predict equipment failure 24 hours in advance.
    • Impact: Reduces downtime by 30%.
  3. eSewa’s Chatbot

    • Idea: RNN (or Transformer) for natural language understanding.
    • How: Processes queries like "Bill payment for NTC Rs. 500" and routes to the correct API.

B. Global Examples

  1. Google’s Image Search (CNN)

    • Idea: Inception-v3 CNN processes 1000+ classes.
    • How: Detects objects in images (e.g., "cat" in a photo of a Siamese cat).
  2. YouTube’s Recommendation System (Deep Learning)

    • Idea: Multi-layer perceptron (MLP) + collaborative filtering.
    • How: Predicts "You might also like" videos based on watch history.
  3. Tesla’s Autopilot (CNN + RNN)

    • Idea: Combines spatial (CNN for camera input) and temporal (RNN for trajectory) data.
    • How: Detects pedestrians and predicts their movement.


7. Exam Tip

What to Expect in PU Exams:

  1. Theory Questions (30%):

    • Define perceptron, backpropagation, and vanishing gradient problem.
    • Compare CNN vs. RNN architectures (use a table like above).
    • Explain why ReLU is preferred over sigmoid in hidden layers.
  2. Numerical Problems (40%):

    • Forward/backward pass: Given weights and inputs, compute outputs or gradients. Example: "For a network with , , and input , compute the hidden layer output using ReLU."
    • Loss calculation: Compute MSE or cross-entropy for given predictions.
    • Weight updates: Perform one step of gradient descent.
  3. Application-Based (30%):

    • Case Study: "How would you design a neural network for NTC’s customer churn prediction?" Structure:
      • Input: Customer data (call duration, complaints, usage).
      • Hidden: 2 layers (128 → 64 neurons, ReLU).
      • Output: 1 neuron (sigmoid) for "churn (1) or not (0)".
      • Loss: Binary cross-entropy.
    • Diagram: Draw a CNN for a given task (e.g., "design a network to classify Nepali handwritten digits").

Common Pitfalls:

  • Forgetting to apply activation functions (e.g., using linear output for classification).
  • Misaligning dimensions in matrix multiplications (e.g., input shape vs. weights).
  • Confusing batch size (number of samples per update) with epochs (full passes over data).

Pro Tip:

  • Memorize the shapes:
    • Input layer size = number of features.
    • Hidden layer size = hyperparameter (tune via validation error).
    • Output layer size = number of classes (for classification) or 1 (for regression).
  • Practice: Use tools like TensorFlow Playground to visualize how layers/learning rates affect training.

Based on the PU BE Computer (PU) syllabus for Artificial Intelligence (CMP346), unit 8.

Discussion

Loading…