CACS410 Artificial Intelligence

Artificial IntelligenceUnit 78 min read

Artificial Neural Networks: Models, Learning & Applications

Unit 7 of Artificial Intelligence explores how artificial neural networks (ANNs) mimic biological neurons to learn patterns, covering architectures (feedforward/recurrent), activation functions, backpropagation, and real-world implementations in image recognition, finance, and recommendation systems.

TAKEAWAYS:

  • Biological Inspiration: ANNs replicate the human brain’s neuron structure (dendrites, axon, synapses) to process information in layers.
  • Key Architectures: Feedforward networks (static) vs. recurrent networks (memory) solve different problems (e.g., image classification vs. time-series prediction).
  • Learning Mechanism: Backpropagation adjusts weights using gradient descent to minimize error, guided by activation functions (sigmoid, ReLU).
  • Real-World Impact: Used in Khalti’s fraud detection (pattern recognition), Daraz’s recommendation systems (collaborative filtering), and NTC’s network traffic prediction (time-series analysis).
  • Math Behind Models: Loss functions (MSE, cross-entropy) and optimization (SGD, Adam) drive training efficiency.
  • Limitations: Require large datasets, struggle with interpretability, and may overfit without regularization.

1. Biological Neurons vs. Artificial Neurons

ANNs are inspired by the human brain’s neurons, which transmit signals via electrical impulses. A single artificial neuron (perceptron) mimics this process:

graph LR
    A["Inputs (x₁, x₂, ..., xₙ)"] --> B["Weights (w₁, w₂, ..., wₙ)"]
    B --> C["Summation: z = Σ(wᵢxᵢ) + b"]
    C --> D["Activation: a = f(z)"]
    D --> E["Output"]
  • Inputs (xᵢ): Features (e.g., pixel values in an image).
  • Weights (wᵢ): Learned parameters adjusted during training.
  • Bias (b): Shifts the activation function.
  • Activation (f): Introduces non-linearity (e.g., sigmoid, ReLU).

2. Neural Network Architectures

A. Feedforward Neural Networks (FNNs)

  • Structure: Layers connected unidirectionally (input → hidden → output).
  • Use Case: Static data (e.g., classifying handwritten digits in eSewa’s OTP verification).
  • Example: A 3-layer FNN for AND gate (weights: w₁=1, w₂=1, b=-1.5; activation: step function).
graph TD
    A["Input Layer"] -->|"x₁, x₂"| B["Hidden Layer (1 neuron)"]
    B -->|"a = step(w₁x₁ + w₂x₂ + b)"| C["Output Layer"]
    C -->|"Output = 1 if a ≥ 0.5"| D["AND Gate"]

Worked Example: For inputs (1, 1): z = 1*1 + 1*1 - 1.5 = 0.5 → a = step(0.5) = 1 (correct AND output).

B. Recurrent Neural Networks (RNNs)

  • Structure: Loops enable memory (e.g., processing sequences like Ncell’s call logs).
  • Variants: LSTM (long-term memory) for time-series (e.g., stock prices in NEPSE).
  • Example: Predicting the next word in a sentence (language models).
graph LR
    A["Input t"] --> B["RNN Cell"]
    B -->|"Hidden State hₜ"| C["Output t"]
    C -->|"Feedback"| B

C. Comparison Table

Feature Feedforward (FNN) Recurrent (RNN/LSTM)
Data Type Static (images, tabular) Sequential (text, time-series)
Memory None Yes (hidden state)
Training Speed Faster Slower (vanishing gradients)
Example Use Khalti’s fraud detection Pathao’s route optimization

3. Activation Functions

Non-linearities enable complex mappings. Common functions:

  1. Sigmoid: Outputs between 0 and 1 (used in binary classification). Graph:
  2. ReLU: Faster training (avoids vanishing gradients).
  3. Tanh: Zero-centered output (better for hidden layers).

Real-World Tie-In:

  • YouTube’s recommendation system uses sigmoid to predict click probability (0 to 1).

4. Training Neural Networks: Backpropagation

Goal: Minimize error (e.g., mean squared error, MSE) between predicted and actual outputs. Steps:

  1. Forward Pass: Compute output and error.
  2. Backward Pass: Propagate error to adjust weights using the chain rule.
  3. Update Weights: Gradient descent: where = learning rate.

Worked Example: Train a neuron to predict y = 2x + 1 (linear regression).

  • Initial weights: w = 0.5, b = 0.
  • Input: x = 1 → Target: y = 3.
  • Forward Pass: z = 0.5*1 + 0 = 0.5 → a = 0.5 (sigmoid) → Error: (3 - 0.5)² = 6.25.
  • Backward Pass: Update: w = 0.5 - 0.1*0.625 = 0.4375.

5. Loss Functions

Measure prediction error:

Function Formula Use Case
MSE Regression (e.g., house pricing)
Cross-Entropy Classification (e.g., spam detection)

Example: NTC’s network latency prediction uses MSE to optimize weights for traffic routing.


6. Optimization Algorithms

  • Gradient Descent (GD): Slow for large datasets.
  • Stochastic GD (SGD): Updates weights per sample (faster but noisy).
  • Adam: Combines momentum and adaptive learning rates (used in Google’s BERT).

Graph:


7. Real-World Applications in Nepal

Company/App ANN Application Key Idea Used
Khalti Fraud detection (credit card transactions) FNN + Sigmoid for anomaly scoring
Daraz Product recommendations Collaborative filtering (RNN/LSTM)
NTC Network traffic prediction LSTM for time-series forecasting
Ncell Call detail record analysis Autoencoders for compression
eSewa OTP verification (image-based) CNN for digit recognition

Worked Example: Daraz’s Recommendation System

  1. Input: User’s past orders (e.g., [laptop, phone, charger]).
  2. RNN Layer: Processes sequence to generate hidden state h.
  3. Output Layer: Predicts next purchase probability (e.g., 80% for headphones).
graph LR
    A["User Orders"] --> B["RNN (LSTM)"]
    B --> C["Hidden State h"]
    C --> D["Output: Probability"]
    D --> E["Recommend: Headphones"]

8. Challenges and Solutions

Challenge Solution
Overfitting Dropout, L2 regularization
Vanishing Gradients ReLU, LSTM, residual connections
Slow Training GPU/TPU acceleration, batch processing
Interpretability SHAP values, LIME for explainability

Example: NEPSE’s stock prediction uses dropout to avoid overfitting on historical data.


9. Exam Tip

  1. Diagrams Are Key:
    • Draw 3-layer FNN, RNN with loops, and backpropagation flow.
    • Label weights, biases, and activation functions.
  2. Math Shortcuts:
    • Memorize sigmoid derivative: .
    • For AND gate, use w₁=w₂=1, b=-1.5 with step activation.
  3. Real-World Links:
    • Tie backpropagation to Khalti’s fraud system (adjusts weights to flag suspicious transactions).
    • Relate LSTM to Pathao’s dynamic pricing (time-series data).
  4. Common Pitfalls:
    • Don’t confuse feedforward (no loops) vs. recurrent (memory).
    • Cross-entropy is for classification; MSE is for regression.
  5. Past Exam Patterns:
    • 3+7 marks: Define ANN (3) + prove a logic statement (7) using resolution (as in past papers).
    • Short notes: Expect Hopefield Network (energy-based learning) and Best-First Search (not ANN, but often paired).

Final Visual Summary:

mindmap
  root((Artificial Neural Networks))
    Biological Inspiration
    Architectures
      Feedforward
      Recurrent
    Activation Functions
      Sigmoid
      ReLU
      Tanh
    Training
      Backpropagation
      Loss Functions
    Applications
      Fraud Detection
      Recommendation Systems
      Time-Series Prediction
    Challenges
      Overfitting
      Vanishing Gradients

Based on the TU BCA syllabus for Artificial Intelligence (CACS410), unit 7.

Discussion

Loading…