Artificial IntelligenceUnit 89 min read

Neural Networks: Layers, Learning & Applications

Unit 8 of Artificial Intelligence explores neural networks—how they mimic biological neurons, process data through layers, and learn via backpropagation, with real-world examples from Nepalese tech (eSewa fraud detection, Ncell chatbots) and global AI (Google Translate, YouTube recommendations).

Neural Networks: The Brain of Modern AI

What is a Neural Network?

A neural network (NN) is a computational model inspired by the human brain’s neural structure. It consists of interconnected nodes (neurons) organized in layers that process input data, learn patterns, and make predictions or decisions. Unlike traditional programming, NNs adapt through training on data rather than following rigid rules.

Biological vs. Artificial Neurons

Inspired byBiological NeuronArtificial Neuron
Comparison of biological and artificial neuron structures (inspired by neural biology)

Structure of a Neural Network

A NN is built in layers:

  1. Input Layer: Receives raw data (e.g., pixel values in an image).
  2. Hidden Layers: Perform computations (can be 1–100+ layers in deep learning).
  3. Output Layer: Produces the final result (e.g., a class label or probability).
Input LayerHidden Layer 1Hidden Layer 2Output Layer
Generic feedforward neural network architecture (3 layers)

Example: Handwritten Digit Recognition (MNIST Dataset)

  • Input: 28×28 pixel image (flattened into 784 values).
  • Hidden Layers: 2 layers with 128 and 64 neurons (ReLU activation).
  • Output: 10 neurons (one for each digit 0–9, softmax activation).
graph LR
    A["Input Layer: 784 neurons"] --> B["Hidden Layer 1: 128 neurons<br/>ReLU"]
    B --> C["Hidden Layer 2: 64 neurons<br/>ReLU"]
    C --> D["Output Layer: 10 neurons<br/>Softmax"]

Key Components

-5-4-3-2-1123450.20.40.60.81xySigmoid (σ(x))
Common activation functions: Sigmoid (S-shaped) vs. ReLU (linear for x > 0)

1. Activation Functions

Convert weighted sums into outputs. Common types:

Function Formula Use Case Graph
Sigmoid Binary classification (0/1) S-shaped curve (0 to 1)
ReLU Hidden layers (avoids vanishing gradients) Linear for , flat at 0
Tanh Normalized outputs (-1 to 1) S-shaped (centered at 0)
Softmax Multi-class classification (probabilities) Peaks at highest

2. Loss Functions

Measure error between predicted and actual outputs. Examples:

  • Mean Squared Error (MSE): For regression.
  • Cross-Entropy: For classification.

Worked Example: Binary Classification Loss Suppose:

  • True label , predicted probability .
  • Loss (binary cross-entropy): .

How Neural Networks Learn: Backpropagation

Neural networks learn by adjusting weights to minimize loss. The process:

  1. Forward Pass: Compute predictions and loss.
  2. Backward Pass: Calculate gradients of loss w.r.t. weights using the chain rule.
  3. Update Weights: Use an optimizer (e.g., Stochastic Gradient Descent, SGD) to adjust weights.
InputNeuron ANeuron BOutput
Backpropagation flow: Error gradients (red arrows) propagate backward through weights

Example: Training a NN to Predict House Prices

Data: 3 features (size, bedrooms, location score) → predict price. Steps:

  1. Initialize random weights and bias .
  2. Forward pass: .
  3. Compute loss (MSE) vs. true price.
  4. Backpropagate to find .
  5. Update weights: (where = learning rate).

Types of Neural Networks

Type Structure Use Case Example (Nepal/Global)
Feedforward NN Input → Hidden → Output (no loops) Classification/regression eSewa fraud detection (binary classification)
Convolutional NN (CNN) Layers: Convolution → Pooling → Fully Connected Image/video processing Ncell face unlock (image recognition)
Recurrent NN (RNN) Loops (memory of past inputs) Time-series, text Google Translate (sequence prediction)
Long Short-Term Memory (LSTM) Special RNN for long sequences Stock price prediction, chatbots Daraz customer review sentiment analysis
Transformer Self-attention mechanism NLP (language models) YouTube comment moderation (NLP)

Real-World Applications in Nepal and Globally

1. eSewa Fraud Detection

  • Idea Used: Feedforward NN with binary classification.
  • How: Trained on transaction data (amount, time, location) to flag fraudulent payments (e.g., unusual high-value transfers at odd hours).
  • Output: Probability score (e.g., 0.92 = "high risk").

2. Ncell Chatbot (Nepali Language Processing)

  • Idea Used: RNN/LSTM for sequence processing.
  • How: Processes customer queries (e.g., "My bill is pending") as sequences of words, predicts intent, and generates responses.
  • Example Trace: Input: ["My", "bill", "is", "pending"] → Encoded as vectors → LSTM hidden states → Output: "Please check your payment method."

3. Daraz Recommendation System

  • Idea Used: Collaborative Filtering + NN Embeddings.
  • How: NN learns user-item interaction patterns (e.g., users who bought X also buy Y) to recommend products.
  • Example: If you view shoes, the NN predicts "Buy these sandals" based on similar users' behavior.

4. NTC Traffic Route Optimization

  • Idea Used: Reinforcement Learning (RL).
  • How: An RL agent (NN) learns to adjust traffic light timings in Kathmandu to minimize congestion by simulating thousands of scenarios.
  • Worked Example:
    • State: Traffic density at 3 intersections.
    • Action: Green light duration for each road.
    • Reward: Negative delay time for vehicles.
    • Policy: NN outputs optimal timing (e.g., "Road A: 45s green, Road B: 30s green").

Advantages and Limitations

Advantages Limitations
Handles complex, non-linear patterns Requires large datasets
Adapts to new data (online learning) Black-box nature (hard to interpret)
Scalable (deep learning for big data) Computationally expensive (GPU/TPU needed)
Used in unstructured data (images, text) Overfitting if not regularized

Worked Example: Predicting Loan Defaults (Nepali Banks)

Problem: A bank wants to predict if a customer will default on a loan using NN. Data: 5 features (income, credit score, loan amount, employment status, age). Steps:

  1. Preprocess:
    • Normalize income (scale to 0–1).
    • Encode employment status (e.g., "employed" → 1, "unemployed" → 0).
  2. Model Architecture:
    Model:
      Input Layer: 5 neurons
      Hidden Layer 1: 16 neurons (ReLU)
      Hidden Layer 2: 8 neurons (ReLU)
      Output Layer: 1 neuron (Sigmoid for probability)
    
  3. Training:
    • Loss: Binary cross-entropy.
    • Optimizer: Adam (adaptive learning rate).
    • Batch size: 32, epochs: 50.
  4. Prediction: Input: [income=0.8, credit_score=0.6, loan_amount=0.5, employed=1, age=0.4] Output: 0.2 → "Low risk of default (78% confidence)."

Exam Tip

  1. Diagrams Are Key: Always draw the NN architecture (layers, activations) for questions on design. Label inputs, hidden layers, and outputs clearly.
  2. Math Shortcuts: For backpropagation, focus on the chain rule and gradient flow. Memorize:
    • if , else 0.
    • .
  3. Real-World Links: Examiners love connections to local tech. For example:
    • "How would you design a NN for NEPSE stock prediction?" → Use LSTM for time-series data.
    • "Explain how Pathao uses CNNs." → Object detection for rider locations in satellite images.
  4. Common Pitfalls:
    • Forgetting to normalize input data (e.g., pixel values 0–255 → 0–1).
    • Confusing activation functions (e.g., using ReLU in the output layer for classification).
  5. Practical Question Tricks:
    • If asked to "explain backpropagation," trace one weight’s gradient through the network.
    • For "advantages of CNNs," mention parameter sharing (same filter applied to all regions of an image).

Final Note: Neural networks are the backbone of modern AI. Master the architecture, math of learning, and real-world mapping to ace this unit. Practice coding a simple NN (even in Python’s numpy) to solidify concepts!

Based on the TU BIM syllabus for Artificial Intelligence (IT228), unit 8.

Discussion

Loading…