Artificial IntelligenceUnit 89 min read
Neural Networks: Layers, Learning & Applications
Unit 8 of Artificial Intelligence explores neural networks—how they mimic biological neurons, process data through layers, and learn via backpropagation, with real-world examples from Nepalese tech (eSewa fraud detection, Ncell chatbots) and global AI (Google Translate, YouTube recommendations).
Neural Networks: The Brain of Modern AI
What is a Neural Network?
A neural network (NN) is a computational model inspired by the human brain’s neural structure. It consists of interconnected nodes (neurons) organized in layers that process input data, learn patterns, and make predictions or decisions. Unlike traditional programming, NNs adapt through training on data rather than following rigid rules.
Biological vs. Artificial Neurons
Structure of a Neural Network
A NN is built in layers:
- Input Layer: Receives raw data (e.g., pixel values in an image).
- Hidden Layers: Perform computations (can be 1–100+ layers in deep learning).
- Output Layer: Produces the final result (e.g., a class label or probability).
Example: Handwritten Digit Recognition (MNIST Dataset)
- Input: 28×28 pixel image (flattened into 784 values).
- Hidden Layers: 2 layers with 128 and 64 neurons (ReLU activation).
- Output: 10 neurons (one for each digit 0–9, softmax activation).
graph LR
A["Input Layer: 784 neurons"] --> B["Hidden Layer 1: 128 neurons<br/>ReLU"]
B --> C["Hidden Layer 2: 64 neurons<br/>ReLU"]
C --> D["Output Layer: 10 neurons<br/>Softmax"]Key Components
1. Activation Functions
Convert weighted sums into outputs. Common types:
| Function | Formula | Use Case | Graph |
|---|---|---|---|
| Sigmoid | Binary classification (0/1) | S-shaped curve (0 to 1) | |
| ReLU | Hidden layers (avoids vanishing gradients) | Linear for , flat at 0 | |
| Tanh | Normalized outputs (-1 to 1) | S-shaped (centered at 0) | |
| Softmax | Multi-class classification (probabilities) | Peaks at highest |
2. Loss Functions
Measure error between predicted and actual outputs. Examples:
- Mean Squared Error (MSE): For regression.
- Cross-Entropy: For classification.
Worked Example: Binary Classification Loss Suppose:
- True label , predicted probability .
- Loss (binary cross-entropy): .
How Neural Networks Learn: Backpropagation
Neural networks learn by adjusting weights to minimize loss. The process:
- Forward Pass: Compute predictions and loss.
- Backward Pass: Calculate gradients of loss w.r.t. weights using the chain rule.
- Update Weights: Use an optimizer (e.g., Stochastic Gradient Descent, SGD) to adjust weights.
Example: Training a NN to Predict House Prices
Data: 3 features (size, bedrooms, location score) → predict price. Steps:
- Initialize random weights and bias .
- Forward pass: .
- Compute loss (MSE) vs. true price.
- Backpropagate to find .
- Update weights: (where = learning rate).
Types of Neural Networks
| Type | Structure | Use Case | Example (Nepal/Global) |
|---|---|---|---|
| Feedforward NN | Input → Hidden → Output (no loops) | Classification/regression | eSewa fraud detection (binary classification) |
| Convolutional NN (CNN) | Layers: Convolution → Pooling → Fully Connected | Image/video processing | Ncell face unlock (image recognition) |
| Recurrent NN (RNN) | Loops (memory of past inputs) | Time-series, text | Google Translate (sequence prediction) |
| Long Short-Term Memory (LSTM) | Special RNN for long sequences | Stock price prediction, chatbots | Daraz customer review sentiment analysis |
| Transformer | Self-attention mechanism | NLP (language models) | YouTube comment moderation (NLP) |
Real-World Applications in Nepal and Globally
1. eSewa Fraud Detection
- Idea Used: Feedforward NN with binary classification.
- How: Trained on transaction data (amount, time, location) to flag fraudulent payments (e.g., unusual high-value transfers at odd hours).
- Output: Probability score (e.g., 0.92 = "high risk").
2. Ncell Chatbot (Nepali Language Processing)
- Idea Used: RNN/LSTM for sequence processing.
- How: Processes customer queries (e.g., "My bill is pending") as sequences of words, predicts intent, and generates responses.
- Example Trace: Input: ["My", "bill", "is", "pending"] → Encoded as vectors → LSTM hidden states → Output: "Please check your payment method."
3. Daraz Recommendation System
- Idea Used: Collaborative Filtering + NN Embeddings.
- How: NN learns user-item interaction patterns (e.g., users who bought X also buy Y) to recommend products.
- Example: If you view shoes, the NN predicts "Buy these sandals" based on similar users' behavior.
4. NTC Traffic Route Optimization
- Idea Used: Reinforcement Learning (RL).
- How: An RL agent (NN) learns to adjust traffic light timings in Kathmandu to minimize congestion by simulating thousands of scenarios.
- Worked Example:
- State: Traffic density at 3 intersections.
- Action: Green light duration for each road.
- Reward: Negative delay time for vehicles.
- Policy: NN outputs optimal timing (e.g., "Road A: 45s green, Road B: 30s green").
Advantages and Limitations
| Advantages | Limitations |
|---|---|
| Handles complex, non-linear patterns | Requires large datasets |
| Adapts to new data (online learning) | Black-box nature (hard to interpret) |
| Scalable (deep learning for big data) | Computationally expensive (GPU/TPU needed) |
| Used in unstructured data (images, text) | Overfitting if not regularized |
Worked Example: Predicting Loan Defaults (Nepali Banks)
Problem: A bank wants to predict if a customer will default on a loan using NN. Data: 5 features (income, credit score, loan amount, employment status, age). Steps:
- Preprocess:
- Normalize income (scale to 0–1).
- Encode employment status (e.g., "employed" → 1, "unemployed" → 0).
- Model Architecture:
Model: Input Layer: 5 neurons Hidden Layer 1: 16 neurons (ReLU) Hidden Layer 2: 8 neurons (ReLU) Output Layer: 1 neuron (Sigmoid for probability) - Training:
- Loss: Binary cross-entropy.
- Optimizer: Adam (adaptive learning rate).
- Batch size: 32, epochs: 50.
- Prediction: Input: [income=0.8, credit_score=0.6, loan_amount=0.5, employed=1, age=0.4] Output: 0.2 → "Low risk of default (78% confidence)."
Exam Tip
- Diagrams Are Key: Always draw the NN architecture (layers, activations) for questions on design. Label inputs, hidden layers, and outputs clearly.
- Math Shortcuts: For backpropagation, focus on the chain rule and gradient flow. Memorize:
- if , else 0.
- .
- Real-World Links: Examiners love connections to local tech. For example:
- "How would you design a NN for NEPSE stock prediction?" → Use LSTM for time-series data.
- "Explain how Pathao uses CNNs." → Object detection for rider locations in satellite images.
- Common Pitfalls:
- Forgetting to normalize input data (e.g., pixel values 0–255 → 0–1).
- Confusing activation functions (e.g., using ReLU in the output layer for classification).
- Practical Question Tricks:
- If asked to "explain backpropagation," trace one weight’s gradient through the network.
- For "advantages of CNNs," mention parameter sharing (same filter applied to all regions of an image).
Final Note: Neural networks are the backbone of modern AI. Master the architecture, math of learning, and real-world mapping to ace this unit. Practice coding a simple NN (even in Python’s numpy) to solidify concepts!
Based on the TU BIM syllabus for Artificial Intelligence (IT228), unit 8.
Discussion
Loading…