Artificial IntelligenceUnit 78 min read
Artificial Neural Networks: Models, Learning & Applications
Unit 7 of Artificial Intelligence explores how artificial neural networks (ANNs) mimic biological neurons to learn patterns, covering architectures (feedforward/recurrent), activation functions, backpropagation, and real-world implementations in image recognition, finance, and recommendation systems.
TAKEAWAYS:
- Biological Inspiration: ANNs replicate the human brain’s neuron structure (dendrites, axon, synapses) to process information in layers.
- Key Architectures: Feedforward networks (static) vs. recurrent networks (memory) solve different problems (e.g., image classification vs. time-series prediction).
- Learning Mechanism: Backpropagation adjusts weights using gradient descent to minimize error, guided by activation functions (sigmoid, ReLU).
- Real-World Impact: Used in Khalti’s fraud detection (pattern recognition), Daraz’s recommendation systems (collaborative filtering), and NTC’s network traffic prediction (time-series analysis).
- Math Behind Models: Loss functions (MSE, cross-entropy) and optimization (SGD, Adam) drive training efficiency.
- Limitations: Require large datasets, struggle with interpretability, and may overfit without regularization.
1. Biological Neurons vs. Artificial Neurons
ANNs are inspired by the human brain’s neurons, which transmit signals via electrical impulses. A single artificial neuron (perceptron) mimics this process:
graph LR
A["Inputs (x₁, x₂, ..., xₙ)"] --> B["Weights (w₁, w₂, ..., wₙ)"]
B --> C["Summation: z = Σ(wᵢxᵢ) + b"]
C --> D["Activation: a = f(z)"]
D --> E["Output"]- Inputs (xᵢ): Features (e.g., pixel values in an image).
- Weights (wᵢ): Learned parameters adjusted during training.
- Bias (b): Shifts the activation function.
- Activation (f): Introduces non-linearity (e.g., sigmoid, ReLU).
2. Neural Network Architectures
A. Feedforward Neural Networks (FNNs)
- Structure: Layers connected unidirectionally (input → hidden → output).
- Use Case: Static data (e.g., classifying handwritten digits in eSewa’s OTP verification).
- Example: A 3-layer FNN for AND gate (weights:
w₁=1, w₂=1, b=-1.5; activation: step function).
graph TD
A["Input Layer"] -->|"x₁, x₂"| B["Hidden Layer (1 neuron)"]
B -->|"a = step(w₁x₁ + w₂x₂ + b)"| C["Output Layer"]
C -->|"Output = 1 if a ≥ 0.5"| D["AND Gate"]Worked Example:
For inputs (1, 1):
z = 1*1 + 1*1 - 1.5 = 0.5 → a = step(0.5) = 1 (correct AND output).
B. Recurrent Neural Networks (RNNs)
- Structure: Loops enable memory (e.g., processing sequences like Ncell’s call logs).
- Variants: LSTM (long-term memory) for time-series (e.g., stock prices in NEPSE).
- Example: Predicting the next word in a sentence (language models).
graph LR
A["Input t"] --> B["RNN Cell"]
B -->|"Hidden State hₜ"| C["Output t"]
C -->|"Feedback"| BC. Comparison Table
| Feature | Feedforward (FNN) | Recurrent (RNN/LSTM) |
|---|---|---|
| Data Type | Static (images, tabular) | Sequential (text, time-series) |
| Memory | None | Yes (hidden state) |
| Training Speed | Faster | Slower (vanishing gradients) |
| Example Use | Khalti’s fraud detection | Pathao’s route optimization |
3. Activation Functions
Non-linearities enable complex mappings. Common functions:
- Sigmoid: Outputs between 0 and 1 (used in binary classification). Graph:
- ReLU: Faster training (avoids vanishing gradients).
- Tanh: Zero-centered output (better for hidden layers).
Real-World Tie-In:
- YouTube’s recommendation system uses sigmoid to predict click probability (0 to 1).
4. Training Neural Networks: Backpropagation
Goal: Minimize error (e.g., mean squared error, MSE) between predicted and actual outputs. Steps:
- Forward Pass: Compute output and error.
- Backward Pass: Propagate error to adjust weights using the chain rule.
- Update Weights: Gradient descent: where = learning rate.
Worked Example: Train a neuron to predict y = 2x + 1 (linear regression).
- Initial weights:
w = 0.5, b = 0. - Input:
x = 1→ Target:y = 3. - Forward Pass:
z = 0.5*1 + 0 = 0.5→a = 0.5(sigmoid) → Error:(3 - 0.5)² = 6.25. - Backward Pass:
Update:
w = 0.5 - 0.1*0.625 = 0.4375.
5. Loss Functions
Measure prediction error:
| Function | Formula | Use Case |
|---|---|---|
| MSE | Regression (e.g., house pricing) | |
| Cross-Entropy | Classification (e.g., spam detection) |
Example: NTC’s network latency prediction uses MSE to optimize weights for traffic routing.
6. Optimization Algorithms
- Gradient Descent (GD): Slow for large datasets.
- Stochastic GD (SGD): Updates weights per sample (faster but noisy).
- Adam: Combines momentum and adaptive learning rates (used in Google’s BERT).
Graph:
7. Real-World Applications in Nepal
| Company/App | ANN Application | Key Idea Used |
|---|---|---|
| Khalti | Fraud detection (credit card transactions) | FNN + Sigmoid for anomaly scoring |
| Daraz | Product recommendations | Collaborative filtering (RNN/LSTM) |
| NTC | Network traffic prediction | LSTM for time-series forecasting |
| Ncell | Call detail record analysis | Autoencoders for compression |
| eSewa | OTP verification (image-based) | CNN for digit recognition |
Worked Example: Daraz’s Recommendation System
- Input: User’s past orders (e.g.,
[laptop, phone, charger]). - RNN Layer: Processes sequence to generate hidden state
h. - Output Layer: Predicts next purchase probability (e.g.,
80% for headphones).
graph LR
A["User Orders"] --> B["RNN (LSTM)"]
B --> C["Hidden State h"]
C --> D["Output: Probability"]
D --> E["Recommend: Headphones"]8. Challenges and Solutions
| Challenge | Solution |
|---|---|
| Overfitting | Dropout, L2 regularization |
| Vanishing Gradients | ReLU, LSTM, residual connections |
| Slow Training | GPU/TPU acceleration, batch processing |
| Interpretability | SHAP values, LIME for explainability |
Example: NEPSE’s stock prediction uses dropout to avoid overfitting on historical data.
9. Exam Tip
- Diagrams Are Key:
- Draw 3-layer FNN, RNN with loops, and backpropagation flow.
- Label weights, biases, and activation functions.
- Math Shortcuts:
- Memorize sigmoid derivative: .
- For AND gate, use
w₁=w₂=1, b=-1.5with step activation.
- Real-World Links:
- Tie backpropagation to Khalti’s fraud system (adjusts weights to flag suspicious transactions).
- Relate LSTM to Pathao’s dynamic pricing (time-series data).
- Common Pitfalls:
- Don’t confuse feedforward (no loops) vs. recurrent (memory).
- Cross-entropy is for classification; MSE is for regression.
- Past Exam Patterns:
- 3+7 marks: Define ANN (3) + prove a logic statement (7) using resolution (as in past papers).
- Short notes: Expect Hopefield Network (energy-based learning) and Best-First Search (not ANN, but often paired).
Final Visual Summary:
mindmap
root((Artificial Neural Networks))
Biological Inspiration
Architectures
Feedforward
Recurrent
Activation Functions
Sigmoid
ReLU
Tanh
Training
Backpropagation
Loss Functions
Applications
Fraud Detection
Recommendation Systems
Time-Series Prediction
Challenges
Overfitting
Vanishing GradientsBased on the TU BCA syllabus for Artificial Intelligence (CACS410), unit 7.
Discussion
Loading…