Machine LearningUnit 59 min read
Neural Networks & Deep Learning Basics
Unit 5 of Machine Learning: Explores artificial neural networks (ANNs), their architecture, activation functions, training (forward/backpropagation), and deep learning basics, with real-world applications in image recognition, NLP, and recommendation systems.
TAKEAWAYS:
- Neural networks are inspired by biological neurons and process data through interconnected layers of nodes (neurons).
- Activation functions introduce non-linearity, enabling networks to learn complex patterns (e.g., ReLU, sigmoid).
- Training involves forward propagation (predictions) and backpropagation (error correction via gradients).
- Deep learning uses multiple hidden layers to model hierarchical features (e.g., CNNs for images, RNNs for sequences).
- Overfitting is mitigated by techniques like dropout, regularization, and early stopping.
- Real-world applications include eSewa fraud detection, Daraz product recommendations, and Pathao route optimization.
1. Introduction to Neural Networks
Neural networks (NNs) are computational models inspired by the human brain’s neural structure. They consist of interconnected layers of artificial neurons that transform input data to produce an output.
1.1 Biological vs. Artificial Neurons
A biological neuron (left) receives signals via dendrites, processes them in the cell body, and sends output via the axon. An artificial neuron (right) takes weighted inputs, applies an activation function, and outputs a value.
- Biological neuron: Processes signals via electrochemical impulses.
- Artificial neuron: Processes data via mathematical operations (weighted sum + activation).
1.2 Architecture of a Neural Network
A typical NN has:
- Input layer: Receives raw data (e.g., pixel values, text embeddings).
- Hidden layers: Perform feature extraction (e.g., edge detection in images).
- Output layer: Produces predictions (e.g., class labels, regression values).
flowchart TD
A["Input Layer"] --> B["Hidden Layer 1"]
B --> C["Hidden Layer 2"]
C --> D["Output Layer"]1.3 Why Neural Networks?
- Non-linearity: Unlike linear models, NNs can model complex relationships.
- Feature learning: Hidden layers automatically extract meaningful patterns.
- Scalability: Deep networks (many layers) excel at high-dimensional data (e.g., images, audio).
2. Key Components of a Neural Network
2.1 Neurons and Weights
Each neuron computes: where:
- = input features,
- = weights (learned during training),
- = bias term.
2.2 Activation Functions
Introduce non-linearity to enable complex decision boundaries. Common functions:
| Function | Formula | Output Range | Use Case |
|---|---|---|---|
| Sigmoid | (0, 1) | Binary classification (e.g., spam detection) | |
| ReLU | [0, ∞) | Deep learning (faster training) | |
| Tanh | (-1, 1) | Hidden layers (centered around 0) |
Plot of ReLU (left) and Sigmoid (right) activation functions. ReLU avoids vanishing gradients, while sigmoid squashes outputs to (0,1).
2.3 Loss Functions
Measure prediction error. Common types:
- Mean Squared Error (MSE): For regression (e.g., house price prediction).
- Cross-Entropy: For classification (e.g., image categorization).
- Binary Cross-Entropy: For binary classification (e.g., fraud detection).
3. Training Neural Networks
3.1 Forward Propagation
Input data flows through the network to produce predictions:
- Input → Hidden Layer 1 → Hidden Layer 2 → Output.
- Each layer applies weights, bias, and activation.
flowchart TD
A["Input: [x₁, x₂]"] --> B["Layer 1: z₁ = w₁x₁ + w₂x₂ + b₁"]
B --> C["Activation: a₁ = ReLU(z₁)"]
C --> D["Layer 2: z₂ = w₃a₁ + b₂"]
D --> E["Output: ŷ = σ(z₂)"]3.2 Backpropagation
Corrects weights using gradients (chain rule):
- Compute loss (e.g., MSE) between prediction and true label.
- Propagate error backward to adjust weights via gradient descent: where = learning rate.
3.3 Example: Training a Simple NN
Problem: Predict if a customer will default on a loan (binary classification). Data: Features = [income, credit_score, loan_amount], Label = default (0/1).
# Pseudocode for one training step
import numpy as np
# Sample data
X = np.array([[50000, 700, 10000], [30000, 500, 5000]])
y = np.array([0, 1]) # Default (0=no, 1=yes)
# Initialize weights and bias
weights = np.random.randn(3, 1)
bias = 0
# Forward pass
z = np.dot(X, weights) + bias
a = 1 / (1 + np.exp(-z)) # Sigmoid activation
# Loss (Binary Cross-Entropy)
loss = -np.mean(y * np.log(a) + (1 - y) * np.log(1 - a))
# Backpropagation (simplified)
gradient = (a - y) / X.shape[0]
weights -= 0.01 * gradient # Update weights
4. Deep Learning Basics
Deep learning uses multiple hidden layers to model hierarchical features.
4.1 Types of Deep Networks
| Network Type | Architecture | Example Use Case |
|---|---|---|
| CNN (Convolutional) | Convolution + Pooling layers | Image recognition (e.g., Daraz product tags) |
| RNN (Recurrent) | Loops for sequential data | Sentiment analysis (e.g., WhatsApp chatbots) |
| Transformer | Attention mechanisms | Machine translation (e.g., Google Translate) |
4.2 Convolutional Neural Networks (CNNs)
Specialized for grid-like data (e.g., images). Key layers:
- Convolutional Layer: Detects edges/patterns via filters.
- Pooling Layer: Reduces spatial dimensions (e.g., max pooling).
flowchart TD
A["Input Image"] --> B["Conv Layer 1: 3×3 Filter (Edge Detection)"]
B --> C["ReLU Activation"]
C --> D["Max Pooling (2×2)"]
D --> E["Conv Layer 2: 5×5 Filter (Textures)"]
E --> F["Flatten → Dense"]
F --> G["Output: [P₁, P₂, P₃] (Softmax)"]Example: eSewa uses CNNs to detect fraudulent transactions by analyzing transaction patterns in images of receipts.
5. Challenges and Solutions
5.1 Overfitting
Occurs when the model memorizes training data instead of generalizing.
| Technique | Description |
|---|---|
| Dropout | Randomly deactivates neurons during training. |
| Regularization | Adds penalty to large weights (e.g., L2). |
| Early Stopping | Halts training when validation error increases. |
5.2 Vanishing/Exploding Gradients
- Vanishing: Gradients become too small (e.g., with sigmoid). Fix: Use ReLU or batch normalization.
- Exploding: Gradients grow uncontrollably. Fix: Gradient clipping or smaller learning rates.
6. Real-World Applications
6.1 eSewa: Fraud Detection
- Idea: Uses a shallow NN to classify transactions as fraudulent/legitimate.
- How: Inputs = transaction amount, time, merchant history → Output = fraud probability.
- Worked Example:
- Transaction: ₹5000 at 3 AM to a new merchant → High fraud risk (output ≈ 0.9).
- Transaction: ₹100 at 10 AM to a trusted store → Low risk (output ≈ 0.1).
6.2 Daraz: Product Recommendations
- Idea: Uses a collaborative filtering NN to suggest products.
- How: Inputs = user purchase history, item features → Output = recommendation scores.
- Worked Example:
- User buys shoes → NN suggests socks, belts (based on co-purchase patterns).
6.3 Pathao: Route Optimization
- Idea: Uses a reinforcement learning NN to optimize driver routes.
- How: Inputs = traffic data, rider locations → Output = fastest path.
- Worked Example:
- Peak hour → NN suggests detour to avoid congestion (reduces travel time by 20%).
7. Exam Tips
- Diagrams: Always draw the architecture of a NN (input → hidden → output layers) with labels for weights, biases, and activation functions.
- Math: Know the formulas for forward propagation, backpropagation, and loss functions (MSE, cross-entropy).
- Applications: Link concepts to real-world examples (e.g., CNNs for images, RNNs for sequences).
- Trade-offs: Discuss bias-variance trade-off and how deep networks reduce bias but risk overfitting.
- Code Snippets: If asked to explain training, use pseudocode (like the loan default example above).
- Visuals: For deep learning, sketch a CNN’s convolutional layers or a transformer’s attention mechanism.
Based on the TU BCA syllabus for Machine Learning (CACS486), unit 5.
Discussion
Loading…