CACS486 Machine Learning

Machine LearningUnit 59 min read

Neural Networks & Deep Learning Basics

Unit 5 of Machine Learning: Explores artificial neural networks (ANNs), their architecture, activation functions, training (forward/backpropagation), and deep learning basics, with real-world applications in image recognition, NLP, and recommendation systems.

TAKEAWAYS:

  • Neural networks are inspired by biological neurons and process data through interconnected layers of nodes (neurons).
  • Activation functions introduce non-linearity, enabling networks to learn complex patterns (e.g., ReLU, sigmoid).
  • Training involves forward propagation (predictions) and backpropagation (error correction via gradients).
  • Deep learning uses multiple hidden layers to model hierarchical features (e.g., CNNs for images, RNNs for sequences).
  • Overfitting is mitigated by techniques like dropout, regularization, and early stopping.
  • Real-world applications include eSewa fraud detection, Daraz product recommendations, and Pathao route optimization.

1. Introduction to Neural Networks

Neural networks (NNs) are computational models inspired by the human brain’s neural structure. They consist of interconnected layers of artificial neurons that transform input data to produce an output.

1.1 Biological vs. Artificial Neurons

A biological neuron (left) receives signals via dendrites, processes them in the cell body, and sends output via the axon. An artificial neuron (right) takes weighted inputs, applies an activation function, and outputs a value.

  • Biological neuron: Processes signals via electrochemical impulses.
  • Artificial neuron: Processes data via mathematical operations (weighted sum + activation).
1: Dendrites → Inputs 2: Axon → Output 3: Synapse → Weights Biological NeuronArtificial Neuron
Comparison of biological and artificial neuron structures and functions.

1.2 Architecture of a Neural Network

A typical NN has:

  • Input layer: Receives raw data (e.g., pixel values, text embeddings).
  • Hidden layers: Perform feature extraction (e.g., edge detection in images).
  • Output layer: Produces predictions (e.g., class labels, regression values).
flowchart TD
    A["Input Layer"] --> B["Hidden Layer 1"]
    B --> C["Hidden Layer 2"]
    C --> D["Output Layer"]

1.3 Why Neural Networks?

  • Non-linearity: Unlike linear models, NNs can model complex relationships.
  • Feature learning: Hidden layers automatically extract meaningful patterns.
  • Scalability: Deep networks (many layers) excel at high-dimensional data (e.g., images, audio).

2. Key Components of a Neural Network

2.1 Neurons and Weights

Each neuron computes: where:

  • = input features,
  • = weights (learned during training),
  • = bias term.

2.2 Activation Functions

Introduce non-linearity to enable complex decision boundaries. Common functions:

-5-4-3-2-112345-6-4-2246xyLinear (f(x) = x)Sigmoid (f(x) = σ(x))x=0ReLUSigmoid
Function Formula Output Range Use Case
Sigmoid (0, 1) Binary classification (e.g., spam detection)
ReLU [0, ∞) Deep learning (faster training)
Tanh (-1, 1) Hidden layers (centered around 0)

Plot of ReLU (left) and Sigmoid (right) activation functions. ReLU avoids vanishing gradients, while sigmoid squashes outputs to (0,1).

2.3 Loss Functions

Measure prediction error. Common types:

  • Mean Squared Error (MSE): For regression (e.g., house price prediction).
  • Cross-Entropy: For classification (e.g., image categorization).
  • Binary Cross-Entropy: For binary classification (e.g., fraud detection).

3. Training Neural Networks

3.1 Forward Propagation

Input data flows through the network to produce predictions:

  1. Input → Hidden Layer 1 → Hidden Layer 2 → Output.
  2. Each layer applies weights, bias, and activation.
flowchart TD
    A["Input: [x₁, x₂]"] --> B["Layer 1: z₁ = w₁x₁ + w₂x₂ + b₁"]
    B --> C["Activation: a₁ = ReLU(z₁)"]
    C --> D["Layer 2: z₂ = w₃a₁ + b₂"]
    D --> E["Output: ŷ = σ(z₂)"]

3.2 Backpropagation

Corrects weights using gradients (chain rule):

  1. Compute loss (e.g., MSE) between prediction and true label.
  2. Propagate error backward to adjust weights via gradient descent: where = learning rate.

3.3 Example: Training a Simple NN

Problem: Predict if a customer will default on a loan (binary classification). Data: Features = [income, credit_score, loan_amount], Label = default (0/1).

# Pseudocode for one training step
import numpy as np

# Sample data
X = np.array([[50000, 700, 10000], [30000, 500, 5000]])
y = np.array([0, 1])  # Default (0=no, 1=yes)

# Initialize weights and bias
weights = np.random.randn(3, 1)
bias = 0

# Forward pass
z = np.dot(X, weights) + bias
a = 1 / (1 + np.exp(-z))  # Sigmoid activation

# Loss (Binary Cross-Entropy)
loss = -np.mean(y * np.log(a) + (1 - y) * np.log(1 - a))

# Backpropagation (simplified)
gradient = (a - y) / X.shape[0]
weights -= 0.01 * gradient  # Update weights

4. Deep Learning Basics

Deep learning uses multiple hidden layers to model hierarchical features.

4.1 Types of Deep Networks

Network Type Architecture Example Use Case
CNN (Convolutional) Convolution + Pooling layers Image recognition (e.g., Daraz product tags)
RNN (Recurrent) Loops for sequential data Sentiment analysis (e.g., WhatsApp chatbots)
Transformer Attention mechanisms Machine translation (e.g., Google Translate)

4.2 Convolutional Neural Networks (CNNs)

Specialized for grid-like data (e.g., images). Key layers:

  • Convolutional Layer: Detects edges/patterns via filters.
  • Pooling Layer: Reduces spatial dimensions (e.g., max pooling).
flowchart TD
    A["Input Image"] --> B["Conv Layer 1: 3×3 Filter (Edge Detection)"]
    B --> C["ReLU Activation"]
    C --> D["Max Pooling (2×2)"]
    D --> E["Conv Layer 2: 5×5 Filter (Textures)"]
    E --> F["Flatten → Dense"]
    F --> G["Output: [P₁, P₂, P₃] (Softmax)"]

Example: eSewa uses CNNs to detect fraudulent transactions by analyzing transaction patterns in images of receipts.


5. Challenges and Solutions

5.1 Overfitting

Occurs when the model memorizes training data instead of generalizing.

Technique Description
Dropout Randomly deactivates neurons during training.
Regularization Adds penalty to large weights (e.g., L2).
Early Stopping Halts training when validation error increases.

5.2 Vanishing/Exploding Gradients

  • Vanishing: Gradients become too small (e.g., with sigmoid). Fix: Use ReLU or batch normalization.
  • Exploding: Gradients grow uncontrollably. Fix: Gradient clipping or smaller learning rates.

6. Real-World Applications

6.1 eSewa: Fraud Detection

  • Idea: Uses a shallow NN to classify transactions as fraudulent/legitimate.
  • How: Inputs = transaction amount, time, merchant history → Output = fraud probability.
  • Worked Example:
    • Transaction: ₹5000 at 3 AM to a new merchant → High fraud risk (output ≈ 0.9).
    • Transaction: ₹100 at 10 AM to a trusted store → Low risk (output ≈ 0.1).

6.2 Daraz: Product Recommendations

  • Idea: Uses a collaborative filtering NN to suggest products.
  • How: Inputs = user purchase history, item features → Output = recommendation scores.
  • Worked Example:
    • User buys shoes → NN suggests socks, belts (based on co-purchase patterns).

6.3 Pathao: Route Optimization

  • Idea: Uses a reinforcement learning NN to optimize driver routes.
  • How: Inputs = traffic data, rider locations → Output = fastest path.
  • Worked Example:
    • Peak hour → NN suggests detour to avoid congestion (reduces travel time by 20%).

7. Exam Tips

  1. Diagrams: Always draw the architecture of a NN (input → hidden → output layers) with labels for weights, biases, and activation functions.
  2. Math: Know the formulas for forward propagation, backpropagation, and loss functions (MSE, cross-entropy).
  3. Applications: Link concepts to real-world examples (e.g., CNNs for images, RNNs for sequences).
  4. Trade-offs: Discuss bias-variance trade-off and how deep networks reduce bias but risk overfitting.
  5. Code Snippets: If asked to explain training, use pseudocode (like the loan default example above).
  6. Visuals: For deep learning, sketch a CNN’s convolutional layers or a transformer’s attention mechanism.

Based on the TU BCA syllabus for Machine Learning (CACS486), unit 5.

Discussion

Loading…