Artificial IntelligenceUnit 88 min read

Neural Networks – Foundations, Training, and Applications

Unit 8 of Artificial Intelligence: introduces neural network concepts, architectures, training algorithms, and real‑world applications.

Key points

  • Neural networks model complex functions using layers of weighted neurons and activation functions.
  • Training relies on gradient descent and back‑propagation to minimise loss functions.
  • Regularisation techniques (dropout, L1/L2, early stopping) combat over‑fitting.
  • Convolutional, recurrent, auto‑encoder and GAN architectures extend basic feed‑forward nets for vision, sequence and generative tasks.
  • Neural nets power everyday services such as fraud detection in e‑wallets, image search in Google, and voice assistants in smartphones.

Introduction

Artificial Neural Networks (ANNs) are computational models inspired by the biological brain. A neuron receives inputs , multiplies them by weights , adds a bias , and applies a non‑linear activation function to produce an output :

A network is a directed acyclic graph of such neurons organized in layers. The simplest form, a feed‑forward neural network (FNN), propagates information from input to output without cycles.

Basic Building Blocks

Neuron and Activation Functions

Function Formula Domain Typical Use
Linear Output layer for regression
Sigmoid Binary classification, hidden layers (historical)
Tanh Hidden layers (zero‑centered)
ReLU Hidden layers (efficient)
Leaky ReLU if else Mitigates dying ReLU
Softmax Multi‑class output
-5-4-3-2-112345-6-4-2246xySigmoid (σ(x))Linear (Identity)
Common activation functions: Sigmoid (saturates at extremes), ReLU (linear for x > 0), Linear (no transformation).

Loss Functions

Task Loss Formula
Regression Mean Squared Error (MSE)
Binary Classification Binary Cross‑Entropy
Multi‑class Categorical Cross‑Entropy

Training a Neural Network

Gradient Descent

Weights are updated iteratively:

-5-4-3-2-112345-6-4-2246xyLoss Surface (J(θ))θ (Parameter)Minimum (Optimal θ)Initial θ
Gradient descent steps: moving from initial θ toward the loss minimum.

where is the learning rate and the loss.

Back‑Propagation

The chain rule propagates gradients from output to input layers. For a single hidden neuron with output and output neuron with :

  1. Compute error at output: .
  2. Propagate to hidden: .
  3. Update weights:

Worked Example

Consider a network with one input , one hidden neuron, one output neuron.
Weights: , ; , .
Activation: ReLU. Target .

  1. Forward pass:
    → .
    → .
  2. Loss (MSE): .
  3. Gradients:
    .
    .
  4. Weight updates ():
    .
    .

After one iteration, weights have moved closer to reducing the error.

Regularisation and Over‑fitting

Technique Idea Effect
L1 / L2 Penalise large weights Encourages sparsity / smoothness
Dropout Randomly deactivate neurons during training Reduces co‑adaptation
Early Stopping Stop training when validation loss stops improving Prevents memorising training data
Data Augmentation Generate synthetic samples Increases effective dataset size

Comparison Table

Model Training Time Generalisation Typical Use
Perceptron Very fast Poor for non‑linear Simple linearly separable tasks
MLP (2‑layer) Moderate Good with regularisation Image classification, regression
CNN Longer (convolutions) Excellent for images Object detection, segmentation
RNN / LSTM Long (sequences) Good for temporal data Speech, language modelling
GAN Very long Generates realistic samples Image synthesis, data augmentation

Advanced Architectures

Convolutional Neural Networks (CNNs)

CNNs replace dense layers with convolutional filters that exploit spatial locality. A typical CNN pipeline:

Input ImageConv1Pool1Conv2Pool2FlattenFCSoftmaxOutput
CNN pipeline showing convolutional layers (Conv), pooling (Max-Pool), flattening, and classification layers (FC + Softmax).

Recurrent Neural Networks (RNNs)

RNNs maintain a hidden state that captures sequence context. LSTM and GRU variants mitigate vanishing gradients.

Autoencoders

Encoder‑decoder pairs learn compact representations. Useful for dimensionality reduction and anomaly detection.

Generative Adversarial Networks (GANs)

Two networks (generator & discriminator) compete, producing realistic synthetic data.

Applications in Nepal and Beyond

Product Neural Network Idea How It Is Used
eSewa Anomaly detection via MLP on transaction features Flags suspicious payments in real time
Google Search CNNs for image indexing + RNNs for query understanding Provides visual search results
WhatsApp RNN (LSTM) for next‑word prediction in typing Auto‑suggestions in chat
Pathao CNN for driver‑vehicle detection in video streams Enhances safety monitoring
NEPSE Time‑series forecasting with RNN (LSTM) Predicts stock price trends

In the real world

  • eSewa employs a shallow MLP trained on historical transaction data to compute a fraud risk score. The network outputs a probability between 0 and 1; transactions above a threshold trigger manual review.
  • Google Photos uses a deep CNN (Inception‑V3) to extract feature vectors from images, enabling quick similarity search.
  • WhatsApp’s on‑device language model is a small LSTM that predicts the next word, reducing typing effort for Nepali users.

Advantages and Disadvantages

Advantage Disadvantage
Learns complex non‑linear mappings Requires large labeled data
Flexible architecture (CNN, RNN, etc.) Training can be computationally expensive
Parallelizable on GPUs Hyper‑parameter tuning is non‑trivial
Continual learning possible Prone to over‑fitting without regularisation

Real‑World Images

NVIDIA Tesla GPUHigh‑performance GPU used for training deep neural networks (Image: Mahogny, Public domain, via Wikimedia Commons) Smartphone cameraInput source for mobile CNN applications (Image: LR.127, CC BY 4.0, via Wikimedia Commons) CCTV cameraVideo feed processed by CNNs for surveillance (Image: Tamasflex, CC BY-SA 3.0, via Wikimedia Commons)

Exam tip

  • Diagram questions: Be ready to draw a 3‑layer MLP, label inputs, weights, biases, activations, and loss.
  • Back‑prop calculations: Practice computing values for a small network; remember the chain rule and activation derivatives.
  • Comparisons: Know the key differences between perceptron, MLP, CNN, RNN, and GAN in terms of architecture, training, and use‑cases.
  • Applications: Relate each architecture to a real product (e.g., CNN → image search, RNN → language model).

Good luck!

Based on the TU BITM syllabus for Artificial Intelligence (IT228), unit 8.

Discussion

Loading…