Neural NetworksUnit 19 min read

Neural Networks: Basics, Models & Biological Inspiration

Unit 1 of Neural Networks introduces the foundational concepts of artificial neural networks (ANNs), their biological inspiration, key components (neurons, layers, weights), and simple models like the McCulloch-Pitts neuron. It contrasts ANNs with traditional computing, explains activation functions, and demonstrates a

Biological Inspiration: The Brain vs. Artificial Neurons

The human brain contains ~86 billion neurons, each connected via ~100 trillion synapses. These biological neurons communicate via electrochemical signals (action potentials) and exhibit plasticity (adapting to new information). Artificial neural networks (ANNs) mimic this structure but simplify it:

Biological NeuronArtificial Neuron
Comparison of biological neuron (left) and artificial neuron (right) components

Key Differences:

Feature Biological Neuron Artificial Neuron
Signal Type Electrochemical Numerical (floating-point)
Learning Hebbian, dopamine-based Backpropagation, gradient descent
Plasticity Lifelong adaptation Fixed architecture (unless dynamic)
Energy Use ~20W (whole brain) ~100W (GPU cluster)

Artificial Neuron: The McCulloch-Pitts Model (1943)

The simplest ANN unit is the McCulloch-Pitts neuron, a binary classifier:

  1. Inputs: (e.g., pixel intensities in an image).
  2. Weights: (learned during training).
  3. Bias: (threshold adjustment).
  4. Activation: Step function , where .
Input 1Input 2Threshold UnitOutput
McCulloch-Pitts neuron model: binary inputs, threshold, and binary output

Worked Example: Handwritten Digit Recognition Suppose a neuron classifies whether a pixel is part of a "1" (output=1) or "0" (output=0). For a 3-pixel input:

  • Inputs: (gray-scale values).
  • Weights: , bias .
  • Calculation: Since , output ("1" detected).

Activation Functions: Nonlinearity in ANNs

Linear neurons (no activation) cannot model complex patterns. Activation functions introduce nonlinearity:

  • Step Function: Binary threshold (used in McCulloch-Pitts).
  • Sigmoid: (outputs between 0 and 1).
  • ReLU: (common in deep learning).
  • Tanh: (outputs between -1 and 1).

Graph Comparison:

-5-4-3-2-112345-1-0.50.51xySigmoid (σ(z))TanhSigmoidReLUTanh
Graphs of common activation functions (Sigmoid, ReLU, Tanh) with key points labeled

Why ReLU Dominates:

  • Computationally efficient (no exponentiation).
  • Mitigates vanishing gradients in deep networks.
  • Sparse activation (many neurons output 0).

Neural Network Architecture: Layers and Connectivity

ANNs are composed of layers:

  1. Input Layer: Raw features (e.g., pixel values).
  2. Hidden Layers: Intermediate computations (can be multiple).
  3. Output Layer: Final prediction (e.g., class probabilities).

Fully Connected (Dense) Network:

Weights (W₁)Weights (W₂)Input Layer (3 neurons)Hidden Layer (4 neurons)Output Layer (2 neurons)
Fully connected (dense) neural network with 3 input, 4 hidden, and 2 output neurons

Sparse Connectivity (e.g., CNNs):

  • Neurons connect only to local regions (e.g., adjacent pixels).
  • Reduces parameters and computational cost.

Training a Neural Network: Supervised Learning Basics

ANNs learn via gradient descent and backpropagation:

  1. Forward Pass: Compute output for given inputs.
  2. Loss Calculation: Compare output to true label (e.g., Mean Squared Error for regression).
  3. Backward Pass: Adjust weights using the chain rule to minimize loss.
0.511.522.533.544.552468xyLoss Function (MSE)Gradient Descent UpdateMinimum Loss
Loss function minimization via gradient descent (simplified example)

Loss Functions:

  • MSE: (regression).
  • Cross-Entropy: (classification).

Worked Example: Predicting House Prices Suppose a single neuron predicts price from size :

  • True price , predicted .
  • MSE loss:
  • Gradient for weight : .

In the Real World

  1. eSewa (Nepal):

    • Idea Used: Multilayer Perceptrons (MLPs) for fraud detection.
    • How: ANNs analyze transaction patterns (time, amount, location) to flag suspicious activities (e.g., sudden large transfers). A hidden layer with ReLU activation learns non-linear relationships between features like "user’s usual spending" and "transaction amount."
  2. Pathao (Ride-Hailing):

    • Idea Used: Dynamic Routing via Neural Networks.
    • How: A deep neural network predicts optimal driver routes in real-time by processing inputs like traffic data (from GPS), driver availability, and historical demand. The network’s output is a probability distribution over possible routes, and the highest-probability path is chosen. For example, during Kathmandu’s evening rush, the ANN might reroute drivers away from Thapathali via a hidden layer that detects congestion patterns.
  3. Ncell’s Customer Churn Prediction:

    • Idea Used: Binary Classification with Sigmoid Output.
    • How: Ncell uses a neural network to predict whether a subscriber will cancel service. Inputs include call duration, data usage, and customer service interactions. The output layer has a single neuron with a sigmoid activation, outputting a probability (e.g., 0.87) that the customer will churn. If this exceeds a threshold (e.g., 0.8), the system triggers a retention offer.

Exam Tip

  1. Define Clearly: Start with definitions (e.g., "An artificial neuron is a mathematical model inspired by biological neurons..."). Examiners check if you distinguish between biological neurons and ANNs.
  2. Draw Diagrams: Sketch a 3-layer ANN (input, hidden, output) and label weights/biases. For partial credit, even a rough diagram with arrows is better than none.
  3. Worked Examples: Always show step-by-step calculations for neuron outputs or loss functions. For example:
    • Given inputs/weights, compute and activation.
    • For MSE, write the formula and plug in numbers.
  4. Compare Activation Functions: In short-answer questions, list one advantage and one disadvantage of ReLU vs. Sigmoid (e.g., ReLU is faster but can cause "dying ReLU" if weights are too large).
  5. Real-World Links: If asked about applications, mention eSewa (fraud detection), Pathao (routing), or Ncell (churn prediction). Avoid vague answers like "ANNs are used in AI."
  6. Common Pitfalls:
    • Forgetting the bias term in neuron calculations.
    • Confusing forward pass (prediction) with backward pass (learning).
    • Using the wrong activation for the task (e.g., ReLU for output in classification).

Summary Table: Key Concepts

Concept Description Example Use Case
McCulloch-Pitts Neuron Binary threshold unit (step activation). Early AI, simple classifiers.
Sigmoid Activation Smooth output between 0 and 1. Probability outputs in classification.
ReLU Activation Linear for positive inputs, else zero. Deep learning (e.g., CNNs).
Loss Function Measures prediction error (MSE, Cross-Entropy). Training ANNs for regression/classification.
Backpropagation Algorithm to update weights via gradient descent. All supervised learning tasks.

Based on the TU BSc CSIT syllabus for Neural Networks, unit 1.

Discussion

Loading…