Neural NetworksUnit 19 min read
Neural Networks: Basics, Models & Biological Inspiration
Unit 1 of Neural Networks introduces the foundational concepts of artificial neural networks (ANNs), their biological inspiration, key components (neurons, layers, weights), and simple models like the McCulloch-Pitts neuron. It contrasts ANNs with traditional computing, explains activation functions, and demonstrates a
Biological Inspiration: The Brain vs. Artificial Neurons
The human brain contains ~86 billion neurons, each connected via ~100 trillion synapses. These biological neurons communicate via electrochemical signals (action potentials) and exhibit plasticity (adapting to new information). Artificial neural networks (ANNs) mimic this structure but simplify it:
Key Differences:
| Feature | Biological Neuron | Artificial Neuron |
|---|---|---|
| Signal Type | Electrochemical | Numerical (floating-point) |
| Learning | Hebbian, dopamine-based | Backpropagation, gradient descent |
| Plasticity | Lifelong adaptation | Fixed architecture (unless dynamic) |
| Energy Use | ~20W (whole brain) | ~100W (GPU cluster) |
Artificial Neuron: The McCulloch-Pitts Model (1943)
The simplest ANN unit is the McCulloch-Pitts neuron, a binary classifier:
- Inputs: (e.g., pixel intensities in an image).
- Weights: (learned during training).
- Bias: (threshold adjustment).
- Activation: Step function , where .
Worked Example: Handwritten Digit Recognition Suppose a neuron classifies whether a pixel is part of a "1" (output=1) or "0" (output=0). For a 3-pixel input:
- Inputs: (gray-scale values).
- Weights: , bias .
- Calculation: Since , output ("1" detected).
Activation Functions: Nonlinearity in ANNs
Linear neurons (no activation) cannot model complex patterns. Activation functions introduce nonlinearity:
- Step Function: Binary threshold (used in McCulloch-Pitts).
- Sigmoid: (outputs between 0 and 1).
- ReLU: (common in deep learning).
- Tanh: (outputs between -1 and 1).
Graph Comparison:
Why ReLU Dominates:
- Computationally efficient (no exponentiation).
- Mitigates vanishing gradients in deep networks.
- Sparse activation (many neurons output 0).
Neural Network Architecture: Layers and Connectivity
ANNs are composed of layers:
- Input Layer: Raw features (e.g., pixel values).
- Hidden Layers: Intermediate computations (can be multiple).
- Output Layer: Final prediction (e.g., class probabilities).
Fully Connected (Dense) Network:
Sparse Connectivity (e.g., CNNs):
- Neurons connect only to local regions (e.g., adjacent pixels).
- Reduces parameters and computational cost.
Training a Neural Network: Supervised Learning Basics
ANNs learn via gradient descent and backpropagation:
- Forward Pass: Compute output for given inputs.
- Loss Calculation: Compare output to true label (e.g., Mean Squared Error for regression).
- Backward Pass: Adjust weights using the chain rule to minimize loss.
Loss Functions:
- MSE: (regression).
- Cross-Entropy: (classification).
Worked Example: Predicting House Prices Suppose a single neuron predicts price from size :
- True price , predicted .
- MSE loss:
- Gradient for weight : .
In the Real World
eSewa (Nepal):
- Idea Used: Multilayer Perceptrons (MLPs) for fraud detection.
- How: ANNs analyze transaction patterns (time, amount, location) to flag suspicious activities (e.g., sudden large transfers). A hidden layer with ReLU activation learns non-linear relationships between features like "user’s usual spending" and "transaction amount."
Pathao (Ride-Hailing):
- Idea Used: Dynamic Routing via Neural Networks.
- How: A deep neural network predicts optimal driver routes in real-time by processing inputs like traffic data (from GPS), driver availability, and historical demand. The network’s output is a probability distribution over possible routes, and the highest-probability path is chosen. For example, during Kathmandu’s evening rush, the ANN might reroute drivers away from Thapathali via a hidden layer that detects congestion patterns.
Ncell’s Customer Churn Prediction:
- Idea Used: Binary Classification with Sigmoid Output.
- How: Ncell uses a neural network to predict whether a subscriber will cancel service. Inputs include call duration, data usage, and customer service interactions. The output layer has a single neuron with a sigmoid activation, outputting a probability (e.g., 0.87) that the customer will churn. If this exceeds a threshold (e.g., 0.8), the system triggers a retention offer.
Exam Tip
- Define Clearly: Start with definitions (e.g., "An artificial neuron is a mathematical model inspired by biological neurons..."). Examiners check if you distinguish between biological neurons and ANNs.
- Draw Diagrams: Sketch a 3-layer ANN (input, hidden, output) and label weights/biases. For partial credit, even a rough diagram with arrows is better than none.
- Worked Examples: Always show step-by-step calculations for neuron outputs or loss functions. For example:
- Given inputs/weights, compute and activation.
- For MSE, write the formula and plug in numbers.
- Compare Activation Functions: In short-answer questions, list one advantage and one disadvantage of ReLU vs. Sigmoid (e.g., ReLU is faster but can cause "dying ReLU" if weights are too large).
- Real-World Links: If asked about applications, mention eSewa (fraud detection), Pathao (routing), or Ncell (churn prediction). Avoid vague answers like "ANNs are used in AI."
- Common Pitfalls:
- Forgetting the bias term in neuron calculations.
- Confusing forward pass (prediction) with backward pass (learning).
- Using the wrong activation for the task (e.g., ReLU for output in classification).
Summary Table: Key Concepts
| Concept | Description | Example Use Case |
|---|---|---|
| McCulloch-Pitts Neuron | Binary threshold unit (step activation). | Early AI, simple classifiers. |
| Sigmoid Activation | Smooth output between 0 and 1. | Probability outputs in classification. |
| ReLU Activation | Linear for positive inputs, else zero. | Deep learning (e.g., CNNs). |
| Loss Function | Measures prediction error (MSE, Cross-Entropy). | Training ANNs for regression/classification. |
| Backpropagation | Algorithm to update weights via gradient descent. | All supervised learning tasks. |
Based on the TU BSc CSIT syllabus for Neural Networks, unit 1.
Discussion
Loading…