Cognitive ScienceUnit 810 min read

Neural Networks & Distributed Info Processing: Models, Learning, and Brain-Like Systems

Unit 8 of Cognitive Science explores how neural networks mimic biological brains to process information in parallel, covering artificial neurons, activation functions, learning rules (backpropagation), deep learning architectures, and distributed representation—with real-world ties to apps like Khalti’s fraud detection

TAKEAWAYS:

  • Artificial neurons are the building blocks of neural networks, combining weighted inputs with activation functions to produce outputs (like a simplified brain cell).
  • Distributed representation encodes information across many neurons (e.g., a concept like "cat" activates multiple neurons differently than "dog").
  • Backpropagation is the core algorithm for training neural networks by adjusting weights to minimize error (used in every modern AI system).
  • Deep learning architectures (CNNs, RNNs) specialize in processing spatial/temporal data (e.g., Khalti’s handwritten signature verification).
  • Biological plausibility compares artificial networks to real brains, highlighting strengths (scalability) and gaps (energy efficiency).
  • Applications range from recommendation systems (Daraz) to medical diagnosis (Nepal’s AI-driven tuberculosis detection).

1. Artificial Neurons: The Basic Unit

Neural networks are inspired by biological neurons but simplified for computation. A single artificial neuron (or perceptron) works as follows:

![biological neuron diagram](/media/20d4e9246491845d5ec0.gif "A real neuron with dendrites, soma, and axon (left) vs. an artificial neuron (right). (Image: Geetika saini, CC BY-SA 4.0, via Wikimedia Commons)")

How an Artificial Neuron Works

  1. Inputs: Receive signals (e.g., pixel intensities in an image).
  2. Weights: Each input has a weight (learned during training).
  3. Bias: A threshold term (like a neuron’s resting potential).
  4. Weighted Sum: Compute .
  5. Activation Function: Apply a nonlinear function (e.g., sigmoid, ReLU) to to produce output .

Example: Classifying a handwritten digit (0–9) from Khalti’s signature verification system.

  • Input: 28×28 pixel image (flattened into 784 values).
  • Weights: Learned during training to detect edges/patterns.
  • Activation: Sigmoid squashes output to [0,1] (probability of being a "5").
graph LR
    A["Inputs (x₁, x₂, ..., xₙ)"] --> B["Weighted Sum (z = Σwᵢxᵢ + b)"]
    B --> C["Activation Function (a = f(z))"]
    C --> D["Output"]

Activation Functions

Function Formula Output Range Use Case
Sigmoid (0, 1) Binary classification (e.g., spam detection).
ReLU [0, ∞) Deep networks (avoids vanishing gradients).
Tanh (-1, 1) Normalized outputs (e.g., Ncell’s signal strength prediction).

Worked Example: Train a neuron to recognize "1" vs. "7" in NEPSE stock charts.

  • Input: Pixel values of a digit image.
  • Weights: Start with random values (e.g., ).
  • Bias: .
  • For input (a "1"): Sigmoid output: (close to 0.5 → uncertain).

2. Learning in Neural Networks: Backpropagation

Neural networks learn by adjusting weights to minimize error. The backpropagation algorithm does this efficiently.

Key Steps

  1. Forward Pass: Compute output for all layers.
  2. Loss Function: Measure error (e.g., mean squared error for regression).
  3. Backward Pass: Propagate error backward to adjust weights using the chain rule.

Loss Functions:

Function Formula Use Case
Mean Squared Error Regression (e.g., predicting Daraz delivery times).
Cross-Entropy Classification (e.g., Khalti transaction fraud).

Worked Example: Train a network to predict Pathao’s driver wait time.

  • Input: [traffic density, time of day, distance].
  • Output: Wait time (continuous value).
  • Loss: MSE between predicted and actual wait times.
  • Backpropagation adjusts weights to reduce error iteratively.
flowchart LR
    A["Input Layer"] --> B["Hidden Layer 1"]
    B --> C["Hidden Layer 2"]
    C --> D["Output Layer"]
    D --> E["Loss Function"]
    E --> F["Backpropagation"]
    F -->|"Adjust"| B
    F -->|"Adjust"| C

3. Distributed Representation

Unlike symbolic AI (where "cat" is a single token), neural networks use distributed representation: a concept is encoded across many neurons.

Example: Word embeddings in Ncell’s chatbot.

  • The word "call" might activate neurons for: [0.3, -0.1, 0.8, ...] (distributed across 100s of neurons).
  • Similar words (e.g., "message") have similar activation patterns.

Advantages:

  • Robust to noise (e.g., handwritten "5" vs. printed "5").
  • Captures semantic relationships (e.g., "king" – "man" + "woman" ≈ "queen").

Comparison: Local vs. Distributed Representation

Feature Local (Symbolic) Distributed (Neural)
Representation Single neuron/bit Many neurons
Fault Tolerance Fragile (one error breaks it) Robust (damage spreads)
Learning Rule-based Data-driven
Example "cat" = 1, "dog" = 0 "cat" = [0.2, -0.5, 0.8,...]

4. Deep Learning Architectures

Specialized networks for specific tasks:

A. Convolutional Neural Networks (CNNs)

  • Use Case: Image/pattern recognition (e.g., Khalti’s signature verification).
  • Key Idea: Use filters (kernels) to detect edges, textures.
  • Example: Detecting potholes in Kathmandu traffic routes for Pathao drivers.
graph LR
    A["Input Image"] --> B["Convolution Layer 1"]
    B --> C["ReLU"]
    C --> D["Pooling"]
    D --> E["Convolution Layer 2"]
    E --> F["Fully Connected"]
    F --> G["Output"]

B. Recurrent Neural Networks (RNNs)

  • Use Case: Sequential data (e.g., Ncell’s call duration prediction).
  • Key Idea: Memory of past inputs via hidden states.
  • Problem: Vanishing gradients → solved by LSTMs/GRUs.

C. Transformers

  • Use Case: Natural language (e.g., NEPSE stock analysis chatbots).
  • Key Idea: Self-attention mechanisms to weigh input importance.

5. Biological Plausibility vs. Artificial Networks

Feature Biological Brain Artificial Neural Networks
Energy Use ~20W (highly efficient) ~1000W (data centers)
Learning Lifelong, unsupervised Supervised, batch-based
Plasticity High (neurogenesis) Low (fixed architecture)
Parallelism Massive (100B neurons) Limited by hardware

Real-World Gap: Artificial networks lack homeostasis (self-regulation) seen in brains.


6. Applications in Nepal

Company/App Neural Network Use Case Example
Khalti Fraud detection (anomaly detection in transactions) Flags unusual spending patterns.
Ncell Predictive maintenance (antenna failures) Reduces downtime by 30%.
Daraz Recommendation systems (collaborative filtering) Suggests products based on browsing.
NTC Network traffic prediction Optimizes bandwidth allocation.
Nepal Rastra Bank Loan default risk assessment Uses credit history + neural nets.

Worked Example: Khalti’s Fraud Detection

  1. Input: Transaction amount, time, location, user history.
  2. Hidden Layers: Detect patterns (e.g., sudden large transfer to unknown account).
  3. Output: Fraud probability score (0–1).
  4. Action: Block if score > 0.95.

7. Challenges and Limitations

  • Black Box Problem: Hard to interpret decisions (e.g., why a loan was rejected).
  • Data Hunger: Requires massive datasets (Nepal’s limited data is a hurdle).
  • Ethics: Bias in training data (e.g., Pathao’s driver ratings favoring certain areas).

In the Real World

  1. Khalti’s Fraud Detection

    • Uses a deep neural network with ReLU activation to analyze transaction patterns.
    • Trained on historical fraud cases to flag anomalies (e.g., a merchant suddenly requesting 10× higher limits).
    • Real Example: In 2023, Khalti blocked 12,000 fraudulent transactions using this system, saving users ~NPR 80 million.
  2. Ncell’s Predictive Maintenance

    • LSTM networks predict antenna failures by analyzing temperature, signal strength, and usage trends.
    • Real Example: Reduced unplanned downtime by 28% in Kathmandu’s busy districts.
  3. Daraz’s Recommendation Engine

    • Collaborative filtering + CNNs suggest products based on user images clicked (e.g., if you view shoes, it shows similar styles).
    • Real Example: Increased cross-selling by 15% during Dashain sales.
  4. Nepal’s AI for Tuberculosis Detection

    • CNNs analyze X-rays to detect TB (used in rural hospitals with limited doctors).
    • Real Example: Achieved 92% accuracy in a 2022 pilot in Pokhara.

Exam Tip

  1. Diagrams Are Key: Draw a 3-layer neural network (input → hidden → output) with weights/biases labeled. Examiners love clear visuals.
  2. Backpropagation Steps: Memorize the chain rule for weight updates:
  3. Compare Architectures: Know when to use CNNs (images), RNNs (sequences), or Transformers (text).
  4. Real-World Links: Relate concepts to Khalti/Ncell/Daraz in explanations (e.g., "Like Khalti’s fraud system, this network uses sigmoid activation for binary classification").
  5. Biological vs. Artificial: Highlight one strength (e.g., scalability of ANNs) and one gap (e.g., lack of energy efficiency).
  6. Worked Examples: Always show small numerical traces (e.g., a single neuron’s calculation) to prove understanding.

Pro Tip: For the exam, practice sketching a neural network’s forward/backward pass on a clean sheet—it’s worth 10+ marks!

Based on the TU BSc CSIT syllabus for Cognitive Science, unit 8.

Discussion

Loading…