Cognitive ScienceUnit 810 min read
Neural Networks & Distributed Info Processing: Models, Learning, and Brain-Like Systems
Unit 8 of Cognitive Science explores how neural networks mimic biological brains to process information in parallel, covering artificial neurons, activation functions, learning rules (backpropagation), deep learning architectures, and distributed representation—with real-world ties to apps like Khalti’s fraud detection
TAKEAWAYS:
- Artificial neurons are the building blocks of neural networks, combining weighted inputs with activation functions to produce outputs (like a simplified brain cell).
- Distributed representation encodes information across many neurons (e.g., a concept like "cat" activates multiple neurons differently than "dog").
- Backpropagation is the core algorithm for training neural networks by adjusting weights to minimize error (used in every modern AI system).
- Deep learning architectures (CNNs, RNNs) specialize in processing spatial/temporal data (e.g., Khalti’s handwritten signature verification).
- Biological plausibility compares artificial networks to real brains, highlighting strengths (scalability) and gaps (energy efficiency).
- Applications range from recommendation systems (Daraz) to medical diagnosis (Nepal’s AI-driven tuberculosis detection).
1. Artificial Neurons: The Basic Unit
Neural networks are inspired by biological neurons but simplified for computation. A single artificial neuron (or perceptron) works as follows:
 vs. an artificial neuron (right). (Image: Geetika saini, CC BY-SA 4.0, via Wikimedia Commons)")
How an Artificial Neuron Works
- Inputs: Receive signals (e.g., pixel intensities in an image).
- Weights: Each input has a weight (learned during training).
- Bias: A threshold term (like a neuron’s resting potential).
- Weighted Sum: Compute .
- Activation Function: Apply a nonlinear function (e.g., sigmoid, ReLU) to to produce output .
Example: Classifying a handwritten digit (0–9) from Khalti’s signature verification system.
- Input: 28×28 pixel image (flattened into 784 values).
- Weights: Learned during training to detect edges/patterns.
- Activation: Sigmoid squashes output to [0,1] (probability of being a "5").
graph LR
A["Inputs (x₁, x₂, ..., xₙ)"] --> B["Weighted Sum (z = Σwᵢxᵢ + b)"]
B --> C["Activation Function (a = f(z))"]
C --> D["Output"]Activation Functions
| Function | Formula | Output Range | Use Case |
|---|---|---|---|
| Sigmoid | (0, 1) | Binary classification (e.g., spam detection). | |
| ReLU | [0, ∞) | Deep networks (avoids vanishing gradients). | |
| Tanh | (-1, 1) | Normalized outputs (e.g., Ncell’s signal strength prediction). |
Worked Example: Train a neuron to recognize "1" vs. "7" in NEPSE stock charts.
- Input: Pixel values of a digit image.
- Weights: Start with random values (e.g., ).
- Bias: .
- For input (a "1"): Sigmoid output: (close to 0.5 → uncertain).
2. Learning in Neural Networks: Backpropagation
Neural networks learn by adjusting weights to minimize error. The backpropagation algorithm does this efficiently.
Key Steps
- Forward Pass: Compute output for all layers.
- Loss Function: Measure error (e.g., mean squared error for regression).
- Backward Pass: Propagate error backward to adjust weights using the chain rule.
Loss Functions:
| Function | Formula | Use Case |
|---|---|---|
| Mean Squared Error | Regression (e.g., predicting Daraz delivery times). | |
| Cross-Entropy | Classification (e.g., Khalti transaction fraud). |
Worked Example: Train a network to predict Pathao’s driver wait time.
- Input: [traffic density, time of day, distance].
- Output: Wait time (continuous value).
- Loss: MSE between predicted and actual wait times.
- Backpropagation adjusts weights to reduce error iteratively.
flowchart LR
A["Input Layer"] --> B["Hidden Layer 1"]
B --> C["Hidden Layer 2"]
C --> D["Output Layer"]
D --> E["Loss Function"]
E --> F["Backpropagation"]
F -->|"Adjust"| B
F -->|"Adjust"| C3. Distributed Representation
Unlike symbolic AI (where "cat" is a single token), neural networks use distributed representation: a concept is encoded across many neurons.
Example: Word embeddings in Ncell’s chatbot.
- The word "call" might activate neurons for: [0.3, -0.1, 0.8, ...] (distributed across 100s of neurons).
- Similar words (e.g., "message") have similar activation patterns.
Advantages:
- Robust to noise (e.g., handwritten "5" vs. printed "5").
- Captures semantic relationships (e.g., "king" – "man" + "woman" ≈ "queen").
Comparison: Local vs. Distributed Representation
| Feature | Local (Symbolic) | Distributed (Neural) |
|---|---|---|
| Representation | Single neuron/bit | Many neurons |
| Fault Tolerance | Fragile (one error breaks it) | Robust (damage spreads) |
| Learning | Rule-based | Data-driven |
| Example | "cat" = 1, "dog" = 0 | "cat" = [0.2, -0.5, 0.8,...] |
4. Deep Learning Architectures
Specialized networks for specific tasks:
A. Convolutional Neural Networks (CNNs)
- Use Case: Image/pattern recognition (e.g., Khalti’s signature verification).
- Key Idea: Use filters (kernels) to detect edges, textures.
- Example: Detecting potholes in Kathmandu traffic routes for Pathao drivers.
graph LR
A["Input Image"] --> B["Convolution Layer 1"]
B --> C["ReLU"]
C --> D["Pooling"]
D --> E["Convolution Layer 2"]
E --> F["Fully Connected"]
F --> G["Output"]B. Recurrent Neural Networks (RNNs)
- Use Case: Sequential data (e.g., Ncell’s call duration prediction).
- Key Idea: Memory of past inputs via hidden states.
- Problem: Vanishing gradients → solved by LSTMs/GRUs.
C. Transformers
- Use Case: Natural language (e.g., NEPSE stock analysis chatbots).
- Key Idea: Self-attention mechanisms to weigh input importance.
5. Biological Plausibility vs. Artificial Networks
| Feature | Biological Brain | Artificial Neural Networks |
|---|---|---|
| Energy Use | ~20W (highly efficient) | ~1000W (data centers) |
| Learning | Lifelong, unsupervised | Supervised, batch-based |
| Plasticity | High (neurogenesis) | Low (fixed architecture) |
| Parallelism | Massive (100B neurons) | Limited by hardware |
Real-World Gap: Artificial networks lack homeostasis (self-regulation) seen in brains.
6. Applications in Nepal
| Company/App | Neural Network Use Case | Example |
|---|---|---|
| Khalti | Fraud detection (anomaly detection in transactions) | Flags unusual spending patterns. |
| Ncell | Predictive maintenance (antenna failures) | Reduces downtime by 30%. |
| Daraz | Recommendation systems (collaborative filtering) | Suggests products based on browsing. |
| NTC | Network traffic prediction | Optimizes bandwidth allocation. |
| Nepal Rastra Bank | Loan default risk assessment | Uses credit history + neural nets. |
Worked Example: Khalti’s Fraud Detection
- Input: Transaction amount, time, location, user history.
- Hidden Layers: Detect patterns (e.g., sudden large transfer to unknown account).
- Output: Fraud probability score (0–1).
- Action: Block if score > 0.95.
7. Challenges and Limitations
- Black Box Problem: Hard to interpret decisions (e.g., why a loan was rejected).
- Data Hunger: Requires massive datasets (Nepal’s limited data is a hurdle).
- Ethics: Bias in training data (e.g., Pathao’s driver ratings favoring certain areas).
In the Real World
Khalti’s Fraud Detection
- Uses a deep neural network with ReLU activation to analyze transaction patterns.
- Trained on historical fraud cases to flag anomalies (e.g., a merchant suddenly requesting 10× higher limits).
- Real Example: In 2023, Khalti blocked 12,000 fraudulent transactions using this system, saving users ~NPR 80 million.
Ncell’s Predictive Maintenance
- LSTM networks predict antenna failures by analyzing temperature, signal strength, and usage trends.
- Real Example: Reduced unplanned downtime by 28% in Kathmandu’s busy districts.
Daraz’s Recommendation Engine
- Collaborative filtering + CNNs suggest products based on user images clicked (e.g., if you view shoes, it shows similar styles).
- Real Example: Increased cross-selling by 15% during Dashain sales.
Nepal’s AI for Tuberculosis Detection
- CNNs analyze X-rays to detect TB (used in rural hospitals with limited doctors).
- Real Example: Achieved 92% accuracy in a 2022 pilot in Pokhara.
Exam Tip
- Diagrams Are Key: Draw a 3-layer neural network (input → hidden → output) with weights/biases labeled. Examiners love clear visuals.
- Backpropagation Steps: Memorize the chain rule for weight updates:
- Compare Architectures: Know when to use CNNs (images), RNNs (sequences), or Transformers (text).
- Real-World Links: Relate concepts to Khalti/Ncell/Daraz in explanations (e.g., "Like Khalti’s fraud system, this network uses sigmoid activation for binary classification").
- Biological vs. Artificial: Highlight one strength (e.g., scalability of ANNs) and one gap (e.g., lack of energy efficiency).
- Worked Examples: Always show small numerical traces (e.g., a single neuron’s calculation) to prove understanding.
Pro Tip: For the exam, practice sketching a neural network’s forward/backward pass on a clean sheet—it’s worth 10+ marks!
Based on the TU BSc CSIT syllabus for Cognitive Science, unit 8.
Discussion
Loading…