Neural NetworksUnit 210 min read
Rosenblatt’s Perceptron: Binary Classification, Learning Rules & Limitations
Unit 2 of Neural Networks covers the foundational Perceptron algorithm, its mathematical formulation, training via the Perceptron Learning Rule (PLR), geometric interpretation as a linear classifier, and its limitations (e.g., XOR problem). Includes real-world applications, worked examples, and comparisons with modern
TAKEAWAYS:
- The Perceptron is a single-layer neural network that classifies linearly separable binary data using a weighted sum and step activation.
- It learns via the Perceptron Learning Rule (PLR), adjusting weights iteratively to minimize classification errors.
- Geometrically, the Perceptron finds the optimal hyperplane that separates two classes in the feature space.
- It fails on non-linearly separable problems (e.g., XOR), motivating later architectures like MLPs.
- Real-world use: Spam detection (e.g., eSewa transaction alerts), simple fraud detection (Khalti), and binary decision systems (NTC network routing).
- Limitations: No hidden layers, restricted to linear boundaries, and requires fully separable training data.
1. What is a Perceptron?
A Perceptron is the simplest feedforward neural network for binary classification. It mimics a biological neuron:
- Takes input features (e.g., pixel intensities, transaction amounts).
- Computes a weighted sum (where is the bias).
- Applies a step (threshold) activation function:
- Outputs 1 (class A) or 0 (class B).
Visual: Perceptron Structure
2. Mathematical Formulation
Input-Output Relationship
For inputs and weights : y = \begin{cases} 1 & \text{if } \mathbf{w}^T \mathbf{x} + b \geq 0 \\ 0 & \text{otherwise} \end \end{cases} Key Idea: The Perceptron finds a decision boundary (hyperplane) that separates the two classes.
Geometric Interpretation
In 2D, the decision boundary is a line: In 3D, it’s a plane, and in -dimensions, a hyperplane.
3. Perceptron Learning Rule (PLR)
The Perceptron learns by adjusting weights to minimize classification errors. The update rule: where:
- = learning rate (typically ),
- = true label (0 or 1),
- = predicted output.
Steps:
- Compute .
- Predict .
- If (misclassification), update weights:
- Repeat for all training examples (epochs).
Worked Example: AND Gate as a Perceptron
Train a Perceptron to implement the AND logic gate:
| (AND) | ||
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 0 |
| 1 | 0 | 0 |
| 1 | 1 | 1 |
Initial weights: , , .
Epoch 1:
Input (1,1), : → (error). Update: .
Input (1,0), : → (error). Update: .
Input (0,1), : → (error). Update: .
Epoch 2: Continue until all inputs are classified correctly. Final weights might converge to , (varies by ).
Visual: Convergence Path
flowchart TD
A["Epoch 1: w=[0.1,0.1], b=0.1"] --> B["Epoch 2: w=[0.5,0.5], b=-0.1"]
B --> C["Epoch 3: w=[0.9,0.9], b=-1.0"]
C --> D["Converged: w≈[1,1], b=-1.5"]4. Real-World Applications
1. eSewa Transaction Fraud Detection
- Problem: Classify transactions as fraudulent (1) or legitimate (0).
- How Perceptron Helps:
- Inputs: Transaction amount, time, location, user history.
- Trained on historical data to flag suspicious patterns (e.g., high amount + unusual time).
- Limitation: Fails if fraudsters use non-linear patterns (e.g., encoded amounts).
2. Khalti Payment Routing
- Problem: Route payments to the correct bank or wallet (binary: success/failure).
- How Perceptron Helps:
- Inputs: Account number format, balance, transaction type.
- Learns to reject invalid routes (e.g., insufficient balance + wrong format).
- Real Example: Khalti’s early fraud detection used Perceptrons before switching to deeper networks.
3. NTC Network Packet Classification
- Problem: Classify network packets as legitimate (0) or malicious (1).
- How Perceptron Helps:
- Inputs: Packet size, source IP, port number.
- Acts as a simple firewall rule (e.g., block packets from known malicious IPs).
- Limitation: Cannot detect encrypted malware (requires MLPs or transformers).
5. Limitations of the Perceptron
| Limitation | Explanation | Example |
|---|---|---|
| Linearly Separable Only | Fails if classes cannot be separated by a straight line/hyperplane. | XOR problem (see below). |
| No Hidden Layers | Cannot model complex, nested patterns. | Handwritten digit recognition (requires MLP). |
| Slow Convergence | May take many epochs for non-trivial data. | High-dimensional medical diagnosis data. |
| Sensitive to Features | Performance degrades if features are not scaled or normalized. | Pixel values [0,255] vs. normalized [0,1]. |
The XOR Problem: Why Perceptrons Fail
The XOR truth table is not linearly separable:
| (XOR) | ||
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
Visual: XOR Data in 2D
Solution: Use a Multilayer Perceptron (MLP) (covered in Unit 5) with hidden layers to model non-linear boundaries.
6. Comparison: Perceptron vs. Other Models
| Feature | Perceptron | Logistic Regression | Multilayer Perceptron (MLP) |
|---|---|---|---|
| Activation | Step function (binary) | Sigmoid (probabilistic) | Sigmoid/ReLU (multi-class) |
| Output | 0 or 1 | Probability | Multiple classes/outputs |
| Training | Perceptron Learning Rule (PLR) | Gradient Descent (log loss) | Backpropagation |
| Handles Non-Linearity | ❌ No | ❌ No (unless features engineered) | ✅ Yes (hidden layers) |
| Use Case | Simple binary classification | Probability estimation | Complex patterns (images, text) |
7. Exam Tip: How to Score Full Marks
Define Clearly:
- Start with: "A Perceptron is a single-layer neural network for binary classification using a weighted sum and step activation."
- Mention inputs, weights, bias, and output in your answer.
Show the Math:
- Write the weight update rule and decision boundary equation.
- For worked examples, show all steps (e.g., AND gate training).
Draw the Geometry:
- Sketch a 2D scatter plot with the decision boundary line.
- Label axes as features and classes.
Discuss Limitations:
- Mention XOR problem and linear separability.
- Compare with MLPs or kernel methods (Unit 6).
Real-World Tie-In:
- Link to eSewa/Khalti fraud detection or NTC routing.
- Example: "A Perceptron in eSewa could flag transactions where amount > threshold AND time = night."
Common Pitfalls:
- ❌ Forgetting to normalize inputs (affects convergence).
- ❌ Assuming Perceptrons work for all problems (they don’t!).
- ❌ Confusing PLR with gradient descent (they’re different).
Final Note: The Perceptron is the building block of neural networks. Master its math, geometry, and limitations to ace Unit 2—and prepare for Units 5 (MLPs) and 6 (kernels) where these ideas evolve!
Based on the TU BSc CSIT syllabus for Neural Networks, unit 2.
Discussion
Loading…