Artificial IntelligenceUnit 610 min read
Learning in AI: Types, Methods & Neural Learning
Unit 6 of Artificial Intelligence explores how AI systems acquire knowledge—from supervised/unsupervised learning to inductive and analogy-based methods—with real-world applications in recommendation systems, fraud detection, and neural network training.
TAKEAWAYS:
- Learning in AI is the ability to improve performance from experience, using data or feedback.
- Supervised learning uses labeled data (e.g., spam detection), while unsupervised learning finds hidden patterns (e.g., customer segmentation).
- Inductive learning generalizes from examples (e.g., medical diagnosis), while analogy-based learning solves problems by mapping known solutions (e.g., Daraz’s route optimization).
- Neural networks learn via backpropagation (adjusting weights to minimize error) and activation functions (e.g., ReLU, sigmoid).
- Overfitting (memorizing training data) and underfitting (failing to capture patterns) are key pitfalls in learning models.
1. What is Learning in AI?
Learning in AI is the process by which a system improves its performance on a task by analyzing data, identifying patterns, and making decisions with minimal human intervention. Unlike traditional rule-based systems, learning systems adapt over time.
Key Characteristics of Learning in AI:
- Experience: The system improves based on input data or feedback.
- Performance: The system’s output becomes more accurate or efficient.
- Generalization: The system applies learned knowledge to unseen data.
Types of Learning in AI
Learning can be classified into three broad categories:
| Type | Description | Example |
|---|---|---|
| Supervised Learning | Uses labeled data (input-output pairs) to train the model. | Spam detection (email labeled as spam/ham). |
| Unsupervised Learning | Finds hidden patterns in unlabeled data. | Customer segmentation (grouping users by behavior). |
| Reinforcement Learning | Learns by interacting with an environment and receiving rewards/penalties. | Pathao’s delivery route optimization (reward = faster delivery). |
2. Supervised Learning
Supervised learning is the most common type, where the model is trained on a dataset with known inputs and outputs.
How It Works:
- Training Data: A dataset with input features (
X) and corresponding labels (Y). Example:X = [email text],Y = [spam/ham]. - Model Training: The algorithm learns a mapping function
f(X) → Y. - Prediction: The trained model predicts
Yfor newX.
Example: Email Spam Detection
- Dataset: 1000 emails labeled as spam/ham.
- Features: Words like "free," "offer," "win."
- Model: Logistic Regression or Naive Bayes.
- Prediction: New email with "free offer" → classified as spam.
Real-World Application: Khalti’s Fraud Detection
Khalti uses supervised learning to detect fraudulent transactions:
- Input: Transaction amount, time, location.
- Output: Fraud (1) or Legitimate (0).
- Model: Random Forest or SVM.
- Result: Reduces false positives in payments.
3. Unsupervised Learning
Unsupervised learning works with unlabeled data to discover hidden structures.
Common Techniques:
- Clustering: Groups similar data points (e.g., K-Means).
- Dimensionality Reduction: Reduces feature space (e.g., PCA).
- Association Rule Learning: Finds relationships (e.g., Market Basket Analysis).
Example: Customer Segmentation for Daraz
Daraz uses K-Means clustering to group customers:
- Data: Purchase history, browsing behavior.
- Clusters: High-value buyers, occasional shoppers, bargain hunters.
- Marketing: Target promotions based on cluster.
Visual: K-Means Clustering
graph TD
A["Data Points"] --> B["Cluster 1: High Spenders"]
A --> C["Cluster 2: Occasional Buyers"]
A --> D["Cluster 3: Bargain Hunters"]
B -->|"Promotions"| E["Personalized Offers"]
C -->|"Discounts"| E
D -->|"Flash Sales"| E4. Inductive Learning
Inductive learning generalizes from specific examples to broader rules.
How It Works:
- Observations: Specific instances (e.g., "All observed swans are white").
- Generalization: Derives a rule (e.g., "All swans are white").
- Prediction: Applies the rule to new cases.
Example: Medical Diagnosis
- Data: Symptoms of 1000 patients with/without diabetes.
- Rule: "If blood sugar > 126 mg/dL, predict diabetes."
- Prediction: New patient with high blood sugar → diagnosed.
Limitations:
- Overfitting: Model memorizes training data but fails on new data.
- Bias: Assumes all swans are white (ignores black swans in Australia).
5. Learning by Analogy
Analogy-based learning solves new problems by mapping them to known solutions.
Types of Analogy:
| Type | Description | Example |
|---|---|---|
| Derivational Analogy | Solves a problem by transforming a known solution. | Daraz’s route optimization (like solving a maze). |
| Structural Analogy | Maps the structure of one problem to another. | Medical diagnosis (like comparing symptoms to known diseases). |
| Transformational Analogy | Applies a transformation to a known solution to fit a new problem. | NTC’s traffic prediction (like weather forecasting). |
Example: Pathao’s Delivery Route Optimization
- Known Problem: Shortest path in a graph (Dijkstra’s algorithm).
- New Problem: Real-time traffic delays.
- Analogy: Adjusts Dijkstra’s algorithm with live traffic data.
6. Learning in Neural Networks
Neural networks learn by adjusting weights to minimize prediction error.
Key Concepts:
- Activation Functions: Introduce non-linearity (e.g., ReLU, Sigmoid).
- Backpropagation: Adjusts weights using gradient descent.
- Loss Function: Measures prediction error (e.g., Mean Squared Error).
Example: Neural Network for Handwritten Digit Recognition
- Input Layer: 28×28 pixels (MNIST dataset).
- Hidden Layers: Extract features (edges, curves).
- Output Layer: Predicts digit (0–9).
- Training: Adjusts weights to minimize misclassification.
Visual: Neural Network Layers
graph TD
A["Input Layer\n(784 neurons)"] --> B["Hidden Layer 1\n(256 neurons, ReLU)"]
B --> C["Hidden Layer 2\n(128 neurons, ReLU)"]
C --> D["Output Layer\n(10 neurons, Softmax)"]
D -->|"Prediction"| E["Digit: 5"]Real-World Application: Ncell’s Churn Prediction
Ncell uses neural networks to predict customer churn:
- Input: Call duration, data usage, complaints.
- Output: Churn probability (0–1).
- Action: Retention offers for high-risk users.
7. Challenges in Learning
| Challenge | Description | Solution |
|---|---|---|
| Overfitting | Model memorizes training data but fails on new data. | Use cross-validation, regularization. |
| Underfitting | Model is too simple to capture patterns. | Add more features, increase model complexity. |
| Bias-Variance Tradeoff | High bias (underfitting) vs. high variance (overfitting). | Use ensemble methods (e.g., Random Forest). |
| Data Quality | Noisy or incomplete data degrades performance. | Clean data, use imputation techniques. |
In the Real World
Khalti’s Fraud Detection
- Uses supervised learning (SVM/Random Forest) to classify transactions as fraudulent or legitimate.
- Example: A transaction from Kathmandu to Pokhara at 3 AM → flagged as high-risk.
Daraz’s Recommendation System
- Uses collaborative filtering (unsupervised learning) to suggest products.
- Example: If User A buys a phone case, Daraz recommends similar cases to User B.
NTC’s Traffic Prediction
- Uses time-series forecasting (neural networks) to predict congestion.
- Example: During Dashain, NTC predicts delays on Ring Road and reroutes buses.
Pathao’s Delivery Optimization
- Uses reinforcement learning to adjust delivery routes dynamically.
- Example: Avoids traffic jams by learning from past deliveries.
Exam Tip
- Define clearly: Always start with definitions (e.g., "Supervised learning is...").
- Use examples: Relate concepts to real-world apps (e.g., Khalti for fraud detection).
- Compare methods: Draw tables for supervised vs. unsupervised learning.
- Diagrams: Sketch neural network layers or decision trees in exams.
- Pitfalls: Mention overfitting/underfitting when discussing learning challenges.
Worked Example for Exam: Question: Explain supervised learning with an example from NEPSE. Answer: Supervised learning uses labeled data to train a model. For NEPSE:
- Data: Historical stock prices (input) and market trends (output).
- Model: Linear Regression or Decision Tree.
- Prediction: Forecasts stock prices for the next trading day. Visual:
graph TD
A["Historical Data\n(2010-2023)"] --> B["Train Model\n(Linear Regression)"]
B --> C["Predict 2024\nStock Price: NRS 5000]"]Based on the TU BCA syllabus for Artificial Intelligence (CACS410), unit 6.
Discussion
Loading…