Artificial IntelligenceUnit 77 min read
Machine Learning Basics: Models, Algorithms & Applications
Unit 7 of Artificial Intelligence explores supervised/unsupervised learning, regression/classification, model evaluation, and real-world ML pipelines, with visuals of decision trees, neural network layers, and loss functions.
TAKEAWAYS:
- Machine learning trains models on data to make predictions (supervised) or find patterns (unsupervised).
- Key algorithms include linear regression, k-nearest neighbors, decision trees, and clustering (k-means).
- Model performance is measured by metrics like accuracy, precision, recall, and loss functions (e.g., MSE).
- Real-world ML is used in recommendation systems (e.g., Daraz), fraud detection (banks), and route optimization (Pathao).
- Overfitting/underfitting must be avoided via techniques like cross-validation and regularization.
1. What is Machine Learning?
Machine learning (ML) is a subset of AI where systems learn from data to improve performance over time without explicit programming. It relies on three key components:
- Data: Input features (e.g., user age, purchase history) and labels (e.g., "buy" or "not buy").
- Model: A mathematical function (e.g., linear equation, decision tree) that maps inputs to outputs.
- Learning Algorithm: Optimizes the model using training data (e.g., gradient descent).
Types of Machine Learning
mindmap
root((Machine Learning))
Supervised Learning
Classification
Example: Spam detection (email → "spam"/"not spam")
Regression
Example: House price prediction (area → price)
Unsupervised Learning
Clustering
Example: Customer segmentation (K-means)
Dimensionality Reduction
Example: PCA for compressing images
Reinforcement Learning
Example: Pathao’s dynamic pricing (reward: profit)2. Supervised Learning: Regression & Classification
Supervised learning uses labeled data to train models. Two main tasks:
A. Regression (Predicting Continuous Values)
Definition: Predicts numeric outputs (e.g., temperature, stock price). Example: Predicting Daraz’s daily order volume based on past sales.
Algorithm: Linear Regression
- Fits a line to minimize error (loss).
- Loss Function: Mean Squared Error (MSE): where = true value, = predicted value.
Worked Example: Predicting NEPSE index using past 5 days.
| Day | Index (y) | Predicted (ŷ) | Error (y - ŷ) | Squared Error |
|---|---|---|---|---|
| 1 | 1500 | 1480 | 20 | 400 |
| 2 | 1520 | 1510 | 10 | 100 |
| MSE = (400 + 100 + ...) / 5 = 200 |
Visual: Gradient descent updates weights to minimize MSE.
B. Classification (Predicting Categories)
Definition: Predicts discrete labels (e.g., "fraud"/"no fraud"). Algorithms:
- Decision Trees: Splits data based on features (e.g., "age > 30?").
- k-Nearest Neighbors (k-NN): Classifies based on majority vote of nearest neighbors.
Worked Example: Bank loan approval (features: income, credit score; label: "approve"/"reject").
graph TD A["Income > 50k?"] -->|"Yes"| B["Credit Score > 700?"] A -->|"No"| C["Reject"] B -->|"Yes"| D["Approve"] B -->|"No"| C
Real-World Tie: Ncell uses decision trees to predict customer churn (leaving the network).
3. Unsupervised Learning: Clustering & Dimensionality Reduction
No labels are provided; the model finds hidden patterns.
A. Clustering (k-Means)
Goal: Group similar data points (e.g., customer segments). Steps:
- Randomly assign centroids.
- Assign each point to the nearest centroid.
- Recalculate centroids until convergence.
Example: Daraz groups users by purchase behavior (3 clusters: high-spenders, mid-spenders, low-spenders).
B. Dimensionality Reduction (PCA)
Goal: Reduce features while retaining most information (e.g., compressing images). Example: Reducing 1000-pixel images to 100 features for faster processing.
4. Model Evaluation Metrics
| Metric | Formula | Use Case |
|---|---|---|
| Accuracy | Overall correctness (e.g., spam filter) | |
| Precision | "Of predicted frauds, how many are real?" | |
| Recall | "Of actual frauds, how many did we catch?" | |
| F1-Score | Balance between precision/recall |
Confusion Matrix for Fraud Detection:
| Predicted No Fraud | Predicted Fraud
----------|-------------------|-----------------
Actual No Fraud | TN = 900 | FP = 50
Actual Fraud | FN = 20 | TP = 30
5. Overfitting & Underfitting
- Overfitting: Model memorizes training data but fails on new data (e.g., a tree with 100 levels for 100 samples).
- Underfitting: Model is too simple (e.g., a straight line for nonlinear data).
Solutions:
- Regularization: Penalize large weights (e.g., L1/L2).
- Cross-Validation: Split data into training/validation sets.
- Pruning: Simplify decision trees.
Visual: Overfitting vs. good fit.
6. Machine Learning Pipeline
flowchart LR A["Data Collection"] --> B["Data Preprocessing"] B --> C["Feature Selection"] C --> D["Model Training"] D --> E["Hyperparameter Tuning"] E --> F["Evaluation"] F -->|"If Poor"| G["Retrain Model"] F -->|"If Good"| H["Deployment"]
Example: NTC’s traffic prediction system:
- Collects GPS data from vehicles.
- Cleans missing values (e.g., removes 0-speed records).
- Trains a random forest model.
- Evaluates using RMSE (Root Mean Squared Error).
In the Real World
Pathao’s Dynamic Pricing:
- Uses reinforcement learning to adjust fares based on demand/supply (reward: maximizing rides).
- Example: During Dash Hour, prices surge to balance supply.
Khalti’s Fraud Detection:
- Employs supervised learning (logistic regression) to flag suspicious transactions.
- Features: transaction amount, time, user location.
Daraz’s Recommendation System:
- Uses collaborative filtering (unsupervised) to suggest products based on similar users’ purchases.
- Example: "Customers who bought X also bought Y."
Exam Tip
- Define clearly: Differentiate supervised/unsupervised learning with examples.
- Draw diagrams: Decision trees, confusion matrices, and loss curves are high-scoring.
- Apply concepts: Relate algorithms to real cases (e.g., "How would you predict NEPSE trends?").
- Metric calculations: Practice computing accuracy/precision from confusion matrices.
- Avoid memorization: Focus on why an algorithm works (e.g., "k-NN is lazy because it doesn’t train a model").
Key Formulae to Remember:
- Linear Regression Loss:
- Decision Tree Gini Impurity: (for binary classification)
- k-Means Objective: Minimize (where is centroid).
Based on the TU BITM syllabus for Artificial Intelligence (IT228), unit 7.
Discussion
Loading…