Artificial IntelligenceUnit 77 min read

Machine Learning Basics: Models, Algorithms & Applications

Unit 7 of Artificial Intelligence explores supervised/unsupervised learning, regression/classification, model evaluation, and real-world ML pipelines, with visuals of decision trees, neural network layers, and loss functions.

TAKEAWAYS:

  • Machine learning trains models on data to make predictions (supervised) or find patterns (unsupervised).
  • Key algorithms include linear regression, k-nearest neighbors, decision trees, and clustering (k-means).
  • Model performance is measured by metrics like accuracy, precision, recall, and loss functions (e.g., MSE).
  • Real-world ML is used in recommendation systems (e.g., Daraz), fraud detection (banks), and route optimization (Pathao).
  • Overfitting/underfitting must be avoided via techniques like cross-validation and regularization.

1. What is Machine Learning?

Machine learning (ML) is a subset of AI where systems learn from data to improve performance over time without explicit programming. It relies on three key components:

  • Data: Input features (e.g., user age, purchase history) and labels (e.g., "buy" or "not buy").
  • Model: A mathematical function (e.g., linear equation, decision tree) that maps inputs to outputs.
  • Learning Algorithm: Optimizes the model using training data (e.g., gradient descent).

Types of Machine Learning

mindmap
  root((Machine Learning))
    Supervised Learning
      Classification
        Example: Spam detection (email → "spam"/"not spam")
      Regression
        Example: House price prediction (area → price)
    Unsupervised Learning
      Clustering
        Example: Customer segmentation (K-means)
      Dimensionality Reduction
        Example: PCA for compressing images
    Reinforcement Learning
      Example: Pathao’s dynamic pricing (reward: profit)

2. Supervised Learning: Regression & Classification

Supervised learning uses labeled data to train models. Two main tasks:

A. Regression (Predicting Continuous Values)

Definition: Predicts numeric outputs (e.g., temperature, stock price). Example: Predicting Daraz’s daily order volume based on past sales.

Algorithm: Linear Regression

  • Fits a line to minimize error (loss).
  • Loss Function: Mean Squared Error (MSE): where = true value, = predicted value.

Worked Example: Predicting NEPSE index using past 5 days.

Day Index (y) Predicted (ŷ) Error (y - ŷ) Squared Error
1 1500 1480 20 400
2 1520 1510 10 100
MSE = (400 + 100 + ...) / 5 = 200

Visual: Gradient descent updates weights to minimize MSE.

B. Classification (Predicting Categories)

Definition: Predicts discrete labels (e.g., "fraud"/"no fraud"). Algorithms:

  • Decision Trees: Splits data based on features (e.g., "age > 30?").
  • k-Nearest Neighbors (k-NN): Classifies based on majority vote of nearest neighbors.

Worked Example: Bank loan approval (features: income, credit score; label: "approve"/"reject").

graph TD
  A["Income > 50k?"] -->|"Yes"| B["Credit Score > 700?"]
  A -->|"No"| C["Reject"]
  B -->|"Yes"| D["Approve"]
  B -->|"No"| C

Real-World Tie: Ncell uses decision trees to predict customer churn (leaving the network).


3. Unsupervised Learning: Clustering & Dimensionality Reduction

No labels are provided; the model finds hidden patterns.

A. Clustering (k-Means)

Goal: Group similar data points (e.g., customer segments). Steps:

  1. Randomly assign centroids.
  2. Assign each point to the nearest centroid.
  3. Recalculate centroids until convergence.

Example: Daraz groups users by purchase behavior (3 clusters: high-spenders, mid-spenders, low-spenders).

B. Dimensionality Reduction (PCA)

Goal: Reduce features while retaining most information (e.g., compressing images). Example: Reducing 1000-pixel images to 100 features for faster processing.


4. Model Evaluation Metrics

Metric Formula Use Case
Accuracy Overall correctness (e.g., spam filter)
Precision "Of predicted frauds, how many are real?"
Recall "Of actual frauds, how many did we catch?"
F1-Score Balance between precision/recall

Confusion Matrix for Fraud Detection:

          | Predicted No Fraud | Predicted Fraud
----------|-------------------|-----------------
Actual No Fraud | TN = 900          | FP = 50
Actual Fraud   | FN = 20           | TP = 30

5. Overfitting & Underfitting

  • Overfitting: Model memorizes training data but fails on new data (e.g., a tree with 100 levels for 100 samples).
  • Underfitting: Model is too simple (e.g., a straight line for nonlinear data).

Solutions:

  • Regularization: Penalize large weights (e.g., L1/L2).
  • Cross-Validation: Split data into training/validation sets.
  • Pruning: Simplify decision trees.

Visual: Overfitting vs. good fit.


6. Machine Learning Pipeline

flowchart LR
  A["Data Collection"] --> B["Data Preprocessing"]
  B --> C["Feature Selection"]
  C --> D["Model Training"]
  D --> E["Hyperparameter Tuning"]
  E --> F["Evaluation"]
  F -->|"If Poor"| G["Retrain Model"]
  F -->|"If Good"| H["Deployment"]

Example: NTC’s traffic prediction system:

  1. Collects GPS data from vehicles.
  2. Cleans missing values (e.g., removes 0-speed records).
  3. Trains a random forest model.
  4. Evaluates using RMSE (Root Mean Squared Error).

In the Real World

  1. Pathao’s Dynamic Pricing:

    • Uses reinforcement learning to adjust fares based on demand/supply (reward: maximizing rides).
    • Example: During Dash Hour, prices surge to balance supply.
  2. Khalti’s Fraud Detection:

    • Employs supervised learning (logistic regression) to flag suspicious transactions.
    • Features: transaction amount, time, user location.
  3. Daraz’s Recommendation System:

    • Uses collaborative filtering (unsupervised) to suggest products based on similar users’ purchases.
    • Example: "Customers who bought X also bought Y."

Exam Tip

  • Define clearly: Differentiate supervised/unsupervised learning with examples.
  • Draw diagrams: Decision trees, confusion matrices, and loss curves are high-scoring.
  • Apply concepts: Relate algorithms to real cases (e.g., "How would you predict NEPSE trends?").
  • Metric calculations: Practice computing accuracy/precision from confusion matrices.
  • Avoid memorization: Focus on why an algorithm works (e.g., "k-NN is lazy because it doesn’t train a model").

Key Formulae to Remember:

  1. Linear Regression Loss:
  2. Decision Tree Gini Impurity: (for binary classification)
  3. k-Means Objective: Minimize (where is centroid).

Based on the TU BITM syllabus for Artificial Intelligence (IT228), unit 7.

Discussion

Loading…