IT228 Artificial Intelligence

Artificial IntelligenceUnit 77 min read

Machine Learning Basics: Models, Algorithms & Applications

Unit 7 of Artificial Intelligence explores foundational concepts of machine learning, including supervised/unsupervised learning, model evaluation, bias-variance tradeoff, and real-world applications in Nepalese and global tech ecosystems. This note covers definitions, algorithms, worked examples (e.g., loan approval,

Core Concepts of Machine Learning

Machine Learning (ML) is a subset of AI where systems learn from data to make predictions or decisions without explicit programming. It is broadly classified into three types:

1. Types of Machine Learning

classDiagram
    class ML {
        <<abstract>>
        +learns from data
    }
    class Supervised {
        +labeled data
        +predicts output
    }
    class Unsupervised {
        +unlabeled data
        +finds patterns
    }
    class Reinforcement {
        +learns by rewards
        +trial-and-error
    }
    ML <|-- Supervised
    ML <|-- Unsupervised
    ML <|-- Reinforcement

Supervised Learning

  • Definition: Uses labeled data (input-output pairs) to train a model to predict outputs.
  • Examples: Spam detection (email labeled as spam/ham), house price prediction.
  • Algorithms:
    • Regression: Predicts continuous values (e.g., temperature, stock prices).
      • Example: Predicting Daraz product demand based on past sales.
    • Classification: Predicts discrete labels (e.g., loan approval: yes/no).
      • Example: Ncell predicting customer churn (will they leave?).

Unsupervised Learning

  • Definition: Finds hidden patterns in unlabeled data.
  • Examples: Customer segmentation (Khalti grouping users by spending), anomaly detection (NTC detecting fraudulent calls).
  • Algorithms:
    • Clustering: Groups similar data (e.g., K-means for market segmentation).
    • Dimensionality Reduction: Simplifies data (e.g., PCA for compressing images).

Reinforcement Learning

  • Definition: Learns by interacting with an environment (rewards/punishments).
  • Examples: Pathao’s dynamic pricing, robotics (e.g., warehouse automation).
  • Key Idea: Agent learns optimal actions via trial-and-error (e.g., Q-learning).

Key ML Components

1. Training, Validation, and Test Sets

  • Split Data: 70% training, 15% validation, 15% test.
  • Why?
    • Training: Model learns.
    • Validation: Tunes hyperparameters (e.g., learning rate).
    • Test: Evaluates final performance.
  • Example: For a loan approval model (Nepal Bank), use 3 years of past data:
    • Train: 2018–2020 (approved/rejected loans).
    • Validate: 2021 (adjust model).
    • Test: 2022 (check accuracy).

2. Model Evaluation Metrics

Metric Supervised Use Case Formula
Accuracy Overall correctness (classification) (TP + TN) / (TP + TN + FP + FN)
Precision How many predicted positives are correct? TP / (TP + FP)
Recall (Sensitivity) How many actual positives are captured? TP / (TP + FN)
F1-Score Balance of precision/recall 2 * (Precision * Recall) / (Precision + Recall)
RMSE Error magnitude (regression) sqrt(mean((predicted - actual)²))

Worked Example: Loan Approval Model

  • Data: 1000 past loans (500 approved, 500 rejected).
  • Model Predicts:
    • True Positives (TP): 450 (correctly approved).
    • False Positives (FP): 50 (wrongly approved).
    • False Negatives (FN): 30 (wrongly rejected).
  • Calculations:
    • Accuracy = (450 + 470) / 1000 = 92%.
    • Precision = 450 / (450 + 50) = 90%.
    • Recall = 450 / (450 + 30) = 94%.

Bias-Variance Tradeoff

graph LR
    A["High Bias"] --> B["Underfitting"]
    C["High Variance"] --> D["Overfitting"]
    E["Optimal Model"] -->|"Balanced"| F["Good Performance"]
    A -->|"Simplistic"| E
    C -->|"Complex"| E
  • Bias: Error due to overly simplistic assumptions (e.g., linear model for nonlinear data).
  • Variance: Error due to excessive sensitivity to training data (e.g., memorizing noise).
  • Solution: Use cross-validation or regularization (e.g., L1/L2 penalties).

Real-World Tie-In:

  • eSewa’s Payment Model:
    • High Bias: Assumes all users pay on time → misses fraud.
    • High Variance: Memorizes exact transaction patterns → fails on new users.
    • Fix: Use ensemble methods (e.g., Random Forest) to balance bias/variance.

Feature Engineering

1. Feature Selection vs. Feature Extraction

Feature Selection Feature Extraction
Chooses relevant features from existing data. Creates new features (e.g., PCA).
Example: For traffic prediction, select "time of day," "holiday," but ignore "weather" if irrelevant. Example: Convert raw sensor data (NTC traffic cameras) into "average speed per lane."

2. Data Preprocessing Steps

flowchart LR
    A["Raw Data"] --> B["Handle Missing Values"]
    B --> C["Normalization: Scale to [0,1] or Z-score"]
    C --> D["Encoding: Categorical → Numerical"]
    D --> E["Split into Train/Validation/Test"]
    E --> F["Model Training"]

Worked Example: Kathmandu Traffic Prediction

  • Raw Data: GPS coordinates, timestamps, speed (missing values for 10% of data).
  • Steps:
    1. Missing Values: Fill gaps with average speed for that route.
    2. Normalization: Scale speed from 0–120 km/h to [0,1].
    3. Encoding: Convert "route type" (e.g., "ring road," "arterial") to numerical codes.
    4. Train-Test Split: 70% 2022 data, 15% 2023, 15% 2024 (future prediction).

In the Real World

  1. Khalti’s Fraud Detection

    • Idea: Supervised learning (classification).
    • How: Uses transaction history (labeled as fraud/legit) to train a model that flags suspicious payments in real time.
    • Impact: Reduces false positives (e.g., blocking legitimate transfers).
  2. Pathao’s Dynamic Pricing

    • Idea: Reinforcement learning.
    • How: Adjusts fares based on demand/supply (reward = maximizing driver earnings and passenger satisfaction).
    • Example: During Dashain, prices surge in Lalitpur but drop in Bhaktapur.
  3. Nepal Stock Exchange (NEPSE) Predictions

    • Idea: Time-series forecasting (regression).
    • How: Uses past stock prices (e.g., NABIL, NMB) to predict trends. Investors use this to decide buy/sell.
    • Challenge: High variance (market crashes) requires robust models.

Exam Tip

  1. Define Clearly:
    • Differentiate supervised vs. unsupervised with examples (e.g., "Khalti uses supervised learning for fraud detection").
  2. Show Calculations:
    • For metrics (accuracy, precision), always write the formula and plug in numbers.
  3. Compare Models:
    • Use tables to contrast algorithms (e.g., "K-means vs. DBSCAN for customer segmentation").
  4. Real-World Links:
    • Tie theory to Nepalese apps (e.g., "eSewa’s underfitting problem → need more features").
  5. Visuals:
    • Draw confusion matrices for classification problems.
    • Sketch bias-variance curves to explain tradeoffs.

Practice Questions

  1. Short Answer:
    • How would you preprocess data for a model predicting Daraz delivery delays?
  2. Calculation:
    • Given a spam filter with TP=80, FP=10, FN=20, calculate precision and recall.
  3. Application:
    • Suggest an ML type for NTC’s task of optimizing call routing during festivals. Justify your choice.

Based on the TU BIM syllabus for Artificial Intelligence (IT228), unit 7.

Discussion

Loading…