Machine LearningUnit 1012 min read

Model Evaluation: Metrics, Validation, Bias-Variance Tradeoff

Unit 10 of Machine Learning explores how to rigorously assess ML models—key metrics (accuracy, precision, recall, F1, ROC), validation strategies (train-test split, cross-validation), bias-variance decomposition, and real-world tradeoffs like precision vs. recall in fraud detection or recall vs. false positives in medi

Core Concepts

1. Why Evaluate Models?

Machine learning models are not inherently "correct"—they make predictions based on patterns in data. Evaluation quantifies how well a model generalizes to unseen data. Without evaluation, you risk deploying models that perform poorly in production (e.g., a spam filter that misclassifies legitimate emails or a loan approval system that unfairly rejects applicants).

Key Idea:

A model’s performance on training data ≠ its performance on real-world data.


2. Evaluation Metrics

Metrics depend on the type of problem (classification, regression, clustering) and the cost of errors. Below are the most critical metrics for classification tasks, visualized for clarity.

Confusion Matrix

For binary classification, the confusion matrix summarizes true/false positives/negatives. It is the foundation for all other metrics.

graph LR
    A["True Label"] --> B["Positive"] & C["Negative"]
    B --> D["Predicted Positive"] & E["Predicted Negative"]
    C --> F["Predicted Positive"] & G["Predicted Negative"]
    D["TP"] -->|"True Positive"| H["Correct"]
    E["TN"] -->|"True Negative"| H
    F["FP"] -->|"False Positive"| I["Error"]
    G["FN"] -->|"False Negative"| I
  • True Positive (TP): Correctly predicted positive (e.g., a fraudulent transaction flagged as fraud).
  • True Negative (TN): Correctly predicted negative (e.g., a legitimate transaction not flagged).
  • False Positive (FP): Incorrectly predicted positive (e.g., a legitimate transaction flagged as fraud → Type I error).
  • False Negative (FN): Incorrectly predicted negative (e.g., a fraudulent transaction not flagged → Type II error).

Accuracy

Accuracy =

  • Pros: Simple to understand.
  • Cons: Misleading for imbalanced datasets (e.g., 99% accuracy in a dataset where 99% are negatives).

Example: Suppose a model predicts loan defaults:

  • TP = 50 (correctly identified defaulters)
  • TN = 900 (correctly identified non-defaulters)
  • FP = 20 (legitimate borrowers flagged as defaulters)
  • FN = 30 (actual defaulters missed) Accuracy = . But if only 1% of loans default, this model is not reliable for risk assessment.

Precision, Recall, and F1-Score

For imbalanced data, precision and recall are more informative.

Metric Formula Interpretation When to Use
Precision Of all predicted positives, how many are correct? High FP cost (e.g., spam detection)
Recall Of all actual positives, how many did we catch? High FN cost (e.g., cancer detection)
F1-Score Harmonic mean of precision and recall (balances both). Imbalanced data, need tradeoff.

Worked Example: Fraud Detection (Khalti)

  • Scenario: Khalti wants to flag fraudulent transactions. FP (legitimate transaction blocked) is costly for users, but FN (fraud missed) is costly for Khalti.
  • Data: 100 frauds (P) and 9900 legitimate transactions (N).
  • Model Output:
    • TP = 80, FP = 100, FN = 20, TN = 9800.
  • Calculations:
    • Precision =
    • Recall =
    • F1-Score =

Interpretation:

  • The model is too aggressive in flagging (high FP). Khalti might adjust the threshold to reduce FP at the cost of recall.

ROC Curve and AUC

The Receiver Operating Characteristic (ROC) curve plots True Positive Rate (TPR = Recall) vs. False Positive Rate (FPR = FP / (FP + TN)) at different classification thresholds.

graph TD
    A["FPR (x-axis)"] --> B["0"] --> C["0.2"] --> D["0.4"] --> E["0.6"] --> F["0.8"] --> G["1"]
    B -->|"TPR"| H["0.1"]
    C -->|"TPR"| I["0.3"]
    D -->|"TPR"| J["0.6"]
    E -->|"TPR"| K["0.8"]
    F -->|"TPR"| L["0.9"]
    G -->|"TPR"| M["1"]
  • AUC (Area Under the Curve): Measures the model’s ability to distinguish classes. AUC = 1 (perfect), AUC = 0.5 (random guessing).
  • Example: A medical test for diabetes with AUC = 0.92 is excellent; AUC = 0.6 is poor.

Real-World Tie-In:

  • Nepal Rastra Bank (NRB) uses AUC to evaluate models for detecting money laundering. A high AUC ensures the model is not just memorizing training data but generalizing to new fraud patterns.

3. Model Validation Strategies

How do you ensure your model generalizes? Train-test split and cross-validation are the gold standards.

Train-Test Split

  • Split data into training (70-80%) and test (20-30%) sets.
  • Pros: Simple, fast.
  • Cons: Single split may not capture data variability. Risk of high variance in performance.

Example:

  • Dataset: 1000 customer churn predictions.
  • Train: 800 samples, Test: 200 samples.
  • If the test set is not representative (e.g., only high-value customers), the model may fail in production.

K-Fold Cross-Validation

Split data into K folds, train on K-1 folds, test on the remaining fold. Repeat K times and average performance.

flowchart LR
    A["Dataset"] --> B["Fold 1"] & C["Fold 2"] & D["Fold 3"]
    B --> E["Train on Fold 2 + 3, Test on Fold 1"]
    C --> F["Train on Fold 1 + 3, Test on Fold 2"]
    D --> G["Train on Fold 1 + 2, Test on Fold 3"]
    H["Average Accuracy"] --> I["Final Model Evaluation"]
  • K=5 or K=10 are common choices.
  • Pros: Reduces variance, better estimate of generalization.
  • Cons: Computationally expensive.

Worked Example: Daraz Customer Segmentation

  • Problem: Daraz wants to predict which customers will buy electronics.
  • Data: 5000 customers, 5 features (age, location, past purchases, etc.).
  • Approach: 5-fold CV.
    • Fold 1: Train on 4000, Test on 1000 → Accuracy = 82%
    • Fold 2: Train on 4000 (different 1000), Test on 1000 → Accuracy = 85%
    • Average accuracy = 83.6% (more reliable than a single split).

Stratified K-Fold

For imbalanced datasets, ensure each fold has the same class distribution as the original dataset.

Example:

  • Dataset: 90% non-churn, 10% churn.
  • Stratified 5-fold ensures each fold has ~9% churners.

4. Bias-Variance Tradeoff

No model is perfect. The bias-variance decomposition explains why models underfit or overfit.

graph LR
    A["Model Error"] --> B["Bias"] & C["Variance"] & D["Irreducible Error"]
    B["High Bias"] --> E["Underfitting: Model too simple"]
    C["High Variance"] --> F["Overfitting: Model too complex"]
    D["Irreducible Error"] --> G["Noise in data"]
Term Definition Symptoms Solution
High Bias Model is too simple to capture patterns. High error on train and test data. Use more complex models, add features.
High Variance Model fits noise in training data. Low train error, high test error. Regularization, more data, simpler models.
Irreducible Error Noise in the data itself (e.g., measurement errors). Cannot be reduced. Collect better data.

Visualizing Bias-Variance:

graph TD
    A["Complexity"] --> B["Low"] --> C["High Bias, Low Variance"] --> D["Underfitting"]
    A --> E["Medium"] --> F["Balanced Bias-Variance"]
    A --> G["High"] --> H["Low Bias, High Variance"] --> I["Overfitting"]

Real-World Example: Traffic Prediction (NTC)

  • High Bias: A linear model predicting traffic speed in Kathmandu may ignore rush-hour patterns → underfits.
  • High Variance: A complex model memorizes every traffic jam location but fails in new areas → overfits.
  • Solution: Use random forests (medium complexity) with cross-validation.

5. Learning Curves

Plot training and validation error vs. training set size to diagnose bias/variance.

graph LR
    A["Training Set Size"] --> B["Small"] --> C["High Train Error, High Val Error"] --> D["High Bias"]
    A --> E["Medium"] --> F["Low Train Error, Low Val Error"] --> G["Balanced"]
    A --> H["Large"] --> I["Low Train Error, High Val Error"] --> J["High Variance"]

Interpretation:

  • Both errors high? → Increase model complexity.
  • Train error low, val error high? → Reduce complexity or get more data.
  • Both errors low? → Model is well-balanced.

Example: Loan Approval Model (Nabil Bank)

  • Observation: Training error = 5%, validation error = 20%.
  • Diagnosis: High variance (overfitting). Solutions:
    1. Add regularization (L1/L2).
    2. Collect more loan data.
    3. Use ensemble methods (e.g., bagging).

In the Real World

  1. Khalti’s Fraud Detection

    • Idea Used: Precision-Recall Tradeoff + AUC.
    • How: Khalti’s ML model flags transactions. A high precision (low FP) is critical to avoid blocking legitimate payments, while high recall ensures most frauds are caught. The team uses AUC-ROC to tune the threshold dynamically based on transaction volume.
  2. Pathao’s Driver Demand Prediction

    • Idea Used: K-Fold Cross-Validation + RMSE (Regression Metric).
    • How: Pathao uses historical ride data to predict driver demand in different Kathmandu zones. They employ 5-fold CV to ensure the model generalizes across seasons (monsoon vs. winter). The RMSE (Root Mean Squared Error) helps them quantify how many drivers to allocate per zone.
  3. Nepal Stock Exchange (NEPSE) Index Prediction

    • Idea Used: Bias-Variance Tradeoff + Walk-Forward Validation.
    • How: Algorithmic traders use ML to predict NEPSE movements. A simple linear model (high bias) may miss trends, while a complex neural net (high variance) overfits to past crashes. Traders use walk-forward validation (rolling train-test windows) to simulate real-time performance.

Exam Tip

What Examiners Look For

  1. Metric Selection:

    • Know when to use accuracy vs. precision/recall/F1.
    • For imbalanced data, never use accuracy alone.
    • Example Question: "A model predicts disease X (1% prevalence). Accuracy is 99%. Is this model reliable? Explain using precision and recall."
    • Expected Answer: Discuss how accuracy is misleading; calculate precision/recall to show the model may have high FP/FN.
  2. Validation Strategies:

    • Train-test split vs. cross-validation: When to use each.
    • Stratified K-fold for imbalanced data.
    • Example Question: "How would you evaluate a model for detecting fake news on social media?"
    • Expected Answer: Use stratified 5-fold CV because fake news is rare (imbalanced). Report precision (to avoid flagging real news) and recall (to catch most fakes).
  3. Bias-Variance Tradeoff:

    • Diagnose from learning curves.
    • Fix overfitting/underfitting with the right techniques.
    • Example Question: "A model has 95% train accuracy but 70% test accuracy. What’s wrong? How to fix?"
    • Expected Answer: High variance (overfitting). Solutions: regularization, pruning (for trees), or more data.
  4. Real-World Applications:

    • Always tie theory to practice. For example:
      • Bank loan approval: Discuss precision (avoid false rejections) vs. recall (catch defaulters).
      • Medical diagnosis: Emphasize recall (catch all diseases) even if it means more FP tests.
    • Example Question: "How would you evaluate a model for predicting patient readmission?"
    • Expected Answer: Use AUC-ROC (since cost of FN is high) and stratified CV (imbalanced data).
  5. Common Pitfalls:

    • Data Leakage: Never use test data to tune hyperparameters (e.g., scaling before train-test split).
    • Ignoring Class Imbalance: Always check class distribution before choosing metrics.
    • Overfitting to a Single Split: Always use cross-validation for robust evaluation.

Quick Revision Checklist

Before the exam, ensure you can: ✅ Define TP, TN, FP, FN and derive precision, recall, F1, accuracy. ✅ Explain why accuracy is insufficient for imbalanced data (give a numerical example). ✅ Draw and interpret a confusion matrix and ROC curve. ✅ Describe train-test split vs. K-fold CV and when to use each. ✅ Diagnose high bias vs. high variance from learning curves. ✅ Apply bias-variance tradeoff to a real scenario (e.g., traffic prediction, fraud detection). ✅ Calculate AUC from a given ROC curve (or interpret it given a value).

Based on the PU BE Computer (PU) syllabus for Machine Learning (CMP364), unit 10.

Discussion

Loading…