CSC317 Simulation and Modeling

Simulation and ModelingUnit 814 min read

Model Validation, Verification & Output Analysis: Methods, Tests & Real-World Checks

Unit 8 of Simulation and Modeling covers the critical final steps of simulation: verification (building the model correctly), validation (building the right model), calibration (adjusting parameters to match reality), and output analysis (extracting meaningful results). Learn statistical tests (Poker, Runs, Chi-Square)

Key Concepts and Definitions

1. Verification vs. Validation: The Core Distinction

VerificationValidationCalibrationAccreditation
Core distinctions in V&V: Verification ensures correctness of model implementation, while Validation checks real-world applicability.
  • Verification: Ensuring the model implements the theory correctly (e.g., no coding errors, correct probability distributions).
    • Example: If you simulate a queue for a Daraz delivery center, verification checks that the arrival/exit logic matches the theoretical queueing model (M/M/1, M/D/1, etc.).
  • Validation: Ensuring the model represents the real system (e.g., does the simulated queue behave like the real Daraz warehouse?).
    • Example: Compare simulated wait times with real Daraz delivery delays during peak sales (Dashain/Eid).

2. The 4-Step V&V Process

Step-by-Step Workflow

  1. Conceptual Model: Define what the simulation will represent.

    • Example: Simulate NTC’s call center queue to predict wait times during power outages.
    • Key Questions:
      • Are all relevant factors included (e.g., call arrival rates, agent skills)?
      • Are assumptions explicit (e.g., Poisson arrivals)?
  2. Verification: Check if the model works as designed.

    • Methods:
      • Code Review: Manual inspection or automated tools (e.g., unit tests in Python/GPSS).
      • Tracing: Follow a single entity (e.g., a customer) through the simulation to ensure correct logic.
      • Extreme-Case Testing: Simulate edge cases (e.g., zero arrivals, infinite queue length).
    • Example: For a Pokhara University exam hall simulation, verify that:
      • Students arrive at the correct rate (λ = 10/hour).
      • The exam duration (μ = 3 hours) is enforced.
  3. Validation: Compare model outputs to real-world data.

    • Methods:
      • Historical Data: Use past records (e.g., Ncell’s call logs).
      • Expert Judgment: Consult domain experts (e.g., a Daraz logistics manager).
      • Statistical Tests: Compare distributions (e.g., Chi-Square for arrival times).
    • Example: Validate a Kathmandu traffic simulation by comparing simulated vs. real congestion at Thapathali during rush hour.
  4. Calibration: Adjust model parameters to match reality.

    • Example: If simulated NTC call wait times are 20% lower than real data, increase the arrival rate (λ) or decrease agent service rate (μ).

3. Statistical Tests for Validation

A. Poker Test (for Random Number Generation)

Purpose: Check if random numbers are uniformly distributed (no patterns). Steps:

  1. Divide numbers into groups (e.g., for 4-digit numbers: 0000–0999, 1000–1999, ..., 9000–9999).
  2. Count numbers in each group. For 1000 numbers, expect ~100 per group.
  3. Use Chi-Square test to compare observed vs. expected counts.

Worked Example:

  • Data: 1000 four-digit random numbers. Observed counts per group:

    Group 0000–0999 1000–1999 2000–2999 3000–3999 4000–4999 5000–5999 6000–6999 7000–7999 8000–8999 9000–9999
    Count 95 110 88 105 120 90 102 85 115 90
  • Expected Count: 100 per group.

  • Chi-Square Statistic:

  • Critical Value (α = 0.05, df = 9): 16.92.

  • Conclusion: Since 12.5 < 16.92, fail to reject H₀ (numbers are uniformly distributed).

Real-World Tie-In:

  • eSewa’s Transaction IDs: If eSewa generates 10,000 transaction IDs daily, a Poker test ensures no ID ranges are over/under-represented (e.g., no bias toward IDs starting with "1" or "9").

B. Runs Test (for Independence)

Purpose: Check if random numbers are independent (no sequences like 1234 or 4321). Steps:

  1. Convert numbers to binary (e.g., 1234 → 0010 0110 0011 0100).
  2. Count "runs" (switches from 0→1 or 1→0).
  3. Compare to expected runs for a given sequence length.

Worked Example:

  • Data: 20 binary digits: 1 0 1 1 0 0 1 0 1 1 1 0 0 0 1 1 0 1 0 1
  • Runs: 1→0, 0→1, 1→0, 0→1, 1→1 (no run), 1→0, 0→0 (no run), 0→1, 1→1 (no run), 1→0, 0→1, 1→0, 0→1 → Total runs = 10.
  • Expected Runs (for 20 digits, n₁=10 ones, n₂=10 zeros):
  • Standard Deviation:
  • Z-Score:
  • Conclusion: |Z| < 1.96 → fail to reject H₀ (numbers are independent).

Real-World Tie-In:

  • Khalti’s OTP Generation: If Khalti sends OTPs like 1234 or 4321, the Runs test would flag it as non-random. Independent OTPs prevent brute-force attacks.

C. Chi-Square Goodness-of-Fit Test

Purpose: Compare simulated vs. real distributions (e.g., arrival times, service times). Steps:

  1. Define bins (e.g., for service times: 0–1 min, 1–2 min, etc.).
  2. Count real vs. simulated observations in each bin.
  3. Compute Chi-Square statistic.

Worked Example:

  • Scenario: Simulate a Pathao driver’s trip time (real data vs. model).
  • Real Data (Nepalese drivers, minutes):
    Bin 0–5 5–10 10–15 15–20 20+
    Count 30 50 40 20 10
  • Simulated Data:
    Bin 0–5 5–10 10–15 15–20 20+
    Count 25 45 45 25 10
  • Chi-Square Calculation:
  • Critical Value (α = 0.05, df = 4): 9.49.
  • Conclusion: 2.25 < 9.49 → model is valid (simulated times match real data).

4. Output Analysis: Extracting Meaningful Results

A. Confidence Intervals

Purpose: Estimate the range of a metric (e.g., average wait time) with uncertainty. Formula:

  • : Sample mean (e.g., average Ncell call wait time = 2.5 minutes).
  • : Sample standard deviation.
  • : Number of replications.

Worked Example:

  • Scenario: Simulate Ncell’s call center for 10 replications.
  • Results:
    • minutes, minutes, .
    • (from t-table).
  • CI:
  • Interpretation: We are 95% confident the true average wait time is between 2.29 and 2.71 minutes.

B. Hypothesis Testing

Example: Test if a new Daraz warehouse layout reduces order processing time.

  • Null Hypothesis (H₀): .
  • Alternative Hypothesis (H₁): .
  • Test Statistic: Use a t-test if data is normal, or Mann-Whitney U if not.

5. Model Calibration: Iterative Refinement

flowchart LR
    A["Start with initial parameters"] --> B["Run simulation"]
    B --> C["Compare outputs to real data"]
    C -->|"Discrepancy found?"| D["Yes"]
    D --> E["Adjust parameters"]
    E --> B
    C -->|"No"| F["Stop"]

Example: Calibrating a Kathmandu traffic model.

  1. Initial Run: Simulate traffic flow with λ = 500 cars/hour, μ = 30 cars/hour.
  2. Validation: Real data shows average speed = 20 km/h, but simulation gives 25 km/h.
  3. Adjustment: Increase λ to 600 cars/hour (more congestion).
  4. Repeat: Now simulation matches real speed (20 km/h).
1234563.844.24.44.64.8ySimulated Wait Time (minutes)Real-World Data (minutes)
Calibration progress: Simulated vs. real wait times after 5 iterations.

6. Accreditation: Stakeholder Approval

Purpose: Ensure the model is acceptable to users (e.g., NTC, banks, eSewa). Steps:

  1. Present findings to domain experts.
  2. Document assumptions and limitations.
  3. Obtain sign-off (e.g., from a university or government body).

Example:

  • A TU-developed model for NEPSE stock price prediction must be reviewed by SEBON before use.

In the Real World

  1. eSewa’s Transaction Queue Simulation

    • Idea Used: Queueing Theory (M/M/1) + Validation via Chi-Square.
    • How: eSewa simulates transaction processing queues to predict peak-hour delays. They validate the model by comparing simulated vs. real transaction success rates during Dashain (when usage spikes). If the model predicts 5% failures but reality shows 8%, they recalibrate arrival rates (λ) or server processing times (μ).
  2. Ncell’s Network Load Modeling

    • Idea Used: Discrete Event Simulation (DES) + Poker Test for Randomness.
    • How: Ncell simulates call arrival patterns to optimize tower placements. They use the Poker test to ensure call IDs are randomly assigned (no bias toward certain prefixes). Output analysis helps predict call drop rates during festivals (e.g., Tihar).
  3. Daraz’s Warehouse Order Fulfillment

    • Idea Used: Confidence Intervals + Hypothesis Testing.
    • How: Daraz simulates order processing times to test if a new warehouse layout reduces delivery times. They run 20 replications, compute a 95% CI for average processing time, and compare it to the old layout’s CI. If the new CI is significantly lower, they adopt the change.

Exam Tip

What Examiners Want to See

  1. Definitions:

    • Clearly distinguish verification (code/logic) vs. validation (real-world fit).
    • Know calibration is parameter tuning, while accreditation is stakeholder approval.
  2. Tests:

    • For Poker/Runs/Chi-Square, show:
      • Step-by-step calculations (like the worked examples above).
      • Whether to reject/fail to reject H₀ and why.
    • Example Short Answer:

      "The Poker test checks uniformity by grouping random numbers into bins and comparing observed vs. expected counts using Chi-Square. For 1000 four-digit numbers, if the Chi-Square statistic (e.g., 12.5) is less than the critical value (16.92 at α=0.05), we accept uniformity."

  3. Real-World Applications:

    • Tie examples to Nepalese companies (eSewa, Ncell, Daraz) or everyday scenarios (traffic, queues).
    • Example:

      "A Pokhara University exam hall simulation would validate arrival rates by comparing simulated vs. real student arrival times during the TU entrance exam. If the Chi-Square test shows a p-value > 0.05, the model’s arrival distribution is accepted."

  4. Output Analysis:

    • Always state the confidence level (e.g., 95%) and interpret the CI (e.g., "we are 95% confident the true wait time is between 2.29 and 2.71 minutes").
    • For hypothesis tests, write the null/alternative hypotheses and decision rule.
  5. Common Pitfalls:

    • Mixing verification/validation: Verification is about the model’s internals; validation is about its realism.
    • Ignoring assumptions: Always state assumptions (e.g., "we assume Poisson arrivals for the NTC call center").
    • Overfitting: Calibration should use separate data from validation (e.g., calibrate on 2019 data, validate on 2020).

Quick Revision Table

Concept Definition Method/Tool Example
Verification Is the model built correctly? Code review, tracing, extreme-case tests Checking if a GPSS model’s GENERATE block uses the right λ.
Validation Does the model represent reality? Chi-Square, Runs test, expert judgment Comparing simulated vs. real Ncell call wait times.
Calibration Adjusting parameters to match real data. Iterative testing Increasing λ in a traffic model to match real congestion.
Accreditation Official approval by stakeholders. Documentation, sign-off TU approving a model for NEPSE use.
Poker Test Checks uniformity of random numbers. Chi-Square Testing eSewa’s transaction ID generation.
Runs Test Checks independence of random numbers. Binary conversion, run counting Validating Khalti’s OTP randomness.
Confidence Interval Estimates a metric’s range with uncertainty. t-distribution "95% CI for Ncell wait time: [2.29, 2.71] min."

Final Worked Example: Full V&V Process

Scenario: Simulate a single-server bank ATM queue (e.g., NMB Bank at Thapathali).

  1. Conceptual Model:

    • Arrivals: Poisson (λ = 10/hour).
    • Service: Exponential (μ = 12/hour).
    • Queue discipline: FIFO.
  2. Verification:

    • Check GPSS code:
      GENERATE 6, EXPO, 10  // λ = 10/hour
      QUEUE ATM
      SEIZE ATM_MACHINE
      ADVANCE 5, EXPO, 12   // μ = 12/hour
      RELEASE ATM_MACHINE
      TERMINATE 1
      
    • Trace: Simulate 1 customer → confirm they wait 0.5 hours on average.
  3. Validation:

    • Real data: Average wait time = 4.5 minutes (from bank records).
    • Simulated mean wait time = 5 minutes.
    • Chi-Square Test: Compare observed vs. simulated wait time bins.
      • Result: p-value = 0.12 > 0.05 → model is valid.
  4. Calibration:

    • Simulated wait time > real data → increase μ (faster service).
    • New μ = 15/hour → simulated wait time = 4.2 minutes (matches real data).
  5. Output Analysis:

    • Run 30 replications → 95% CI for wait time: [3.8, 4.6] minutes.
    • Conclusion: The ATM can handle the current load, but adding a second machine would reduce wait times further.

3.63.844.24.44.64.812345xLower 95% CIUpper 95% CIMean Wait Time95% CI Lower95% CI UpperMean
Confidence interval for ATM wait time (30 replications): [3.8, 4.6] minutes.

Based on the TU BSc CSIT syllabus for Simulation and Modeling (CSC317), unit 8.

Discussion

Loading…