Elective Data Analysis and Modeling

Data Analysis and ModelingUnit 515 min read

Hypothesis Testing: Types, Tests & Decision Rules

Unit 5 of Data Analysis and Modeling explores how to make data-driven decisions by testing assumptions (hypotheses) about populations using sample data, covering null/alternative hypotheses, test statistics, p-values, significance levels, and common tests (z-test, t-test, chi-square, ANOVA) with real-world applications

TAKEAWAYS:

  • Hypothesis testing is the foundation for making statistically valid decisions by comparing observed data against a null hypothesis (e.g., "no effect" or "no difference").
  • Type I (α) and Type II (β) errors are inevitable trade-offs: rejecting a true null hypothesis (false positive) vs. failing to reject a false null hypothesis (false negative).
  • Test statistics (z, t, F, χ²) quantify how extreme sample results are under the null hypothesis, while p-values measure the probability of observing data as extreme as—or more extreme than—the sample if the null is true.
  • One-tailed vs. two-tailed tests depend on whether the alternative hypothesis specifies a direction (e.g., "increase" or "decrease") or not.
  • Confidence intervals provide a range of plausible values for a parameter (e.g., mean, proportion) and are directly linked to hypothesis tests (e.g., reject H₀ if the interval excludes the hypothesized value).
  • Real-world applications include A/B testing in apps (e.g., Pathao’s ride-sharing algorithms), quality control in manufacturing (e.g., NTC’s network reliability tests), and financial fraud detection (e.g., bank transaction anomaly checks).

1. Core Concepts: Hypotheses and Errors

Hypothesis testing starts with two competing statements about a population parameter:

  • Null hypothesis (H₀): A default assumption of "no effect" or "no difference" (e.g., "The new marketing campaign has no impact on sales").
  • Alternative hypothesis (H₁ or Ha): The claim we seek evidence for (e.g., "Sales increase after the campaign").

How Decisions Work

We use sample data to compute a test statistic (e.g., z-score, t-score) and compare it to a critical value (from distribution tables) or compute a p-value. The decision rule:

  • If p-value ≤ α (significance level, e.g., 0.05), reject H₀.
  • Otherwise, fail to reject H₀ (not "accept H₀").

Errors in Hypothesis Testing

Two types of errors can occur:

Error Type Definition Probability Notation Real-World Example
Type I (α) Rejecting a true H₀ (false positive) α = P(Type I error) Pathao’s algorithm flags a normal ride as fraudulent, causing customer complaints.
Type II (β) Failing to reject a false H₀ (false negative) β = P(Type II error) A bank fails to detect fraudulent transactions because the anomaly threshold was too strict.
Power (1 − β) Probability of correctly rejecting a false H₀ (true positive rate) Power = 1 − β NTC’s network tests correctly identify a failing tower before outages occur.

Key Insight: You cannot eliminate both errors simultaneously. Reducing α (e.g., from 0.05 to 0.01) decreases Type I errors but increases Type II errors.


graph TD
    A["Start: Define H₀ and H₁"] --> B["Collect sample data"]
    B --> C["Calculate test statistic (e.g., z, t)"]
    C --> D["Compute p-value or compare to critical value"]
    D -->|"p ≤ α"| E["Reject H₀\n(Conclude evidence supports H₁)"]
    D -->|"p > α"| F["Fail to reject H₀\n(Insufficient evidence)"]
    E --> G["Risk: Type I error if H₀ was true"]
    F --> H["Risk: Type II error if H₁ was true"]

Example: Suppose Ncell claims its new network upgrade increases average download speed from 20 Mbps (H₀: μ = 20) to 25 Mbps (H₁: μ > 20). A sample of 50 users shows a mean of 22 Mbps with a standard deviation of 5 Mbps. Using α = 0.05:

  1. Test statistic (t-test):
  2. Critical value (df = 49, one-tailed, α = 0.05): ~1.677.
  3. Decision: Since 2.83 > 1.677, reject H₀. Conclusion: Evidence suggests the upgrade improved speed.

2. Types of Hypothesis Tests

The choice of test depends on the data type, population distribution, and sample size. Below is a decision flowchart:

flowchart TD
    A["Is the data categorical or numerical?"]
    A -->|"Numerical"| B["Is the population standard deviation σ known?"]
    B -->|"Yes"| C["Use z-test"]
    B -->|"No"| D["Use t-test"]
    A -->|"Categorical"| E["Is it a goodness-of-fit or test of independence?"]
    E -->|"Goodness-of-fit"| F["Use χ² test"]
    E -->|"Test of independence"| G["Use χ² test"]
    D --> H["Is the sample size >30 or population normal?"]
    H -->|"Yes"| I["One-sample t-test"]
    H -->|"No"| J["Non-parametric test (e.g., Wilcoxon)"]
    I --> K["Compare means of two groups?"]
    K -->|"Yes"| L["Independent samples?\nUse independent t-test"]
    K -->|"No"| M["Paired samples?\nUse paired t-test"]

Common Tests

Test When to Use Assumptions Formula
z-test Numerical data, σ known, large sample (n ≥ 30) Population normal or n ≥ 30
t-test Numerical data, σ unknown, small sample (n < 30) Population normal or n ≥ 30 (for one-sample t-test)
Chi-square (χ²) test Categorical data (goodness-of-fit or independence) Expected frequencies ≥5 in each category
ANOVA Compare means of ≥3 groups (numerical data) Normality, homogeneity of variance

Example (Chi-square): Daraz wants to test if customer reviews are evenly distributed across 5-star, 4-star, 3-star, 2-star, and 1-star ratings. Observed data: 120, 80, 60, 30, 10. Expected (uniform): 60 each.

  1. Compute χ²:
  2. Critical value (df = 4, α = 0.05): 9.49.
  3. Decision: Since 100 > 9.49, reject H₀. Conclusion: Ratings are not uniformly distributed.

3. p-values vs. Critical Values

  • p-value: Probability of observing data as extreme as—or more extreme than—the sample, assuming H₀ is true.
    • Small p-value (≤ α): Strong evidence against H₀.
    • Large p-value (> α): Weak evidence against H₀.
  • Critical value: Threshold from the test statistic’s distribution (e.g., z-table, t-table) that divides the "reject" and "fail to reject" regions.

Example (p-value): Khalti tests if the average transaction time (H₀: μ = 10 sec) has decreased (H₁: μ < 10). Sample: n = 40, sec, s = 1.2 sec.

  1. Test statistic (t-test):
  2. p-value (one-tailed, df = 39): P(t < −2.04) ≈ 0.024.
  3. Decision: Since 0.024 ≤ 0.05, reject H₀. Conclusion: Transaction time has significantly decreased.

graph LR
    A["p-value Approach"] --> B["Compute p-value from test statistic"]
    B -->|"p ≤ α"| C["Reject H₀"]
    B -->|"p > α"| D["Fail to reject H₀"]
    E["Critical Value Approach"] --> F["Find critical value from table"]
    F -->|"Test statistic > critical value"| G["Reject H₀"]
    F -->|"Test statistic ≤ critical value"| H["Fail to reject H₀"]
    I["Both approaches are equivalent!"]

4. Confidence Intervals and Hypothesis Testing

Confidence intervals (CIs) provide a range of plausible values for a parameter and are directly linked to hypothesis tests:

  • If the CI for a parameter excludes the hypothesized value (H₀ value), reject H₀.
  • If the CI includes the H₀ value, fail to reject H₀.

Example (CI): NEPSE wants to test if the average return on a stock is 8% (H₀: μ = 8). Sample: n = 25, , s = 2.5.

  1. 95% CI for μ:
  2. Decision: Since the CI [6.49, 8.51] includes 8, fail to reject H₀.

graph TD
    A["95% CI: [6.49, 8.51]"]
    B["H₀: μ = 8"]
    C["Is 8 in [6.49, 8.51]?"]
    C -->|"Yes"| D["Fail to reject H₀"]
    C -->|"No"| E["Reject H₀"]

5. One-Tailed vs. Two-Tailed Tests

Test Type Alternative Hypothesis (H₁) When to Use Example
One-tailed μ > μ₀ or μ < μ₀ Directional claim (e.g., "increase" or "decrease") Pathao’s algorithm tests if ride time decreases after a new route.
Two-tailed μ ≠ μ₀ Non-directional claim (e.g., "different") NTC tests if network latency differs from the standard.

Example (One-tailed): A bank tests if the average loan approval time (H₀: μ = 30 days) has decreased (H₁: μ < 30). Sample: n = 36, days, s = 5 days.

  1. Test statistic (t-test):
  2. Critical value (one-tailed, df = 35, α = 0.05): −1.69.
  3. Decision: Since −2.4 < −1.69, reject H₀. Conclusion: Approval time has significantly decreased.

6. Real-World Applications

1. A/B Testing in Apps (Pathao, Daraz)

  • Idea Used: Two-sample t-test to compare user engagement between two versions of an app feature.
  • How It Works:
    • H₀: Version A and Version B have equal click-through rates (CTR).
    • H₁: Version B’s CTR > Version A’s CTR.
    • Example: Pathao tests a new UI button color. After 10,000 rides:
      • Version A CTR: 12% (n = 5,000)
      • Version B CTR: 14% (n = 5,000)
      • Test statistic (z-test):
      • p-value: P(z > 3.46) ≈ 0.0003.
      • Decision: Reject H₀. Deploy Version B.

2. Quality Control in Manufacturing (NTC, Ncell)

  • Idea Used: Chi-square goodness-of-fit test to ensure product defects meet expected rates.
  • How It Works:
    • H₀: Defect rates follow the expected distribution (e.g., 1% major, 3% minor, 96% good).
    • Example: NTC tests 200 network towers for failures:
      • Observed: 5 major, 8 minor, 187 good.
      • Expected: 2 major, 6 minor, 192 good.
      • χ² statistic: 4.5.
      • Critical value (df = 2, α = 0.05): 5.99.
      • Decision: Fail to reject H₀. Defect rates are acceptable.

3. Financial Fraud Detection (Banks)

  • Idea Used: Z-test for proportions to detect unusual transaction patterns.
  • How It Works:
    • H₀: The proportion of fraudulent transactions is ≤ 0.5%.
    • Example: A bank observes 12 fraudulent transactions in 2,000 (p̂ = 0.006).
    • Test statistic:
    • p-value (two-tailed): 0.53.
    • Decision: Fail to reject H₀. No significant increase in fraud (yet).

7. Common Mistakes to Avoid

  1. Assuming "fail to reject H₀" means H₀ is true: It only means insufficient evidence against H₀.
  2. Ignoring assumptions: Non-normal data with small samples require non-parametric tests (e.g., Mann-Whitney U).
  3. Using one-tailed tests without justification: Only use directional tests if theory supports it.
  4. P-hacking: Manipulating data or tests to achieve p ≤ 0.05 (e.g., cherry-picking samples).
  5. Confusing statistical significance (p ≤ 0.05) with practical significance: A result may be "significant" but trivial in real-world terms (e.g., a 0.1% increase in sales).

8. Step-by-Step Hypothesis Testing Workflow

flowchart TD
    A["1. State H₀ and H₁"] --> B["2. Choose significance level (α)"]
    B --> C["3. Select test (z, t, χ², etc.)"]
    C --> D["4. Calculate test statistic"]
    D --> E["5. Determine critical value or p-value"]
    E --> F["6. Make decision (reject/fail to reject H₀)"]
    F --> G["7. Draw conclusion in context"]
    G --> H["8. Check assumptions and report limitations"]

Worked Example (ANOVA): A university compares the average exam scores of students taught by 3 different professors (H₀: μ₁ = μ₂ = μ₃).

  • Data:
    • Professor A: [85, 90, 78, 88] (n₁ = 4, )
    • Professor B: [76, 82, 80, 79] (n₂ = 4, )
    • Professor C: [92, 88, 90, 95] (n₃ = 4, )
  • Step 1: Compute between-group variance (SSB) and within-group variance (SSW).
  • Step 2: Calculate F-statistic: (Assume SSB = 500, SSW = 200, k = 3 groups, N = 12 total students).
  • Step 3: Critical F-value (df₁ = 2, df₂ = 9, α = 0.05): 4.26.
  • Decision: Since 11.25 > 4.26, reject H₀. Conclusion: Not all professors have equal average scores.

Exam Tip

  1. Always state H₀ and H₁ clearly in words and symbols (e.g., "H₀: μ = 50 vs. H₁: μ ≠ 50").
  2. Show all steps: Test statistic → critical value/p-value → decision → conclusion.
  3. Interpret results in context: Avoid vague statements like "reject H₀." Instead, say:
    • "There is sufficient evidence at the 5% level to conclude that [specific effect]."
  4. Watch assumptions: If data is non-normal, mention non-parametric alternatives (e.g., Wilcoxon rank-sum test).
  5. Practice with real data: Use datasets from Nepal’s NEPSE stock prices, NTC’s network latency logs, or Khalti’s transaction times to build intuition.
  6. Common exam traps:
    • Directionality: Two-tailed tests require α/2 in each tail.
    • Degrees of freedom: For t-tests, df = n − 1; for ANOVA, df₁ = k − 1, df₂ = N − k.
    • Effect size: Report confidence intervals or Cohen’s d (for t-tests) to show practical significance.

Final Visual Summary:

mindmap
  root((Hypothesis Testing))
    Concepts
      Null Hypothesis (H₀)
      Alternative Hypothesis (H₁)
      Test Statistic (z, t, χ², F)
      p-value
      Significance Level (α)
    Errors
      Type I (α)
      Type II (β)
      Power (1 − β)
    Tests
      z-test
      t-test
      Chi-square (χ²)
      ANOVA
    Applications
      A/B Testing (Pathao, Daraz)
      Quality Control (NTC, Ncell)
      Fraud Detection (Banks)
    Workflow
      State H₀/H₁
      Choose α
      Select test
      Compute statistic
      Compare to critical value/p-value
      Decide & conclude

Based on the PU BBA (PU) syllabus for Data Analysis and Modeling, unit 5.

Discussion

Loading…