MTH216 Probability and Statistics

Probability and StatisticsUnit 89 min read

Hypothesis Testing: Types, Tests & Decision Rules

Unit 8 of Probability and Statistics covers hypothesis testing—how to make data-driven decisions by comparing sample data to assumptions (null/alternative hypotheses), calculating test statistics, and determining significance using p-values or critical values.

TAKEAWAYS:

  • Hypothesis testing follows a structured process: state hypotheses, choose a test, compute test statistics, and make a decision based on significance.
  • Type I (α) and Type II (β) errors are inevitable trade-offs in decision-making, with α controlled by significance level (e.g., 0.05).
  • Parametric tests (e.g., t-test, z-test) assume normality, while non-parametric tests (e.g., chi-square, Mann-Whitney U) do not.
  • p-values quantify evidence against the null hypothesis; smaller p-values (≤ α) reject .
  • Confidence intervals provide a range for population parameters (e.g., mean) and indirectly test hypotheses.
  • Real-world applications include A/B testing (e.g., Daraz’s ad click rates), quality control (e.g., NTC’s call-drop rates), and medical trials (e.g., vaccine efficacy).

1. Introduction to Hypothesis Testing

Hypothesis testing is a statistical method to make inferences about a population based on sample data. It answers questions like:

  • Is the new drug more effective than the existing one?
  • Does Pathao’s new delivery route reduce delays?
  • Is NEPSE’s stock return significantly different from 10%?

Key Definitions

  • Null Hypothesis (): Default assumption (e.g., "no effect," "no difference"). Tested for evidence against it.
  • Alternative Hypothesis ( or ): Claim we suspect is true (e.g., "the drug works," "the route is faster").
  • Test Statistic: Standardized value (e.g., z-score, t-score) comparing sample data to .
  • Significance Level (α): Probability of rejecting when it’s true (commonly 0.05). Controls Type I error.
  • p-value: Probability of observing data as extreme as the sample, assuming is true. Small p-value (≤ α) → Reject .

Steps in Hypothesis Testing

flowchart TD
    A["1. State Hypotheses"] --> B["2. Choose Significance Level (α)"]
    B --> C["3. Select Test Statistic (z, t, χ², etc.)"]
    C --> D["4. Compute Test Statistic from Sample Data"]
    D --> E["5. Determine Critical Value or p-value"]
    E --> F["6. Make Decision: Reject/Fail to Reject \(H_0\)"]
    F --> G["7. Draw Conclusion"]

Example 1: Testing NEPSE’s Stock Return Assume NEPSE’s historical return is 10%. A broker claims a new fund yields 12%.

  • (one-tailed test)
  • Sample: 30 days, mean return = 11.5%, std dev = 2%.
  • Test: One-sample z-test (since , population std dev known).
  • Calculation:
    • p-value = (from z-table).
    • Decision: Since , reject . The fund’s return is significantly higher.

2. Types of Hypothesis Tests

Tests are classified by:

  1. Parameter Tested: Mean (-test, -test), proportion (z-test), variance (-test).
  2. Sample Size: Large (, use -test) vs. small (, use -test).
  3. Tail of Test:
    • One-tailed: Directional claim (e.g., "increase" or "decrease").
    • Two-tailed: Non-directional claim (e.g., "different").

Comparison Table: Common Tests

Test When to Use Assumptions Test Statistic
z-test Large sample (), known σ Normal population or CLT applies
t-test Small sample (), unknown σ Normal population (df = )
Chi-square () Categorical data (goodness-of-fit, independence) Expected frequencies > 5
ANOVA Compare means of ≥3 groups Normality, homogeneity of variance

Example 2: Daraz’s A/B Test for Ad Click Rates Daraz tests two ad designs:

  • Group A: Old ad → 5% click rate ().
  • Group B: New ad → 6% click rate ().
  • Question: Is the new ad significantly better?
  • Test: Two-proportion z-test (since samples are large and independent).
  • Calculation:
    • p-value (two-tailed) ≈ 0.159.
    • Decision: Fail to reject (p > 0.05). No significant difference.

3. Errors in Hypothesis Testing

Two types of errors arise from incorrect decisions:

Error Type Definition Probability Notation Consequence
Type I (α) Reject when it’s true False alarm (e.g., banning a safe drug)
Type II (β) Fail to reject when it’s false Missed opportunity (e.g., ignoring a better ad)
  • Power of a Test = (probability of correctly rejecting when false).
  • Trade-off: Reducing α increases β, and vice versa.

Example 3: NTC’s Call-Drop Rate Test NTC claims call-drop rate is <1% in a new tower. A sample of 500 calls shows 3 drops.

  • (one-tailed)
  • Test: One-proportion z-test.
  • Calculation:
    • p-value = (one-tailed).
    • Decision: Fail to reject . Insufficient evidence to claim the rate is higher.

4. Decision Rules: Critical Values vs. p-values

Two methods to decide:

  1. Critical Value Approach:

    • Compare test statistic to critical value (from z-table/t-table).
    • Example: For α = 0.05 (two-tailed), critical z = ±1.96.
    • If , reject .
  2. p-value Approach:

    • Compare p-value to α.
    • If p ≤ α, reject .

Example 4: Pathao’s Delivery Time Reduction Pathao claims a new algorithm reduces delivery time from 45 to 40 minutes.

  • Sample: 40 deliveries, mean = 39 min, std dev = 5 min.
  • , (one-tailed).
  • Test: One-sample t-test (small , unknown σ).
  • Calculation:
    • Critical t-value (α = 0.05, one-tailed) ≈ -1.684.
    • Decision: → Reject . The algorithm works.

5. Confidence Intervals and Hypothesis Testing

Confidence intervals (CIs) provide a range for population parameters and indirectly test hypotheses:

  • If the 95% CI for does not include , reject at α = 0.05.
  • Example: For NEPSE’s return (Example 1), the 95% CI is (10.5%, 12.5%).
    • Since 10% is not in this range, we reject .

Visualization:


6. Real-World Applications

In the Real World

  1. eSewa’s Fraud Detection:

    • Uses chi-square tests to detect unusual transaction patterns (e.g., sudden spikes in small-value payments).
    • Idea: Compare observed vs. expected transaction frequencies across user segments.
  2. Khalti’s Promotional Discount Testing:

    • Runs A/B tests (two-sample t-tests) to compare conversion rates between discount offers (e.g., 10% vs. 15% off).
    • Example: If Group A (10% off) has 5% conversion () and Group B (15% off) has 6% (), a z-test checks if the difference is significant.
  3. NTC’s Network Reliability:

    • Uses hypothesis tests on call-drop rates (Example 3) to validate claims about new tower performance.
    • Idea: Ensure Type I error (false alarm) is minimized to avoid unnecessary maintenance costs.

Exam Tip

  1. Always state hypotheses clearly ( and ) and specify the test type (one-tailed/two-tailed).
  2. Show calculations step-by-step, especially the test statistic formula. Partial credit is given for correct setup.
  3. Interpret p-values correctly:
    • "p = 0.03" means there’s a 3% chance of observing the data if were true.
    • Never say "p is the probability is true."
  4. For non-parametric tests (e.g., Mann-Whitney U), state why parametric tests aren’t appropriate (e.g., non-normal data).
  5. Common pitfalls:
    • Using a t-test when (use z-test).
    • Ignoring assumptions (e.g., normality for t-tests).
    • Misinterpreting p-values (e.g., confusing statistical significance with practical importance).
  6. Practical questions often involve real scenarios (e.g., "A bank claims its loan default rate is 5%. Test this claim given sample data."). Relate your answer to the context.

Key Formula Summary:

mindmap
  root((Hypothesis Testing))
    z-test
      formula["z = (x̄ - μ₀)/(σ/√n)"]
      use["Large n, known σ"]
    t-test
      formula["t = (x̄ - μ₀)/(s/√n)"]
      use["Small n, unknown σ"]
    chi-square
      formula["χ² = Σ(O - E)²/E"]
      use["Categorical data"]
    p-value
      rule["p ≤ α → Reject H₀"]
    errors
      typeI["α: False positive"]
      typeII["β: False negative"]

Based on the PU BE Computer (PU) syllabus for Probability and Statistics (MTH216), unit 8.

Discussion

Loading…