Probability and StatisticsUnit 89 min read
Hypothesis Testing: Types, Tests & Decision Rules
Unit 8 of Probability and Statistics covers hypothesis testing—how to make data-driven decisions by comparing sample data to assumptions (null/alternative hypotheses), calculating test statistics, and determining significance using p-values or critical values.
TAKEAWAYS:
- Hypothesis testing follows a structured process: state hypotheses, choose a test, compute test statistics, and make a decision based on significance.
- Type I (α) and Type II (β) errors are inevitable trade-offs in decision-making, with α controlled by significance level (e.g., 0.05).
- Parametric tests (e.g., t-test, z-test) assume normality, while non-parametric tests (e.g., chi-square, Mann-Whitney U) do not.
- p-values quantify evidence against the null hypothesis; smaller p-values (≤ α) reject .
- Confidence intervals provide a range for population parameters (e.g., mean) and indirectly test hypotheses.
- Real-world applications include A/B testing (e.g., Daraz’s ad click rates), quality control (e.g., NTC’s call-drop rates), and medical trials (e.g., vaccine efficacy).
1. Introduction to Hypothesis Testing
Hypothesis testing is a statistical method to make inferences about a population based on sample data. It answers questions like:
- Is the new drug more effective than the existing one?
- Does Pathao’s new delivery route reduce delays?
- Is NEPSE’s stock return significantly different from 10%?
Key Definitions
- Null Hypothesis (): Default assumption (e.g., "no effect," "no difference"). Tested for evidence against it.
- Alternative Hypothesis ( or ): Claim we suspect is true (e.g., "the drug works," "the route is faster").
- Test Statistic: Standardized value (e.g., z-score, t-score) comparing sample data to .
- Significance Level (α): Probability of rejecting when it’s true (commonly 0.05). Controls Type I error.
- p-value: Probability of observing data as extreme as the sample, assuming is true. Small p-value (≤ α) → Reject .
Steps in Hypothesis Testing
flowchart TD
A["1. State Hypotheses"] --> B["2. Choose Significance Level (α)"]
B --> C["3. Select Test Statistic (z, t, χ², etc.)"]
C --> D["4. Compute Test Statistic from Sample Data"]
D --> E["5. Determine Critical Value or p-value"]
E --> F["6. Make Decision: Reject/Fail to Reject \(H_0\)"]
F --> G["7. Draw Conclusion"]Example 1: Testing NEPSE’s Stock Return Assume NEPSE’s historical return is 10%. A broker claims a new fund yields 12%.
- (one-tailed test)
- Sample: 30 days, mean return = 11.5%, std dev = 2%.
- Test: One-sample z-test (since , population std dev known).
- Calculation:
- p-value = (from z-table).
- Decision: Since , reject . The fund’s return is significantly higher.
2. Types of Hypothesis Tests
Tests are classified by:
- Parameter Tested: Mean (-test, -test), proportion (z-test), variance (-test).
- Sample Size: Large (, use -test) vs. small (, use -test).
- Tail of Test:
- One-tailed: Directional claim (e.g., "increase" or "decrease").
- Two-tailed: Non-directional claim (e.g., "different").
Comparison Table: Common Tests
| Test | When to Use | Assumptions | Test Statistic |
|---|---|---|---|
| z-test | Large sample (), known σ | Normal population or CLT applies | |
| t-test | Small sample (), unknown σ | Normal population | (df = ) |
| Chi-square () | Categorical data (goodness-of-fit, independence) | Expected frequencies > 5 | |
| ANOVA | Compare means of ≥3 groups | Normality, homogeneity of variance |
Example 2: Daraz’s A/B Test for Ad Click Rates Daraz tests two ad designs:
- Group A: Old ad → 5% click rate ().
- Group B: New ad → 6% click rate ().
- Question: Is the new ad significantly better?
- Test: Two-proportion z-test (since samples are large and independent).
- Calculation:
- p-value (two-tailed) ≈ 0.159.
- Decision: Fail to reject (p > 0.05). No significant difference.
3. Errors in Hypothesis Testing
Two types of errors arise from incorrect decisions:
| Error Type | Definition | Probability Notation | Consequence |
|---|---|---|---|
| Type I (α) | Reject when it’s true | False alarm (e.g., banning a safe drug) | |
| Type II (β) | Fail to reject when it’s false | Missed opportunity (e.g., ignoring a better ad) |
- Power of a Test = (probability of correctly rejecting when false).
- Trade-off: Reducing α increases β, and vice versa.
Example 3: NTC’s Call-Drop Rate Test NTC claims call-drop rate is <1% in a new tower. A sample of 500 calls shows 3 drops.
- (one-tailed)
- Test: One-proportion z-test.
- Calculation:
- p-value = (one-tailed).
- Decision: Fail to reject . Insufficient evidence to claim the rate is higher.
4. Decision Rules: Critical Values vs. p-values
Two methods to decide:
Critical Value Approach:
- Compare test statistic to critical value (from z-table/t-table).
- Example: For α = 0.05 (two-tailed), critical z = ±1.96.
- If , reject .
p-value Approach:
- Compare p-value to α.
- If p ≤ α, reject .
Example 4: Pathao’s Delivery Time Reduction Pathao claims a new algorithm reduces delivery time from 45 to 40 minutes.
- Sample: 40 deliveries, mean = 39 min, std dev = 5 min.
- , (one-tailed).
- Test: One-sample t-test (small , unknown σ).
- Calculation:
- Critical t-value (α = 0.05, one-tailed) ≈ -1.684.
- Decision: → Reject . The algorithm works.
5. Confidence Intervals and Hypothesis Testing
Confidence intervals (CIs) provide a range for population parameters and indirectly test hypotheses:
- If the 95% CI for does not include , reject at α = 0.05.
- Example: For NEPSE’s return (Example 1), the 95% CI is (10.5%, 12.5%).
- Since 10% is not in this range, we reject .
Visualization:
6. Real-World Applications
In the Real World
eSewa’s Fraud Detection:
- Uses chi-square tests to detect unusual transaction patterns (e.g., sudden spikes in small-value payments).
- Idea: Compare observed vs. expected transaction frequencies across user segments.
Khalti’s Promotional Discount Testing:
- Runs A/B tests (two-sample t-tests) to compare conversion rates between discount offers (e.g., 10% vs. 15% off).
- Example: If Group A (10% off) has 5% conversion () and Group B (15% off) has 6% (), a z-test checks if the difference is significant.
NTC’s Network Reliability:
- Uses hypothesis tests on call-drop rates (Example 3) to validate claims about new tower performance.
- Idea: Ensure Type I error (false alarm) is minimized to avoid unnecessary maintenance costs.
Exam Tip
- Always state hypotheses clearly ( and ) and specify the test type (one-tailed/two-tailed).
- Show calculations step-by-step, especially the test statistic formula. Partial credit is given for correct setup.
- Interpret p-values correctly:
- "p = 0.03" means there’s a 3% chance of observing the data if were true.
- Never say "p is the probability is true."
- For non-parametric tests (e.g., Mann-Whitney U), state why parametric tests aren’t appropriate (e.g., non-normal data).
- Common pitfalls:
- Using a t-test when (use z-test).
- Ignoring assumptions (e.g., normality for t-tests).
- Misinterpreting p-values (e.g., confusing statistical significance with practical importance).
- Practical questions often involve real scenarios (e.g., "A bank claims its loan default rate is 5%. Test this claim given sample data."). Relate your answer to the context.
Key Formula Summary:
mindmap
root((Hypothesis Testing))
z-test
formula["z = (x̄ - μ₀)/(σ/√n)"]
use["Large n, known σ"]
t-test
formula["t = (x̄ - μ₀)/(s/√n)"]
use["Small n, unknown σ"]
chi-square
formula["χ² = Σ(O - E)²/E"]
use["Categorical data"]
p-value
rule["p ≤ α → Reject H₀"]
errors
typeI["α: False positive"]
typeII["β: False negative"]Based on the PU BE Computer (PU) syllabus for Probability and Statistics (MTH216), unit 8.
Discussion
Loading…