Data Analysis and ModelingUnit 515 min read
Hypothesis Testing: Types, Tests & Decision Rules
Unit 5 of Data Analysis and Modeling explores how to make data-driven decisions by testing assumptions (hypotheses) about populations using sample data, covering null/alternative hypotheses, test statistics, p-values, significance levels, and common tests (z-test, t-test, chi-square, ANOVA) with real-world applications
TAKEAWAYS:
- Hypothesis testing is the foundation for making statistically valid decisions by comparing observed data against a null hypothesis (e.g., "no effect" or "no difference").
- Type I (α) and Type II (β) errors are inevitable trade-offs: rejecting a true null hypothesis (false positive) vs. failing to reject a false null hypothesis (false negative).
- Test statistics (z, t, F, χ²) quantify how extreme sample results are under the null hypothesis, while p-values measure the probability of observing data as extreme as—or more extreme than—the sample if the null is true.
- One-tailed vs. two-tailed tests depend on whether the alternative hypothesis specifies a direction (e.g., "increase" or "decrease") or not.
- Confidence intervals provide a range of plausible values for a parameter (e.g., mean, proportion) and are directly linked to hypothesis tests (e.g., reject H₀ if the interval excludes the hypothesized value).
- Real-world applications include A/B testing in apps (e.g., Pathao’s ride-sharing algorithms), quality control in manufacturing (e.g., NTC’s network reliability tests), and financial fraud detection (e.g., bank transaction anomaly checks).
1. Core Concepts: Hypotheses and Errors
Hypothesis testing starts with two competing statements about a population parameter:
- Null hypothesis (H₀): A default assumption of "no effect" or "no difference" (e.g., "The new marketing campaign has no impact on sales").
- Alternative hypothesis (H₁ or Ha): The claim we seek evidence for (e.g., "Sales increase after the campaign").
How Decisions Work
We use sample data to compute a test statistic (e.g., z-score, t-score) and compare it to a critical value (from distribution tables) or compute a p-value. The decision rule:
- If p-value ≤ α (significance level, e.g., 0.05), reject H₀.
- Otherwise, fail to reject H₀ (not "accept H₀").
Errors in Hypothesis Testing
Two types of errors can occur:
| Error Type | Definition | Probability Notation | Real-World Example |
|---|---|---|---|
| Type I (α) | Rejecting a true H₀ (false positive) | α = P(Type I error) | Pathao’s algorithm flags a normal ride as fraudulent, causing customer complaints. |
| Type II (β) | Failing to reject a false H₀ (false negative) | β = P(Type II error) | A bank fails to detect fraudulent transactions because the anomaly threshold was too strict. |
| Power (1 − β) | Probability of correctly rejecting a false H₀ (true positive rate) | Power = 1 − β | NTC’s network tests correctly identify a failing tower before outages occur. |
Key Insight: You cannot eliminate both errors simultaneously. Reducing α (e.g., from 0.05 to 0.01) decreases Type I errors but increases Type II errors.
graph TD
A["Start: Define H₀ and H₁"] --> B["Collect sample data"]
B --> C["Calculate test statistic (e.g., z, t)"]
C --> D["Compute p-value or compare to critical value"]
D -->|"p ≤ α"| E["Reject H₀\n(Conclude evidence supports H₁)"]
D -->|"p > α"| F["Fail to reject H₀\n(Insufficient evidence)"]
E --> G["Risk: Type I error if H₀ was true"]
F --> H["Risk: Type II error if H₁ was true"]Example: Suppose Ncell claims its new network upgrade increases average download speed from 20 Mbps (H₀: μ = 20) to 25 Mbps (H₁: μ > 20). A sample of 50 users shows a mean of 22 Mbps with a standard deviation of 5 Mbps. Using α = 0.05:
- Test statistic (t-test):
- Critical value (df = 49, one-tailed, α = 0.05): ~1.677.
- Decision: Since 2.83 > 1.677, reject H₀. Conclusion: Evidence suggests the upgrade improved speed.
2. Types of Hypothesis Tests
The choice of test depends on the data type, population distribution, and sample size. Below is a decision flowchart:
flowchart TD
A["Is the data categorical or numerical?"]
A -->|"Numerical"| B["Is the population standard deviation σ known?"]
B -->|"Yes"| C["Use z-test"]
B -->|"No"| D["Use t-test"]
A -->|"Categorical"| E["Is it a goodness-of-fit or test of independence?"]
E -->|"Goodness-of-fit"| F["Use χ² test"]
E -->|"Test of independence"| G["Use χ² test"]
D --> H["Is the sample size >30 or population normal?"]
H -->|"Yes"| I["One-sample t-test"]
H -->|"No"| J["Non-parametric test (e.g., Wilcoxon)"]
I --> K["Compare means of two groups?"]
K -->|"Yes"| L["Independent samples?\nUse independent t-test"]
K -->|"No"| M["Paired samples?\nUse paired t-test"]Common Tests
| Test | When to Use | Assumptions | Formula |
|---|---|---|---|
| z-test | Numerical data, σ known, large sample (n ≥ 30) | Population normal or n ≥ 30 | |
| t-test | Numerical data, σ unknown, small sample (n < 30) | Population normal or n ≥ 30 (for one-sample t-test) | |
| Chi-square (χ²) test | Categorical data (goodness-of-fit or independence) | Expected frequencies ≥5 in each category | |
| ANOVA | Compare means of ≥3 groups (numerical data) | Normality, homogeneity of variance |
Example (Chi-square): Daraz wants to test if customer reviews are evenly distributed across 5-star, 4-star, 3-star, 2-star, and 1-star ratings. Observed data: 120, 80, 60, 30, 10. Expected (uniform): 60 each.
- Compute χ²:
- Critical value (df = 4, α = 0.05): 9.49.
- Decision: Since 100 > 9.49, reject H₀. Conclusion: Ratings are not uniformly distributed.
3. p-values vs. Critical Values
- p-value: Probability of observing data as extreme as—or more extreme than—the sample, assuming H₀ is true.
- Small p-value (≤ α): Strong evidence against H₀.
- Large p-value (> α): Weak evidence against H₀.
- Critical value: Threshold from the test statistic’s distribution (e.g., z-table, t-table) that divides the "reject" and "fail to reject" regions.
Example (p-value): Khalti tests if the average transaction time (H₀: μ = 10 sec) has decreased (H₁: μ < 10). Sample: n = 40, sec, s = 1.2 sec.
- Test statistic (t-test):
- p-value (one-tailed, df = 39): P(t < −2.04) ≈ 0.024.
- Decision: Since 0.024 ≤ 0.05, reject H₀. Conclusion: Transaction time has significantly decreased.
graph LR
A["p-value Approach"] --> B["Compute p-value from test statistic"]
B -->|"p ≤ α"| C["Reject H₀"]
B -->|"p > α"| D["Fail to reject H₀"]
E["Critical Value Approach"] --> F["Find critical value from table"]
F -->|"Test statistic > critical value"| G["Reject H₀"]
F -->|"Test statistic ≤ critical value"| H["Fail to reject H₀"]
I["Both approaches are equivalent!"]4. Confidence Intervals and Hypothesis Testing
Confidence intervals (CIs) provide a range of plausible values for a parameter and are directly linked to hypothesis tests:
- If the CI for a parameter excludes the hypothesized value (H₀ value), reject H₀.
- If the CI includes the H₀ value, fail to reject H₀.
Example (CI): NEPSE wants to test if the average return on a stock is 8% (H₀: μ = 8). Sample: n = 25, , s = 2.5.
- 95% CI for μ:
- Decision: Since the CI [6.49, 8.51] includes 8, fail to reject H₀.
graph TD
A["95% CI: [6.49, 8.51]"]
B["H₀: μ = 8"]
C["Is 8 in [6.49, 8.51]?"]
C -->|"Yes"| D["Fail to reject H₀"]
C -->|"No"| E["Reject H₀"]5. One-Tailed vs. Two-Tailed Tests
| Test Type | Alternative Hypothesis (H₁) | When to Use | Example |
|---|---|---|---|
| One-tailed | μ > μ₀ or μ < μ₀ | Directional claim (e.g., "increase" or "decrease") | Pathao’s algorithm tests if ride time decreases after a new route. |
| Two-tailed | μ ≠ μ₀ | Non-directional claim (e.g., "different") | NTC tests if network latency differs from the standard. |
Example (One-tailed): A bank tests if the average loan approval time (H₀: μ = 30 days) has decreased (H₁: μ < 30). Sample: n = 36, days, s = 5 days.
- Test statistic (t-test):
- Critical value (one-tailed, df = 35, α = 0.05): −1.69.
- Decision: Since −2.4 < −1.69, reject H₀. Conclusion: Approval time has significantly decreased.
6. Real-World Applications
1. A/B Testing in Apps (Pathao, Daraz)
- Idea Used: Two-sample t-test to compare user engagement between two versions of an app feature.
- How It Works:
- H₀: Version A and Version B have equal click-through rates (CTR).
- H₁: Version B’s CTR > Version A’s CTR.
- Example: Pathao tests a new UI button color. After 10,000 rides:
- Version A CTR: 12% (n = 5,000)
- Version B CTR: 14% (n = 5,000)
- Test statistic (z-test):
- p-value: P(z > 3.46) ≈ 0.0003.
- Decision: Reject H₀. Deploy Version B.
2. Quality Control in Manufacturing (NTC, Ncell)
- Idea Used: Chi-square goodness-of-fit test to ensure product defects meet expected rates.
- How It Works:
- H₀: Defect rates follow the expected distribution (e.g., 1% major, 3% minor, 96% good).
- Example: NTC tests 200 network towers for failures:
- Observed: 5 major, 8 minor, 187 good.
- Expected: 2 major, 6 minor, 192 good.
- χ² statistic: 4.5.
- Critical value (df = 2, α = 0.05): 5.99.
- Decision: Fail to reject H₀. Defect rates are acceptable.
3. Financial Fraud Detection (Banks)
- Idea Used: Z-test for proportions to detect unusual transaction patterns.
- How It Works:
- H₀: The proportion of fraudulent transactions is ≤ 0.5%.
- Example: A bank observes 12 fraudulent transactions in 2,000 (p̂ = 0.006).
- Test statistic:
- p-value (two-tailed): 0.53.
- Decision: Fail to reject H₀. No significant increase in fraud (yet).
7. Common Mistakes to Avoid
- Assuming "fail to reject H₀" means H₀ is true: It only means insufficient evidence against H₀.
- Ignoring assumptions: Non-normal data with small samples require non-parametric tests (e.g., Mann-Whitney U).
- Using one-tailed tests without justification: Only use directional tests if theory supports it.
- P-hacking: Manipulating data or tests to achieve p ≤ 0.05 (e.g., cherry-picking samples).
- Confusing statistical significance (p ≤ 0.05) with practical significance: A result may be "significant" but trivial in real-world terms (e.g., a 0.1% increase in sales).
8. Step-by-Step Hypothesis Testing Workflow
flowchart TD
A["1. State H₀ and H₁"] --> B["2. Choose significance level (α)"]
B --> C["3. Select test (z, t, χ², etc.)"]
C --> D["4. Calculate test statistic"]
D --> E["5. Determine critical value or p-value"]
E --> F["6. Make decision (reject/fail to reject H₀)"]
F --> G["7. Draw conclusion in context"]
G --> H["8. Check assumptions and report limitations"]Worked Example (ANOVA): A university compares the average exam scores of students taught by 3 different professors (H₀: μ₁ = μ₂ = μ₃).
- Data:
- Professor A: [85, 90, 78, 88] (n₁ = 4, )
- Professor B: [76, 82, 80, 79] (n₂ = 4, )
- Professor C: [92, 88, 90, 95] (n₃ = 4, )
- Step 1: Compute between-group variance (SSB) and within-group variance (SSW).
- Step 2: Calculate F-statistic: (Assume SSB = 500, SSW = 200, k = 3 groups, N = 12 total students).
- Step 3: Critical F-value (df₁ = 2, df₂ = 9, α = 0.05): 4.26.
- Decision: Since 11.25 > 4.26, reject H₀. Conclusion: Not all professors have equal average scores.
Exam Tip
- Always state H₀ and H₁ clearly in words and symbols (e.g., "H₀: μ = 50 vs. H₁: μ ≠ 50").
- Show all steps: Test statistic → critical value/p-value → decision → conclusion.
- Interpret results in context: Avoid vague statements like "reject H₀." Instead, say:
- "There is sufficient evidence at the 5% level to conclude that [specific effect]."
- Watch assumptions: If data is non-normal, mention non-parametric alternatives (e.g., Wilcoxon rank-sum test).
- Practice with real data: Use datasets from Nepal’s NEPSE stock prices, NTC’s network latency logs, or Khalti’s transaction times to build intuition.
- Common exam traps:
- Directionality: Two-tailed tests require α/2 in each tail.
- Degrees of freedom: For t-tests, df = n − 1; for ANOVA, df₁ = k − 1, df₂ = N − k.
- Effect size: Report confidence intervals or Cohen’s d (for t-tests) to show practical significance.
Final Visual Summary:
mindmap
root((Hypothesis Testing))
Concepts
Null Hypothesis (H₀)
Alternative Hypothesis (H₁)
Test Statistic (z, t, χ², F)
p-value
Significance Level (α)
Errors
Type I (α)
Type II (β)
Power (1 − β)
Tests
z-test
t-test
Chi-square (χ²)
ANOVA
Applications
A/B Testing (Pathao, Daraz)
Quality Control (NTC, Ncell)
Fraud Detection (Banks)
Workflow
State H₀/H₁
Choose α
Select test
Compute statistic
Compare to critical value/p-value
Decide & concludeBased on the PU BBA (PU) syllabus for Data Analysis and Modeling, unit 5.
Discussion
Loading…