Business StatisticsUnit 99 min read
Hypothesis Testing: Tests of Significance, Z-Tests, t-Tests, p-Values
Unit 9 of Business Statistics introduces hypothesis testing—the scientific method for making data-driven decisions. Learn how to frame null/alternative hypotheses, choose significance levels (α), perform Z-tests and t-tests, interpret p-values, and apply these to real-world scenarios like comparing product performance
TAKEAWAYS:
- Hypothesis testing uses sample data to evaluate claims about populations (e.g., "Is Product X better than Product Y?").
- The null hypothesis (H₀) assumes no effect, while the alternative hypothesis (H₁) proposes a change (e.g., "mean sales increased").
- p-values measure evidence against H₀: if p < α (e.g., 0.05), reject H₀ (statistically significant result).
- Z-tests use population standard deviation (known σ) for large samples (n ≥ 30); t-tests use sample standard deviation (unknown σ) for small samples.
- Type I/II errors are critical: α controls Type I (false positives), β controls Type II (false negatives).
- Real-world applications include A/B testing (e.g., Daraz’s ad campaigns), quality control (e.g., NTC’s network reliability), and medical trials (e.g., vaccine efficacy).
1. Core Concepts: Hypotheses and Significance
Hypothesis testing is the backbone of inferential statistics. It answers: "Is the observed difference in data due to random chance or a real effect?"
Key Definitions
- Null Hypothesis (H₀): Default assumption of no effect or no difference. Example: "The new teaching method does not improve exam scores."
- Alternative Hypothesis (H₁): Claims there is an effect. Example: "The new method increases scores by >5%."
- Significance Level (α): Threshold for rejecting H₀ (commonly 0.05 or 5%).
- p-value: Probability of observing data as extreme as yours, if H₀ is true.
- p < α → Reject H₀ (significant result).
- p ≥ α → Fail to reject H₀ (insufficient evidence).
Types of Hypotheses
classDiagram
class Hypothesis {
+isDirectional()
+isNonDirectional()
}
class H0 {
<<null>>
"No effect/difference"
}
class H1 {
<<alternative>>
"Effect/difference exists"
}
H0 --> H1 : "vs."
H1 : "Can be one-tailed or two-tailed"
note for H1
One-tailed: "μ > 50" or "μ < 50"
Two-tailed: "μ ≠ 50"Example: eSewa’s User Growth
Claim: "eSewa’s new referral program increased daily active users by 15%."
- H₀: μ = 0% (no change).
- H₁: μ > 0% (one-tailed test, since we expect an increase).
- α = 0.05.
2. Test Statistics: Z-Test vs. t-Test
Choose the test based on sample size and population standard deviation (σ).
Comparison Table
| Feature | Z-Test | t-Test |
|---|---|---|
| Population σ | Known | Unknown |
| Sample Size | Large (n ≥ 30) | Small (n < 30) |
| Distribution | Normal (Z-table) | t-distribution (t-table) |
| Formula | ||
| When to Use | Quality control (e.g., NTC’s call wait times) | A/B testing (e.g., Daraz’s ad clicks) |
Worked Example: Ncell’s Call Drop Rate
Scenario: Ncell claims its call drop rate is ≤1%. A sample of 50 calls shows 3 drops.
- H₀: p = 1% (no increase in drops).
- H₁: p > 1% (one-tailed).
- α = 0.05, n = 50 (small sample → t-test for proportions).
Steps:
- Calculate sample proportion (p̂): (6%).
- Standard error (SE): .
- Test statistic (Z): .
- p-value: For Z = 3.57 (one-tailed), p ≈ 0.0002 < 0.05 → Reject H₀. Conclusion: Ncell’s drop rate has significantly increased.
3. Decision Rules and Errors
Decision Flowchart
flowchart TD
A["Start"] --> B{"Is p < α?"}
B -->|"Yes"| C["Reject H₀<br/>'Significant result'"]
B -->|"No"| D["Fail to reject H₀<br/>'No significant result'"]
C --> E["Conclude: Effect exists"]
D --> F["Conclude: No evidence of effect"]Types of Errors
| Error | Definition | Consequence | How to Reduce |
|---|---|---|---|
| Type I (α) | Reject H₀ when true | False alarm (e.g., blaming a good product) | Lower α (e.g., 0.01 instead of 0.05) |
| Type II (β) | Fail to reject H₀ when false | Miss real effects (e.g., ignoring a better ad) | Increase sample size or power |
Example: Pathao’s Ride Pricing
- H₀: New dynamic pricing has no effect on demand.
- Type I Error: Pathao raises prices, but demand doesn’t drop (lost revenue).
- Type II Error: Prices stay low, but demand could have been higher (missed profit).
4. Hypothesis Testing Steps (With Visual Trace)
Use this 5-step framework for any problem:
State Hypotheses
- Example: "Does Kathmandu’s traffic speed differ from the national average (60 km/h)?"
- H₀: μ = 60 km/h
- H₁: μ ≠ 60 km/h (two-tailed)
- Example: "Does Kathmandu’s traffic speed differ from the national average (60 km/h)?"
Choose Significance Level (α)
- Typically 0.05.
Calculate Test Statistic
- For a sample of 25 cars with mean speed = 55 km/h, s = 8 km/h: .
Find Critical Value or p-value
- For df = 24, two-tailed t = ±2.064 (from t-table).
- Since |−3.125| > 2.064 → Reject H₀.
Make Decision
- Conclusion: Kathmandu’s traffic is significantly slower than the national average.
5. Real-World Applications
A. Daraz’s A/B Testing for Ad Campaigns
- Idea Used: Two-sample t-test for independent means.
- How:
- H₀: Ad A’s conversion rate = Ad B’s conversion rate.
- H₁: Ad A’s rate > Ad B’s rate (one-tailed).
- Data: 1000 users per ad; Ad A: 8% conversions, Ad B: 6%.
- Result: p = 0.01 → Daraz switches to Ad A.
B. NEPSE Stock Price Analysis
- Idea Used: One-sample Z-test for mean.
- How:
- Claim: "NEPSE’s average daily return is 0.5%."
- H₀: μ = 0.5%.
- Sample: 30 days, mean return = 0.3%, σ = 0.2%.
- Test Statistic: .
- p-value (two-tailed): 0.029 → Reject H₀ (returns are significantly lower).
C. Khalti’s Fraud Detection
- Idea Used: Chi-square goodness-of-fit test.
- How:
- H₀: Fraud transactions follow expected distribution (e.g., 5% of all transactions).
- Observed: 8 frauds in 100 transactions (8%).
- Expected: 5 frauds.
- Chi-square Statistic: .
- p-value: 0.37 → Fail to reject H₀ (no significant increase in fraud).
6. Common Pitfalls and Exam Tips
Mistakes to Avoid
- Ignoring H₀/H₁: Always state both hypotheses explicitly.
- Wrong Test: Using a Z-test when σ is unknown (use t-test).
- Direction Matters: One-tailed vs. two-tailed affects p-value.
- Assuming Causation: Correlation ≠ causation (e.g., ice cream sales and drowning don’t imply causation).
Exam Tip: Structured Answer Format
Use this template for full marks:
- State H₀ and H₁ (1 mark).
- Choose α and test type (1 mark).
- Calculate test statistic (show formula + steps) (3 marks).
- Find critical value/p-value (1 mark).
- Decision + Conclusion (2 marks).
- Interpretation in context (e.g., "This suggests...") (2 marks).
Example Answer Starter:
"For the given data on NTC’s call wait times (μ₀ = 30 sec, n = 40, s = 5 sec, X̄ = 35 sec), we perform a one-sample t-test at α = 0.05 to test if wait times have increased. H₀: μ = 30 sec vs. H₁: μ > 30 sec. Test Statistic: . Critical t (df=39, one-tailed): 1.685. Since 4.90 > 1.685, we reject H₀ at 5% significance. Conclusion: NTC’s wait times have significantly increased."
7. Practice Questions (With Hints)
Bank Loan Defaults:
- A bank claims default rates are ≤2%. In a sample of 200 loans, 6 defaulted. Test at α = 0.01.
- Hint: Use Z-test for proportions; H₀: p = 0.02.
YouTube Ad Engagement:
- YouTube tests a new ad format. Sample A (old): 5% click-through; Sample B (new): 7%, n = 500 each.
- Hint: Two-sample Z-test; H₁: p_B > p_A.
Khalti Transaction Fees:
- Khalti claims fees are ≤1%. A sample of 100 transactions shows 1.5% fees. Test at α = 0.05.
- Hint: One-sample Z-test; H₁: p > 0.01.
8. Key Formulas Summary
| Test | Formula | When to Use |
|---|---|---|
| Z-test (σ known) | Large n, known σ | |
| t-test (σ unknown) | Small n, unknown σ | |
| Two-sample t-test | Compare two independent groups | |
| Chi-square | Categorical data (e.g., fraud rates) |
Exam Tip: Visualizing Results
Always draw the distribution with:
- Mean/Expected value (μ₀).
- Observed statistic (Z or t).
- Critical region (shaded for rejection).
- p-value area.
Example for a t-test:
Based on the TU BITM syllabus for Business Statistics (STT201), unit 9.
Discussion
Loading…