Business StatisticsUnit 914 min read
Hypothesis Testing: Tests of Significance, p-Values & Decision Rules
Unit 9 of Business Statistics introduces hypothesis testing, covering null/alternative hypotheses, test statistics, significance levels, p-values, and decision-making frameworks for statistical inference in business and research.
TAKEAWAYS:
- Hypothesis testing follows a structured 5-step process (state hypotheses, choose test, compute statistic, determine critical value/p-value, make decision).
- Type I and Type II errors are inevitable trade-offs in decision-making, with α (significance level) controlling Type I error risk.
- p-values quantify evidence against the null hypothesis: smaller p-values (≤ α) reject H₀.
- One-tailed vs. two-tailed tests depend on the alternative hypothesis’s directionality.
- Common tests (z-test, t-test, chi-square) are chosen based on data type (continuous/discrete) and sample size.
- Real-world applications include A/B testing in apps (e.g., Pathao’s delivery route optimization) and quality control in manufacturing (e.g., NTC’s network reliability checks).
1. Introduction to Hypothesis Testing
Hypothesis testing is a statistical method used to make inferences about a population based on sample data. It helps businesses and researchers answer questions like:
- Does a new marketing strategy increase sales? (e.g., Daraz’s promotional campaigns)
- Is the average loan default rate higher than 5%? (e.g., Nabil Bank’s risk assessment)
- Does a new drug reduce recovery time? (e.g., pharmaceutical trials)
Key Definitions
- Null Hypothesis (H₀): Assumes no effect or no difference (default position). Example: H₀: μ = 50 (mean salary = 50,000 NPR).
- Alternative Hypothesis (H₁ or Ha): Claims a significant effect or difference. Can be:
- Two-tailed: H₁: μ ≠ 50 (no direction specified).
- One-tailed: H₁: μ > 50 (right-tailed) or H₁: μ < 50 (left-tailed).
- Significance Level (α): Probability of rejecting H₀ when it’s true (e.g., α = 0.05 or 5%).
- Test Statistic: Standardized value (e.g., z-score, t-score) calculated from sample data to compare against critical values.
- p-value: Probability of observing test results as extreme as those seen, assuming H₀ is true. Smaller p-value → stronger evidence against H₀.
Decision Rule
Reject H₀ if:
- Test statistic falls in the critical region (e.g., z > 1.96 for α = 0.05, two-tailed), or
- p-value ≤ α.
2. Steps in Hypothesis Testing
Use this 5-step framework for any test:
flowchart TD
A["Step 1: State Hypotheses"] --> B["Step 2: Choose Test Statistic"]
B --> C["Step 3: Compute Test Statistic"]
C --> D["Step 4: Determine Critical Value or p-value"]
D --> E["Step 5: Make Decision & Conclusion"]Step 1: State Hypotheses
Example: A factory claims its light bulbs last 1000 hours on average. A sample of 50 bulbs has a mean life of 980 hours (σ = 50 hours). Test at α = 0.05 if the claim is false.
- H₀: μ = 1000 (claim is true)
- H₁: μ < 1000 (claim is false; left-tailed test)
Step 2: Choose Test Statistic
Select based on:
| Data Type | Population σ Known? | Sample Size (n) | Test Statistic |
|---|---|---|---|
| Continuous | Yes | Any | z-test |
| Continuous | No | n ≥ 30 | z-test (approx.) |
| Continuous | No | n < 30 | t-test |
| Categorical | — | — | Chi-square test |
For our example:
- σ is known (50 hours), n = 50 (≥ 30) → z-test.
Step 3: Compute Test Statistic
Formula for z-test: Plugging in values:
Step 4: Determine Critical Value or p-value
Option 1: Critical Value Approach For α = 0.05 (left-tailed), critical z = -1.645. Since -2.83 < -1.645, reject H₀.
Option 2: p-value Approach Look up z = -2.83 in standard normal table: p-value ≈ 0.0023 (one-tailed). Since 0.0023 < 0.05, reject H₀.
Step 5: Conclusion
Decision: Reject H₀ (sufficient evidence that μ < 1000). Interpretation: The factory’s claim that bulbs last 1000 hours is false at 5% significance.
3. Types of Errors
No test is perfect. Two possible errors:
| Error Type | Definition | Probability Notation | Example |
|---|---|---|---|
| Type I | Reject H₀ when it’s true (false alarm) | α (significance level) | Firing an employee due to "poor performance" when they’re actually excellent. |
| Type II | Fail to reject H₀ when it’s false | β | Approving a faulty batch of medicines because the test missed the defect. |
Trade-off: Reducing α (e.g., to 0.01) lowers Type I error but increases Type II error.
4. One-Tailed vs. Two-Tailed Tests
| Test Type | Alternative Hypothesis (H₁) | When to Use | Critical Region |
|---|---|---|---|
| Two-tailed | μ ≠ μ₀ | No prior expectation of direction (e.g., "Is there any difference?"). | Both tails (e.g., z < -1.96 or z > 1.96) |
| Right-tailed | μ > μ₀ | Expecting an increase (e.g., "Does training improve scores?"). | Right tail (e.g., z > 1.645) |
| Left-tailed | μ < μ₀ | Expecting a decrease (e.g., "Does a new drug reduce recovery time?"). | Left tail (e.g., z < -1.645) |
Example for Two-Tailed Test: A bank claims its loan approval time is 10 days. A sample of 40 loans has a mean approval time of 11.5 days (σ = 2 days). Test at α = 0.05.
- H₀: μ = 10
- H₁: μ ≠ 10 (two-tailed)
- z-test:
- Critical z: ±1.96
- Decision: Reject H₀ (3.03 > 1.96). The approval time is significantly different from 10 days.
5. Common Hypothesis Tests
| Test | Purpose | When to Use | Formula |
|---|---|---|---|
| z-test | Compare sample mean to population mean (σ known). | Large samples (n ≥ 30) or σ known. | |
| t-test | Compare sample mean to population mean (σ unknown). | Small samples (n < 30) or σ unknown. | (df = n - 1) |
| Chi-square | Test independence between categorical variables. | Contingency tables (e.g., "Is gender independent of product preference?"). | (df = (rows-1)(cols-1)) |
| ANOVA | Compare means across ≥3 groups. | Experimental designs (e.g., "Do 3 fertilizers affect crop yield differently?"). |
6. Real-World Applications
Example 1: Pathao’s Delivery Route Optimization
Scenario: Pathao wants to test if a new algorithm reduces delivery times.
- H₀: μ = 25 minutes (current average).
- H₁: μ < 25 minutes (algorithm improves speed).
- Sample: 100 deliveries with new algorithm → mean = 23 minutes, σ = 4 minutes.
- Test: One-tailed z-test (σ known).
- p-value: < 0.0001 → Reject H₀. The algorithm significantly reduces delivery time.
Example 2: NTC’s Network Reliability Check
Scenario: NTC claims 99.9% network uptime. A sample of 500 hours shows 4 outages.
- H₀: p = 0.999 (99.9% uptime).
- H₁: p < 0.999 (worse than claimed).
- Test: One-tailed z-test for proportion.
- p-value: 0.0057 → Reject H₀. Uptime is significantly worse than claimed.
Example 3: Daraz’s A/B Testing for Sales
Scenario: Daraz tests a new website layout to increase sales.
- H₀: No difference in conversion rates (p₁ = p₂).
- H₁: p₁ ≠ p₂ (two-tailed).
- Sample: Group A (old layout): 500 sales/10,000 visitors. Group B (new layout): 600 sales/10,000 visitors.
- Test: Two-proportion z-test.
- p-value: 0.018 → Reject H₀. The new layout significantly increases sales.
7. Visualizing Hypothesis Testing
Figure 1: Critical Regions for α = 0.05
Figure 2: p-value for z = -2.83 (Left-Tailed Test)
Figure 3: Type I and Type II Errors
pie
title Errors in Hypothesis Testing
"Type I Error (α)" : 5
"Type II Error (β)" : 15
"Correct Decisions" : 808. Exam Tip
- Always state H₀ and H₁ clearly (include directionality for one-tailed tests).
- Justify your test choice (z-test vs. t-test vs. chi-square) based on data type and sample size.
- Show calculations step-by-step for test statistics (e.g., z, t, χ²). Partial credit is common for correct formulas.
- Interpret p-values correctly:
- p ≤ α → "Reject H₀; sufficient evidence to support H₁."
- p > α → "Fail to reject H₀; insufficient evidence."
- Avoid common mistakes:
- Confusing one-tailed vs. two-tailed tests (check H₁).
- Misusing σ vs. s (population vs. sample standard deviation).
- Forgetting degrees of freedom (df = n - 1 for t-test).
- Real-world connection: In exams, relate examples to business scenarios (e.g., "A company claims its product lasts X hours—test their claim").
9. Worked Example: Kathmandu Traffic Congestion Study
Scenario: The Kathmandu Metropolitan City claims the average daily traffic delay is ≤ 30 minutes. A survey of 64 commuters finds a mean delay of 32 minutes (s = 8 minutes). Test at α = 0.01.
Step 1: Hypotheses
- H₀: μ ≤ 30
- H₁: μ > 30 (right-tailed; claim is false)
Step 2: Test Choice
- σ unknown, n = 64 (≥ 30) → z-test (approximation).
Step 3: Compute z
Step 4: Critical Value
For α = 0.01 (right-tailed), critical z = 2.326. Since 2 < 2.326, fail to reject H₀.
Step 5: Conclusion
Decision: Fail to reject H₀. Interpretation: At 1% significance, there is insufficient evidence that the average delay exceeds 30 minutes. The city’s claim cannot be rejected.
10. Summary Table: Hypothesis Testing Checklist
| Step | Action Items |
|---|---|
| 1. State Hypotheses | Write H₀ and H₁ (include =, ≠, >, or <). |
| 2. Choose Test | Select z-test, t-test, or chi-square based on data type and sample size. |
| 3. Compute Statistic | Plug values into the correct formula (show all steps). |
| 4. Determine Critical Value/p-value | Use tables or software (e.g., Excel, calculator). |
| 5. Make Decision | Compare test statistic to critical value or p-value to α. |
| 6. Interpret Result | Relate conclusion to the real-world scenario (e.g., "Reject H₀: The drug is effective."). |
11. Common Pitfalls
- Ignoring directionality: Using a two-tailed test when H₁ is one-tailed (or vice versa) leads to incorrect critical values.
- Wrong test selection: Using a t-test when σ is known (should be z-test).
- Misinterpreting p-values: A p-value of 0.06 is not "almost significant" at α = 0.05—it’s still not significant.
- Overlooking assumptions: Hypothesis tests assume normality (for small samples) and independence of observations.
In the real world
- Pathao’s Delivery Route Optimization: Uses one-tailed z-tests to determine if new algorithms reduce delivery times (H₁: μ < current mean). Example: Testing if a 10% route adjustment lowers delivery time from 45 to 40 minutes (α=0.01).
- NTC’s Network Reliability Checks: Employs two-tailed t-tests to verify if call drop rates (H₀: μ=5%) differ after infrastructure upgrades (sample: 200 calls, σ unknown).
- Daraz’s A/B Testing: Compares conversion rates between two ad creatives using chi-square tests (H₀: no difference in click-through rates) to decide which to scale.
Based on the TU BIM syllabus for Business Statistics (STT201), unit 9.
Discussion
Loading…