Business StatisticsUnit 814 min read
Statistical Inference & Hypothesis Testing: Tests, Confidence, p-Values
Unit 8 of Business Statistics: covers statistical inference (point/interval estimation), hypothesis testing (Z, t, χ², F tests), Type I/II errors, significance levels, and real-world applications like quality control, market research, and financial risk assessment.
TAKEAWAYS:
- Statistical inference lets us estimate population parameters from sample data using confidence intervals and point estimators.
- Hypothesis testing uses p-values and critical values to decide whether observed data supports a null hypothesis (H₀).
- Z-tests and t-tests compare means; χ²-tests compare distributions or goodness-of-fit; F-tests compare variances.
- Type I error (rejecting H₀ when true) and Type II error (failing to reject H₀ when false) trade off with sample size and significance level (α).
- Business uses these tools for A/B testing (e.g., Daraz promotions), quality checks (NTC network reliability), and financial risk (NEPSE stock performance).
- Always assume H₀ is true unless evidence overwhelmingly contradicts it (proof by contradiction).
1. Statistical Inference: Estimating Population Parameters
Statistical inference bridges the gap between sample data and population conclusions. It includes:
- Point estimation: Single-value estimates (e.g., sample mean for population mean ).
- Interval estimation: Confidence intervals (e.g., for 95% CI).
Key Concepts
Confidence Interval (CI): A range where the true parameter lies with confidence (e.g., 95% CI).
- Formula for mean ( known):
- For small samples () or unknown , use t-distribution:
Margin of Error (ME): Half the width of the CI:
Worked Example: Daraz’s Promotional Effectiveness
Scenario: Daraz runs a discount campaign and collects a sample of 100 orders. Before the campaign, the average order value was Rs 1,200 (σ = Rs 300). After the campaign, the sample mean is Rs 1,300. Test if the campaign increased order value at 95% confidence.
Steps:
- Hypotheses:
- (no effect)
- (campaign helped)
- Test Statistic (Z-test for large ):
- Critical Value: For 95% CI, .
- Since , reject .
- 95% Confidence Interval: The entire interval is above 1200, confirming the campaign’s effect.
Visual: Interpretation: Daraz can claim the campaign increased order values with 95% confidence.
2. Hypothesis Testing: Framework and Steps
Hypothesis testing evaluates claims about populations using sample data. Steps:
- State hypotheses:
- : Null hypothesis (status quo, e.g., "no effect").
- : Alternative hypothesis (research claim, e.g., "there is an effect").
- Choose significance level (): Commonly 0.05 (5%).
- Select test statistic: Z, t, χ², or F based on data and hypotheses.
- Compute test statistic and p-value (probability of observing data if is true).
- Decision rule:
- Reject if p-value < or test statistic > critical value.
- Fail to reject otherwise.
Types of Errors
| Error Type | Definition | Consequence | Example |
|---|---|---|---|
| Type I | Reject when true | False alarm | Accusing a Daraz seller of fraud when innocent |
| Type II | Fail to reject when false | Missed opportunity | Missing a price drop in NEPSE |
| α | Probability of Type I error | Set by researcher (e.g., 0.05) | |
| β | Probability of Type II error | Depends on effect size and sample size |
Trade-off: Lowering reduces Type I errors but increases Type II errors (and vice versa).
Visual:
flowchart TD
A["Start: Assume H₀ true"] --> B["Collect sample data"]
B --> C["Calculate test statistic"]
C --> D{"Is p-value < α?"}
D -->|"Yes"| E["Reject H₀: Evidence against H₀"]
D -->|"No"| F["Fail to reject H₀: Insufficient evidence"]
E --> G["Conclude effect exists"]
F --> G3. Types of Hypothesis Tests
| Test | Use Case | Test Statistic | Assumptions | Example in Nepal |
|---|---|---|---|---|
| Z-test | Compare population mean (σ known) | Normal distribution, large (≥30) | Testing if Ncell’s call drop rate improved after network upgrade | |
| t-test | Compare population mean (σ unknown) | Normal distribution, small | Comparing average loan defaults between two banks (e.g., NMB vs. Global IME) | |
| χ²-test | Goodness-of-fit or independence | Categorical data, expected frequencies >5 | Testing if Pathao’s surge pricing affects rider distribution by time of day | |
| F-test | Compare variances (e.g., two populations) | Normal distributions, independent samples | Comparing variability in eSewa transaction times vs. Khalti transaction times |
Worked Example: NEPSE Stock Performance
Scenario: An investor claims NEPSE’s average daily return is 0.5% (σ = 1.2%). Over 30 trading days, the sample mean return is 0.7%. Test at 95% confidence.
Steps:
- Hypotheses:
- (two-tailed test)
- Test Statistic (Z-test):
- Critical Values: For 95% CI, .
- Since , fail to reject .
- p-value: (from Z-table).
- p-value > 0.05 → no significant evidence to reject .
Visual:
Interpretation: The investor’s claim lacks statistical support. The data does not prove the average return differs from 0.5%.
4. Comparing Two Populations
Independent Samples (Two-Sample t-test)
Tests if two population means are equal:
- Equal variances: Pooled variance
- Unequal variances: Welch’s t-test (uses separate variances).
Worked Example: NTC vs. Ncell Call Quality Scenario: NTC claims its call drop rate is lower than Ncell’s. Samples:
- NTC: , ,
- Ncell: , ,
Steps:
- Hypotheses:
- (one-tailed)
- Pooled Variance:
- t-statistic:
- Critical Value: For 95% CI, one-tailed, (from t-table).
- Since , reject .
Visual:
Interpretation: NTC’s call drop rate is significantly lower than Ncell’s at 95% confidence.
Paired Samples (Dependent t-test)
Used when data is paired (e.g., before/after measurements).
- Test statistic: where (differences).
Worked Example: Kathmandu Traffic Before/After Flyover Scenario: Traffic speeds (km/h) before and after a flyover were recorded for 20 vehicles:
| Vehicle | Before | After | Difference () |
|---|---|---|---|
| 1 | 20 | 25 | -5 |
| 2 | 15 | 20 | -5 |
| ... | ... | ... | ... |
| Mean = -4 km/h, km/h |
Steps:
- Hypotheses:
- (no effect)
- (flyover improved speed)
- t-statistic:
- Critical Value: For 95% CI, one-tailed, .
- Since , reject .
Visual:
5. Chi-Square (χ²) Tests
Goodness-of-Fit Test
Tests if observed frequencies match expected frequencies (e.g., die rolls, survey responses).
- Test statistic: where = observed, = expected.
Worked Example: Pathao Rider Preferences Scenario: Pathao surveyed 200 riders on preferred pickup times:
| Time Slot | Observed () | Expected () |
|---|---|---|
| 6–10 AM | 50 | 40 |
| 10 AM–2 PM | 30 | 50 |
| 2–6 PM | 80 | 50 |
| 6–10 PM | 40 | 40 |
| 10 PM–6 AM | 0 | 20 |
Steps:
- Hypotheses:
- : Rider preferences follow uniform distribution.
- : Preferences are non-uniform.
- Expected Frequencies: for each slot.
- χ²-statistic:
- Critical Value: For 4 df (k-1) and α=0.05, .
- Since , reject .
Visual:
Interpretation: Rider preferences are not uniform; Pathao should adjust surge pricing accordingly.
Test of Independence
Tests if two categorical variables are independent (e.g., gender vs. product choice).
Contingency table:
Product A Product B Total Male 40 60 100 Female 30 70 100 Total 70 130 200 Test statistic same as above; degrees of freedom = .
6. F-Test: Comparing Variances
Tests if two populations have equal variances (e.g., consistency in loan processing times).
- Test statistic:
- Critical values from F-distribution tables.
Worked Example: Loan Processing Times Scenario: Two banks report:
- Bank X: , days
- Bank Y: , days
Steps:
- Hypotheses:
- F-statistic:
- Critical Values: For 95% CI, two-tailed, and .
- Since , fail to reject .
Interpretation: No significant difference in processing time variability between the banks.
In the Real World
eSewa’s Transaction Fraud Detection
- Idea: Uses hypothesis testing to flag unusual transaction patterns (e.g., sudden high-volume transfers).
- How: Compares observed transaction frequencies to expected distributions (χ²-test). If p-value < 0.05, alerts fraud teams.
- Example: A user’s transactions spike from Rs 5,000/month to Rs 50,000 in one day. The χ²-test rejects the null hypothesis of "normal usage," triggering a review.
Daraz’s A/B Testing for Discounts
- Idea: Tests if a discount offer increases conversion rates using two-sample t-tests.
- How: Randomly assigns users to control (no discount) and treatment (20% off) groups. If the mean order value in the treatment group is significantly higher (p < 0.05), the discount is rolled out.
- Worked Example: Control group mean = Rs 1,200 (, σ=Rs 300); treatment group mean = Rs 1,350 (, σ=Rs 350).
- Z-test:
- p-value ≈ 0 → reject . Daraz concludes the discount works.
NEPSE’s Stock Performance Monitoring
- Idea: Uses confidence intervals to track if a stock’s return deviates from historical norms.
- How: Monthly, NEPSE calculates a 95% CI for daily returns. If a stock’s return falls outside the CI (e.g., 0.5% ± 0.5%), analysts investigate anomalies.
- Example: If a stock’s return is 1.2% and the CI is [0.0%, 1.0%], the upper bound is exceeded, prompting further analysis.
Exam Tip
- Master the Steps: Always follow the 5-step hypothesis testing framework (state hypotheses → choose α → select test → compute statistic → decide). Partial credit is lost if steps are skipped.
- Know When to Use Which Test:
- Z-test: Large samples, σ known.
- t-test: Small samples or σ unknown.
- χ²-test: Categorical data (goodness-of-fit or independence).
- F-test: Compare variances.
- Paired t-test: Before/after or matched pairs.
- Calculate p-values or Critical Values:
- For Z-tests, use standard normal tables.
- For t-tests, use t-distribution tables (df = ).
- For χ², use χ² tables (df = categories - 1).
- For F-tests, use F-distribution tables (df₁, df₂).
- Interpret Rejection Regions:
- Two-tailed: Reject if or .
- One-tailed: Reject if or .
- Real-World Context:
- Always tie your answer to a business scenario (e.g., "This test helps Daraz decide whether to continue a discount campaign").
- Show Work:
- For calculations, write out the test statistic formula and substitute values clearly. Partial credit is given for correct setup even if the final answer is wrong.
- Common Pitfalls:
- Direction of Alternative Hypothesis: (one-tailed) vs. (two-tailed) changes critical values.
- Degrees of Freedom: For t-tests, df = ; for χ², df = (rows-1)(cols-1).
- Assumptions: Always state assumptions (e.g., "Data is normally distributed" for t-tests).
Final Visual Summary:
mindmap
root((Statistical Inference & Hypothesis Testing))
Statistical_Inference
Point_Estimation
Interval_Estimation
Confidence_Intervals
Z-Intervals
t-Intervals
Hypothesis_Testing
Framework
Hypotheses
Test_Statistics
Decision_Rules
Types
Z-Tests
t-Tests
Independent
Paired
Chi-Square
Goodness-of-Fit
Independence
F-Tests
Real-World_Applications
eSewa_Fraud_Detection
Daraz_A_B_Testing
NEPSE_Stock_Monitoring
Exam_Tips
Steps
Test_Selection
p-values_vs_Critical_Values
AssumptionsBased on the TU BBA syllabus for Business Statistics (STT201), unit 8.
Discussion
Loading…