Business StatisticsUnit 1313 min read
Statistical Inference & Hypothesis Testing: Tests, Confidence, and Decision-Making
Unit 13 of Business Statistics: Covers statistical inference (estimation, confidence intervals), hypothesis testing (z, t, chi-square tests), p-values, Type I/II errors, and real-world applications like quality control and market research.
TAKEAWAYS:
- Statistical inference lets us estimate population parameters from sample data (e.g., average wage, market share).
- Hypothesis testing uses p-values and significance levels to decide whether observed data supports a claim (e.g., "Does eSewa’s transaction volume increase with promotions?").
- Confidence intervals quantify uncertainty in estimates (e.g., "Ncell’s 95% CI for call drop rate is 0.02–0.05%").
- Z-tests and t-tests compare sample means to known values or between groups (e.g., Daraz’s order delivery times before/after warehouse upgrades).
- Chi-square tests check if categorical data fits expected distributions (e.g., "Does Pathao’s surge pricing correlate with traffic density?").
- Type I/II errors have real costs: false alarms (e.g., recalling safe products) vs. missed defects (e.g., approving unsafe loans).
1. Statistical Inference: From Samples to Populations
Statistical inference is the process of making conclusions about a population based on data from a sample. Since we rarely survey every worker, customer, or stock (population), we use samples to generalize.
Key Concepts
- Population: The entire group of interest (e.g., all Ncell subscribers in Kathmandu).
- Sample: A subset of the population (e.g., 500 randomly selected Ncell users).
- Parameter: A fixed value describing the population (e.g., mean call duration).
- Statistic: A calculated value from the sample (e.g., sample mean call duration).
Types of Inference
Estimation: Guessing a population parameter (e.g., "What’s the average loan default rate in Nepal?").
- Point estimate: Single value (e.g., sample mean = 3.2%).
- Interval estimate: Range with confidence (e.g., 95% CI: 2.8%–3.6%).
Hypothesis Testing: Testing claims about the population (e.g., "Does Khalti’s transaction fee decrease after a rate cut?").
FIGURE 1: Population vs. Sample
figure:
graph TD
A["Population: All Ncell users in Kathmandu (1M)"] -->|"Sample"| B["Sample: 500 users surveyed"]
B --> C["Statistic: Sample mean call duration = 4.1 min"]
C --> D["Parameter: True population mean (unknown)"]Confidence Intervals
A confidence interval (CI) gives a range of values likely to contain the population parameter, with a given confidence level (e.g., 95%).
Formula for Mean (Large Sample, σ known): Where:
- = sample mean
- = critical z-value (e.g., 1.96 for 95% CI)
- = population standard deviation
- = sample size
Worked Example: Daraz’s Order Delivery Time Daraz claims its average order delivery time is ≤48 hours. A sample of 100 orders has a mean of 50 hours and σ=12 hours. Test at 95% CI.
Calculate CI: Since 48 hours is within the CI, we cannot reject Daraz’s claim at 95% confidence.
Interpretation:
- 95% CI: We’re 95% sure the true mean delivery time is between 47.65 and 52.35 hours.
- Business implication: Daraz’s claim is plausible, but the upper bound suggests delays exceed 48 hours.
FIGURE 2: Confidence Interval for Daraz’s Delivery Time
figure:
graph LR
A["Sample Mean: 50 hours"] --> B["± Margin of Error: ±2.352"]
B --> C["95% CI: [47.65, 52.35] hours"]
C --> D["Claim: ≤48 hours"]
D -->|"Within CI"| E["Fail to reject claim"]Advantages/Disadvantages
| Advantages | Disadvantages |
|---|---|
| Provides a range, not just a point. | Wider intervals = less precision. |
| Quantifies uncertainty. | Requires random sampling. |
| Useful for decision-making (e.g., risk assessment). | Assumes normality (for small samples). |
2. Hypothesis Testing: Testing Claims
Hypothesis testing answers: "Is there enough evidence to support a claim?" Steps:
- State hypotheses:
- : Null hypothesis (default claim, e.g., "No effect").
- : Alternative hypothesis (what we test, e.g., "There is an effect").
- Choose significance level (, e.g., 0.05).
- Calculate test statistic (z, t, chi-square).
- Compare to critical value or p-value.
- Make a decision: Reject or fail to reject.
FIGURE 3: Hypothesis Testing Workflow
figure:
graph TD
A["State Hypotheses"] --> B["Choose α (e.g., 0.05)"]
B --> C["Calculate Test Statistic"]
C --> D["Compare to Critical Value or p-value"]
D --> E["Decision: Reject/Fail to Reject H₀"]Types of Tests
| Test | When to Use | Formula | Example |
|---|---|---|---|
| Z-test | Large sample (), σ known. | Test if Ncell’s call drop rate changed after a network upgrade. | |
| t-test | Small sample () or σ unknown. | Compare average wages of TU vs. PU students. | |
| Chi-square test | Categorical data (e.g., goodness-of-fit). | Check if Pathao’s surge pricing tiers fit a normal distribution. |
Worked Example: eSewa Transaction Volume
eSewa claims its daily transaction volume increased after a new app update. Before: mean = 50,000 transactions/day (, σ=5,000). After: mean = 52,000 (, σ=5,000). Test at α=0.05.
State hypotheses:
- : (no increase).
- : (increase).
Calculate z-test statistic:
Critical value: For α=0.05 (one-tailed), .
Decision:
- Since , reject .
- p-value: (from z-table).
- Since , reject .
Conclusion: There is statistically significant evidence that eSewa’s transaction volume increased after the update.
FIGURE 4: Z-Test for eSewa Transactions
figure:
graph LR
A["Sample Mean (After): 52,000"] --> B["Population Mean (H₀): 50,000"]
B --> C["Standard Error: 913.42"]
C --> D["z = 2.19"]
D --> E["Critical z: 1.645"]
E -->|"z > critical"| F["Reject H₀"]P-values and Decision Rules
- P-value: Probability of observing data as extreme as the sample, assuming is true.
- If , reject .
- If , fail to reject .
Example Interpretation:
- For eSewa’s test, → reject .
- For Daraz’s CI example, since 48 hours is within the CI, we fail to reject the claim at 95% confidence.
FIGURE 5: P-value Decision Rule
figure:
graph LR
A["Calculate p-value"] --> B["Compare to α"]
B --> C["If p ≤ α"]
C --> D["Reject H₀"]
B --> E["If p > α"]
E --> F["Fail to reject H₀"]3. Type I and Type II Errors
| Error | Definition | Consequence | Example |
|---|---|---|---|
| Type I (α) | Reject when true. | False alarm. | Recall a product that’s actually safe. |
| Type II (β) | Fail to reject when false. | Missed defect. | Approve a loan for a risky borrower. |
| Power | : Probability of correctly rejecting . | Higher power = better test. |
Trade-off:
- Lowering (e.g., 0.01) reduces Type I errors but increases Type II errors.
- Increasing sample size () reduces both errors.
FIGURE 6: Type I vs. Type II Errors
figure:
graph TD
A["H₀ is True"] --> B["Correct Decision: Fail to Reject H₀"]
A --> C["Type I Error: Reject H₀"]
D["H₀ is False"] --> E["Type II Error: Fail to Reject H₀"]
D --> F["Correct Decision: Reject H₀"]Worked Example: NEPSE Stock Returns
NEPSE claims its stock returns are normally distributed with μ=12%. A sample of 40 stocks has a mean return of 10.5% and σ=4%. Test if the true mean is 12% at α=0.05.
State hypotheses:
- : .
- : (two-tailed).
Calculate t-test statistic (small sample, σ unknown):
Critical t-value (df=39, α=0.05/2): .
Decision:
- Since , reject .
Conclusion: The sample evidence suggests NEPSE’s stock returns are not normally distributed with μ=12% at 95% confidence.
FIGURE 7: T-test for NEPSE Returns
figure:
graph LR
A["Sample Mean: 10.5%"] --> B["H₀ Mean: 12%"]
B --> C["Standard Error: 0.632"]
C --> D["t = -2.37"]
D --> E["Critical t: ±2.023"]
E -->|"t < -2.023"| F["Reject H₀"]4. Chi-Square Tests: Categorical Data
Used to test:
- Goodness-of-fit: Does observed data fit an expected distribution?
- Independence: Are two categorical variables independent?
Example: Does Pathao’s surge pricing tier usage fit a uniform distribution?
FIGURE 8: Chi-Square Goodness-of-Fit
figure:
graph TD
A["Observed Frequencies"] --> B["Expected Frequencies (Uniform)"]
B --> C["Calculate χ² = Σ (O - E)²/E"]
C --> D["Compare to Critical χ²"]
D --> E["Decision: Fit or Not Fit"]Worked Example: Daraz Order Categories
*Daraz’s orders are categorized as:
- Electronics (30%)
- Clothing (40%)
- Groceries (30%). A sample of 100 orders has: Electronics: 35, Clothing: 38, Groceries: 27. Test if the distribution fits Daraz’s claim at α=0.05.*
State hypotheses:
- : Observed distribution matches claimed (30%, 40%, 30%).
- : Distribution does not match.
Expected counts:
- Electronics:
- Clothing:
- Groceries:
Calculate χ²:
Critical χ² (df=2, α=0.05): .
Decision:
- Since , fail to reject .
Conclusion: The sample distribution does not significantly differ from Daraz’s claimed categories.
FIGURE 9: Chi-Square Calculation for Daraz
figure:
mermaid table
| Category | Observed | Expected | (O-E)²/E |
|---|---|---|---|
| Electronics | 35 | 30 | 0.33 |
| Clothing | 38 | 40 | 0.05 |
| Groceries | 27 | 30 | 0.33 |
| Total | 100 | 100 | 0.71 |
5. Real-World Applications
In the Real World
eSewa/Khalti: Transaction Fraud Detection
- Idea: Hypothesis testing checks if fraud rates increase after new payment methods.
- Example: eSewa tests if the mean fraud rate rises from 0.5% to 1% after adding UPI. A z-test compares the new sample mean to the old rate.
Ncell/NTC: Network Performance
- Idea: Confidence intervals estimate call drop rates after network upgrades.
- Example: Ncell claims its call drop rate is ≤2%. A 95% CI of [1.8%, 2.5%] includes 2%, so the claim is plausible.
Daraz: Inventory Management
- Idea: Chi-square tests if product categories align with demand forecasts.
- Example: Daraz checks if observed sales (Electronics: 35%, Clothing: 38%) match expected (30%, 40%). The χ²=0.71 suggests no significant deviation.
NEPSE: Stock Market Stability
- Idea: t-tests verify if stock returns deviate from historical norms.
- Example: If NEPSE’s sample mean return is 10.5% (vs. historical 12%), a t-test may reject the null hypothesis, signaling instability.
FIGURE 10: Real-World Examples Summary
figure:
mindmap
root((Statistical Inference))
eSewa: Hypothesis Testing for Fraud Rates
Ncell: Confidence Intervals for Call Drop Rates
Daraz: Chi-Square for Category Fit
NEPSE: t-Tests for Stock Return StabilityExam Tips
Master the Steps:
- Always state and clearly.
- For z/t-tests, show the formula and plug in numbers.
- For chi-square, build the table with O, E, and .
Interpret Results Correctly:
- "Reject " means evidence supports the alternative.
- "Fail to reject" does not mean "prove "; it means insufficient evidence.
Choose the Right Test:
- Z-test: Large , σ known.
- t-test: Small or σ unknown.
- Chi-square: Categorical data.
Show Work for Full Marks:
- Write out calculations (e.g., standard error, test statistic).
- Include p-values or critical values with decisions.
Real-World Linkage:
- Connect tests to business scenarios (e.g., "How would a bank use a t-test for loan defaults?").
- Use examples from the unit (e.g., Daraz, eSewa) to explain concepts.
Common Pitfalls:
- Directionality: One-tailed vs. two-tailed tests.
- Assumptions: Normality for t-tests, independence for chi-square.
- Sample Size: Use z for , t otherwise.
FIGURE 11: Exam Checklist
figure:
graph LR
A["State H₀ and H₀"] --> B["Choose α"]
B --> C["Calculate Test Statistic"]
C --> D["Compare to Critical Value or p-value"]
D --> E["Decision: Reject/Fail to Reject"]
E --> F["Interpret in Context"]
F --> G["Show All Work"]Based on the TU BBS syllabus for Business Statistics (MGT207), unit 13.
Discussion
Loading…