MGT207 Business Statistics

Business StatisticsUnit 1313 min read

Statistical Inference & Hypothesis Testing: Tests, Confidence, and Decision-Making

Unit 13 of Business Statistics: Covers statistical inference (estimation, confidence intervals), hypothesis testing (z, t, chi-square tests), p-values, Type I/II errors, and real-world applications like quality control and market research.

TAKEAWAYS:

  • Statistical inference lets us estimate population parameters from sample data (e.g., average wage, market share).
  • Hypothesis testing uses p-values and significance levels to decide whether observed data supports a claim (e.g., "Does eSewa’s transaction volume increase with promotions?").
  • Confidence intervals quantify uncertainty in estimates (e.g., "Ncell’s 95% CI for call drop rate is 0.02–0.05%").
  • Z-tests and t-tests compare sample means to known values or between groups (e.g., Daraz’s order delivery times before/after warehouse upgrades).
  • Chi-square tests check if categorical data fits expected distributions (e.g., "Does Pathao’s surge pricing correlate with traffic density?").
  • Type I/II errors have real costs: false alarms (e.g., recalling safe products) vs. missed defects (e.g., approving unsafe loans).

1. Statistical Inference: From Samples to Populations

Statistical inference is the process of making conclusions about a population based on data from a sample. Since we rarely survey every worker, customer, or stock (population), we use samples to generalize.

Key Concepts

  • Population: The entire group of interest (e.g., all Ncell subscribers in Kathmandu).
  • Sample: A subset of the population (e.g., 500 randomly selected Ncell users).
  • Parameter: A fixed value describing the population (e.g., mean call duration).
  • Statistic: A calculated value from the sample (e.g., sample mean call duration).

Types of Inference

  1. Estimation: Guessing a population parameter (e.g., "What’s the average loan default rate in Nepal?").

    • Point estimate: Single value (e.g., sample mean = 3.2%).
    • Interval estimate: Range with confidence (e.g., 95% CI: 2.8%–3.6%).
  2. Hypothesis Testing: Testing claims about the population (e.g., "Does Khalti’s transaction fee decrease after a rate cut?").


FIGURE 1: Population vs. Sample

figure:
graph TD
    A["Population: All Ncell users in Kathmandu (1M)"] -->|"Sample"| B["Sample: 500 users surveyed"]
    B --> C["Statistic: Sample mean call duration = 4.1 min"]
    C --> D["Parameter: True population mean (unknown)"]

Confidence Intervals

A confidence interval (CI) gives a range of values likely to contain the population parameter, with a given confidence level (e.g., 95%).

Formula for Mean (Large Sample, σ known): Where:

  • = sample mean
  • = critical z-value (e.g., 1.96 for 95% CI)
  • = population standard deviation
  • = sample size

Worked Example: Daraz’s Order Delivery Time Daraz claims its average order delivery time is ≤48 hours. A sample of 100 orders has a mean of 50 hours and σ=12 hours. Test at 95% CI.

  1. Calculate CI: Since 48 hours is within the CI, we cannot reject Daraz’s claim at 95% confidence.

  2. Interpretation:

    • 95% CI: We’re 95% sure the true mean delivery time is between 47.65 and 52.35 hours.
    • Business implication: Daraz’s claim is plausible, but the upper bound suggests delays exceed 48 hours.

FIGURE 2: Confidence Interval for Daraz’s Delivery Time

figure:
graph LR
    A["Sample Mean: 50 hours"] --> B["± Margin of Error: ±2.352"]
    B --> C["95% CI: [47.65, 52.35] hours"]
    C --> D["Claim: ≤48 hours"]
    D -->|"Within CI"| E["Fail to reject claim"]

Advantages/Disadvantages

Advantages Disadvantages
Provides a range, not just a point. Wider intervals = less precision.
Quantifies uncertainty. Requires random sampling.
Useful for decision-making (e.g., risk assessment). Assumes normality (for small samples).

2. Hypothesis Testing: Testing Claims

Hypothesis testing answers: "Is there enough evidence to support a claim?" Steps:

  1. State hypotheses:
    • : Null hypothesis (default claim, e.g., "No effect").
    • : Alternative hypothesis (what we test, e.g., "There is an effect").
  2. Choose significance level (, e.g., 0.05).
  3. Calculate test statistic (z, t, chi-square).
  4. Compare to critical value or p-value.
  5. Make a decision: Reject or fail to reject.

FIGURE 3: Hypothesis Testing Workflow

figure:
graph TD
    A["State Hypotheses"] --> B["Choose α (e.g., 0.05)"]
    B --> C["Calculate Test Statistic"]
    C --> D["Compare to Critical Value or p-value"]
    D --> E["Decision: Reject/Fail to Reject H₀"]

Types of Tests

Test When to Use Formula Example
Z-test Large sample (), σ known. Test if Ncell’s call drop rate changed after a network upgrade.
t-test Small sample () or σ unknown. Compare average wages of TU vs. PU students.
Chi-square test Categorical data (e.g., goodness-of-fit). Check if Pathao’s surge pricing tiers fit a normal distribution.

Worked Example: eSewa Transaction Volume

eSewa claims its daily transaction volume increased after a new app update. Before: mean = 50,000 transactions/day (, σ=5,000). After: mean = 52,000 (, σ=5,000). Test at α=0.05.

  1. State hypotheses:

    • : (no increase).
    • : (increase).
  2. Calculate z-test statistic:

  3. Critical value: For α=0.05 (one-tailed), .

  4. Decision:

    • Since , reject .
    • p-value: (from z-table).
    • Since , reject .

Conclusion: There is statistically significant evidence that eSewa’s transaction volume increased after the update.


FIGURE 4: Z-Test for eSewa Transactions

figure:
graph LR
    A["Sample Mean (After): 52,000"] --> B["Population Mean (H₀): 50,000"]
    B --> C["Standard Error: 913.42"]
    C --> D["z = 2.19"]
    D --> E["Critical z: 1.645"]
    E -->|"z > critical"| F["Reject H₀"]

P-values and Decision Rules

  • P-value: Probability of observing data as extreme as the sample, assuming is true.
    • If , reject .
    • If , fail to reject .

Example Interpretation:

  • For eSewa’s test, → reject .
  • For Daraz’s CI example, since 48 hours is within the CI, we fail to reject the claim at 95% confidence.

FIGURE 5: P-value Decision Rule

figure:
graph LR
    A["Calculate p-value"] --> B["Compare to α"]
    B --> C["If p ≤ α"]
    C --> D["Reject H₀"]
    B --> E["If p > α"]
    E --> F["Fail to reject H₀"]

3. Type I and Type II Errors

Error Definition Consequence Example
Type I (α) Reject when true. False alarm. Recall a product that’s actually safe.
Type II (β) Fail to reject when false. Missed defect. Approve a loan for a risky borrower.
Power : Probability of correctly rejecting . Higher power = better test.

Trade-off:

  • Lowering (e.g., 0.01) reduces Type I errors but increases Type II errors.
  • Increasing sample size () reduces both errors.

FIGURE 6: Type I vs. Type II Errors

figure:
graph TD
    A["H₀ is True"] --> B["Correct Decision: Fail to Reject H₀"]
    A --> C["Type I Error: Reject H₀"]
    D["H₀ is False"] --> E["Type II Error: Fail to Reject H₀"]
    D --> F["Correct Decision: Reject H₀"]

Worked Example: NEPSE Stock Returns

NEPSE claims its stock returns are normally distributed with μ=12%. A sample of 40 stocks has a mean return of 10.5% and σ=4%. Test if the true mean is 12% at α=0.05.

  1. State hypotheses:

    • : .
    • : (two-tailed).
  2. Calculate t-test statistic (small sample, σ unknown):

  3. Critical t-value (df=39, α=0.05/2): .

  4. Decision:

    • Since , reject .

Conclusion: The sample evidence suggests NEPSE’s stock returns are not normally distributed with μ=12% at 95% confidence.


FIGURE 7: T-test for NEPSE Returns

figure:
graph LR
    A["Sample Mean: 10.5%"] --> B["H₀ Mean: 12%"]
    B --> C["Standard Error: 0.632"]
    C --> D["t = -2.37"]
    D --> E["Critical t: ±2.023"]
    E -->|"t < -2.023"| F["Reject H₀"]

4. Chi-Square Tests: Categorical Data

Used to test:

  • Goodness-of-fit: Does observed data fit an expected distribution?
  • Independence: Are two categorical variables independent?

Example: Does Pathao’s surge pricing tier usage fit a uniform distribution?


FIGURE 8: Chi-Square Goodness-of-Fit

figure:
graph TD
    A["Observed Frequencies"] --> B["Expected Frequencies (Uniform)"]
    B --> C["Calculate χ² = Σ (O - E)²/E"]
    C --> D["Compare to Critical χ²"]
    D --> E["Decision: Fit or Not Fit"]

Worked Example: Daraz Order Categories

*Daraz’s orders are categorized as:

  • Electronics (30%)
  • Clothing (40%)
  • Groceries (30%). A sample of 100 orders has: Electronics: 35, Clothing: 38, Groceries: 27. Test if the distribution fits Daraz’s claim at α=0.05.*
  1. State hypotheses:

    • : Observed distribution matches claimed (30%, 40%, 30%).
    • : Distribution does not match.
  2. Expected counts:

    • Electronics:
    • Clothing:
    • Groceries:
  3. Calculate χ²:

  4. Critical χ² (df=2, α=0.05): .

  5. Decision:

    • Since , fail to reject .

Conclusion: The sample distribution does not significantly differ from Daraz’s claimed categories.


FIGURE 9: Chi-Square Calculation for Daraz

figure:

mermaid table

Category Observed Expected (O-E)²/E
Electronics 35 30 0.33
Clothing 38 40 0.05
Groceries 27 30 0.33
Total 100 100 0.71

5. Real-World Applications

In the Real World

  1. eSewa/Khalti: Transaction Fraud Detection

    • Idea: Hypothesis testing checks if fraud rates increase after new payment methods.
    • Example: eSewa tests if the mean fraud rate rises from 0.5% to 1% after adding UPI. A z-test compares the new sample mean to the old rate.
  2. Ncell/NTC: Network Performance

    • Idea: Confidence intervals estimate call drop rates after network upgrades.
    • Example: Ncell claims its call drop rate is ≤2%. A 95% CI of [1.8%, 2.5%] includes 2%, so the claim is plausible.
  3. Daraz: Inventory Management

    • Idea: Chi-square tests if product categories align with demand forecasts.
    • Example: Daraz checks if observed sales (Electronics: 35%, Clothing: 38%) match expected (30%, 40%). The χ²=0.71 suggests no significant deviation.
  4. NEPSE: Stock Market Stability

    • Idea: t-tests verify if stock returns deviate from historical norms.
    • Example: If NEPSE’s sample mean return is 10.5% (vs. historical 12%), a t-test may reject the null hypothesis, signaling instability.

FIGURE 10: Real-World Examples Summary

figure:
mindmap
  root((Statistical Inference))
    eSewa: Hypothesis Testing for Fraud Rates
    Ncell: Confidence Intervals for Call Drop Rates
    Daraz: Chi-Square for Category Fit
    NEPSE: t-Tests for Stock Return Stability

Exam Tips

  1. Master the Steps:

    • Always state and clearly.
    • For z/t-tests, show the formula and plug in numbers.
    • For chi-square, build the table with O, E, and .
  2. Interpret Results Correctly:

    • "Reject " means evidence supports the alternative.
    • "Fail to reject" does not mean "prove "; it means insufficient evidence.
  3. Choose the Right Test:

    • Z-test: Large , σ known.
    • t-test: Small or σ unknown.
    • Chi-square: Categorical data.
  4. Show Work for Full Marks:

    • Write out calculations (e.g., standard error, test statistic).
    • Include p-values or critical values with decisions.
  5. Real-World Linkage:

    • Connect tests to business scenarios (e.g., "How would a bank use a t-test for loan defaults?").
    • Use examples from the unit (e.g., Daraz, eSewa) to explain concepts.
  6. Common Pitfalls:

    • Directionality: One-tailed vs. two-tailed tests.
    • Assumptions: Normality for t-tests, independence for chi-square.
    • Sample Size: Use z for , t otherwise.

FIGURE 11: Exam Checklist

figure:
graph LR
    A["State H₀ and H₀"] --> B["Choose α"]
    B --> C["Calculate Test Statistic"]
    C --> D["Compare to Critical Value or p-value"]
    D --> E["Decision: Reject/Fail to Reject"]
    E --> F["Interpret in Context"]
    F --> G["Show All Work"]

Based on the TU BBS syllabus for Business Statistics (MGT207), unit 13.

Discussion

Loading…