Business StatisticsUnit 814 min read

Statistical Inference & Hypothesis Testing: Tests, Confidence, p-Values

Unit 8 of Business Statistics: covers statistical inference (point/interval estimation), hypothesis testing (Z, t, χ², F tests), Type I/II errors, significance levels, and real-world applications like quality control, market research, and financial risk assessment.

TAKEAWAYS:

  • Statistical inference lets us estimate population parameters from sample data using confidence intervals and point estimators.
  • Hypothesis testing uses p-values and critical values to decide whether observed data supports a null hypothesis (H₀).
  • Z-tests and t-tests compare means; χ²-tests compare distributions or goodness-of-fit; F-tests compare variances.
  • Type I error (rejecting H₀ when true) and Type II error (failing to reject H₀ when false) trade off with sample size and significance level (α).
  • Business uses these tools for A/B testing (e.g., Daraz promotions), quality checks (NTC network reliability), and financial risk (NEPSE stock performance).
  • Always assume H₀ is true unless evidence overwhelmingly contradicts it (proof by contradiction).

1. Statistical Inference: Estimating Population Parameters

Statistical inference bridges the gap between sample data and population conclusions. It includes:

  • Point estimation: Single-value estimates (e.g., sample mean for population mean ).
  • Interval estimation: Confidence intervals (e.g., for 95% CI).

Key Concepts

  • Confidence Interval (CI): A range where the true parameter lies with confidence (e.g., 95% CI).

    • Formula for mean ( known):
    • For small samples () or unknown , use t-distribution:
  • Margin of Error (ME): Half the width of the CI:

Worked Example: Daraz’s Promotional Effectiveness

Scenario: Daraz runs a discount campaign and collects a sample of 100 orders. Before the campaign, the average order value was Rs 1,200 (σ = Rs 300). After the campaign, the sample mean is Rs 1,300. Test if the campaign increased order value at 95% confidence.

Steps:

  1. Hypotheses:
    • (no effect)
    • (campaign helped)
  2. Test Statistic (Z-test for large ):
  3. Critical Value: For 95% CI, .
    • Since , reject .
  4. 95% Confidence Interval: The entire interval is above 1200, confirming the campaign’s effect.

Visual: Interpretation: Daraz can claim the campaign increased order values with 95% confidence.


2. Hypothesis Testing: Framework and Steps

Hypothesis testing evaluates claims about populations using sample data. Steps:

  1. State hypotheses:
    • : Null hypothesis (status quo, e.g., "no effect").
    • : Alternative hypothesis (research claim, e.g., "there is an effect").
  2. Choose significance level (): Commonly 0.05 (5%).
  3. Select test statistic: Z, t, χ², or F based on data and hypotheses.
  4. Compute test statistic and p-value (probability of observing data if is true).
  5. Decision rule:
    • Reject if p-value < or test statistic > critical value.
    • Fail to reject otherwise.

Types of Errors

Error Type Definition Consequence Example
Type I Reject when true False alarm Accusing a Daraz seller of fraud when innocent
Type II Fail to reject when false Missed opportunity Missing a price drop in NEPSE
α Probability of Type I error Set by researcher (e.g., 0.05)
β Probability of Type II error Depends on effect size and sample size

Trade-off: Lowering reduces Type I errors but increases Type II errors (and vice versa).

Visual:

flowchart TD
    A["Start: Assume H₀ true"] --> B["Collect sample data"]
    B --> C["Calculate test statistic"]
    C --> D{"Is p-value < α?"}
    D -->|"Yes"| E["Reject H₀: Evidence against H₀"]
    D -->|"No"| F["Fail to reject H₀: Insufficient evidence"]
    E --> G["Conclude effect exists"]
    F --> G

3. Types of Hypothesis Tests

Test Use Case Test Statistic Assumptions Example in Nepal
Z-test Compare population mean (σ known) Normal distribution, large (≥30) Testing if Ncell’s call drop rate improved after network upgrade
t-test Compare population mean (σ unknown) Normal distribution, small Comparing average loan defaults between two banks (e.g., NMB vs. Global IME)
χ²-test Goodness-of-fit or independence Categorical data, expected frequencies >5 Testing if Pathao’s surge pricing affects rider distribution by time of day
F-test Compare variances (e.g., two populations) Normal distributions, independent samples Comparing variability in eSewa transaction times vs. Khalti transaction times

Worked Example: NEPSE Stock Performance

Scenario: An investor claims NEPSE’s average daily return is 0.5% (σ = 1.2%). Over 30 trading days, the sample mean return is 0.7%. Test at 95% confidence.

Steps:

  1. Hypotheses:
    • (two-tailed test)
  2. Test Statistic (Z-test):
  3. Critical Values: For 95% CI, .
    • Since , fail to reject .
  4. p-value: (from Z-table).
    • p-value > 0.05 → no significant evidence to reject .

Visual:

Interpretation: The investor’s claim lacks statistical support. The data does not prove the average return differs from 0.5%.


4. Comparing Two Populations

Independent Samples (Two-Sample t-test)

Tests if two population means are equal:

  • Equal variances: Pooled variance
  • Unequal variances: Welch’s t-test (uses separate variances).

Worked Example: NTC vs. Ncell Call Quality Scenario: NTC claims its call drop rate is lower than Ncell’s. Samples:

  • NTC: , ,
  • Ncell: , ,

Steps:

  1. Hypotheses:
    • (one-tailed)
  2. Pooled Variance:
  3. t-statistic:
  4. Critical Value: For 95% CI, one-tailed, (from t-table).
    • Since , reject .

Visual:

Interpretation: NTC’s call drop rate is significantly lower than Ncell’s at 95% confidence.

Paired Samples (Dependent t-test)

Used when data is paired (e.g., before/after measurements).

  • Test statistic: where (differences).

Worked Example: Kathmandu Traffic Before/After Flyover Scenario: Traffic speeds (km/h) before and after a flyover were recorded for 20 vehicles:

Vehicle Before After Difference ()
1 20 25 -5
2 15 20 -5
... ... ... ...
Mean = -4 km/h, km/h

Steps:

  1. Hypotheses:
    • (no effect)
    • (flyover improved speed)
  2. t-statistic:
  3. Critical Value: For 95% CI, one-tailed, .
    • Since , reject .

Visual:


5. Chi-Square (χ²) Tests

Goodness-of-Fit Test

Tests if observed frequencies match expected frequencies (e.g., die rolls, survey responses).

  • Test statistic: where = observed, = expected.

Worked Example: Pathao Rider Preferences Scenario: Pathao surveyed 200 riders on preferred pickup times:

Time Slot Observed () Expected ()
6–10 AM 50 40
10 AM–2 PM 30 50
2–6 PM 80 50
6–10 PM 40 40
10 PM–6 AM 0 20

Steps:

  1. Hypotheses:
    • : Rider preferences follow uniform distribution.
    • : Preferences are non-uniform.
  2. Expected Frequencies: for each slot.
  3. χ²-statistic:
  4. Critical Value: For 4 df (k-1) and α=0.05, .
    • Since , reject .

Visual:

Interpretation: Rider preferences are not uniform; Pathao should adjust surge pricing accordingly.

Test of Independence

Tests if two categorical variables are independent (e.g., gender vs. product choice).

  • Contingency table:

    Product A Product B Total
    Male 40 60 100
    Female 30 70 100
    Total 70 130 200
  • Test statistic same as above; degrees of freedom = .


6. F-Test: Comparing Variances

Tests if two populations have equal variances (e.g., consistency in loan processing times).

  • Test statistic:
  • Critical values from F-distribution tables.

Worked Example: Loan Processing Times Scenario: Two banks report:

  • Bank X: , days
  • Bank Y: , days

Steps:

  1. Hypotheses:
  2. F-statistic:
  3. Critical Values: For 95% CI, two-tailed, and .
    • Since , fail to reject .

Interpretation: No significant difference in processing time variability between the banks.


In the Real World

  1. eSewa’s Transaction Fraud Detection

    • Idea: Uses hypothesis testing to flag unusual transaction patterns (e.g., sudden high-volume transfers).
    • How: Compares observed transaction frequencies to expected distributions (χ²-test). If p-value < 0.05, alerts fraud teams.
    • Example: A user’s transactions spike from Rs 5,000/month to Rs 50,000 in one day. The χ²-test rejects the null hypothesis of "normal usage," triggering a review.
  2. Daraz’s A/B Testing for Discounts

    • Idea: Tests if a discount offer increases conversion rates using two-sample t-tests.
    • How: Randomly assigns users to control (no discount) and treatment (20% off) groups. If the mean order value in the treatment group is significantly higher (p < 0.05), the discount is rolled out.
    • Worked Example: Control group mean = Rs 1,200 (, σ=Rs 300); treatment group mean = Rs 1,350 (, σ=Rs 350).
      • Z-test:
      • p-value ≈ 0 → reject . Daraz concludes the discount works.
  3. NEPSE’s Stock Performance Monitoring

    • Idea: Uses confidence intervals to track if a stock’s return deviates from historical norms.
    • How: Monthly, NEPSE calculates a 95% CI for daily returns. If a stock’s return falls outside the CI (e.g., 0.5% ± 0.5%), analysts investigate anomalies.
    • Example: If a stock’s return is 1.2% and the CI is [0.0%, 1.0%], the upper bound is exceeded, prompting further analysis.

Exam Tip

  1. Master the Steps: Always follow the 5-step hypothesis testing framework (state hypotheses → choose α → select test → compute statistic → decide). Partial credit is lost if steps are skipped.
  2. Know When to Use Which Test:
    • Z-test: Large samples, σ known.
    • t-test: Small samples or σ unknown.
    • χ²-test: Categorical data (goodness-of-fit or independence).
    • F-test: Compare variances.
    • Paired t-test: Before/after or matched pairs.
  3. Calculate p-values or Critical Values:
    • For Z-tests, use standard normal tables.
    • For t-tests, use t-distribution tables (df = ).
    • For χ², use χ² tables (df = categories - 1).
    • For F-tests, use F-distribution tables (df₁, df₂).
  4. Interpret Rejection Regions:
    • Two-tailed: Reject if or .
    • One-tailed: Reject if or .
  5. Real-World Context:
    • Always tie your answer to a business scenario (e.g., "This test helps Daraz decide whether to continue a discount campaign").
  6. Show Work:
    • For calculations, write out the test statistic formula and substitute values clearly. Partial credit is given for correct setup even if the final answer is wrong.
  7. Common Pitfalls:
    • Direction of Alternative Hypothesis: (one-tailed) vs. (two-tailed) changes critical values.
    • Degrees of Freedom: For t-tests, df = ; for χ², df = (rows-1)(cols-1).
    • Assumptions: Always state assumptions (e.g., "Data is normally distributed" for t-tests).

Final Visual Summary:

mindmap
  root((Statistical Inference & Hypothesis Testing))
    Statistical_Inference
      Point_Estimation
      Interval_Estimation
        Confidence_Intervals
          Z-Intervals
          t-Intervals
    Hypothesis_Testing
      Framework
        Hypotheses
        Test_Statistics
        Decision_Rules
      Types
        Z-Tests
        t-Tests
          Independent
          Paired
        Chi-Square
          Goodness-of-Fit
          Independence
        F-Tests
    Real-World_Applications
      eSewa_Fraud_Detection
      Daraz_A_B_Testing
      NEPSE_Stock_Monitoring
    Exam_Tips
      Steps
      Test_Selection
      p-values_vs_Critical_Values
      Assumptions

Based on the TU BBA syllabus for Business Statistics (STT201), unit 8.

Discussion

Loading…