STT201 Business Statistics

Business StatisticsUnit 117 min read

ANOVA & Chi-Square Tests: Hypothesis, F-Test, Goodness-of-Fit

Unit 11 of Business Statistics introduces ANOVA (Analysis of Variance) for comparing means across ≥3 groups and Chi-Square tests for categorical data analysis, including goodness-of-fit and independence tests. Learn F-distribution, test statistics, assumptions, and real-world applications in quality control, market seg

Key Concepts & Definitions

1. Analysis of Variance (ANOVA)

ANOVA is a statistical method used to compare the means of three or more groups to determine if at least one group mean is different from the others. It partitions total variability into:

  • Between-group variability (due to differences between group means)
  • Within-group variability (due to natural variation within each group)

Assumptions of ANOVA

mindmap
  root((ANOVA Assumptions))
    Normality["Data in each group is normally distributed"]
    Homogeneity["Variances (σ²) are equal across groups (homoscedasticity)"]
    Independence["Observations are independent"]
    Random["Samples are randomly selected"]

Steps in ANOVA

  1. State Hypotheses:
    • : All group means are equal ()
    • : At least one group mean is different.
  2. Calculate Sum of Squares (SS):
    • Total SS (SST): Total variation in data.
    • Between-group SS (SSB): Variation due to group differences.
    • Within-group SS (SSW): Variation within groups.
  3. Compute Mean Squares (MS):
    • (Between-group mean square)
    • (Within-group mean square)
  4. Calculate F-statistic:
  5. Compare with Critical F-value (from F-distribution table) or compute p-value.

2. Chi-Square Tests

Chi-Square tests analyze categorical data (frequencies) and include:

  • Goodness-of-Fit Test: Checks if observed frequencies match expected frequencies.
  • Test of Independence: Tests if two categorical variables are independent.

Assumptions of Chi-Square Tests

mindmap
  root((Chi-Square Assumptions))
    Independence["Observations are independent"]
    Expected["Expected frequency ≥5 in ≥80% of cells"]
    Categorical["Data is categorical (counts/frequencies)"]

Steps in Chi-Square Test

  1. State Hypotheses:
    • : Observed frequencies = Expected frequencies (or variables are independent).
    • : They are not equal (or variables are dependent).
  2. Compute Chi-Square Statistic: where = observed frequency, = expected frequency.
  3. Determine Critical Value (from Chi-Square table) or compute p-value.

Worked Examples

Example 1: One-Way ANOVA (Comparing 3 Groups)

Scenario: A company tests the effectiveness of 3 different advertising methods (TV, Radio, Social Media) on sales. Sample data (in Rs. '000):

Group TV Radio Social Media
Sample Size 10 10 10
Mean Sales 50 45 60
Variance 100 80 120

Step 1: Calculate SSB (Between-group SS) where

Step 2: Calculate SSW (Within-group SS)

Step 3: Degrees of Freedom

Step 4: Mean Squares

Step 5: F-statistic

Step 6: Compare with Critical F-value From F-table (): Critical .

Since , fail to reject . No significant difference in sales across advertising methods.



Example 2: Chi-Square Goodness-of-Fit Test

Scenario: A factory claims 20% of its products are defective. In a sample of 200, 50 were defective. Test the claim at .

Step 1: State Hypotheses

  • : (20% defective)
  • :

Step 2: Expected Frequencies

Step 3: Chi-Square Statistic

Step 4: Critical Value From Chi-Square table (): Critical .

Since , fail to reject . The claim is not rejected.



In the Real World

1. ANOVA in E-Sewa & Khalti (Nepal)

  • Scenario: E-Sewa tests 3 payment methods (Mobile Banking, Credit Card, UPI) to see which drives higher transaction success rates.
  • How ANOVA Helps:
    • Compares mean success rates across methods.
    • If , at least one method performs significantly better.
    • Example: If Mobile Banking has a higher mean success rate, E-Sewa can prioritize it.

2. Chi-Square in Daraz & Pathao (Market Segmentation)

  • Scenario: Daraz wants to check if customer preferences for product categories (Electronics, Fashion, Grocery) are independent of age groups (18-25, 26-35, 36+).
  • How Chi-Square Helps:
    • Tests if age group affects product choice.
    • If , preferences are not independent (e.g., younger users prefer Electronics).
    • Example: Pathao uses this to tailor promotions (e.g., discounts on food for 26-35 age group).

3. ANOVA in NTC (Network Performance)

  • Scenario: NTC tests 3 internet service providers (Ncell, NTC, Smart) for download speeds in Kathmandu.
  • How ANOVA Helps:
    • Compares mean speeds across providers.
    • If is significant, NTC can identify the best-performing ISP for upgrades.

Comparison Table: ANOVA vs. Chi-Square

Feature ANOVA Chi-Square Test
Data Type Continuous (means) Categorical (frequencies)
Groups Compared ≥3 groups 2 variables (rows × columns)
Test Statistic F-distribution Chi-Square ()
Assumptions Normality, Homogeneity, Independence Expected frequency ≥5, Independence
Use Case Compare group means (e.g., test scores, sales) Test proportions (e.g., survey responses, defect rates)

Exam Tip

  1. ANOVA Questions:

    • Always state hypotheses clearly.
    • Show all calculations (SSB, SSW, MS, F).
    • Use F-table to compare critical values.
    • Interpret results: "Reject " means at least one group is different.
  2. Chi-Square Questions:

    • Expected frequencies must be calculated carefully ().
    • Check assumptions (e.g., no expected frequency <5).
    • Degrees of freedom for independence test: (rows × columns).
  3. Common Mistakes to Avoid:

    • Forgetting to subtract 1 for degrees of freedom in ANOVA.
    • Misapplying Chi-Square to continuous data (use ANOVA instead).
    • Ignoring assumptions (e.g., non-normal data in ANOVA).

Based on the TU BITM syllabus for Business Statistics (STT201), unit 11.

Discussion

Loading…