Business StatisticsUnit 117 min read
ANOVA & Chi-Square Tests: Hypothesis, F-Test, Goodness-of-Fit
Unit 11 of Business Statistics introduces ANOVA (Analysis of Variance) for comparing means across ≥3 groups and Chi-Square tests for categorical data analysis, including goodness-of-fit and independence tests. Learn F-distribution, test statistics, assumptions, and real-world applications in quality control, market seg
Key Concepts & Definitions
1. Analysis of Variance (ANOVA)
ANOVA is a statistical method used to compare the means of three or more groups to determine if at least one group mean is different from the others. It partitions total variability into:
- Between-group variability (due to differences between group means)
- Within-group variability (due to natural variation within each group)
Assumptions of ANOVA
mindmap
root((ANOVA Assumptions))
Normality["Data in each group is normally distributed"]
Homogeneity["Variances (σ²) are equal across groups (homoscedasticity)"]
Independence["Observations are independent"]
Random["Samples are randomly selected"]Steps in ANOVA
- State Hypotheses:
- : All group means are equal ()
- : At least one group mean is different.
- Calculate Sum of Squares (SS):
- Total SS (SST): Total variation in data.
- Between-group SS (SSB): Variation due to group differences.
- Within-group SS (SSW): Variation within groups.
- Compute Mean Squares (MS):
- (Between-group mean square)
- (Within-group mean square)
- Calculate F-statistic:
- Compare with Critical F-value (from F-distribution table) or compute p-value.
2. Chi-Square Tests
Chi-Square tests analyze categorical data (frequencies) and include:
- Goodness-of-Fit Test: Checks if observed frequencies match expected frequencies.
- Test of Independence: Tests if two categorical variables are independent.
Assumptions of Chi-Square Tests
mindmap
root((Chi-Square Assumptions))
Independence["Observations are independent"]
Expected["Expected frequency ≥5 in ≥80% of cells"]
Categorical["Data is categorical (counts/frequencies)"]Steps in Chi-Square Test
- State Hypotheses:
- : Observed frequencies = Expected frequencies (or variables are independent).
- : They are not equal (or variables are dependent).
- Compute Chi-Square Statistic: where = observed frequency, = expected frequency.
- Determine Critical Value (from Chi-Square table) or compute p-value.
Worked Examples
Example 1: One-Way ANOVA (Comparing 3 Groups)
Scenario: A company tests the effectiveness of 3 different advertising methods (TV, Radio, Social Media) on sales. Sample data (in Rs. '000):
| Group | TV | Radio | Social Media |
|---|---|---|---|
| Sample Size | 10 | 10 | 10 |
| Mean Sales | 50 | 45 | 60 |
| Variance | 100 | 80 | 120 |
Step 1: Calculate SSB (Between-group SS) where
Step 2: Calculate SSW (Within-group SS)
Step 3: Degrees of Freedom
Step 4: Mean Squares
Step 5: F-statistic
Step 6: Compare with Critical F-value From F-table (): Critical .
Since , fail to reject . No significant difference in sales across advertising methods.
Example 2: Chi-Square Goodness-of-Fit Test
Scenario: A factory claims 20% of its products are defective. In a sample of 200, 50 were defective. Test the claim at .
Step 1: State Hypotheses
- : (20% defective)
- :
Step 2: Expected Frequencies
Step 3: Chi-Square Statistic
Step 4: Critical Value From Chi-Square table (): Critical .
Since , fail to reject . The claim is not rejected.
In the Real World
1. ANOVA in E-Sewa & Khalti (Nepal)
- Scenario: E-Sewa tests 3 payment methods (Mobile Banking, Credit Card, UPI) to see which drives higher transaction success rates.
- How ANOVA Helps:
- Compares mean success rates across methods.
- If , at least one method performs significantly better.
- Example: If Mobile Banking has a higher mean success rate, E-Sewa can prioritize it.
2. Chi-Square in Daraz & Pathao (Market Segmentation)
- Scenario: Daraz wants to check if customer preferences for product categories (Electronics, Fashion, Grocery) are independent of age groups (18-25, 26-35, 36+).
- How Chi-Square Helps:
- Tests if age group affects product choice.
- If , preferences are not independent (e.g., younger users prefer Electronics).
- Example: Pathao uses this to tailor promotions (e.g., discounts on food for 26-35 age group).
3. ANOVA in NTC (Network Performance)
- Scenario: NTC tests 3 internet service providers (Ncell, NTC, Smart) for download speeds in Kathmandu.
- How ANOVA Helps:
- Compares mean speeds across providers.
- If is significant, NTC can identify the best-performing ISP for upgrades.
Comparison Table: ANOVA vs. Chi-Square
| Feature | ANOVA | Chi-Square Test |
|---|---|---|
| Data Type | Continuous (means) | Categorical (frequencies) |
| Groups Compared | ≥3 groups | 2 variables (rows × columns) |
| Test Statistic | F-distribution | Chi-Square () |
| Assumptions | Normality, Homogeneity, Independence | Expected frequency ≥5, Independence |
| Use Case | Compare group means (e.g., test scores, sales) | Test proportions (e.g., survey responses, defect rates) |
Exam Tip
ANOVA Questions:
- Always state hypotheses clearly.
- Show all calculations (SSB, SSW, MS, F).
- Use F-table to compare critical values.
- Interpret results: "Reject " means at least one group is different.
Chi-Square Questions:
- Expected frequencies must be calculated carefully ().
- Check assumptions (e.g., no expected frequency <5).
- Degrees of freedom for independence test: (rows × columns).
Common Mistakes to Avoid:
- Forgetting to subtract 1 for degrees of freedom in ANOVA.
- Misapplying Chi-Square to continuous data (use ANOVA instead).
- Ignoring assumptions (e.g., non-normal data in ANOVA).
Based on the TU BITM syllabus for Business Statistics (STT201), unit 11.
Discussion
Loading…