Business Research MethodsUnit 913 min read
Hypothesis Testing: Types, Tests & Decision Rules
Unit 9 of Business Research Methods explores hypothesis testing—how to formulate, test, and interpret hypotheses using statistical tools, significance levels, and decision-making frameworks to draw valid business conclusions.
TAKEAWAYS:
- Hypothesis testing is the process of making data-driven decisions by comparing sample statistics to population parameters using null and alternative hypotheses.
- Key steps include formulating hypotheses, selecting a test statistic, determining significance level (α), and making a decision (reject/fail to reject H₀).
- Common tests include Z-test (large samples), t-test (small samples), Chi-square (categorical data), and ANOVA (multiple groups).
- Type I (false positive) and Type II (false negative) errors guide risk management in business decisions.
- Real-world applications include A/B testing in eSewa’s app updates, market segmentation by Daraz, and credit risk assessment by Nabil Bank.
- The p-value and critical value approaches are two methods to evaluate hypothesis test results.
1. What is Hypothesis Testing?
Hypothesis testing is a statistical method used to make inferences about a population based on sample data. It helps researchers answer questions like:
- "Does the new marketing campaign increase sales?"
- "Is there a significant difference in customer satisfaction between two product versions?"
- "Does employee training improve productivity?"
Key Definitions
| Term | Definition | Example |
|---|---|---|
| Null Hypothesis (H₀) | Assumes no effect or no difference (default position). | "The new ad campaign does not increase sales." |
| Alternative Hypothesis (H₁ or Ha) | Claims an effect exists (what the researcher aims to prove). | "The new ad campaign increases sales." |
| Significance Level (α) | Probability of rejecting H₀ when it’s true (Type I error). Typically 0.05 (5%). | If α = 0.05, there’s a 5% chance of a false positive. |
| Test Statistic | A standardized value (e.g., Z, t, Chi-square) calculated from sample data. | For a sample mean, Z = (X̄ - μ) / (σ/√n). |
| p-value | Probability of observing the test statistic if H₀ is true. Lower p-value → stronger evidence against H₀. | p = 0.03 < α → Reject H₀. |
| Critical Value | Threshold value from statistical tables (e.g., Z = ±1.96 for α = 0.05). | If test statistic > critical value, reject H₀. |
2. Steps in Hypothesis Testing
flowchart TD
A["1. Formulate Hypotheses"] --> B["2. Choose Significance Level (α)"]
B --> C["3. Select Test Statistic"]
C --> D["4. Collect & Analyze Data"]
D --> E["5. Calculate Test Statistic"]
E --> F["6. Determine p-value or Compare to Critical Value"]
F --> G["7. Make Decision (Reject/Fail to Reject *H₀*)"]
G --> H["8. Draw Conclusion"]Step-by-Step Worked Example: Daraz’s A/B Test
Scenario: Daraz wants to test if a new checkout button color (green vs. red) increases conversion rates.
- Population: All Daraz users in Nepal.
- Sample: 1,000 users randomly assigned to two groups.
- Hypotheses:
- H₀: μ_red = μ_green (no difference in conversion rates).
- H₁: μ_red ≠ μ_green (conversion rates differ).
Data:
- Red button group: 12% conversion (n₁ = 500).
- Green button group: 14% conversion (n₂ = 500).
Test: Two-sample Z-test (since sample sizes are large). Calculation:
- Calculate pooled standard deviation (s_p).
- Compute Z-statistic:
- Compare Z to critical value (Z = ±1.96 for α = 0.05) or find p-value.
Result: Z = 2.10 > 1.96 → Reject H₀. Conclusion: The green button significantly improves conversion rates (p < 0.05).
3. Types of Hypothesis Tests
| Test Type | When to Use | Example in Nepal |
|---|---|---|
| Z-test | Large sample size (n > 30), known population σ. | NTC testing if average call drop rate exceeds 2% (historical data available). |
| t-test | Small sample size (n ≤ 30), unknown σ. | Nabil Bank comparing loan default rates before/after new policy (n = 25 loans). |
| Chi-square Test | Categorical data (e.g., survey responses, contingency tables). | NEPSE analyzing if investor preferences differ by age group (e.g., young vs. old). |
| ANOVA | Comparing 3+ groups (e.g., product preferences across regions). | Daraz testing sales performance across Kathmandu, Pokhara, and Biratnagar. |
| Correlation (Pearson/Spearman) | Testing relationship between two variables. | Pathao studying if ride duration correlates with customer ratings. |
4. Errors in Hypothesis Testing
Two types of errors can occur:
mindmap
root((Errors in Hypothesis Testing))
Type I Error["False Positive\n(α error)\nReject *H₀* when true"]
Example["NEPSE approves a stock IPO that later fails."]
Cost["Wasted resources, lost investor trust."]
Type II Error["False Negative\n(β error)\nFail to reject *H₀* when false"]
Example["Nabil Bank rejects a low-risk loan applicant."]
Cost["Missed business opportunities."]
Trade-off["α ↑ → β ↓\nα ↓ → β ↑"]Real-World Impact:
- Type I Error (False Alarm): WhatsApp banning a user’s account for "suspicious activity" when they’re innocent.
- Type II Error (Missed Opportunity): Google not detecting a bug in an app update because the test sample was too small.
5. Decision Rules: p-value vs. Critical Value
| Method | Process | Example |
|---|---|---|
| p-value Approach | Compare p-value to α. If p ≤ α, reject H₀. | p = 0.02 ≤ 0.05 → Reject H₀. |
| Critical Value Approach | Compare test statistic to critical value (from Z/t tables). | Z = 2.3 > 1.96 → Reject H₀. |
Worked Example: Ncell’s Customer Satisfaction Survey Question: Does Ncell’s new customer service training improve satisfaction scores?
- H₀: μ_before = μ_after (no improvement).
- H₁: μ_after > μ_before (one-tailed test).
- Sample: 40 customers before/after training.
- Test: Paired t-test (same customers surveyed twice).
- Result: t = 2.5, p = 0.01 < 0.05 → Reject H₀. Conclusion: Training significantly improved satisfaction (p < 0.05).
6. Hypothesis Testing in Business Research
Applications in Nepali Companies
| Company | Research Question | Hypothesis Test Used | Outcome |
|---|---|---|---|
| eSewa | Does the new OTP verification reduce fraud? | Chi-square test | Rejected H₀: Fraud cases dropped by 30% (p < 0.01). |
| Nabil Bank | Does credit score predict loan defaults? | Logistic Regression (Chi-square) | Accepted H₀: Score > 650 → 90% repayment rate. |
| Daraz | Are sales higher on weekends? | One-way ANOVA | Rejected H₀: Weekend sales significantly higher (F = 4.2, p = 0.03). |
| NTC | Is call quality worse in monsoon season? | Two-sample t-test | Failed to reject H₀: No significant difference (p = 0.12). |
Case Study: Himalayan Java’s Market Expansion
Problem: Himalayan Java wants to test if organic coffee sales differ by region (Kathmandu vs. Pokhara). Hypotheses:
- H₀: μ_Kathmandu = μ_Pokhara.
- H₁: μ_Kathmandu ≠ μ_Pokhara.
Data:
- Kathmandu (n = 100): Mean sales = $500, σ = $50.
- Pokhara (n = 80): Mean sales = $450, σ = $40.
Test: Two-sample t-test (unequal variances). Calculation: Critical t-value (df = 178, α = 0.05): ±1.97. Decision: |3.16| > 1.97 → Reject H₀. Conclusion: Sales differ significantly (p < 0.05). Action: Targeted marketing in Pokhara.
7. Common Mistakes to Avoid
- Ignoring Assumptions: Using a t-test when data is not normally distributed.
- Incorrect Hypothesis Formulation: Testing H₀: "Sales increase" (should be H₀: "Sales do not increase").
- P-hacking: Adjusting α or sample size to force significance.
- Overlooking Effect Size: A significant p-value doesn’t always mean a practical difference (e.g., 1% vs. 0.9% conversion rates).
- Multiple Testing: Running 10 tests increases Type I error risk (use Bonferroni correction).
## In the Real World
eSewa’s Fraud Detection
- Idea Used: Chi-square test for categorical data.
- How: eSewa tests if transaction patterns (time, amount) differ between fraudulent and legitimate users. A significant Chi-square result triggers alerts.
- Example: If fraudsters use transactions > $500 at night more often (p < 0.05), eSewa flags such transactions for manual review.
Nabil Bank’s Loan Approval
- Idea Used: Logistic Regression (Chi-square-based).
- How: The bank tests whether factors like income, credit score, and employment status predict loan defaults. A rejected H₀ (e.g., "Credit score doesn’t matter") leads to stricter approval criteria.
- Example: Data shows applicants with scores < 600 default 40% of the time (p < 0.001), so Nabil Bank raises the minimum score to 650.
Daraz’s Dynamic Pricing
- Idea Used: ANOVA for multi-group comparisons.
- How: Daraz tests if product prices should vary by region (Kathmandu vs. rural areas). If ANOVA shows significant differences (p < 0.05), prices are adjusted dynamically.
- Example: A laptop sells for $800 in Kathmandu but $750 in Pokhara (F = 5.2, p = 0.02), so Daraz sets regional price tiers.
Pathao’s Driver Ratings
- Idea Used: Correlation analysis (Pearson’s r).
- How: Pathao tests if driver ratings correlate with ride duration or cancellation rates. A strong negative correlation (r = -0.7, p < 0.05) might lead to bonuses for high-rated drivers.
- Example: Drivers with >4.5 ratings have 20% fewer cancellations (p = 0.01), so Pathao promotes them.
NEPSE’s Investor Sentiment
- Idea Used: Z-test for proportions.
- How: NEPSE tests if investor confidence differs before/after policy changes. A rejected H₀ (e.g., "Confidence is unchanged") triggers communication strategies.
- Example: 60% of investors were bullish before a policy change vs. 70% after (Z = 2.5, p = 0.01), indicating improved sentiment.
## Exam Tip
How This Unit is Examined in PU (Pokhara University):
Theoretical Questions (30%):
- Define H₀, H₁, p-value, and Type I/II errors.
- Differentiate between parametric (Z, t) and non-parametric (Chi-square) tests.
- Explain when to use one-tailed vs. two-tailed tests.
Problem-Solving (50%):
- Given: Hypotheses, sample data, and significance level.
- Do:
- Identify the correct test (Z, t, Chi-square, ANOVA).
- Calculate test statistic (show formulas).
- Compare to critical value or find p-value.
- State decision (reject/fail to reject H₀) and conclusion.
- Example Question:
"A sample of 50 Ncell users has an average call drop rate of 3% (σ = 0.5). Test at α = 0.05 if drops exceed the industry standard of 2%."
Case Studies (20%):
- Apply hypothesis testing to real business scenarios (e.g., bank loan defaults, e-commerce sales).
- Tip: Always link your answer to business decisions (e.g., "Reject H₀ → Launch the campaign").
Common Pitfalls in Exams:
- Forgetting to state assumptions (e.g., normality, independence).
- Misinterpreting one-tailed vs. two-tailed tests (e.g., using Z = ±1.645 for one-tailed).
- Incorrectly calculating degrees of freedom (e.g., t-test df = n - 1, ANOVA df = groups - 1).
- Ignoring effect size: Even if p < 0.05, ask if the difference is meaningful (e.g., 51% vs. 50% conversion).
Quick Revision Checklist:
Final Note: Hypothesis testing is the bridge between data and decision-making. Master it, and you’ll ace business research—and help companies like Nabil Bank, Daraz, and NTC make data-driven choices!
Based on the PU BBA (PU) syllabus for Business Research Methods, unit 9.
Discussion
Loading…