Business Research MethodsUnit 916 min read
Hypothesis Testing: Types, Steps, and Statistical Tools
Unit 9 of Business Research Methods explores hypothesis testing—how to formulate, test, and interpret hypotheses using statistical tools, significance levels, and decision rules. Learn about null/alternative hypotheses, test statistics, p-values, and real-world applications in business decisions.
TAKEAWAYS:
- Hypothesis testing is a structured method to make data-driven decisions by comparing a null hypothesis () against an alternative ().
- Key steps include formulating hypotheses, selecting a significance level (α), choosing a test statistic, calculating p-values, and making a decision.
- Types of errors (Type I and Type II) and power of a test determine the reliability of conclusions.
- Businesses use hypothesis testing for A/B testing (e.g., Daraz’s ad campaigns), customer satisfaction surveys (e.g., Ncell’s Net Promoter Score), and financial forecasting (e.g., Nabil Bank’s loan approval models).
- Parametric tests (e.g., t-tests, ANOVA) assume normal distribution, while non-parametric tests (e.g., Chi-square, Mann-Whitney U) do not.
- The p-value and critical value approach are two methods to reject or fail to reject , with α (e.g., 0.05) defining the threshold for significance.
1. What is Hypothesis Testing?
Hypothesis testing is a statistical method used to make inferences about a population based on sample data. It helps researchers or businesses test assumptions (hypotheses) to determine if observed effects are statistically significant or due to random chance.
Key Definitions
- Null Hypothesis (): The default assumption (e.g., "There is no effect" or "The new marketing strategy does not increase sales").
- Alternative Hypothesis ( or ): The claim we want to test (e.g., "The new strategy increases sales").
- Significance Level (α): The probability of rejecting when it is true (e.g., α = 0.05 means a 5% risk of a false positive).
- Test Statistic: A standardized value (e.g., z-score, t-score) calculated from sample data to compare against a critical value.
- p-value: The probability of observing the test statistic (or more extreme) if is true. A low p-value (≤ α) suggests rejecting .
2. Steps in Hypothesis Testing
flowchart TD
A["1. State Hypotheses"] --> B["2. Choose Significance Level (α)"]
B --> C["3. Select Test Statistic"]
C --> D["4. Collect Data & Calculate Test Statistic"]
D --> E["5. Determine Critical Value or p-value"]
E --> F["6. Make Decision: Reject or Fail to Reject \(H_0\)"]
F --> G["7. Draw Conclusion"]Step-by-Step Explanation
Formulate Hypotheses
- Example: A bank (e.g., Nabil Bank) wants to test if a new loan approval process reduces processing time.
- : The new process does not reduce time (μ ≥ 10 days).
- : The new process reduces time (μ < 10 days).
- Example: A bank (e.g., Nabil Bank) wants to test if a new loan approval process reduces processing time.
Choose α (Significance Level)
- Common values: 0.01 (strict), 0.05 (standard), or 0.10 (less strict).
- Trade-off: Lower α reduces false positives but increases false negatives.
Select Test Statistic
- Depends on data type and assumptions:
- Parametric tests (normal distribution assumed):
- t-test (small samples, population variance unknown).
- z-test (large samples, known population variance).
- ANOVA (compare means across >2 groups).
- Non-parametric tests (no distribution assumption):
- Chi-square test (categorical data, e.g., customer preferences).
- Mann-Whitney U test (ordinal data, e.g., survey rankings).
- Parametric tests (normal distribution assumed):
- Depends on data type and assumptions:
Calculate Test Statistic
- Example: Daraz tests if a new checkout button increases conversion rates.
- Sample: 1000 users with old button (3% conversion) vs. 1000 with new button (4% conversion).
- Use a two-proportion z-test: where , , .
- Example: Daraz tests if a new checkout button increases conversion rates.
Determine Critical Value or p-value
- Critical Value Approach: Compare test statistic to a critical value from a table (e.g., z = ±1.96 for α = 0.05).
- p-value Approach: Find the probability of the test statistic under . If p ≤ α, reject .
Make Decision
- Reject : Strong evidence supports (e.g., the new checkout button works).
- Fail to Reject : Insufficient evidence (e.g., no significant difference in sales).
Draw Conclusion
- Example: If p = 0.02 < α = 0.05, conclude the new button significantly increases conversions.
3. Types of Hypothesis Tests
| Test Type | When to Use | Example in Nepal | Formula/Key Idea |
|---|---|---|---|
| One-sample t-test | Compare sample mean to a known value. | Test if NTC’s average customer wait time (8 min) differs from the target (5 min). | |
| Two-sample t-test | Compare means of two independent groups. | Compare Khalti’s vs. eSewa’s transaction success rates. | |
| Paired t-test | Compare means of the same group before/after. | Test if Pathao’s driver earnings increase after a bonus scheme. | |
| Chi-square test | Test independence in categorical data. | Check if NEPSE’s stock performance correlates with political events. | |
| ANOVA | Compare means across >2 groups. | Compare Daraz’s sales across 3 regions (Kathmandu, Pokhara, Biratnagar). | |
| Mann-Whitney U test | Non-parametric alternative to t-test. | Compare Himalayan Java’s customer satisfaction ratings (ordinal data). | Rank-transformed data, calculate U statistic. |
4. Errors in Hypothesis Testing
Two types of errors can occur:
| Error Type | Definition | Example | Consequence |
|---|---|---|---|
| Type I Error (False Positive) | Reject when it’s true. | Nabil Bank approves a risky loan due to a false signal in data. | Financial loss, reputational damage. |
| Type II Error (False Negative) | Fail to reject when it’s false. | Ncell misses a decline in customer satisfaction due to weak survey data. | Lost revenue, poor customer retention. |
- Power of a Test: Probability of correctly rejecting (1 − β). Higher power reduces Type II errors.
- How to Increase Power:
- Increase sample size.
- Use a higher α (but increases Type I error).
- Reduce variability in data.
5. Real-World Applications in Nepal
Case 1: Daraz’s A/B Testing for Ad Campaigns
- Problem: Daraz wants to test if a new ad banner increases click-through rates (CTR).
- Hypotheses:
- : CTR with new banner = CTR with old banner (5%).
- : CTR with new banner > 5%.
- Method: Two-sample z-test on 5000 users per group.
- Result: New banner shows a CTR of 5.8% (p = 0.03 < 0.05). Conclusion: Roll out the new banner.
Case 2: Ncell’s Net Promoter Score (NPS) Analysis
- Problem: Ncell wants to check if a new customer service training improves NPS.
- Hypotheses:
- : NPS after training = NPS before training (50).
- : NPS after training > 50.
- Method: Paired t-test on 200 customers surveyed before/after training.
- Result: Mean NPS increases from 50 to 58 (p = 0.001). Conclusion: Training is effective.
Case 3: Nabil Bank’s Loan Default Prediction
- Problem: Predict if a borrower will default using credit history.
- Method: Logistic regression (a form of hypothesis testing for classification).
- Hypotheses:
- : Credit history has no predictive power.
- : Credit history predicts default risk.
- Result: p < 0.01 for credit score variable. Conclusion: Use credit scores to approve/reject loans.
6. Parametric vs. Non-Parametric Tests
| Feature | Parametric Tests | Non-Parametric Tests |
|---|---|---|
| Assumptions | Normal distribution, equal variances. | No distribution assumption. |
| Data Type | Interval/ratio data. | Ordinal/categorical data. |
| Examples | t-test, ANOVA, z-test. | Chi-square, Mann-Whitney U, Kruskal-Wallis. |
| When to Use | Large samples, symmetric data. | Small samples, skewed data. |
| Power | More powerful if assumptions hold. | Less powerful but robust to violations. |
Example:
- Parametric: Test if Khalti’s average transaction time (in seconds) differs from eSewa’s (use t-test).
- Non-parametric: Test if Pathao’s driver ratings (1-5 stars) differ by region (use Kruskal-Wallis).
7. Decision Rules: p-value vs. Critical Value
| Approach | Steps | Example |
|---|---|---|
| p-value | 1. Calculate p-value. | For a t-test, p = 0.02. |
| 2. Compare p to α. | If α = 0.05, p < α → reject . | |
| Critical Value | 1. Find critical value (e.g., z = ±1.96). | For α = 0.05, two-tailed test. |
| 2. Compare test statistic to critical value. | If test statistic > 1.96, reject . |
Visual Decision Rule:
mindmap
root((Decision Rule))
p-value
"p ≤ α → Reject \(H_0\)"
"p > α → Fail to reject \(H_0\)"
Critical Value
"Test Statistic > Critical Value → Reject \(H_0\)"
"Test Statistic ≤ Critical Value → Fail to reject \(H_0\)"8. Common Mistakes to Avoid
- Ignoring Assumptions: Using a t-test when data is not normally distributed (→ use Mann-Whitney U).
- Multiple Testing: Running 10 tests increases Type I error risk (use Bonferroni correction).
- P-hacking: Manipulating data to achieve p < 0.05 (unethical and invalidates results).
- Confusing Correlation and Causation: Hypothesis testing shows association, not necessarily cause (e.g., ice cream sales and drowning deaths both rise in summer, but neither causes the other).
9. Software Tools for Hypothesis Testing
| Tool | Use Case | Example Command |
|---|---|---|
| Excel | Basic t-tests, z-tests, Chi-square. | =T.TEST(array1, array2, tails, type) |
| SPSS | Advanced statistical tests, ANOVA. | Analyze → Compare Means → Independent-Samples T Test |
| R | Customizable tests, large datasets. | t.test(x, y, paired = TRUE) |
| Python (SciPy) | Automated testing. | from scipy.stats import ttest_ind |
| JASP | User-friendly alternative to SPSS. | Drag-and-drop interface for tests. |
Example in Python:
from scipy.stats import ttest_ind
group1 = [12, 15, 14, 10, 13] # Old checkout time (Daraz)
group2 = [9, 11, 10, 8, 12] # New checkout time
t_stat, p_val = ttest_ind(group2, group1)
print(f"p-value: {p_val}") # Output: p-value = 0.04 → Reject H0
In the Real World
eSewa and Khalti: Fraud Detection
- Idea Used: Chi-square goodness-of-fit test to detect unusual transaction patterns (e.g., sudden spikes in refunds).
- How: Compare observed transaction frequencies to expected frequencies under normal conditions. If p < 0.05, flag as potential fraud.
Daraz: A/B Testing for UI Changes
- Idea Used: Two-sample t-test to compare conversion rates between old and new website layouts.
- How: Split users randomly into two groups. If the new layout shows a significantly higher conversion rate (p < 0.05), roll it out globally.
Ncell: Customer Churn Prediction
- Idea Used: Logistic regression (a hypothesis-testing framework) to predict churn based on call drop rates, bill payments, and usage data.
- How: Test if "high call drop rate" () predicts churn (: no relationship). If p < 0.01, include the variable in the model.
Nabil Bank: Loan Approval Models
- Idea Used: ANOVA to compare default rates across income groups.
- How: Test if borrowers with income < Rs. 50,000 have higher default rates than those earning > Rs. 100,000. If p < 0.05, adjust loan terms by income bracket.
NTC: Network Performance Testing
- Idea Used: One-sample t-test to check if average internet speed meets regulatory standards (e.g., 50 Mbps).
- How: Sample 50 users’ speeds. If the mean is significantly below 50 Mbps (p < 0.05), investigate network issues.
Exam Tip
How This Unit is Examined
Theory Questions (30%):
- Define null/alternative hypotheses, Type I/II errors, and p-value.
- Explain the difference between parametric and non-parametric tests.
- Common Pitfalls: Students often confuse and or misapply tests (e.g., using ANOVA for two groups instead of a t-test).
Problem-Solving (50%):
- Given: A scenario (e.g., "A company tests two advertising methods").
- Task: Formulate hypotheses, choose a test, calculate test statistic, and interpret results.
- Example Question:
"A sample of 30 students has an average exam score of 75 with a standard deviation of 10. Test if the true mean is 80 at α = 0.05." Solution: Use a one-sample t-test: Critical t-value (df = 29, α = 0.05) = ±2.045. Since |−2.74| > 2.045, reject .
Case Studies (20%):
- Given: A real-world scenario (e.g., "Pathao wants to test if a new incentive increases driver earnings").
- Task: Design a hypothesis test, justify your choice of test, and explain implications.
- Tip: Always link your answer to business decisions (e.g., "If p < 0.05, Pathao should scale the incentive").
Marks Distribution in Exams
| Component | Weightage | What to Focus On |
|---|---|---|
| Definitions (e.g., p-value) | 10% | Memorize key terms and their meanings. |
| Test Selection | 15% | Know when to use t-test, Chi-square, ANOVA. |
| Calculations | 35% | Practice formulas and critical values. |
| Interpretation | 20% | Explain results in business context. |
| Errors and Power | 10% | Understand Type I/II errors and how to minimize them. |
Quick Revision Checklist
- Can you formulate and for a given scenario?
- Do you know when to use parametric vs. non-parametric tests?
- Can you calculate a t-test or z-test from scratch?
- Do you understand what a p-value means and how to interpret it?
- Can you explain the consequences of Type I/II errors in a business context?
Final Worked Example: Kathmandu Traffic Congestion Study
Scenario: The Kathmandu Metropolitan City wants to test if a new traffic signal system reduces average wait times at a busy intersection.
Step 1: Formulate Hypotheses
- : The new system does not reduce wait time (μ ≥ 2 minutes).
- : The new system reduces wait time (μ < 2 minutes).
Step 2: Choose Test
- One-sample t-test (since we compare a sample mean to a known value).
- Assumptions: Wait times are approximately normally distributed (checked via histogram).
Step 3: Collect Data
- Sample: 40 drivers timed at the intersection.
- Sample mean () = 1.8 minutes.
- Sample standard deviation () = 0.5 minutes.
- α = 0.05.
Step 4: Calculate Test Statistic
Step 5: Determine Critical Value
- Degrees of freedom (df) = n − 1 = 39.
- Critical t-value (one-tailed, α = 0.05) ≈ −1.685 (from t-table).
Step 6: Make Decision
- Since |−2.53| > 1.685, reject .
Conclusion: The new traffic signal system significantly reduces wait times (p < 0.05). The city should implement it citywide.
Based on the TU BITM syllabus for Business Research Methods (RCH201), unit 9.
Discussion
Loading…