RCH311 Business Research Methods

Business Research MethodsUnit 1010 min read

Hypothesis Testing & Statistical Analysis: Tests, Errors, Validity, Reliability

Unit 10 of Business Research Methods: Explores how researchers use statistical tests to validate hypotheses, differentiate parametric/non-parametric methods, and ensure reliability/validity in data—with real-world applications in finance, marketing, and operations.

TAKEAWAYS:

  • Hypothesis testing uses null (H₀) and alternative (H₁) hypotheses to decide whether observed data supports a claim, with Type I (false positive) and Type II (false negative) errors as trade-offs.
  • Parametric tests (e.g., t-tests, ANOVA) assume normal distribution and specific data types, while non-parametric tests (e.g., chi-square, Mann-Whitney U) do not.
  • Validity ensures a test measures what it claims (e.g., content, construct, external validity), while reliability checks consistency (e.g., test-retest, split-half).
  • Chi-square tests (e.g., independence, goodness-of-fit) analyze categorical data, like survey responses or market segment distributions.
  • P-value approach rejects H₀ if P < significance level (e.g., 0.01 for 99% confidence), balancing risk of Type I/II errors.
  • Real-world tools: Banks use t-tests for loan default risk; Daraz applies chi-square to optimize inventory placement.

1. Hypothesis Testing: The Core Process

Hypothesis testing is the scientific method for deciding whether sample data supports a claim about a population. It starts with two hypotheses:

  • Null hypothesis (H₀): Default assumption (e.g., "No effect," "No difference").
  • Alternative hypothesis (H₁): Researcher’s claim (e.g., "There is an effect," "Mean income differs by gender").

The process:

  1. State hypotheses (directional or non-directional).
  2. Choose significance level (α) (e.g., 0.05 for 95% confidence).
  3. Select test statistic (e.g., z-test, t-test, chi-square).
  4. Compute test statistic from sample data.
  5. Compare to critical value or use P-value approach:
    • If P < α, reject H₀ (evidence supports H₁).
    • Else, fail to reject H₀.
flowchart TD
    A["State H₀ and H₁"] --> B["Set significance level α (e.g., 0.05)"]
    B --> C["Select appropriate test (z-test, t-test, chi-square, etc.)"]
    C --> D["Compute test statistic (e.g., z = (x̄ - μ₀)/(σ/√n))"]
    D --> E["Compare to critical value OR
     check P-value"]
    E -->|"P < α"| F["Reject H₀: Evidence supports H₁"]
    E -->|"P ≥ α"| G["Fail to reject H₀: Insufficient evidence"]
    F -->|"Type I error"| H["α risk: False positive"]
    G -->|"Type II error"| I["β risk: False negative"]

| A flowchart showing the steps of hypothesis testing: state hypotheses → set significance level → select test → compute statistic → compare to critical value/P-value → decide. |


2. Types of Errors in Hypothesis Testing

Errors occur when we misinterpret data:

  • Type I Error (α-error): False positive—rejecting H₀ when it’s true. Example: A bank approves a loan (rejects H₀: "borrower is low-risk") but the borrower defaults (H₀ was true).
  • Type II Error (β-error): False negative—failing to reject H₀ when it’s false. Example: A hospital misses a disease (fails to reject H₀: "patient is healthy") when they’re sick (H₀ is false).
Reject H₀ when true (False Positive)Risk: α (e.g., 5%)Type I Error (α)Fail to reject H₀ when false (False Negative)Risk: β (depends on sample size, effect size)Type II Error (β)Types of Errors
Classification of errors in hypothesis testing with their consequences.

Trade-off: Lowering α reduces Type I errors but increases Type II errors, and vice versa. Power of a test (1 − β): Probability of correctly rejecting H₀ when it’s false. Higher power = better test.


3. Parametric vs. Non-Parametric Tests

Tests differ by data assumptions:

Feature Parametric Tests Non-Parametric Tests
Data type Interval/ratio, normally distributed Ordinal/nominal, non-normal
Examples t-test, ANOVA, regression Chi-square, Mann-Whitney U, Kruskal-Wallis
Assumptions Normality, homogeneity of variance No assumptions on distribution
When to use Continuous data, large samples Categorical data, small/non-normal samples
Example in Nepal NEPSE tests stock returns (parametric) Daraz tests product category preferences (non-parametric)

| A comparison table listing key differences between parametric and non-parametric tests, with Nepali business examples. |


4. Validity in Research

Validity ensures a test measures what it claims. Types:

  1. Content Validity: Does the test cover the topic? (e.g., A survey on customer satisfaction includes all key aspects.)
  2. Construct Validity: Does the test measure the theoretical construct? (e.g., A "stress test" for employees measures psychological stress.)
  3. External Validity: Can results generalize? (e.g., A study on Kathmandu traffic applies to other cities.)
  4. Internal Validity: Are causal conclusions valid? (e.g., Does a training program cause sales increases, or is it correlation?)

How to check validity:

  • Face validity: Expert review (e.g., a professor checks a questionnaire).
  • Criterion validity: Compare to a gold-standard measure (e.g., correlate a new IQ test with an established one).

5. Reliability in Research

Reliability checks consistency. Types:

  1. Test-Retest Reliability: Same test on two occasions (e.g., Re-administering a customer satisfaction survey after 2 weeks).
  2. Equivalent Form Reliability: Two parallel versions (e.g., Two forms of a math test for students).
  3. Split-Half Reliability: Split test into halves and correlate (e.g., Odd vs. even questions on a personality test).
  4. Inter-Rater Reliability: Agreement between raters (e.g., Two judges scoring Pathao driver performance).

6. Chi-Square Tests: Analyzing Categorical Data

Chi-square tests compare observed vs. expected frequencies in categorical data. Types:

  1. Chi-Square Goodness-of-Fit: Tests if observed data fits a distribution (e.g., Does a coin flip have a 50-50 bias?).
  2. Chi-Square Test of Independence: Tests if two categorical variables are associated (e.g., Is Ncell usage independent of age group?).

Worked Example: Daraz Inventory Placement Scenario: Daraz wants to know if product placement affects sales. They categorize products by shelf position (A, B, C) and sales (high/low). Hypothesis:

  • H₀: Placement has no effect on sales (independent).
  • H₁: Placement affects sales (dependent).
Shelf High Sales Low Sales Total
A 40 10 50
B 20 30 50
C 10 40 50
Total 70 80 150

Steps:

  1. Calculate expected frequencies (e.g., for A & High Sales: (50 × 70)/150 = 23.33).
  2. Compute chi-square statistic: (Plug in values for all cells.)
  3. Compare to critical value (df = (rows−1)(cols−1) = 2; α=0.05 → critical value ≈ 5.99).
    • If χ² > 5.99, reject H₀ (placement affects sales).

Real-world tie: Daraz uses this to optimize shelf space, reducing stockouts and increasing revenue.


7. P-Value Approach vs. Critical Value Approach

Method P-Value Approach Critical Value Approach
Decision Rule Reject H₀ if P < α Reject H₀ if test statistic > critical value
Flexibility Adapts to sample size Fixed critical values (e.g., z₀.₀₅ = 1.96)
Example For α=0.01, reject if P < 0.01 For z-test, reject if z > 2.33
When to Use Large samples, complex tests Small samples, standard tests

| A comparison table showing the decision rules, flexibility, and examples of P-value and critical value approaches. |


8. In the Real World

  1. Ncell’s Customer Segmentation:
    • Uses chi-square tests to analyze if usage patterns (e.g., data vs. voice) differ by age group. If χ² shows dependence, Ncell tailors marketing (e.g., discounts for heavy data users).
  2. Nabil Bank’s Loan Approvals:
    • Applies t-tests to compare default rates between loan applicants with high/low credit scores. Rejecting H₀ (no difference) justifies stricter criteria for low-score applicants.
  3. Pathao’s Driver Scheduling:
    • Uses reliability checks (e.g., test-retest) to validate survey data on driver availability. High reliability ensures accurate demand forecasting.

Exam Tip

  • Focus on definitions: Clearly state null/alternative hypotheses, Type I/II errors, and validity/reliability types.
  • Show calculations: For chi-square or t-tests, write the formula and plug in numbers (even if hypothetical). Examiners reward step-by-step logic.
  • Link to real data: Use Nepali examples (e.g., NEPSE stock analysis, Daraz inventory) to explain parametric/non-parametric choices.
  • P-value vs. critical value: Know when to use each (e.g., P-value for complex tests, critical value for z/t-tests).
  • Common pitfalls:
    • Confusing validity (what you measure) with reliability (consistency).
    • Forgetting degrees of freedom for chi-square (df = (rows−1)(cols−1)).
    • Misinterpreting rejecting H₀ as "proving H₁" (it only means enough evidence exists).

Sample Exam Answer Structure:

  1. Define hypothesis testing and errors (1 mark each).
  2. Compare parametric/non-parametric tests with a Nepali example (2 marks).
  3. Calculate chi-square for a given table (3 marks: steps + formula + conclusion).
  4. Discuss reliability methods in a survey (2 marks).

Final Note: This unit is 50% calculation-heavy. Practice chi-square, t-tests, and P-value interpretations with real datasets (e.g., NEB’s past papers). For validity/reliability, explain with examples—examiners love concrete Nepali cases!

Based on the TU BBM syllabus for Business Research Methods (RCH311), unit 10.

Discussion

Loading…