STT201 Business Statistics

Business StatisticsUnit 129 min read

Non-Parametric Tests & Decision Theory: Tests, Choices & Applications

Unit 12 of Business Statistics covers non-parametric tests (when to use them, how they differ from parametric tests) and decision theory (Bayes’ theorem, payoff matrices, minimax/maximin rules). Learn the Wilcoxon, Mann-Whitney U, Kruskal-Wallis, and Chi-Square tests, their assumptions, and real-world applications in q

1. Why Non-Parametric Tests? When to Use Them

Non-parametric tests are used when data does not meet the assumptions of parametric tests (e.g., normality, homogeneity of variance). They rely on ranks or frequencies rather than raw data values.

Key Differences: Parametric vs. Non-Parametric Tests

classDiagram
    class ParametricTests {
        +Assumes normality
        +Uses raw data (means, variances)
        +Examples: t-test, ANOVA, Regression
    }
    class NonParametricTests {
        +No normality assumption
        +Uses ranks/frequencies
        +Examples: Wilcoxon, Mann-Whitney U, Kruskal-Wallis
    }
    ParametricTests --> "Requires" Assumptions
    NonParametricTests --> "Used when" AssumptionsFail

When to Choose Non-Parametric Tests?

  • Data is ordinal or nominal (e.g., survey ratings, categorical outcomes).
  • Sample size is small (<30) and data is not normally distributed.
  • Data has outliers or skewed distributions.

Example: If you test whether Pathao drivers’ earnings differ by city (Kathmandu vs. Pokhara) but the earnings data is highly skewed, use the Mann-Whitney U test instead of an independent t-test.


2. Common Non-Parametric Tests & Their Uses

05101520Wilcoxon Signed-Rank15Mann-Whitney U20Kruskal-Wallis10Chi-Square12Frequency of Use in Studies (Nepal, 2023)
Relative Popularity of Non-Parametric Tests in Business Statistics Research

(A) Wilcoxon Signed-Rank Test

  • Purpose: Compare paired samples (e.g., before/after treatment).
  • Assumptions:
    • Data is ordinal or continuous.
    • Differences between pairs are symmetrically distributed.

Worked Example: A Daraz customer service team tests if response times improve after training. Before training, average response time = 12 min; after training = 9 min. Test at 5% significance.

Steps:

  1. Calculate differences (D) between paired observations.
  2. Rank absolute differences (ignore signs).
  3. Assign signs to ranks.
  4. Calculate T = sum of smaller ranks.

Test Statistic (T):

  • Sum of positive ranks = 3 + 1 + 2 + 2 = 8
  • Sum of negative ranks = 1
  • T = min(8, 1) = 1
  • Critical value (n=5, α=0.05, two-tailed): 0
  • Decision: Reject H₀ (p < 0.05). Response times improved significantly.

(B) Mann-Whitney U Test

  • Purpose: Compare two independent samples (e.g., sales performance of two teams).
  • Assumptions:
    • Independent samples.
    • Ordinal or continuous data.

Worked Example: NTC vs. Ncell customer satisfaction scores (1-10 scale). Test if there’s a difference.

Steps:

  1. Rank all data together (ties get average rank).

  2. Calculate U₁ and U₂ using: Where:

    • (NTC), (Ncell)
    • (sum of ranks for NTC)
    • (sum of ranks for Ncell)
  3. U = min(U₁, U₂) = min(16, 40) = 16

  4. Critical value (α=0.05): 8

  5. Decision: Fail to reject H₀ (U > critical value). No significant difference.


(C) Kruskal-Wallis Test

  • Purpose: Compare three or more independent groups (e.g., sales across three regions).
  • Assumptions:
    • Independent samples.
    • Ordinal or continuous data.

Worked Example: Khalti transaction success rates in three cities (Kathmandu, Pokhara, Bharatpur). Test for differences.

Steps:

  1. Rank all data (ties get average rank).

  2. Calculate H-statistic: Where:

    • (total observations)
    • sum of ranks for each group
    • (per group)
  3. H = 7.8 (calculated)

  4. Critical value (df=2, α=0.05): 5.99

  5. Decision: Reject H₀ (H > critical value). Significant differences exist.


(D) Chi-Square Test (Goodness-of-Fit & Independence)

  • Purpose:
    • Goodness-of-fit: Test if observed frequencies match expected (e.g., coin fairness).
    • Independence: Test if two categorical variables are related (e.g., gender vs. product preference).

Worked Example (Independence): NEPSE stock performance vs. economic news sentiment (positive/negative). Test independence.

Steps:

  1. Calculate expected frequencies (E) for each cell.
  2. Compute Chi-Square (χ²):
  3. χ² = 12.5 (calculated)
  4. Critical value (df=1, α=0.05): 3.84
  5. Decision: Reject H₀. News sentiment affects stock performance.

3. Decision Theory: Making Optimal Choices

Decision theory helps choose the best action under uncertainty using probabilities and payoffs.

(A) Payoff Matrix

A table showing outcomes for different decisions under uncertain conditions.

Example: Pathao’s decision to expand to a new city (Lalitpur) based on demand.

(B) Decision Rules

  1. Maximax (Optimist): Choose the best possible outcome.

    • Decision: Expand (max payoff = 50).
  2. Maximin (Pessimist): Choose the least worst outcome.

    • Decision: Don’t expand (min payoff = -20 vs. 0).
  3. Minimax Regret: Choose to minimize maximum regret.

    • Regret table:
    • Decision: Expand (minimizes max regret = 30).
  4. Expected Value (Bayesian Approach):

    • Use probabilities of states (e.g., P(High Demand) = 0.6, P(Low Demand) = 0.4).
    • EV(Expand) = 50×0.6 + (-20)×0.4 = 22
    • EV(Don’t Expand) = 0×0.6 + 10×0.4 = 4
    • Decision: Expand (higher EV).

(C) Decision Trees

Visualize sequential decisions and their outcomes.

Example: Daraz’s decision to launch a new product.

flowchart TD
    A["Launch Product"] --> B{"Market Response"}
    B -->|"High (0.7)"| C["Profit: Rs. 50 lakhs"]
    B -->|"Low (0.3)"| D["Loss: Rs. 10 lakhs"]
    A --> E["Don't Launch"] --> F["Profit: Rs. 0"]
    G["EV(Launch)"] -->|"0.7*50 + 0.3*(-10) = 32"| H["Choose Launch"]

4. In the Real World

Company/App Non-Parametric Test Used How It’s Applied
Khalti Mann-Whitney U Test Compares transaction success rates across regions (urban vs. rural).
Daraz Wilcoxon Signed-Rank Test Tests if customer satisfaction improves after a UI redesign (paired before/after).
NTC/Ncell Kruskal-Wallis Test Compares network speeds in three cities (Kathmandu, Pokhara, Bharatpur).
NEPSE Chi-Square Test Checks if stock price movements correlate with economic news sentiment.
Pathao Decision Trees Decides whether to expand to a new city based on demand probabilities.
Banks (e.g., NMB) Mann-Whitney U Test Compares loan default rates between two demographic groups.

Worked Example (Real Tie-In): NMB Bank’s Loan Default Risk

  • Scenario: Compare default rates of urban vs. rural loans (non-normal data).
  • Test: Mann-Whitney U.
  • Data:
  • Result: U = 0 (significant difference). Rural loans have higher default risk.

5. Exam Tip: How to Score Full Marks

  1. State Assumptions Clearly:
    • Always write: "This test assumes independent samples, ordinal data, and no normality."
  2. Show All Steps:
    • Rank data, calculate test statistics, and compare with critical values.
  3. Interpret Results Properly:
    • "Reject H₀ at 5% significance. There is significant evidence that..."
  4. Decision Theory:
    • For payoff matrices, always calculate expected values if probabilities are given.
  5. Common Mistakes to Avoid:
    • Using parametric tests when data is skewed.
    • Forgetting to rank ties correctly in Mann-Whitney U.
    • Misinterpreting p-values (e.g., "p > 0.05 means accept H₀" is incorrect; say "fail to reject").

Quick Checklist for Non-Parametric Tests:

flowchart LR
    A["Check Data Type"] --> B{"Normal?"}
    B -->|"Yes"| C["Use Parametric Test"]
    B -->|"No"| D["Use Non-Parametric Test"]
    D --> E{"Paired Data?"}
    E -->|"Yes"| F["Wilcoxon Signed-Rank"]
    E -->|"No"| G{"Independent Groups?"}
    G -->|"2 Groups"| H["Mann-Whitney U"]
    G -->|">2 Groups"| I["Kruskal-Wallis"]
    G -->|"Categorical"| J["Chi-Square"]
    C --> K["Examples: t-test, ANOVA"]
    F --> L["Example: Paired sample comparison"]
    H --> M["Example: Two independent groups"]
    I --> N["Example: >2 independent groups"]
    J --> O["Example: Goodness-of-fit/Independence"]

Final Note: Non-parametric tests are powerful for real-world data where assumptions fail. Decision theory helps businesses minimize risk (e.g., Pathao’s expansion, NMB’s loans). Practice ranking data and interpreting test statistics—this is where marks are lost in exams!

Based on the TU BBM syllabus for Business Statistics (STT201), unit 12.

Discussion

Loading…