STA154 Basic Statistics

Basic StatisticsUnit 810 min read

Non-Parametric Statistics & Data Analysis

Unit 8 of Basic Statistics: Covers non-parametric tests (rank-based), data analysis tools (box plots, chi-square), and real-world applications like ranking systems and categorical data analysis—with step-by-step examples and visuals.

TAKEAWAYS

  • Non-parametric tests do not assume data follows a specific distribution (e.g., normal), making them ideal for ordinal or non-normal data.
  • Rank-based methods (e.g., Mann-Whitney U, Kruskal-Wallis) compare medians instead of means, reducing sensitivity to outliers.
  • Box plots visually summarize median, quartiles, and outliers—critical for detecting skewed distributions.
  • Chi-square tests analyze categorical data (e.g., preferences, survey responses) to check independence or goodness-of-fit.
  • Secondary data (e.g., NEPSE stock trends, NTC call records) is cheaper/faster but requires validation for biases.
  • Real-world tie-ins: Pathao’s driver ratings (rank correlation), Daraz’s product reviews (chi-square), Ncell’s survey data (non-parametric trends).

1. Introduction to Non-Parametric Statistics

Non-parametric methods do not rely on population parameters (e.g., mean, variance) or strict distributional assumptions. They are robust for:

  • Ordinal data (e.g., survey rankings: "Poor/Medium/Good").
  • Non-normal data (e.g., skewed income distributions).
  • Small sample sizes where parametric tests (e.g., t-tests) fail.

Key Advantages

flowchart TD
  A["Non-parametric tests"] --> B["No assumption of **normality** (works with skewed data)"]
  A --> C["Handles **ordinal/categorical** data (e.g., Likert scales)"]
  A --> D["Robust to **outliers** (unlike t-tests)"]
  A --> E["Flexible for **small samples** or **non-normal distributions**"]

When to Use?

Scenario Parametric Test Non-parametric Test
Normal data, continuous t-test, ANOVA Mann-Whitney U, Kruskal-Wallis
Ordinal data ❌ Not applicable Spearman’s rank correlation
Categorical data ❌ Not applicable Chi-square test
Small samples Risky Wilcoxon signed-rank test

2. Rank-Based Tests

01234567891011121314151 (L)1 (P)2 (L), 2 (P)2 (P)3 (L), 3 (P)3 (P)4 (L), 4 (P)4 (P)
Ranked data for Lalitpur (L) and Pokhara (P) driver ratings (1–5). Ties receive average rank.
UNon-parametric TestsParametric TestsNormal data, interval/ratio scales (e.g., t-test)Non-normal data, ordinal/categorical (e.g., Mann-Whitney)
When to choose non-parametric tests: Non-normality or non-interval data.

(A) Mann-Whitney U Test

Purpose: Compare medians of two independent groups (non-parametric alternative to independent t-test). Example: Comparing Pathao driver ratings (1–5 stars) between two cities.

Worked Example: Data: Ratings from Lalitpur (n=8) and Pokhara (n=7).

Lalitpur: 4, 5, 3, 2, 5, 4, 1, 3
Pokhara:  5, 4, 2, 3, 1, 4, 2

Steps:

  1. Rank all 15 values (ties get average rank):
    Ranked data: 1, 1, 2, 2, 2, 3, 3, 3, 4, 4, 4, 5, 5, 5, 5
    
  2. Sum ranks for each group:
    • Lalitpur:
    • Pokhara:
  3. Calculate U: Where , , :
  4. Compare to critical U (from tables) at α=0.05 → Reject H₀ if U < 20 (one-tailed).

Visual:

(B) Kruskal-Wallis Test

Purpose: Extend Mann-Whitney to ≥3 independent groups (non-parametric ANOVA). Example: Comparing Ncell, NTC, Smartphone call drop rates across 3 cities.


3. Correlation and Rank Correlation

12345678910246810xDELL Rank (X)HP Rank (Y)Student 1Student 2Student 5Student 10
Spearman’s rank correlation (ρ ≈ -0.82): Inverse relationship between DELL and HP preference ranks.

(A) Spearman’s Rank Correlation (ρ)

Purpose: Measure monotonic relationship between two ordinal/ranked variables (non-parametric alternative to Pearson’s r). Formula: Where .

Worked Example: Data: 10 students’ preferences for DELL vs HP (ranks 1–10).

Student | DELL Rank | HP Rank | \(d_i\) | \(d_i^2\)
--------|-----------|---------|---------|--------
1      | 5         | 10      | -5      | 25
2      | 2         | 9       | -7      | 49
...    | ...       | ...     | ...     | ...
10     | 10        | 5       | +5      | 25

Steps:

  1. Compute .
  2. Plug into formula:
  3. Interpret: Strong negative rank correlation (ρ ≈ -0.82).

Visual:


4. Chi-Square Tests

010203040Electronics40Clothing30Books20Others10Number of Shoppers (n=100)
Chi-square goodness-of-fit test: Observed vs. expected shopper categories (χ² = 54, p < 0.05 → reject H₀).

(A) Chi-Square Goodness-of-Fit

Purpose: Test if observed frequencies match expected frequencies (e.g., Daraz product categories). Test Statistic: Example: Do 100 Daraz shoppers prefer categories equally?

Category   | Observed (O) | Expected (E=25)
-----------|--------------|-----------------
Electronics| 40           | 25
Clothing    | 30           | 25
Books       | 20           | 25
Others      | 10           | 25

Steps:

  1. Compute .
  2. Calculate :
  3. Compare to critical χ² (df=3, α=0.05) → Reject H₀ (χ²₀.₀₅₃=7.81 < 54).

Visual:

(B) Chi-Square Test of Independence

Purpose: Check if two categorical variables are independent (e.g., Ncell plan choice vs gender). Example: Do men and women choose Ncell plans equally?

          | Plan A | Plan B | Total
----------|--------|--------|-------
Male      | 30     | 20     | 50
Female    | 25     | 25     | 50
Total     | 55     | 45     | 100

Steps:

  1. Compute expected counts (e.g., for Male/Plan A: ).
  2. Calculate :
  3. Compare to critical χ² (df=1, α=0.05) → Fail to reject H₀ (χ²₀.₀₅₁=3.84 > 2.67).

5. Box-and-Whisker Plots

Purpose: Visualize median, quartiles, and outliers in skewed data. Components:

graph LR
  A["Whisker (Q1 - 1.5*IQR)"] --> B["Q1"]
  B --> C["Box (Q1 to Q3)"]
  C --> D["Median (line)"]
  C --> E["Q3"]
  E --> F["Whisker (Q3 + 1.5*IQR)"]
UQ3 (72)Q1 (52)44, 48, 5268, 72, 76, 80
Student A’s test marks: Q1=52, Q3=72, IQR=20 (no outliers).

Worked Example: Data: Marks of Student A (10 tests).

44, 80, 76, 48, 52, 72, 68, 56, 60, 54

Steps:

  1. Order data: 44, 48, 52, 54, 56, 60, 68, 72, 76, 80.
  2. Find Q1, Q3, IQR:
    • Q1 (median of first 5): 52
    • Q3 (median of last 5): 72
    • IQR = 72 − 52 = 20
  3. Outliers: Lower bound = 52 − 1.5×20 = 22 (none below 44). Upper bound = 72 + 30 = 102 (none above 80).
  4. Plot:

6. Secondary Data Analysis

Definition: Data collected by others (e.g., NEPSE stock prices, NTC call records). Sources in Nepal:

Source Example Data Caution
Government Census, NEPSE listings Bias: Underreporting (e.g., income)
Banks Loan defaults, interest rates Timeliness: Delayed reporting
Private Companies Daraz sales, Pathao ratings Purpose: May hide flaws
Media News polls, surveys Sampling: Non-random respondents

Worked Example: Data: NEPSE’s top 5 stocks (2023) with monthly returns.

Stock   | Jan | Feb | Mar | Mean | Std Dev
--------|-----|-----|-----|------|--------
ABC     | 2% | 5% | -1%| 2.33%| 2.52%
XYZ     | 3% | 0% | 4% | 2.67%| 2.08%

Analysis:

  • Median return: (2% + 3% + 4%)/3 = 3% (better than mean for skewed data).
  • Box plot would show XYZ’s lower variability (smaller whiskers).

In the Real World

  1. Pathao’s Driver Ratings:

    • Uses Spearman’s rank correlation to compare driver performance across cities.
    • Example: If Lalitpur drivers rank higher in "Punctuality" (ρ=0.75) than Pokhara (ρ=0.45), Pathao may train Pokhara drivers.
  2. Daraz’s Product Categories:

    • Chi-square test checks if shoppers’ preferences match Daraz’s stock distribution.
    • Example: If observed Electronics sales (40%) > expected (25%), Daraz may increase inventory.
  3. Ncell’s Survey Data:

    • Kruskal-Wallis test compares call drop rates by plan type (Prepaid vs Postpaid) across regions.
    • Example: If Postpaid users in Kathmandu have significantly higher drops (U=12 < critical U), Ncell may upgrade infrastructure.

Exam Tip

  • Focus on 3 key tests:
    1. Mann-Whitney U (2 groups) → Compare medians.
    2. Spearman’s ρ (rank correlation) → Monotonic trends.
    3. Chi-square (categorical data) → Independence/goodness-of-fit.
  • Always show calculations for χ² and U tests (partial marks for steps).
  • Box plots are high-weight in exams—practice labeling Q1, Q3, outliers.
  • Secondary data: Cite real Nepalese sources (NEPSE, NTC) and discuss bias risks.
  • Rank correlation vs Pearson: Use Spearman for ordinal data, Pearson for interval/ratio.

Key Formulae to Memorize:

  1. Spearman’s ρ:
  2. Mann-Whitney U:
  3. Chi-square:

Based on the TU BIT syllabus for Basic Statistics (STA154), unit 8.

Discussion

Loading…