Basic StatisticsUnit 810 min read
Non-Parametric Statistics & Data Analysis
Unit 8 of Basic Statistics: Covers non-parametric tests (rank-based), data analysis tools (box plots, chi-square), and real-world applications like ranking systems and categorical data analysis—with step-by-step examples and visuals.
TAKEAWAYS
- Non-parametric tests do not assume data follows a specific distribution (e.g., normal), making them ideal for ordinal or non-normal data.
- Rank-based methods (e.g., Mann-Whitney U, Kruskal-Wallis) compare medians instead of means, reducing sensitivity to outliers.
- Box plots visually summarize median, quartiles, and outliers—critical for detecting skewed distributions.
- Chi-square tests analyze categorical data (e.g., preferences, survey responses) to check independence or goodness-of-fit.
- Secondary data (e.g., NEPSE stock trends, NTC call records) is cheaper/faster but requires validation for biases.
- Real-world tie-ins: Pathao’s driver ratings (rank correlation), Daraz’s product reviews (chi-square), Ncell’s survey data (non-parametric trends).
1. Introduction to Non-Parametric Statistics
Non-parametric methods do not rely on population parameters (e.g., mean, variance) or strict distributional assumptions. They are robust for:
- Ordinal data (e.g., survey rankings: "Poor/Medium/Good").
- Non-normal data (e.g., skewed income distributions).
- Small sample sizes where parametric tests (e.g., t-tests) fail.
Key Advantages
flowchart TD A["Non-parametric tests"] --> B["No assumption of **normality** (works with skewed data)"] A --> C["Handles **ordinal/categorical** data (e.g., Likert scales)"] A --> D["Robust to **outliers** (unlike t-tests)"] A --> E["Flexible for **small samples** or **non-normal distributions**"]
When to Use?
| Scenario | Parametric Test | Non-parametric Test |
|---|---|---|
| Normal data, continuous | t-test, ANOVA | Mann-Whitney U, Kruskal-Wallis |
| Ordinal data | ❌ Not applicable | Spearman’s rank correlation |
| Categorical data | ❌ Not applicable | Chi-square test |
| Small samples | Risky | Wilcoxon signed-rank test |
2. Rank-Based Tests
(A) Mann-Whitney U Test
Purpose: Compare medians of two independent groups (non-parametric alternative to independent t-test). Example: Comparing Pathao driver ratings (1–5 stars) between two cities.
Worked Example: Data: Ratings from Lalitpur (n=8) and Pokhara (n=7).
Lalitpur: 4, 5, 3, 2, 5, 4, 1, 3
Pokhara: 5, 4, 2, 3, 1, 4, 2
Steps:
- Rank all 15 values (ties get average rank):
Ranked data: 1, 1, 2, 2, 2, 3, 3, 3, 4, 4, 4, 5, 5, 5, 5 - Sum ranks for each group:
- Lalitpur:
- Pokhara:
- Calculate U: Where , , :
- Compare to critical U (from tables) at α=0.05 → Reject H₀ if U < 20 (one-tailed).
Visual:
(B) Kruskal-Wallis Test
Purpose: Extend Mann-Whitney to ≥3 independent groups (non-parametric ANOVA). Example: Comparing Ncell, NTC, Smartphone call drop rates across 3 cities.
3. Correlation and Rank Correlation
(A) Spearman’s Rank Correlation (ρ)
Purpose: Measure monotonic relationship between two ordinal/ranked variables (non-parametric alternative to Pearson’s r). Formula: Where .
Worked Example: Data: 10 students’ preferences for DELL vs HP (ranks 1–10).
Student | DELL Rank | HP Rank | \(d_i\) | \(d_i^2\)
--------|-----------|---------|---------|--------
1 | 5 | 10 | -5 | 25
2 | 2 | 9 | -7 | 49
... | ... | ... | ... | ...
10 | 10 | 5 | +5 | 25
Steps:
- Compute .
- Plug into formula:
- Interpret: Strong negative rank correlation (ρ ≈ -0.82).
Visual:
4. Chi-Square Tests
(A) Chi-Square Goodness-of-Fit
Purpose: Test if observed frequencies match expected frequencies (e.g., Daraz product categories). Test Statistic: Example: Do 100 Daraz shoppers prefer categories equally?
Category | Observed (O) | Expected (E=25)
-----------|--------------|-----------------
Electronics| 40 | 25
Clothing | 30 | 25
Books | 20 | 25
Others | 10 | 25
Steps:
- Compute .
- Calculate :
- Compare to critical χ² (df=3, α=0.05) → Reject H₀ (χ²₀.₀₅₃=7.81 < 54).
Visual:
(B) Chi-Square Test of Independence
Purpose: Check if two categorical variables are independent (e.g., Ncell plan choice vs gender). Example: Do men and women choose Ncell plans equally?
| Plan A | Plan B | Total
----------|--------|--------|-------
Male | 30 | 20 | 50
Female | 25 | 25 | 50
Total | 55 | 45 | 100
Steps:
- Compute expected counts (e.g., for Male/Plan A: ).
- Calculate :
- Compare to critical χ² (df=1, α=0.05) → Fail to reject H₀ (χ²₀.₀₅₁=3.84 > 2.67).
5. Box-and-Whisker Plots
Purpose: Visualize median, quartiles, and outliers in skewed data. Components:
graph LR A["Whisker (Q1 - 1.5*IQR)"] --> B["Q1"] B --> C["Box (Q1 to Q3)"] C --> D["Median (line)"] C --> E["Q3"] E --> F["Whisker (Q3 + 1.5*IQR)"]
Worked Example: Data: Marks of Student A (10 tests).
44, 80, 76, 48, 52, 72, 68, 56, 60, 54
Steps:
- Order data: 44, 48, 52, 54, 56, 60, 68, 72, 76, 80.
- Find Q1, Q3, IQR:
- Q1 (median of first 5): 52
- Q3 (median of last 5): 72
- IQR = 72 − 52 = 20
- Outliers: Lower bound = 52 − 1.5×20 = 22 (none below 44). Upper bound = 72 + 30 = 102 (none above 80).
- Plot:
6. Secondary Data Analysis
Definition: Data collected by others (e.g., NEPSE stock prices, NTC call records). Sources in Nepal:
| Source | Example Data | Caution |
|---|---|---|
| Government | Census, NEPSE listings | Bias: Underreporting (e.g., income) |
| Banks | Loan defaults, interest rates | Timeliness: Delayed reporting |
| Private Companies | Daraz sales, Pathao ratings | Purpose: May hide flaws |
| Media | News polls, surveys | Sampling: Non-random respondents |
Worked Example: Data: NEPSE’s top 5 stocks (2023) with monthly returns.
Stock | Jan | Feb | Mar | Mean | Std Dev
--------|-----|-----|-----|------|--------
ABC | 2% | 5% | -1%| 2.33%| 2.52%
XYZ | 3% | 0% | 4% | 2.67%| 2.08%
Analysis:
- Median return: (2% + 3% + 4%)/3 = 3% (better than mean for skewed data).
- Box plot would show XYZ’s lower variability (smaller whiskers).
In the Real World
Pathao’s Driver Ratings:
- Uses Spearman’s rank correlation to compare driver performance across cities.
- Example: If Lalitpur drivers rank higher in "Punctuality" (ρ=0.75) than Pokhara (ρ=0.45), Pathao may train Pokhara drivers.
Daraz’s Product Categories:
- Chi-square test checks if shoppers’ preferences match Daraz’s stock distribution.
- Example: If observed Electronics sales (40%) > expected (25%), Daraz may increase inventory.
Ncell’s Survey Data:
- Kruskal-Wallis test compares call drop rates by plan type (Prepaid vs Postpaid) across regions.
- Example: If Postpaid users in Kathmandu have significantly higher drops (U=12 < critical U), Ncell may upgrade infrastructure.
Exam Tip
- Focus on 3 key tests:
- Mann-Whitney U (2 groups) → Compare medians.
- Spearman’s ρ (rank correlation) → Monotonic trends.
- Chi-square (categorical data) → Independence/goodness-of-fit.
- Always show calculations for χ² and U tests (partial marks for steps).
- Box plots are high-weight in exams—practice labeling Q1, Q3, outliers.
- Secondary data: Cite real Nepalese sources (NEPSE, NTC) and discuss bias risks.
- Rank correlation vs Pearson: Use Spearman for ordinal data, Pearson for interval/ratio.
Key Formulae to Memorize:
- Spearman’s ρ:
- Mann-Whitney U:
- Chi-square:
Based on the TU BIT syllabus for Basic Statistics (STA154), unit 8.
Discussion
Loading…