STA169 Statistics I

Statistics IUnit 712 min read

Correlation & Regression: Rank Correlation & Consistency Analysis

Unit 7 of Statistics I: Explores how to measure and analyze relationships between variables using Spearman’s rank correlation, Pearson’s correlation assumptions, regression models, and real-world consistency checks—with step-by-step calculations, visual graphs, and exam-focused applications.

TAKEAWAYS:

  • Learn how Spearman’s rank correlation measures monotonic relationships between ranked data (e.g., judge rankings in contests).
  • Understand Pearson’s correlation assumptions (linearity, normality, homoscedasticity) and when to use rank vs. Pearson methods.
  • Apply regression analysis to predict outcomes (e.g., website traffic vs. server downtime) using least-squares fitting.
  • Compare correlation vs. regression: correlation describes strength/direction; regression models the relationship.
  • Use consistency analysis (e.g., checking judge fairness in rankings) to validate data reliability.
  • Solve worked examples with step-by-step traces, including real datasets (e.g., athlete vigor vs. anger, file transmission times).

1. Correlation Analysis: Measuring Relationships

Correlation quantifies how two variables move together. It ranges from -1 (perfect negative) to +1 (perfect positive), with 0 meaning no linear relationship.

1.1 Types of Correlation

Type Description Example (Real-World)
Positive Variables increase/decrease together. Daraz order volume ↑ → delivery time ↑
Negative One variable increases as the other decreases. NTC data speed ↑ → latency ↓
Perfect Variables follow a straight-line relationship (ρ = ±1). Pathao ride distance vs. fare (fixed ratio)
Zero No linear relationship (ρ = 0). WhatsApp messages sent vs. weather temperature

1.2 Pearson vs. Spearman’s Rank Correlation

Feature Pearson’s r Spearman’s ρ
Data Type Interval/ratio scale Ordinal/ranked data
Assumptions Linear relationship, normality Monotonic (not necessarily linear)
Formula Covariance/(σ₁σ₂) 1 − (6Σd²)/(n(n²−1))
Use Case Continuous data (e.g., height vs. weight) Ranked data (e.g., judge scores)

Why Spearman? When data is ranked (e.g., music contest judges), Pearson’s method fails. Spearman’s ρ uses ranks to detect monotonic trends (e.g., higher rank in Judge 1 → higher rank in Judge 2).


2. Worked Example: Spearman’s Rank Correlation

Scenario: Three judges rank 10 contestants. Calculate Spearman’s ρ to check consistency.

Contestant Judge 1 Rank Judge 2 Rank Judge 3 Rank
A 2 1 4
B 1 2 6
... ... ... ...

Step 1: Assign ranks (ties get average rank). Step 2: Compute differences (d) between Judge 1 and Judge 2. Step 3: Square differences (d²) and sum them. Step 4: Apply formula:

Visual Trace:

Contestant Judge 1 Judge 2 d = R₁ − R₂ d²
A 2 1 1 1
B 1 2 -1 1
... ... ... ... ...
Sum Σd² = 20

Calculation: Interpretation: Strong positive correlation (ρ ≈ 0.88) → Judges agree on rankings.


3. Regression Analysis: Modeling Relationships

Regression predicts a dependent variable (Y) from an independent variable (X) using a best-fit line: where:

  • a = y-intercept,
  • b = slope (rise/run).
50100150200250300-8000-6000-4000-2000200040006000800010000xyBest-fit line (Ŷ = 0.03X + 2)Error distribution (SSE)Observed (100, 5)Observed (200, 7)Predicted (100, 4.2)Predicted (200, 6.5)
Regression Line for Server Downtime vs. Website Traffic (SSE = 12.3).

3.1 Least-Squares Method

Minimizes the sum of squared errors (SSE) between observed and predicted values.

Example: Predict server downtime (Y) from website traffic (X).

X (Traffic/day) | Y (Downtime/min) | Predicted Ŷ | Error (Y − Ŷ) | Error²
----------------|-------------------|---------------|----------------|--------
100             | 5                 | 4.2           | 0.8            | 0.64
200             | 7                 | 6.5           | 0.5            | 0.25
...             | ...               | ...           | ...            | ...    |
**SSE**         |                   |               | **Total = 12.3**|

Slope (b) and intercept (a) are calculated as:


4. Correlation vs. Regression: Key Differences

flowchart TD
    A["Correlation"] -->|"Measures strength/direction"| B["Pearson/Spearman"]
    A -->|"No prediction"| C["No equation"]
    D["Regression"] -->|"Models relationship"| E["Y = a + bX"]
    D -->|"Predicts Y"| F["Uses least-squares"]
flowchart TD
    A["Correlation"] -->|"Measures strength/direction"| B["Pearson/Spearman"]
    A -->|"No prediction"| C["No equation"]
    D["Regression"] -->|"Models relationship"| E["Y = a + bX"]
    D -->|"Predicts Y"| F["Uses least-squares"]
    G["Correlation"] -->|"ρ or r"| H["Value between -1 and +1"]
    I["Regression"] -->|"R²"| J["Coefficient of determination"]
Correlation vs. Regression: Conceptual Flowchart.
Feature Correlation Regression
Goal Describe relationship Predict Y from X
Output Coefficient (ρ or r) Equation (Ŷ = a + bX)
Direction Positive/negative Slope (b)
Strength Value between -1 and +1 R² (coefficient of determination)

5. In the Real World

  1. eSewa/Khalti Payment Processing

    • Idea: Negative correlation between transaction success rate (Y) and network latency (X).
    • How: Faster networks (lower X) → higher success rates (higher Y).
    • Example: If latency increases by 100ms, success rate drops by 3% (measured via regression).
  2. Daraz Order Fulfillment

    • Idea: Positive correlation between order volume (X) and delivery time (Y).
    • How: Spearman’s ρ checks if higher-order days consistently delay shipments.
    • Worked Example:
      | Day | Orders (X) | Delivery Time (Y) | Rank X | Rank Y |
      |-----|------------|-------------------|--------|--------|
      | 1   | 50         | 2 days           | 1      | 1      |
      | 2   | 120        | 4 days           | 3      | 3      |
      | 3   | 80         | 3 days           | 2      | 2      |
      
      ρ = 1 → Perfect positive correlation → More orders → Longer delivery.
  3. NTC Network Optimization

    • Idea: Regression to predict data speed (Y) from signal strength (X).
    • Equation: Ŷ = 5 + 0.8X (if X = signal strength in dBm).
    • Real Impact: NTC uses this to deploy towers where speed drops sharply.

6. Assumptions for Correlation/Regression

Analysis Assumptions
Pearson’s r Linear relationship, normal distribution, homoscedasticity (constant variance).
Spearman’s ρ Monotonic trend (not necessarily linear).
Regression Independent errors, no multicollinearity.

Violation Example:

  • If data is non-normal (e.g., skewed anger scores in athletes), Pearson’s r is unreliable. Use Spearman’s ρ instead.

7. Worked Example: Pearson’s Correlation (with Assumptions Check)

Scenario: Does ambient temperature (X) affect electric power (Y) in a chemical plant?

-551015202468101214xyPearson Line (Y = 0.5X + 3)Normality Check (Bell Curve)Observed (5, 5.5)Observed (10, 8)Observed (15, 10.5)Outlier (2, 4)
Pearson’s Correlation Check: Data points (n=10) with normality assumption (bell curve overlay). Outlier at (2, 4) may violate linearity.
Temp (°F) Power (kWh) (X − X̄) (Y − Ȳ) (X − X̄)(Y − Ȳ) (X − X̄)² (Y − Ȳ)²
27 120 -13 -20 260 169 400
45 150 +15 -10 -150 225 100
... ... ... ... ... ... ...

Step 1: Calculate means (X̄ = 36, Ȳ = 130). Step 2: Compute covariances and standard deviations. Step 3: Apply Pearson’s formula: Check Assumptions:

  • Linearity: Plot shows a straight trend (✅).
  • Normality: Power data is roughly symmetric (✅).
  • Homoscedasticity: Variance is constant (✅).

Conclusion: Strong positive correlation (r = 0.98) → Temperature and power are linearly related.


8. Consistency Analysis: Checking Judge Fairness

Example: Are three judges consistent in ranking athletes’ vigor (X) vs. anger (Y)?

Athlete Vigor (X) Anger (Y) Rank X Rank Y
1 30 6 1 1
2 23 7 2 2
... ... ... ... ...

Spearman’s ρ Calculation:

  • If ρ ≈ 1 → Judges agree on vigor rankings.
  • If ρ ≈ -1 → High vigor ranks correlate with low anger ranks (unlikely).

Real-World Tie: Nepal’s football team selection uses Spearman’s ρ to ensure coaches rank players consistently across speed, endurance, and aggression.


9. Exam Tip: How to Score Full Marks

  1. Definitions: Always define correlation (measure of association) and regression (prediction model) clearly.
  2. Formulas: Memorize Pearson’s r and Spearman’s ρ formulas. Show every step in calculations (e.g., covariance, standard deviation).
  3. Assumptions: List 3 key assumptions for Pearson’s r (linearity, normality, homoscedasticity) when asked.
  4. Graphs: Plot scatter plots for correlation and regression lines for prediction. Label axes properly.
  5. Real-World Link: Connect examples to Nepali apps (eSewa, Daraz) or global tools (Google Trends, WhatsApp metrics).
  6. Critical Thinking:
    • If data is ranked, use Spearman’s ρ.
    • If data is continuous and linear, use Pearson’s r + regression.
    • Always check assumptions (e.g., "Is the relationship linear?").

Common Pitfalls:

  • ❌ Mixing up correlation (descriptive) and regression (predictive).
  • ❌ Forgetting to rank data before Spearman’s ρ.
  • ❌ Ignoring assumptions (e.g., assuming normality without checking).

Final Visual Summary:

mindmap
  root((Correlation & Regression))
    Correlation
      Pearson's r
      Spearman's ρ
    Regression
      Least-Squares Line
      Assumptions Check
    Real-World Apps
      eSewa: Latency vs. Success
      Daraz: Orders vs. Delivery Time
      NTC: Signal Strength vs. Speed

Based on the TU BSc CSIT syllabus for Statistics I (STA169), unit 7.

Discussion

Loading…