Statistics IUnit 712 min read
Correlation & Regression: Rank Correlation & Consistency Analysis
Unit 7 of Statistics I: Explores how to measure and analyze relationships between variables using Spearman’s rank correlation, Pearson’s correlation assumptions, regression models, and real-world consistency checks—with step-by-step calculations, visual graphs, and exam-focused applications.
TAKEAWAYS:
- Learn how Spearman’s rank correlation measures monotonic relationships between ranked data (e.g., judge rankings in contests).
- Understand Pearson’s correlation assumptions (linearity, normality, homoscedasticity) and when to use rank vs. Pearson methods.
- Apply regression analysis to predict outcomes (e.g., website traffic vs. server downtime) using least-squares fitting.
- Compare correlation vs. regression: correlation describes strength/direction; regression models the relationship.
- Use consistency analysis (e.g., checking judge fairness in rankings) to validate data reliability.
- Solve worked examples with step-by-step traces, including real datasets (e.g., athlete vigor vs. anger, file transmission times).
1. Correlation Analysis: Measuring Relationships
Correlation quantifies how two variables move together. It ranges from -1 (perfect negative) to +1 (perfect positive), with 0 meaning no linear relationship.
1.1 Types of Correlation
| Type | Description | Example (Real-World) |
|---|---|---|
| Positive | Variables increase/decrease together. | Daraz order volume ↑ → delivery time ↑ |
| Negative | One variable increases as the other decreases. | NTC data speed ↑ → latency ↓ |
| Perfect | Variables follow a straight-line relationship (ρ = ±1). | Pathao ride distance vs. fare (fixed ratio) |
| Zero | No linear relationship (ρ = 0). | WhatsApp messages sent vs. weather temperature |
1.2 Pearson vs. Spearman’s Rank Correlation
| Feature | Pearson’s r | Spearman’s ρ |
|---|---|---|
| Data Type | Interval/ratio scale | Ordinal/ranked data |
| Assumptions | Linear relationship, normality | Monotonic (not necessarily linear) |
| Formula | Covariance/(σ₁σ₂) | 1 − (6Σd²)/(n(n²−1)) |
| Use Case | Continuous data (e.g., height vs. weight) | Ranked data (e.g., judge scores) |
Why Spearman? When data is ranked (e.g., music contest judges), Pearson’s method fails. Spearman’s ρ uses ranks to detect monotonic trends (e.g., higher rank in Judge 1 → higher rank in Judge 2).
2. Worked Example: Spearman’s Rank Correlation
Scenario: Three judges rank 10 contestants. Calculate Spearman’s ρ to check consistency.
| Contestant | Judge 1 Rank | Judge 2 Rank | Judge 3 Rank |
|---|---|---|---|
| A | 2 | 1 | 4 |
| B | 1 | 2 | 6 |
| ... | ... | ... | ... |
Step 1: Assign ranks (ties get average rank). Step 2: Compute differences (d) between Judge 1 and Judge 2. Step 3: Square differences (d²) and sum them. Step 4: Apply formula:
Visual Trace:
| Contestant | Judge 1 | Judge 2 | d = R₁ − R₂ | d² |
|---|---|---|---|---|
| A | 2 | 1 | 1 | 1 |
| B | 1 | 2 | -1 | 1 |
| ... | ... | ... | ... | ... |
| Sum | Σd² = 20 |
Calculation: Interpretation: Strong positive correlation (ρ ≈ 0.88) → Judges agree on rankings.
3. Regression Analysis: Modeling Relationships
Regression predicts a dependent variable (Y) from an independent variable (X) using a best-fit line: where:
- a = y-intercept,
- b = slope (rise/run).
3.1 Least-Squares Method
Minimizes the sum of squared errors (SSE) between observed and predicted values.
Example: Predict server downtime (Y) from website traffic (X).
X (Traffic/day) | Y (Downtime/min) | Predicted Ŷ | Error (Y − Ŷ) | Error²
----------------|-------------------|---------------|----------------|--------
100 | 5 | 4.2 | 0.8 | 0.64
200 | 7 | 6.5 | 0.5 | 0.25
... | ... | ... | ... | ... |
**SSE** | | | **Total = 12.3**|
Slope (b) and intercept (a) are calculated as:
4. Correlation vs. Regression: Key Differences
flowchart TD
A["Correlation"] -->|"Measures strength/direction"| B["Pearson/Spearman"]
A -->|"No prediction"| C["No equation"]
D["Regression"] -->|"Models relationship"| E["Y = a + bX"]
D -->|"Predicts Y"| F["Uses least-squares"]flowchart TD
A["Correlation"] -->|"Measures strength/direction"| B["Pearson/Spearman"]
A -->|"No prediction"| C["No equation"]
D["Regression"] -->|"Models relationship"| E["Y = a + bX"]
D -->|"Predicts Y"| F["Uses least-squares"]
G["Correlation"] -->|"ρ or r"| H["Value between -1 and +1"]
I["Regression"] -->|"R²"| J["Coefficient of determination"]Correlation vs. Regression: Conceptual Flowchart.| Feature | Correlation | Regression |
|---|---|---|
| Goal | Describe relationship | Predict Y from X |
| Output | Coefficient (ρ or r) | Equation (Ŷ = a + bX) |
| Direction | Positive/negative | Slope (b) |
| Strength | Value between -1 and +1 | R² (coefficient of determination) |
5. In the Real World
eSewa/Khalti Payment Processing
- Idea: Negative correlation between transaction success rate (Y) and network latency (X).
- How: Faster networks (lower X) → higher success rates (higher Y).
- Example: If latency increases by 100ms, success rate drops by 3% (measured via regression).
Daraz Order Fulfillment
- Idea: Positive correlation between order volume (X) and delivery time (Y).
- How: Spearman’s ρ checks if higher-order days consistently delay shipments.
- Worked Example:
ρ = 1 → Perfect positive correlation → More orders → Longer delivery.| Day | Orders (X) | Delivery Time (Y) | Rank X | Rank Y | |-----|------------|-------------------|--------|--------| | 1 | 50 | 2 days | 1 | 1 | | 2 | 120 | 4 days | 3 | 3 | | 3 | 80 | 3 days | 2 | 2 |
NTC Network Optimization
- Idea: Regression to predict data speed (Y) from signal strength (X).
- Equation: Ŷ = 5 + 0.8X (if X = signal strength in dBm).
- Real Impact: NTC uses this to deploy towers where speed drops sharply.
6. Assumptions for Correlation/Regression
| Analysis | Assumptions |
|---|---|
| Pearson’s r | Linear relationship, normal distribution, homoscedasticity (constant variance). |
| Spearman’s ρ | Monotonic trend (not necessarily linear). |
| Regression | Independent errors, no multicollinearity. |
Violation Example:
- If data is non-normal (e.g., skewed anger scores in athletes), Pearson’s r is unreliable. Use Spearman’s ρ instead.
7. Worked Example: Pearson’s Correlation (with Assumptions Check)
Scenario: Does ambient temperature (X) affect electric power (Y) in a chemical plant?
| Temp (°F) | Power (kWh) | (X − X̄) | (Y − Ȳ) | (X − X̄)(Y − Ȳ) | (X − X̄)² | (Y − Ȳ)² |
|---|---|---|---|---|---|---|
| 27 | 120 | -13 | -20 | 260 | 169 | 400 |
| 45 | 150 | +15 | -10 | -150 | 225 | 100 |
| ... | ... | ... | ... | ... | ... | ... |
Step 1: Calculate means (X̄ = 36, Ȳ = 130). Step 2: Compute covariances and standard deviations. Step 3: Apply Pearson’s formula: Check Assumptions:
- Linearity: Plot shows a straight trend (✅).
- Normality: Power data is roughly symmetric (✅).
- Homoscedasticity: Variance is constant (✅).
Conclusion: Strong positive correlation (r = 0.98) → Temperature and power are linearly related.
8. Consistency Analysis: Checking Judge Fairness
Example: Are three judges consistent in ranking athletes’ vigor (X) vs. anger (Y)?
| Athlete | Vigor (X) | Anger (Y) | Rank X | Rank Y |
|---|---|---|---|---|
| 1 | 30 | 6 | 1 | 1 |
| 2 | 23 | 7 | 2 | 2 |
| ... | ... | ... | ... | ... |
Spearman’s ρ Calculation:
- If ρ ≈ 1 → Judges agree on vigor rankings.
- If ρ ≈ -1 → High vigor ranks correlate with low anger ranks (unlikely).
Real-World Tie: Nepal’s football team selection uses Spearman’s ρ to ensure coaches rank players consistently across speed, endurance, and aggression.
9. Exam Tip: How to Score Full Marks
- Definitions: Always define correlation (measure of association) and regression (prediction model) clearly.
- Formulas: Memorize Pearson’s r and Spearman’s ρ formulas. Show every step in calculations (e.g., covariance, standard deviation).
- Assumptions: List 3 key assumptions for Pearson’s r (linearity, normality, homoscedasticity) when asked.
- Graphs: Plot scatter plots for correlation and regression lines for prediction. Label axes properly.
- Real-World Link: Connect examples to Nepali apps (eSewa, Daraz) or global tools (Google Trends, WhatsApp metrics).
- Critical Thinking:
- If data is ranked, use Spearman’s ρ.
- If data is continuous and linear, use Pearson’s r + regression.
- Always check assumptions (e.g., "Is the relationship linear?").
Common Pitfalls:
- ❌ Mixing up correlation (descriptive) and regression (predictive).
- ❌ Forgetting to rank data before Spearman’s ρ.
- ❌ Ignoring assumptions (e.g., assuming normality without checking).
Final Visual Summary:
mindmap
root((Correlation & Regression))
Correlation
Pearson's r
Spearman's ρ
Regression
Least-Squares Line
Assumptions Check
Real-World Apps
eSewa: Latency vs. Success
Daraz: Orders vs. Delivery Time
NTC: Signal Strength vs. SpeedBased on the TU BSc CSIT syllabus for Statistics I (STA169), unit 7.
Discussion
Loading…