Basic StatisticsUnit 78 min read
Correlation & Regression: Pearson’s r, Rank Correlation, Least Squares Line
Unit 7 of Basic Statistics covers how to measure the strength and direction of relationships between two variables (correlation) and how to predict one variable from another (regression). Learn Pearson’s r, Spearman’s rho, regression equations, and real-world applications like predicting blood pressure from age or rank
TAKEAWAYS:
- Correlation ≠ Causation: Pearson’s r measures linear relationships between two continuous variables (e.g., age vs. blood pressure), while Spearman’s rho ranks ordinal data (e.g., student preferences for DELL vs. HP).
- Regression predicts: The least-squares line estimates values (e.g., predicting pumpkin weight from volume or factory productivity from aptitude scores).
- Interpret r carefully: Values range from –1 (perfect negative) to +1 (perfect positive); r = 0 means no linear trend.
- Real-world ties: Kathmandu traffic routes use regression to predict congestion; eSewa’s fraud detection relies on correlation between transaction patterns.
- Exam focus: Always show calculations step-by-step (e.g., covariance, standard deviations) and interpret results in context (e.g., "A r = 0.85 suggests strong positive correlation between experience and performance").
1. Correlation: Measuring Relationships
Correlation quantifies how two variables move together. The two key methods are:
- Pearson’s r (for continuous data, e.g., age vs. blood pressure).
- Spearman’s rho (for ranked data, e.g., student preferences).
Pearson’s Correlation Coefficient (r)
Definition: where:
- = covariance (how X and Y vary together).
- = standard deviations of X and Y.
Key Properties:
graph TD
A["Pearson’s *r* Properties"] --> B["Range: -1 to +1"]
A --> C["*r* = +1: Perfect positive linear relationship"]
A --> D["*r* = -1: Perfect negative linear relationship"]
A --> E["*r* = 0: No linear relationship (could be nonlinear!)"]
A --> F["Sensitive to outliers"]
A --> G["Unitless (standardized)"]Worked Example 1: Husband-Wife Ages Question: Calculate r for the ages of 10 husbands and wives. Interpret the result. Data: | Husband (X) | 23 | 27 | 28 | 28 | 29 | 30 | 31 | 33 | 35 | 36 | | Wife (Y) | 22 | 25 | 26 | 27 | 28 | 29 | 30 | 32 | 34 | 35 |
Step-by-Step Calculation:
- Compute means:
- Compute covariance and standard deviations: Sums:
- Plug into r formula:
Interpretation:
- r = 0.9 indicates a strong positive linear correlation between husband and wife ages. As one spouse’s age increases, the other’s tends to increase proportionally.
- Real-world tie: This mirrors how eSewa’s age verification system uses similar correlations to detect fraud (e.g., mismatched ages in transaction pairs).
Spearman’s Rank Correlation (rho)
Used for ordinal data (ranks, preferences). Formula: where = difference in ranks for each pair, = number of items.
Worked Example 2: Student Preferences (DELL vs. HP) Question: Calculate Spearman’s rho for 10 students’ rankings of DELL and HP. Data:
| Student | DELL Rank | HP Rank | ||
|---|---|---|---|---|
| 1 | 5 | 10 | -5 | 25 |
| 2 | 2 | 5 | -3 | 9 |
| ... | ... | ... | ... | ... |
| 10 | 4 | 7 | -3 | 9 |
Calculation:
Interpretation:
- rho = 0.33 suggests a weak positive correlation between student preferences for DELL and HP. Most students don’t strongly prefer one over the other.
- Real-world tie: Pathao’s driver ratings use Spearman’s rho to correlate customer feedback scores with trip completion times.
2. Regression Analysis: Prediction and Modeling
Regression finds the best-fit line to predict from . Key methods:
- Least-squares regression: Minimizes the sum of squared errors.
- Interpretation: = slope (change in per unit ), = intercept.
Regression Line Equation
Worked Example 3: Blood Pressure vs. Age Question: Fit a regression line to predict blood pressure () from age () for 10 individuals. Data:
| Age (X) | 56 | 42 | 27 | 23 | 66 | 34 | 47 |
|---|---|---|---|---|---|---|---|
| BP (Y) | 147 | 125 | 160 | 118 | 149 | 128 | 140 |
Steps:
- Compute means, covariance, and (from earlier: , , , ).
- Calculate slope () and intercept ():
- Regression equation:
Interpretation:
- For every 1-year increase in age, blood pressure rises by 0.75 mmHg.
- Real-world tie: NTC uses similar models to predict network congestion based on call volumes.
3. Comparing Correlation and Regression
| Aspect | Correlation | Regression |
|---|---|---|
| Purpose | Measures strength/direction of relationship | Predicts one variable from another |
| Output | Single value (r or rho) | Equation () |
| Data Types | Continuous (Pearson) or ranked (Spearman) | Continuous (least-squares) |
| Directionality | Symmetric (X vs. Y same r) | Asymmetric (predicts Y from X) |
| Example | r = 0.8 between study hours and exam scores | Predicting Daraz delivery time from distance |
4. Common Pitfalls and Misconceptions
mindmap
root((Correlation/Regression Mistakes))
A1[Correlation ≠ Causation]
A2[Extrapolation Danger]
A3[Outliers Skew Results]
A4[Nonlinear ≠ No Correlation]
A5[Rank vs. Continuous Data]Example:
- Causation vs. Correlation: Ice cream sales and drowning incidents are correlated (r ≈ 0.8), but neither causes the other (both rise in summer).
- Nonlinear Data: If , Pearson’s r may be low even though the relationship is strong.
## In the Real World
eSewa’s Fraud Detection:
- Idea Used: Pearson’s r to correlate transaction amounts with user location/device.
- How: Unusually high r between a user’s typical spending and a sudden large transaction flags fraud (e.g., r > 0.95 triggers review).
Pathao’s Driver Ratings:
- Idea Used: Spearman’s rho to rank driver performance (e.g., trip completion time vs. customer feedback).
- How: Drivers with rho < 0.5 between "on-time" and "clean car" metrics get retrained.
NTC’s Network Prediction:
- Idea Used: Regression to predict call drops from call volume.
- Equation: (where = % drops, = calls/hour).
- Impact: Helps NTC allocate towers during festivals (e.g., Dashain).
Daraz’s Order Fulfillment:
- Idea Used: Regression to estimate delivery time from distance.
- Example: For a 20 km order, hours (intercept = warehouse processing time).
## Exam Tip
Always show calculations:
- For r: Write out covariance and standard deviations explicitly.
- For regression: Derive and step-by-step (examiners deduct for skipping steps).
Interpret results in context:
- Bad: "r = 0.85".
- Good: "r = 0.85 suggests a strong positive correlation between experience (X) and performance (Y), meaning operators with more years of experience tend to produce more defect-free parts."
Watch units:
- If is in years and in mmHg, state the slope as "0.75 mmHg per year."
Rank correlation shortcut:
- For small datasets (), compute ranks manually. For larger sets, use the formula:
Graphs are mandatory:
- Always sketch a scatter plot with the regression line. Label axes and highlight outliers.
Visual Summary:
Based on the TU BIT syllabus for Basic Statistics (STA154), unit 7.
Discussion
Loading…