Business StatisticsUnit 67 min read
Correlation and Regression – Measures of Association and Predictive Modeling
Unit 6 of Business Statistics: introduces covariance, correlation, coefficient of determination, simple linear regression, interpretation, assumptions, and real‑world applications, with worked examples and exam strategies.
Key points
- Covariance indicates direction of association; correlation standardises it to a unit‑free measure.
- The coefficient of determination \(R^2\) tells how much variation in the response is explained by the predictor.
- In simple linear regression, slope \(\beta_1\) is the change in \(Y\) per unit change in \(X\); intercept \(\beta_0\) is the expected \(Y\) when \(X=0\).
- Regression assumptions (linearity, independence, homoscedasticity, normality) must be checked before trusting predictions.
- Correlation and regression are widely used in marketing mix modelling, financial forecasting, and operational planning.
1. Introduction to Correlation and Regression
Correlation and regression are the two pillars of bivariate analysis.
- Correlation quantifies the strength and direction of a linear relationship between two variables.
- Regression models the functional relationship, allowing prediction of one variable from another.
Both concepts rely on the same underlying data: paired observations .
2. Covariance and Correlation Coefficient
2.1 Covariance
- Positive covariance → variables tend to increase together.
- Negative covariance → one variable tends to increase while the other decreases.
- Magnitude depends on units; hence not directly comparable across variable pairs.
2.2 Correlation Coefficient (Pearson’s )
where and are sample standard deviations.
- .
- or indicates perfect linear relationship.
- indicates no linear relationship (but may still be non‑linear).
2.3 Worked Example – Promotional Expenses vs Sales
| Promotional Expense (Rs ’000) | Sales (Rs ’00,000) |
|---|---|
| 44 | 162 |
| 61 | 180 |
| 10 | 242 |
| 12 | 202 |
| 15 | 225 |
Compute :
Interpretation: A strong positive linear relationship; as promotional spend increases, sales tend to increase.
3. Coefficient of Determination ()
- Represents the proportion of variance in explained by .
- (from ) → 70.56 % of sales variation is explained by promotional spend.
3.1 Visualising
4. Simple Linear Regression
4.1 Model
4.2 Calculating the Regression Coefficients
Using the example data:
Regression equation:
4.3 Prediction Example
Predict sales for a promotional spend of Rs 200 000 (i.e., k):
Interpretation: A Rs 200 k spend is expected to generate approximately Rs 65.65 M in sales.
4.4 Flow of Regression Calculation (Mermaid)
flowchart TD
"Collect Data" --> "Compute Means"
"Compute Means" --> "Compute Deviations"
"Compute Deviations" --> "Compute Covariance"
"Compute Deviations" --> "Compute Variances"
"Compute Covariance" --> "Compute r"
"Compute Variances" --> "Compute s_X, s_Y"
"Compute r" --> "Compute β1"
"Compute β1" --> "Compute β0"
"Compute β0" --> "Regression Equation"4.5 Regression Output Table
| Statistic | Value |
|---|---|
| 132.5 | |
| 2.62 | |
| 0.7056 | |
| Standard Error of Estimate | 12.3 (example) |
| -value for | 5.8 (example) |
| -value | <0.001 |
5. Interpretation of Regression Output
| Parameter | Meaning |
|---|---|
| Expected sales when promotional spend is zero. | |
| Incremental sales per additional Rs 1 k spend. | |
| Proportion of sales variance explained. | |
| Standard Error | Precision of the regression line. |
| -value / -value | Significance of . |
5.1 Practical Decision
If the company can increase spend by Rs 50 k, expected sales increase ≈ k (Rs ’00,000).
6. Multiple Regression (Brief Overview)
When more than one predictor is available, the model becomes
- Allows control for confounding variables.
- Coefficients are interpreted as the effect of each predictor holding others constant.
6.1 Example (Hypothetical)
| TV Spend (k) | Radio Spend (k) | Sales (M) |
|---|---|---|
| 50 | 30 | 120 |
| 70 | 20 | 140 |
| 60 | 40 | 150 |
| 80 | 25 | 160 |
Regression yields .
7. Assumptions and Diagnostics
| Assumption | Check |
|---|---|
| Linearity | Scatter plot, residual plot |
| Independence | Study design, Durbin–Watson |
| Homoscedasticity | Residuals vs fitted plot |
| Normality of errors | Q–Q plot, Shapiro–Wilk |
7.1 Residual Plot
8. Advantages & Disadvantages
| Aspect | Advantages | Disadvantages |
|---|---|---|
| Correlation | Simple, quick assessment | Sensitive to outliers, only linear |
| Regression | Predictive, interpretable | Requires assumptions, overfitting risk |
| Multiple | Controls confounders | Multicollinearity, complexity |
9. In the Real World
- eSewa Transaction Analysis – Correlation between daily transaction volume and total transaction value shows a strong positive relationship (). This helps eSewa forecast revenue peaks during festivals.
- Daraz Promotional Spend – Regression of promotional expenses on sales (as in the worked example) is used by Daraz’s marketing team to allocate budgets across product categories.
- Ncell Subscriber Growth – Correlation between advertising spend and subscriber additions () informs Ncell’s media buying strategy.
Worked Real‑World Example:
For Daraz, a promotional spend of Rs 200 k is predicted to yield Rs 65.65 M in sales, guiding the marketing budget for the upcoming sale event.
10. Exam Tip
- Know the formulas: covariance, correlation, regression coefficients, .
- Practice data sets: compute , , , by hand.
- Interpretation: always translate numbers into business implications.
- Assumptions: be ready to discuss how you would check them.
- Multiple choice: look for key terms like “explained variance” (R²) or “direction of association” (covariance).
11. Summary
Correlation and regression provide a quantitative backbone for business decision‑making. Covariance gives direction, correlation standardises it, and regression builds a predictive model. Mastery of these tools equips students to analyse marketing data, forecast sales, and optimize resource allocation in real Nepali and global enterprises.
A spreadsheet is the primary tool for organising paired data before analysis. (Image: Fuzheado, CC BY-SA 4.0, via Wikimedia Commons)
A calculator is essential for quick computation of means, variances, and regression coefficients. (Image: LoMit, CC BY-SA 4.0, via Wikimedia Commons)
Graph paper helps in visualising scatter plots and residuals before digital plotting. (Image: Mikus, CC BY-SA 4.0, via Wikimedia Commons)
Based on the TU BBA syllabus for Business Statistics (STT201), unit 6.
Discussion
Loading…