Business StatisticsUnit 76 min read

Correlation and Regression Analysis – Key Concepts and Applications

Unit 7 of Business Statistics: introduces correlation, covariance, Pearson’s r, coefficient of determination, simple linear regression, assumptions, diagnostics, and real‑world applications for IT management students in Nepal.

Key points

  • Correlation measures linear association; Pearson’s r ranges from –1 to +1.
  • Coefficient of determination \(R^2\) tells the proportion of variance explained by the model.
  • Simple linear regression estimates the best‑fit line \(Y=\beta_0+\beta_1X\).
  • Regression diagnostics (residual plots, normality) validate model assumptions.
  • Correlation and regression are used in app performance tuning, financial forecasting, and logistics optimization.

Introduction

Correlation and regression are the backbone of predictive analytics in business. They allow managers to quantify relationships between variables, forecast future outcomes, and make data‑driven decisions. In IT management, these tools help evaluate system performance, estimate project costs, and optimize resource allocation.

1. Correlation

-5-4-3-2-112345-2-112345678xyPositive correlationNegative correlationNo correlation
Scatter plots illustrating positive, negative and zero correlation

1.1 Definition

Correlation is a statistical measure that describes the strength and direction of a linear relationship between two quantitative variables.

1.2 Covariance

Covariance is the numerator of the correlation coefficient.

1.3 Pearson’s Correlation Coefficient


where and are sample standard deviations.

  • : positive association.
  • : negative association.
  • : perfect linear relationship.

1.4 Properties

Property Description
Symmetry
Scale‑invariant Units of measurement do not affect
Bounded

1.5 Visualizing Correlation

2. Coefficient of Determination

2.1 Definition

is the square of Pearson’s and represents the proportion of variance in explained by .

2.2 Interpretation

  • means 64 % of the variability in is accounted for by .
  • The remaining 36 % is due to other factors or random error.

2.3 Visual Representation

3. Simple Linear Regression

3.1 Model


where is the intercept, the slope, and the error term.

3.2 Estimating Parameters

3.3 Worked Example – Promotional Expenses vs Sales

Promotional Expense (Rs ’000) Sales (Rs ’00,000)
44 162
61 180
10 242
12 202
15 225
  1. Means

  2. Standard Deviations
    , (calculated from deviations)

  3. Correlation
    (given)

  4. Slope

  5. Intercept

  6. Regression Equation

  7. Interpretation
    For every additional Rs 1 000 spent on promotion, sales increase by approximately Rs 162 000.

3.4 Visualizing the Regression Line

4. Assumptions & Diagnostics

Assumption Check Diagnostic
Linearity Scatter plot Residual plot
Independence Study design Durbin–Watson test
Homoscedasticity Residuals vs fitted Plot residuals
Normality QQ‑plot Shapiro–Wilk test

4.1 Residual Plot

1234567-1-0.8-0.6-0.4-0.20.20.40.60.81xy(1, 0.4)(2, -0.2)(3, 0.1)(4, -0.3)(5, 0)(6, 0.2)
Residual plot showing random scatter around the zero line

5. Multiple Regression (Brief Overview)

When more than one predictor is used:

Coefficients are estimated by minimizing the sum of squared residuals.

5.1 Regression Process Flow

6. Applications in IT Management

Product Idea Used How It Works
eSewa Correlation between number of concurrent users and transaction latency Helps scale servers by predicting latency spikes
NEPSE Regression of trading volume on market index Forecasts future volume for risk management
Pathao Correlation of distance and delivery time Optimizes route planning and driver allocation

6.1 Real‑World Worked Example – Daraz Order Queue

Daraz wants to estimate average delivery time based on order size.

  • Data: Order size (items) vs delivery time (minutes).
  • Result: Regression equation .
  • Interpretation: Each additional item adds ~2.3 min to delivery time, aiding staffing decisions.

7. Comparison Table

Feature Covariance Correlation Coefficient of Determination
Formula
Units Same as product of variables Unitless Unitless
Scale sensitivity Yes No No
Interpretation Direction & magnitude Strength & direction Proportion of variance explained

8. Advantages & Disadvantages

Advantage Disadvantage
Easy to compute and interpret Sensitive to outliers
Provides quick insight into linearity Does not capture non‑linear relationships
Basis for regression analysis Requires interval/ratio data
Useful for forecasting Assumes linearity and homoscedasticity

9. In the real world

  • eSewa: Uses correlation between server load and transaction success rate to auto‑scale cloud instances, ensuring 99.9 % uptime.
  • NEPSE: Applies linear regression of daily trading volume on index level to set margin requirements for brokers.
  • Pathao: Employs regression of delivery time on distance and traffic density to optimize driver dispatch and reduce wait times.

10. Exam tip

  • Know the formulas: , , , .
  • Practice data sets: Compute means, deviations, covariance, and correlation quickly.
  • Interpret results: Translate numeric answers into business implications.
  • Check assumptions: Be ready to comment on residual plots or normality tests.
  • Time management: Allocate 5 min for data cleaning, 10 min for calculations, 5 min for interpretation.

Based on the TU BITM syllabus for Business Statistics (STT201), unit 7.

Discussion

Loading…