Business StatisticsUnit 212 min read
Measures of Central Tendency & Dispersion: Mean, Median, Mode, Range, Variance, Standard Deviation
Unit 2 of Business Statistics covers how to summarize data using central tendency (mean, median, mode) and dispersion (range, variance, standard deviation), when to use each measure, and how to calculate them for grouped and ungrouped data—with real-world applications in Nepal’s business sector.
TAKEAWAYS:
- Central tendency (mean, median, mode) identifies the "typical" value in a dataset, but each has unique strengths (e.g., mean uses all data, median resists outliers).
- Dispersion (range, variance, standard deviation) measures how spread out values are—critical for risk assessment (e.g., loan defaults) and quality control (e.g., product consistency).
- Grouped data requires special formulas (e.g., assumed mean method) to calculate measures, while ungrouped data uses direct summation.
- Skewed distributions dictate which central tendency to prioritize (e.g., median for income data skewed by high earners).
- Real-world ties: Banks use standard deviation to assess loan risk; eSewa analyzes transaction dispersion to detect fraud; Daraz optimizes delivery routes based on delivery-time variance.
- Exam focus: Always justify your choice of measure (e.g., "mode for categorical data") and show calculations step-by-step with clear labels.
1. Central Tendency: The "Typical" Value
Central tendency measures summarize a dataset with a single value. Three key measures:
- Mean (Arithmetic Average): Sum of all values divided by the number of values.
- Median: Middle value when data is ordered (or average of two middle values for even n).
- Mode: Most frequently occurring value(s).
When to Use Which?
graph TD
A["Choose Measure"] --> B["Data Type"]
B --> C["Numerical Data"]
B --> D["Categorical Data"]
C --> E["Symmetrical Distribution?"]
E --> F["Yes\nUse **Mean** (most efficient)"]
E --> G["No\nSkewed?"]
G --> H["Left-Skewed\nUse **Median**"]
G --> I["Right-Skewed\nUse **Median**"]
D --> J["Use **Mode**"]Worked Example 1: Mean, Median, Mode for Ungrouped Data
Question: Calculate the mean, median, and mode for the following monthly incomes (in Rs.) of 10 families:
5000, 6000, 4000, 8000, 5000, 7000, 4000, 9000, 3000, 6000.
Solution:
- Mean:
- Median:
- Ordered data:
3000, 4000, 4000, 5000, 5000, 6000, 6000, 7000, 8000, 9000. - Middle values (5th and 6th):
5000, 6000.
- Ordered data:
- Mode:
- Most frequent values:
4000and5000(each appears twice). Mode = 4000 and 5000 (bimodal).
- Most frequent values:
Why the differences?
- Mean (5800) is pulled higher by the
9000outlier. - Median (5500) is less affected by extremes.
- Mode (4000, 5000) shows the most common incomes.
2. Central Tendency for Grouped Data
For frequency distributions (e.g., income classes), use:
- Mean: , where = frequency, = midpoint of class.
- Median: Locate the class containing the th value, then interpolate.
- Mode: Use the formula:
where:
- = lower limit of modal class,
- = frequency of modal class,
- = frequencies of classes before/after,
- = class width.
Worked Example 2: Mean and Median for Grouped Data
Question: Calculate the mean and median for the following income distribution of 50 families:
| Income (Rs.) | No. of Families |
|---|---|
| Below 1000 | 5 |
| 1000–1999 | 5 |
| 2000–2999 | 5 |
| 3000–3999 | 10 |
| 4000–4999 | 25 |
Solution:
Midpoints and Frequencies:
Class Midpoint () Frequency () Below 1000 500 5 2500 1000–1999 1500 5 7500 2000–2999 2500 5 12500 3000–3999 3500 10 35000 4000–4999 4500 25 112500 Total = 50, = 169000. Mean:
Median:
- th value falls in the 4000–4999 class.
- Cumulative frequency before this class: .
- Since the 25th value is the first in this class, Median = 4000 Rs. (lower limit).
Real-World Tie: Nepal Rastra Bank (NRB) uses median income data to set poverty lines, as the median is less skewed by ultra-rich or ultra-poor outliers.
3. Measures of Dispersion: How Spread Out Is the Data?
Dispersion measures quantify variability. Key tools:
- Range: (simplest but sensitive to outliers).
- Variance (): Average of squared deviations from the mean.
- Standard Deviation (): Square root of variance (same units as data).
Formulas
| Measure | Ungrouped Data | Grouped Data |
|---|---|---|
| Range | ||
| Variance | ||
| Std Dev |
Worked Example 3: Variance and Standard Deviation
Question: Calculate the variance and standard deviation for the incomes in Worked Example 1 (5000, 6000, ..., 6000).
Solution:
Mean (): Already calculated as 5800 Rs.
Deviations and Squared Deviations:
Income () 5000 -800 640000 6000 +200 40000 ... ... ... 6000 +200 40000 Total = 32,800,000. Variance:
Standard Deviation:
Interpretation:
- A standard deviation of 1811 Rs. means most incomes fall within 5800 ± 1811 Rs. (i.e., 3989–7611 Rs.).
- Real-World Tie: Pathao uses standard deviation to analyze driver earnings variability. If is high, it may indicate inconsistent demand or driver availability issues.
4. Comparing Central Tendency and Dispersion
| Measure | Definition | When to Use | Limitation |
|---|---|---|---|
| Mean | Average of all values | Symmetrical data, numerical data | Affected by outliers |
| Median | Middle value | Skewed data, ordinal data | Ignores all but middle values |
| Mode | Most frequent value | Categorical data, identifying trends | May not exist or be unique |
| Range | Difference between max and min | Quick estimate of spread | Ignores distribution shape |
| Variance/Std Dev | Average squared deviation from mean | Comparing datasets, risk analysis | Units are squared (std dev fixes this) |
5. Coefficient of Variation (CV)
Compares dispersion relative to the mean: Use Case: Compare variability across datasets with different units (e.g., daily income vs. monthly income).
Worked Example 4: CV for Two Datasets
Question: Compare the dispersion of:
- Dataset A: Mean = 5000 Rs., Rs.
- Dataset B: Mean = 10000 Rs., Rs.
Solution:
- CV for A:
- CV for B: Conclusion: Both datasets have equal relative dispersion (20%), even though Dataset B’s absolute is higher.
In the Real World
eSewa and Khalti (Digital Payments)
- Idea Used: Standard Deviation of Transaction Amounts
- How: These apps flag unusual transactions by calculating the standard deviation of typical user spending. If a user’s transaction exceeds , it triggers fraud alerts.
- Example: If your average monthly Khalti transaction is 10,000 Rs. with Rs., a sudden 20,000 Rs. payment (within 1) might be normal, but a 50,000 Rs. payment (2.5 above) could be flagged.
Daraz and Ncell (Delivery and Network Reliability)
- Idea Used: Range and Median Delivery Times
- How: Daraz uses the median delivery time (not mean, to avoid skewing by delays) to set customer expectations. Ncell reports the interquartile range (IQR = Q3 – Q1) of network speeds to show typical performance (ignoring extreme outliers).
- Example: If Daraz’s median delivery time is 3 days but the range is 1–7 days, they may offer "guaranteed 3-day delivery" while acknowledging some orders take longer.
Nepal Rastra Bank (NRB) and Loan Risk Assessment
- Idea Used: Coefficient of Variation (CV) of Borrower Incomes
- How: Banks calculate the CV of a borrower’s income over 6 months. A high CV (e.g., >30%) signals unstable income, increasing loan default risk.
- Example: A borrower with mean monthly income = 50,000 Rs. and Rs. has: NRB may reject or require collateral for such loans.
Exam Tip
Always justify your choice of measure:
- "I used the median because the data is right-skewed (e.g., income distributions)."
- "Mode is appropriate for categorical data (e.g., most popular product brands)."
Show all steps clearly:
- Label columns (e.g., , , ) in grouped data tables.
- Use the assumed mean method for grouped data to simplify calculations (see below).
Assumed Mean Method (Shortcut for Grouped Mean): For Worked Example 2, assume a mean of 3000 Rs.:
Class Midpoint () 500 -2500 5 -12500 1500 -1500 5 -7500 2500 -500 5 -2500 3500 +500 10 +5000 4500 +1500 25 +37500 , , Rs. (close to earlier 3380; rounding error). Watch for traps:
- Open-ended classes: For "250 and above," assume a midpoint (e.g., 300) or use if no data is given.
- Units: Always label answers (e.g., "Rs.") and ensure has the same units as the data.
Past Exam Question Solved
Question: Calculate the appropriate measure of central tendency for the following income distribution and justify your choice.
| Monthly Income (Rs.) | No. of Families |
|---|---|
| Below 1000 | 5 |
| 1000–1999 | 5 |
| 2000–2999 | 5 |
| 3000–3999 | 10 |
| 4000–4999 | 25 |
Solution:
Choice of Measure:
- Data is grouped and numerical, but the distribution is right-skewed (most families earn <4000 Rs., but the highest class has 25 families).
- Median is the best choice because it is resistant to skewness and represents the "typical" family’s income better than the mean (which would be pulled upward by the 4000–4999 class).
Calculation:
- Total families () = 50.
- Median position: th family.
- Cumulative frequencies:
- Below 1000: 5
- 1000–1999: 5 + 5 = 10
- 2000–2999: 10 + 5 = 15
- 3000–3999: 15 + 10 = 25 ← 25th family falls here.
- Median = Lower limit of median class = 3000 Rs.
Answer: The median income is 3000 Rs., chosen because the data is skewed and the median best represents the central tendency.
Key Formulas Summary
mindmap
root((Measures))
Central Tendency
Mean: \(\bar{X} = \frac{\sum X}{N}\)
Median: Middle value (or average of two middle values)
Mode: Most frequent value
Dispersion
Range: Max - Min
Variance: \(\sigma^2 = \frac{\sum (X - \bar{X})^2}{N}\)
Std Dev: \(\sigma = \sqrt{\sigma^2}\)
CV: \(\left( \frac{\sigma}{\bar{X}} \right) \times 100\%\)Based on the TU BBA syllabus for Business Statistics (STT201), unit 2.
Discussion
Loading…