Business StatisticsUnit 313 min read
Measures of Dispersion: Range, Quartiles, SD, Variance & Real-World Spread
Unit 3 of Business Statistics explores how data varies around central values, covering range, quartile deviation, standard deviation, and variance with formulas, calculations, and real-world applications like loan interest rates and traffic congestion analysis.
TAKEAWAYS:
- Range shows total spread but ignores internal distribution (e.g., Daraz order delays).
- Quartile deviation measures middle 50% spread (used in Pokhara household expenditure analysis).
- Variance and standard deviation quantify average deviation from the mean (critical for Ncell call-drop probability).
- Coefficient of variation compares dispersion across different scales (e.g., NEPSE stock volatility).
- Real-world ties: Bank loan interest (SD), Pathao driver earnings (quartiles), Kathmandu traffic routes (range).
- Exam focus: Formula derivation, interpretation, and confidence interval calculations (e.g., "95% confidence limits for population mean").
1. Why Dispersion Matters: Beyond the Average
Statistics often report a mean (e.g., average salary = Rs 40,000), but this hides how spread out the data is. Dispersion measures answer:
- "Are most salaries close to Rs 40,000, or do some earn Rs 10,000 while others earn Rs 100,000?"
- "Is this dataset reliable for predictions?"
A box plot showing median, quartiles, and outliers for a dataset. (Image: Sergio Avena, Marc Via, Elad Ziv, Eliseo J. Pérez-Stable, Ch, CC BY 2.5, via Wikimedia Commons)
Key Definitions
| Term | Definition | Formula | When to Use |
|---|---|---|---|
| Range | Difference between max and min values. | Quick overview of total spread. | |
| Quartile Deviation | Half the distance between 1st (Q1) and 3rd (Q3) quartiles. | Robust to outliers (e.g., household income). | |
| Variance (σ²) | Average squared deviation from the mean. | Theoretical analysis (e.g., stock returns). | |
| Standard Deviation (σ) | Square root of variance; units match data. | Real-world comparisons (e.g., exam scores). | |
| Coefficient of Variation (CV) | SD as a % of the mean; unitless. | Comparing dispersion across datasets (e.g., NEPSE vs. Daraz sales). |
2. Range: The Simplest (But Crude) Measure
How it works:
- Identify the minimum and maximum values in your dataset.
- Subtract: .
Example 1: Daraz Order Delivery Times (in hours) Data: 24, 36, 48, 60, 72, 84, 96, 120, 144, 168
- Range = 168 – 24 = 144 hours (4 days!).
- Problem: One delayed order skews the range. A better measure? → Quartiles.
Real-World Tie:
- NTC Internet Speeds: Range alone can’t tell if most users get 50 Mbps or if speeds vary wildly. Quartiles or SD are better.
3. Quartile Deviation: Ignoring Outliers
Why? Range is sensitive to extremes. Quartile deviation focuses on the middle 50% of data.
Steps to Calculate Q1, Q2 (Median), Q3
- Order data: .
- Find Q2 (Median):
- If is odd: middle value.
- If even: average of two middle values.
- Find Q1 (1st Quartile): Median of the lower half (excluding Q2 if is odd).
- Find Q3 (3rd Quartile): Median of the upper half.
Example 2: Pokhara Household Expenditure (Rs ’00)
| Expenditure (Rs ’00) | No. of Households |
|---|---|
| 30–50 | 54 |
| 50–70 | 100 |
| 70–90 | 140 |
| 90–110 | 300 |
| 110–130 | 230 |
| 130–150 | 176 |
| Total households: 1000. |
Step-by-Step:
Cumulative frequencies:
- 30–50: 54
- 50–70: 54 + 100 = 154
- 70–90: 154 + 140 = 294
- 90–110: 294 + 300 = 594 (Q2 lies here)
- 110–130: 594 + 230 = 824
- 130–150: 824 + 176 = 1000
Q2 (Median):
- th value falls in 90–110.
- .
Q1 (25th percentile):
- th value falls in 70–90.
- .
Q3 (75th percentile):
- th value falls in 110–130.
- .
Quartile Deviation (QD):
- .
Interpretation:
- The middle 50% of households spend between Rs 83.57k and Rs 125.65k.
- Coefficient of Quartile Deviation: .
4. Variance and Standard Deviation: The Gold Standard
Why? They measure average deviation from the mean, accounting for all data points.
Formulas
| Population | Sample |
|---|---|
Example 3: Ncell Call Drop Probability A sample of 5 calls has drop times (seconds): 2, 4, 4, 5, 7.
- Mean (): sec.
- Variance ():
- Standard Deviation (): sec.
Interpretation:
- Calls drop on average 1.82 seconds from the mean drop time of 4.4 sec.
- High SD → Inconsistent service (some calls drop instantly, others last longer).
5. Coefficient of Variation (CV): Comparing Apples to Oranges
Problem: SD alone can’t compare dispersion across datasets with different units (e.g., Rs vs. hours). Solution: CV standardizes SD as a percentage of the mean.
Formula:
Example 4: NEPSE vs. Daraz Sales Volatility
| Company | Mean Daily Sales (Rs ’00) | SD (Rs ’00) | CV (%) |
|---|---|---|---|
| NEPSE Stocks | 5000 | 800 | |
| Daraz Orders | 200 | 30 |
Insight:
- NEPSE stocks are slightly more volatile (16% vs. 15%) despite higher absolute SD.
- Use CV when comparing risk across industries (e.g., bank loans vs. stock investments).
6. Real-World Applications
A. Bank Loan Interest Rates (Standard Deviation)
Scenario: A bank offers loans with an average interest rate of 10% but varies by customer.
- Low SD (e.g., 1%): Most customers pay 9–11% → Predictable revenue.
- High SD (e.g., 5%): Some pay 5%, others 15% → Higher risk for the bank.
Example 5: Kathmandu Bank Loan Data
| Interest Rate (%) | No. of Customers |
|---|---|
| 8–10 | 100 |
| 10–12 | 200 |
| 12–14 | 150 |
| 14–16 | 50 |
Calculations:
- Mean (): 11% (weighted average).
- Variance ():
- Convert to midpoints: 9, 11, 13, 15.
- .
- .
- SD (): .
Bank’s Decision:
- Low SD (1.76%) → Stable revenue; can plan budgets confidently.
- If SD were 5%, the bank might offer risk buffers or higher default protections.
B. Pathao Driver Earnings (Quartiles)
Scenario: Pathao drivers earn Rs 30,000–50,000/month, but some earn as low as Rs 10,000.
- Range: Rs 40,000 (misleading if outliers exist).
- Quartiles:
- Q1 (25th percentile): Rs 20,000
- Q3 (75th percentile): Rs 40,000
- Interquartile Range (IQR): Rs 20,000 → Shows middle 50% earn Rs 20k–40k.
Why Quartiles?
- Ignores top 25% (high earners) and bottom 25% (struggling drivers).
- Helps Pathao target support to drivers in Q1 (e.g., training programs).
C. Kathmandu Traffic Congestion (Range vs. SD)
Scenario: Traffic police measure vehicle speeds (km/h) on Ring Road.
- Range: 20–80 km/h → 60 km/h spread (seems bad).
- SD: 10 km/h → Most cars travel within 10 km/h of the mean (50 km/h).
- Insight: Range overstates congestion; SD shows most traffic is slow but stable.
7. Comparing Dispersion Measures
| Measure | Advantages | Disadvantages | Best For |
|---|---|---|---|
| Range | Simple, easy to calculate. | Affected by outliers. | Quick overview. |
| Quartile Deviation | Robust to outliers. | Ignores extreme values outside Q1–Q3. | Skewed distributions (e.g., income). |
| Variance/SD | Uses all data points. | Sensitive to outliers; units squared. | Normal distributions (e.g., heights). |
| Coefficient of Variation | Unitless; compares datasets. | Meaningless if mean = 0. | Risk analysis (e.g., investments). |
8. Common Pitfalls
- Using Range for skewed data: A single outlier (e.g., a Rs 1M salary in a Rs 50k dataset) can distort it.
- Ignoring sample vs. population formulas:
- Population: Divide by .
- Sample: Divide by (Bessel’s correction).
- Misinterpreting SD:
- SD = 0 → All values are identical.
- High SD → Data is widely spread (e.g., stock prices).
- CV for zero mean: Undefined (e.g., temperature changes around 0°C).
9. Worked Example: Past Exam Question
Question: A sample of 500 bulbs has an average life of 1400 hours with a standard deviation of 30 hours. Find the 95% confidence limits for the population mean.
Solution:
Given:
- Sample mean () = 1400 hours.
- Sample SD () = 30 hours.
- Sample size () = 500.
- Confidence level = 95% → (from Z-table).
Standard Error (SE):
Confidence Interval:
Interpretation:
- We are 95% confident the true population mean bulb life is between 1397 and 1403 hours.
10. Exam Tip: How to Score Full Marks
- Show all steps: Examiners deduct marks for skipped calculations.
- Example: For variance, write:
- Label units: Always specify units (e.g., "Rs ’00", "seconds", "km/h").
- Interpret results:
- Bad: "SD = 10".
- Good: "The standard deviation of 10 hours means most calls drop within ±10 seconds of the average drop time."
- Use real-world context:
- For quartiles: "This shows 50% of households spend between Rs X and Rs Y."
- Common exam traps:
- Forget Bessel’s correction () for sample variance.
- Mix up population/sample formulas.
- Ignore cumulative frequencies in grouped data (e.g., quartile calculation).
Mermaid Diagram: Exam Checklist
In the Real World
eSewa Transaction Delays:
- Measure: Standard deviation of processing times.
- Why? Low SD means consistent service; high SD means some users wait much longer.
- Example: If SD = 5 seconds, most transactions take ~mean ±5 sec. If SD = 20 sec, some users face delays.
Khalti Loan Approvals:
- Measure: Coefficient of variation of loan amounts.
- Why? Compares risk across different loan sizes (e.g., Rs 50k vs. Rs 500k loans).
- Example: If CV = 20% for Rs 50k loans but 50% for Rs 500k loans, the bank may require stricter checks for larger loans.
Daraz Order Fulfillment:
- Measure: Quartile deviation of delivery times.
- Why? Ignores the fastest/slowest 25% of orders to focus on typical performance.
- Example: If Q1 = 2 days and Q3 = 5 days, Daraz knows half its orders take 2–5 days to deliver.
NTC Internet Speed Ads:
- Measure: Range vs. standard deviation.
- Why? Ads might claim "up to 100 Mbps" (range), but SD shows if most users get 50 Mbps or speeds vary wildly.
NEPSE Stock Volatility:
- Measure: Coefficient of variation of daily returns.
- Why? Investors compare CV across stocks to pick less risky options.
- Example: Stock A (mean return 5%, SD 2%) has CV = 40%; Stock B (mean 10%, SD 8%) has CV = 80%. Stock A is less volatile per unit of return.
Ncell Network Coverage:
- Measure: Standard deviation of signal strength (dBm).
- Why? Low SD means consistent coverage; high SD means some areas have weak signals.
- Example: If SD = 5 dBm, most users experience signal within ±5 dBm of the mean.
Based on the TU BITM syllabus for Business Statistics (STT201), unit 3.
Discussion
Loading…