STT201 Business Statistics

Business StatisticsUnit 313 min read

Measures of Dispersion: Range, Quartiles, SD, Variance & Real-World Spread

Unit 3 of Business Statistics explores how data varies around central values, covering range, quartile deviation, standard deviation, and variance with formulas, calculations, and real-world applications like loan interest rates and traffic congestion analysis.

TAKEAWAYS:

  • Range shows total spread but ignores internal distribution (e.g., Daraz order delays).
  • Quartile deviation measures middle 50% spread (used in Pokhara household expenditure analysis).
  • Variance and standard deviation quantify average deviation from the mean (critical for Ncell call-drop probability).
  • Coefficient of variation compares dispersion across different scales (e.g., NEPSE stock volatility).
  • Real-world ties: Bank loan interest (SD), Pathao driver earnings (quartiles), Kathmandu traffic routes (range).
  • Exam focus: Formula derivation, interpretation, and confidence interval calculations (e.g., "95% confidence limits for population mean").

1. Why Dispersion Matters: Beyond the Average

Statistics often report a mean (e.g., average salary = Rs 40,000), but this hides how spread out the data is. Dispersion measures answer:

  • "Are most salaries close to Rs 40,000, or do some earn Rs 10,000 while others earn Rs 100,000?"
  • "Is this dataset reliable for predictions?"

box plot componentsA box plot showing median, quartiles, and outliers for a dataset. (Image: Sergio Avena, Marc Via, Elad Ziv, Eliseo J. Pérez-Stable, Ch, CC BY 2.5, via Wikimedia Commons)

Key Definitions

Term Definition Formula When to Use
Range Difference between max and min values. Quick overview of total spread.
Quartile Deviation Half the distance between 1st (Q1) and 3rd (Q3) quartiles. Robust to outliers (e.g., household income).
Variance (σ²) Average squared deviation from the mean. Theoretical analysis (e.g., stock returns).
Standard Deviation (σ) Square root of variance; units match data. Real-world comparisons (e.g., exam scores).
Coefficient of Variation (CV) SD as a % of the mean; unitless. Comparing dispersion across datasets (e.g., NEPSE vs. Daraz sales).

2. Range: The Simplest (But Crude) Measure

How it works:

  • Identify the minimum and maximum values in your dataset.
  • Subtract: .

Example 1: Daraz Order Delivery Times (in hours) Data: 24, 36, 48, 60, 72, 84, 96, 120, 144, 168

  • Range = 168 – 24 = 144 hours (4 days!).
  • Problem: One delayed order skews the range. A better measure? → Quartiles.

Real-World Tie:

  • NTC Internet Speeds: Range alone can’t tell if most users get 50 Mbps or if speeds vary wildly. Quartiles or SD are better.

3. Quartile Deviation: Ignoring Outliers

Why? Range is sensitive to extremes. Quartile deviation focuses on the middle 50% of data.

Steps to Calculate Q1, Q2 (Median), Q3

  1. Order data: .
  2. Find Q2 (Median):
    • If is odd: middle value.
    • If even: average of two middle values.
  3. Find Q1 (1st Quartile): Median of the lower half (excluding Q2 if is odd).
  4. Find Q3 (3rd Quartile): Median of the upper half.

Example 2: Pokhara Household Expenditure (Rs ’00)

Expenditure (Rs ’00) No. of Households
30–50 54
50–70 100
70–90 140
90–110 300
110–130 230
130–150 176
Total households: 1000.

Step-by-Step:

  1. Cumulative frequencies:

    • 30–50: 54
    • 50–70: 54 + 100 = 154
    • 70–90: 154 + 140 = 294
    • 90–110: 294 + 300 = 594 (Q2 lies here)
    • 110–130: 594 + 230 = 824
    • 130–150: 824 + 176 = 1000
  2. Q2 (Median):

    • th value falls in 90–110.
    • .
  3. Q1 (25th percentile):

    • th value falls in 70–90.
    • .
  4. Q3 (75th percentile):

    • th value falls in 110–130.
    • .
  5. Quartile Deviation (QD):

    • .

Interpretation:

  • The middle 50% of households spend between Rs 83.57k and Rs 125.65k.
  • Coefficient of Quartile Deviation: .

4. Variance and Standard Deviation: The Gold Standard

Why? They measure average deviation from the mean, accounting for all data points.

33.544.555.566.571.41.451.51.551.61.65σ (Population SD)σ ≈ 1.58
Standard deviation (σ) measures average deviation from the mean (μ = 5).

Formulas

Population Sample

Example 3: Ncell Call Drop Probability A sample of 5 calls has drop times (seconds): 2, 4, 4, 5, 7.

  1. Mean (): sec.
  2. Variance ():
  3. Standard Deviation (): sec.

Interpretation:

  • Calls drop on average 1.82 seconds from the mean drop time of 4.4 sec.
  • High SD → Inconsistent service (some calls drop instantly, others last longer).

5. Coefficient of Variation (CV): Comparing Apples to Oranges

Problem: SD alone can’t compare dispersion across datasets with different units (e.g., Rs vs. hours). Solution: CV standardizes SD as a percentage of the mean.

Formula:

Example 4: NEPSE vs. Daraz Sales Volatility

Company Mean Daily Sales (Rs ’00) SD (Rs ’00) CV (%)
NEPSE Stocks 5000 800
Daraz Orders 200 30

Insight:

  • NEPSE stocks are slightly more volatile (16% vs. 15%) despite higher absolute SD.
  • Use CV when comparing risk across industries (e.g., bank loans vs. stock investments).

6. Real-World Applications

A. Bank Loan Interest Rates (Standard Deviation)

Scenario: A bank offers loans with an average interest rate of 10% but varies by customer.

  • Low SD (e.g., 1%): Most customers pay 9–11% → Predictable revenue.
  • High SD (e.g., 5%): Some pay 5%, others 15% → Higher risk for the bank.

Example 5: Kathmandu Bank Loan Data

Interest Rate (%) No. of Customers
8–10 100
10–12 200
12–14 150
14–16 50

Calculations:

  1. Mean (): 11% (weighted average).
  2. Variance ():
    • Convert to midpoints: 9, 11, 13, 15.
    • .
    • .
  3. SD (): .

Bank’s Decision:

  • Low SD (1.76%) → Stable revenue; can plan budgets confidently.
  • If SD were 5%, the bank might offer risk buffers or higher default protections.

B. Pathao Driver Earnings (Quartiles)

Scenario: Pathao drivers earn Rs 30,000–50,000/month, but some earn as low as Rs 10,000.

  • Range: Rs 40,000 (misleading if outliers exist).
  • Quartiles:
    • Q1 (25th percentile): Rs 20,000
    • Q3 (75th percentile): Rs 40,000
    • Interquartile Range (IQR): Rs 20,000 → Shows middle 50% earn Rs 20k–40k.

Why Quartiles?

  • Ignores top 25% (high earners) and bottom 25% (struggling drivers).
  • Helps Pathao target support to drivers in Q1 (e.g., training programs).

C. Kathmandu Traffic Congestion (Range vs. SD)

Scenario: Traffic police measure vehicle speeds (km/h) on Ring Road.

  • Range: 20–80 km/h → 60 km/h spread (seems bad).
  • SD: 10 km/h → Most cars travel within 10 km/h of the mean (50 km/h).
  • Insight: Range overstates congestion; SD shows most traffic is slow but stable.

7. Comparing Dispersion Measures

Measure Advantages Disadvantages Best For
Range Simple, easy to calculate. Affected by outliers. Quick overview.
Quartile Deviation Robust to outliers. Ignores extreme values outside Q1–Q3. Skewed distributions (e.g., income).
Variance/SD Uses all data points. Sensitive to outliers; units squared. Normal distributions (e.g., heights).
Coefficient of Variation Unitless; compares datasets. Meaningless if mean = 0. Risk analysis (e.g., investments).

8. Common Pitfalls

  1. Using Range for skewed data: A single outlier (e.g., a Rs 1M salary in a Rs 50k dataset) can distort it.
  2. Ignoring sample vs. population formulas:
    • Population: Divide by .
    • Sample: Divide by (Bessel’s correction).
  3. Misinterpreting SD:
    • SD = 0 → All values are identical.
    • High SD → Data is widely spread (e.g., stock prices).
  4. CV for zero mean: Undefined (e.g., temperature changes around 0°C).

9. Worked Example: Past Exam Question

Question: A sample of 500 bulbs has an average life of 1400 hours with a standard deviation of 30 hours. Find the 95% confidence limits for the population mean.

Solution:

  1. Given:

    • Sample mean () = 1400 hours.
    • Sample SD () = 30 hours.
    • Sample size () = 500.
    • Confidence level = 95% → (from Z-table).
  2. Standard Error (SE):

  3. Confidence Interval:

Interpretation:

  • We are 95% confident the true population mean bulb life is between 1397 and 1403 hours.

10. Exam Tip: How to Score Full Marks

  1. Show all steps: Examiners deduct marks for skipped calculations.
    • Example: For variance, write:
  2. Label units: Always specify units (e.g., "Rs ’00", "seconds", "km/h").
  3. Interpret results:
    • Bad: "SD = 10".
    • Good: "The standard deviation of 10 hours means most calls drop within ±10 seconds of the average drop time."
  4. Use real-world context:
    • For quartiles: "This shows 50% of households spend between Rs X and Rs Y."
  5. Common exam traps:
    • Forget Bessel’s correction () for sample variance.
    • Mix up population/sample formulas.
    • Ignore cumulative frequencies in grouped data (e.g., quartile calculation).

Mermaid Diagram: Exam Checklist


In the Real World

  1. eSewa Transaction Delays:

    • Measure: Standard deviation of processing times.
    • Why? Low SD means consistent service; high SD means some users wait much longer.
    • Example: If SD = 5 seconds, most transactions take ~mean ±5 sec. If SD = 20 sec, some users face delays.
  2. Khalti Loan Approvals:

    • Measure: Coefficient of variation of loan amounts.
    • Why? Compares risk across different loan sizes (e.g., Rs 50k vs. Rs 500k loans).
    • Example: If CV = 20% for Rs 50k loans but 50% for Rs 500k loans, the bank may require stricter checks for larger loans.
  3. Daraz Order Fulfillment:

    • Measure: Quartile deviation of delivery times.
    • Why? Ignores the fastest/slowest 25% of orders to focus on typical performance.
    • Example: If Q1 = 2 days and Q3 = 5 days, Daraz knows half its orders take 2–5 days to deliver.
  4. NTC Internet Speed Ads:

    • Measure: Range vs. standard deviation.
    • Why? Ads might claim "up to 100 Mbps" (range), but SD shows if most users get 50 Mbps or speeds vary wildly.
  5. NEPSE Stock Volatility:

    • Measure: Coefficient of variation of daily returns.
    • Why? Investors compare CV across stocks to pick less risky options.
    • Example: Stock A (mean return 5%, SD 2%) has CV = 40%; Stock B (mean 10%, SD 8%) has CV = 80%. Stock A is less volatile per unit of return.
  6. Ncell Network Coverage:

    • Measure: Standard deviation of signal strength (dBm).
    • Why? Low SD means consistent coverage; high SD means some areas have weak signals.
    • Example: If SD = 5 dBm, most users experience signal within ±5 dBm of the mean.

Based on the TU BITM syllabus for Business Statistics (STT201), unit 3.

Discussion

Loading…