Statistics IUnit 314 min read
Measures of Central Tendency & Dispersion: Concepts, Calculations & Applications
Unit 3 of Statistics I covers the core tools for summarizing data: mean, median, mode, range, variance, and standard deviation. Learn when to use each measure, how to calculate them for raw and grouped data, and how they reveal the true story behind numbers—with real-world examples from Nepal’s tech and finance sectors
TAKEAWAYS:
- Central tendency (mean, median, mode) tells you the "typical" value, but dispersion (range, variance, standard deviation) shows how spread out the data is—always report both to avoid misleading conclusions.
- Mean is sensitive to outliers; median is robust; mode is for categorical data or most frequent values—choose based on data type and skewness.
- Grouped data requires midpoints and frequency weights; ungrouped data uses raw values—always check if data is grouped before calculating.
- Variance measures squared deviation; standard deviation is its square root—use standard deviation for interpretation (it’s in the same units as the data).
- Coefficient of variation (CV) compares dispersion across datasets with different units—critical for benchmarking (e.g., comparing stock volatility).
- Exam focus: 30% of questions test calculation (raw/grouped data), 40% test conceptual choice (why mean vs. median?), and 30% test real-world application (e.g., interpreting bus wait times or loan interest rates).
1. Measures of Central Tendency: The "Typical" Value
Central tendency summarizes a dataset with a single value. The choice depends on data type (nominal, ordinal, interval, ratio) and skewness (symmetry).
1.1 Definitions and When to Use Each
classDiagram
class Measure {
+Name: String
+DataType: String
+Skewness: String
+Formula: String
+UseCase: String
}
class Mean {
+Name: "Mean (Average)"
+DataType: "Interval/Ratio"
+Skewness: "Symmetric or mild skew"
+Formula: "ΣX / N"
+UseCase: "Most common for quantitative data"
}
class Median {
+Name: "Median"
+DataType: "Ordinal/Interval/Ratio"
+Skewness: "Skewed data (robust to outliers)"
+Formula: "(n+1)/2th value (ungrouped)"
+UseCase: "Income, house prices (outliers distort mean)"
}
class Mode {
+Name: "Mode"
+DataType: "Nominal/Ordinal"
+Skewness: "Any"
+Formula: "Most frequent value"
+UseCase: "Categorical data (e.g., bus routes used)"
}
Measure <|-- Mean
Measure <|-- Median
Measure <|-- ModeKey Idea:
- Mean = Total sum / Number of values (sensitive to outliers).
- Median = Middle value (50th percentile; robust to outliers).
- Mode = Most frequent value (can be multiple; used for categories).
Visual: Data Skewness and Measure Choice
Worked Example 1: Choosing the Right Measure
Problem: The following are the waiting times (in minutes) for a bus over 20 days:
15, 10, 2, 17, 5, 8, 3, 10, 12, 18, 4, 6, 9, 11, 7, 14, 1, 20, 13, 5
Step 1: Check for outliers.
Observation: The data is right-skewed (one extreme value: 20 minutes). The mean will be higher than the median.
Step 2: Calculate all three measures.
- Mean = minutes.
- Median = Average of 10th and 11th values = minutes.
- Mode = 5 and 10 (bimodal).
Conclusion: The median (10.5) best represents the "typical" wait time because the mean is inflated by the 20-minute outlier.
1.2 Calculations for Grouped Data
When data is presented in classes (e.g., age groups), use midpoints and frequency weights.
Formula:
- Mean =
Where:
- = frequency of the class
- = midpoint of the class =
Worked Example 2: Grouped Data Mean Problem: Calculate the mean age from the following table:
| Age in Years | Below 20 | 20-30 | 30-40 | 40-50 | 50 and more |
|---|---|---|---|---|---|
| No. of People | 15 | 25 | 30 | 20 | 10 |
Step 1: Assign midpoints.
| Class | Midpoint (m) | Frequency (f) | f × m |
|---|---|---|---|
| Below 20 | 10 | 15 | 150 |
| 20-30 | 25 | 25 | 625 |
| 30-40 | 35 | 30 | 1050 |
| 40-50 | 45 | 20 | 900 |
| 50+ | 55 | 10 | 550 |
| Total | 100 | 3275 |
Step 2: Calculate mean.
2. Measures of Dispersion: How Spread Out the Data Is
Dispersion tells you how much the data varies. High dispersion = data is spread out; low dispersion = data is clustered.
2.1 Absolute vs. Relative Measures
| Measure | Formula | When to Use | Units |
|---|---|---|---|
| Range | Quick estimate of spread | Original | |
| Interquartile Range (IQR) | Robust to outliers (used with median) | Original | |
| Variance | Statistical analysis (squared units) | Squared | |
| Standard Deviation (SD) | Most common (same units as data) | Original | |
| Coefficient of Variation (CV) | Compare dispersion across datasets | Percentage |
Key Idea:
- Absolute measures (range, IQR, SD) depend on the data’s units.
- Relative measures (CV) are unitless and allow comparison (e.g., stock volatility).
Visual: Dispersion in Action
Worked Example 3: Variance and Standard Deviation Problem: Calculate the variance and standard deviation for the bus wait times from Worked Example 1. Step 1: Compute the mean (). Step 2: Calculate squared deviations from the mean. Step 3: Compute variance and SD.
Interpretation: The typical wait time varies by ±7.14 minutes from the mean.
2.2 Coefficient of Variation (CV)
Formula: Worked Example 4: Comparing Dispersion Problem: Two companies, A and B, have the following monthly salaries (in thousands):
- Company A: Mean = 40, SD = 5
- Company B: Mean = 60, SD = 10
Which has more consistent salaries? Step 1: Calculate CV for both. Conclusion: Company A has more consistent salaries (lower CV).
3. Real-World Applications in Nepal
3.1 eSewa and Khalti: Transaction Times
- Idea Used: Central tendency (median) and dispersion (IQR).
- How:
- eSewa reports the "typical" transaction time as the median (robust to slow outliers).
- The IQR shows how much times vary (e.g., "Most transactions take 5–10 seconds, but some take up to 20").
- Why not mean? A few very slow transactions (e.g., due to network issues) would skew the average.
Visual: eSewa Transaction Times
3.2 Ncell and NTC: Customer Wait Times
- Idea Used: Mean vs. median for skewed data.
- How:
- Ncell’s customer service reports the average wait time (mean) for calls, but the median is more reliable because some calls take much longer (e.g., technical issues).
- Dispersion (SD) helps them identify high-variability periods (e.g., weekends vs. weekdays).
Worked Example 5: Ncell Call Center
Problem: Ncell records call wait times (minutes) for 30 customers:
2, 3, 4, 5, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 35, 40
Step 1: Check skewness.
Observation: Right-skewed (long waits are outliers).
Step 2: Calculate measures.
- Mean = 12.5 minutes (inflated by 35 and 40).
- Median = 10 minutes (better "typical" value).
- SD = 11.2 minutes (high dispersion).
Conclusion: Ncell should advertise the median (10 minutes) and work to reduce the SD (e.g., hire more agents during peak hours).
3.3 Daraz and Pathao: Order Fulfillment Times
- Idea Used: Standard deviation for performance benchmarking.
- How:
- Daraz tracks the SD of delivery times to ensure consistency. A high SD means some orders take much longer than others (e.g., due to traffic or weather).
- Pathao uses IQR to set "guaranteed delivery windows" (e.g., "90% of orders arrive within 15–30 minutes").
Worked Example 6: Pathao Delivery Times
Problem: Pathao records delivery times (minutes) for 20 orders:
12, 15, 18, 20, 22, 25, 28, 30, 32, 35, 14, 16, 19, 21, 24, 26, 29, 31, 33, 40
Step 1: Compute five-number summary.
- Min = 12
- Q1 = 18th value = 19
- Median (Q2) = Average of 10th and 11th = 22.5
- Q3 = 29
- Max = 40 Step 2: Calculate IQR. Interpretation: Pathao can promise "Most deliveries take between 19 and 29 minutes" (Q1 to Q3), with only 25% of orders outside this range.
4. Common Pitfalls and Exam Traps
Ignoring Data Type:
- ❌ Using mean for nominal data (e.g., bus routes: "Route 1", "Route 2").
- ✅ Use mode for categories.
Forgetting Midpoints for Grouped Data:
- ❌ Calculating mean without converting classes to midpoints.
- ✅ Always use .
Mixing Population and Sample Variance:
- Population variance: Divide by .
- Sample variance: Divide by (for unbiased estimate).
Overlooking Skewness:
- ❌ Reporting mean for highly skewed data.
- ✅ Use median for skewed distributions.
Units in Dispersion:
- ❌ Comparing SD of salary (₹) and height (cm) directly.
- ✅ Use CV for unitless comparison.
Exam Tip: How to Score Full Marks
Always State the Formula:
- Write the formula before plugging in numbers. Example:
"Mean = . Here, and , so Mean = ."
- Write the formula before plugging in numbers. Example:
Justify Your Choice:
- If asked to pick a measure, explain why. Example:
"The median is more appropriate here because the data is right-skewed (presence of extreme values like 20 and 40), and the median is robust to outliers."
- If asked to pick a measure, explain why. Example:
Show All Steps for Grouped Data:
- Exam markers deduct marks if you skip midpoint calculations. Always show a table like this:
| Class | Midpoint (m) | Frequency (f) | f × m | |-------|--------------|---------------|-------| | ... | ... | ... | ... |
- Exam markers deduct marks if you skip midpoint calculations. Always show a table like this:
Interpret Results:
- Don’t just compute; explain what it means. Example:
"A standard deviation of 7.14 minutes means that most bus wait times fall within ±7 minutes of the mean (9.15 minutes), i.e., between 2 and 16 minutes."
- Don’t just compute; explain what it means. Example:
Watch for Tricks:
- Open-ended classes: For "50 and more", assume midpoint = 55 (or use upper limit if specified).
- Sample vs. population: If the problem says "sample," use for variance.
Quick Revision Table
| Concept | Formula | When to Use | Example |
|---|---|---|---|
| Mean | Symmetric data, interval/ratio | Average salary in a company | |
| Median | th value | Skewed data, ordinal data | Median house price in Kathmandu |
| Mode | Most frequent value | Nominal data, unimodal distributions | Most popular Daraz product |
| Range | Quick spread estimate | Temperature range in a day | |
| Variance | Statistical analysis | Calculating risk in investments | |
| Standard Deviation | Most common dispersion measure | Ncell call wait time variability | |
| Coefficient of Variation | Comparing datasets with different units | Comparing stock volatility |
Final Worked Example: Comprehensive Problem
Problem: The following table shows the marks obtained by 130 students in a Statistics exam:
| Marks | less than 40 | 40-50 | 50-60 | 60-70 | 70-80 | 80 or above |
|---|---|---|---|---|---|---|
| No. of students | 14 | 25 | 30 | 40 | 15 | 6 |
i) Compute the mean marks. ii) Compute the standard deviation. iii) Interpret the results.
Solution: i) Mean Calculation:
| Class | Midpoint (m) | Frequency (f) | f × m |
|---|---|---|---|
| <40 | 20 | 14 | 280 |
| 40-50 | 45 | 25 | 1125 |
| 50-60 | 55 | 30 | 1650 |
| 60-70 | 65 | 40 | 2600 |
| 70-80 | 75 | 15 | 1125 |
| ≥80 | 85 | 6 | 510 |
| Total | 130 | 7290 |
ii) Standard Deviation: Step 1: Compute for each class. Step 2: Compute variance and SD.
iii) Interpretation:
- The average mark is 56.08, but the SD of 40.16 shows high variability. This suggests:
- Some students scored very low (<40), while others scored near perfect (80+).
- The exam may have been difficult for some or had uneven question distribution.
- Action: Check if the exam had easy and hard sections or if teaching was inconsistent.
Summary Checklist for Exams
Before submitting, ensure you’ve covered:
- Defined all terms (mean, median, mode, range, SD, CV).
- Calculated all required measures (show tables for grouped data).
- Justified your choice of measure (e.g., "median for skewed data").
- Interpreted results in real-world terms (e.g., "high SD means inconsistent performance").
- Avoided common mistakes (e.g., using mean for nominal data, forgetting midpoints).
Based on the TU BSc CSIT syllabus for Statistics I (STA169), unit 3.
Discussion
Loading…