Business StatisticsUnit 138 min read
Combined Data Analysis – pooled mean, variance, five‑number summary & box plot
Unit 13 of Business Statistics explains how to merge two or more data sets, compute combined measures (mean, variance, SD, CV), construct five‑number summaries and box plots, and apply these techniques to real‑world business problems.
Key points
- Pooled (combined) mean is a weighted average of individual sample means.
- Combined variance/standard deviation requires both means and variances of the component samples.
- The five‑number summary (min, Q1, median, Q3, max) summarises shape and spread; a box plot visualises it.
- Coefficient of variation standardises dispersion for comparing different scales.
- Combined statistics are essential for aggregating sales, transaction amounts, or performance metrics across branches or periods.
1. What is “Combined Data Analysis”?
Combined data analysis (also called pooled or aggregate analysis) deals with merging two or more independent samples to obtain overall descriptive statistics. It is used when the underlying variable is the same (e.g., height, sales amount) but data are collected separately (different regions, time periods, or product lines).
Key concepts:
| Concept | Symbol | Formula (for two groups) |
|---|---|---|
| Combined (pooled) mean | ||
| Combined variance (unbiased) | ||
| Pooled variance (assuming equal population variance) | ||
| Combined standard deviation | ||
| Coefficient of variation | ||
| Five‑number summary | – |
Note: When more than two groups are involved, extend the formulas by summing over all groups.
2. Computing the Combined Mean
Worked Example 1 (Exam style)
A sample of heights of 6 400 Indians has a mean of 67.85 inches, SD = 2.56 in. A sample of 1 600 British has a mean of 68.55 inches, SD = 2.52 in. Find the combined mean.
Interpretation: The overall average height of the combined sample (8 000 people) is 67.99 inches.
3. Combined Variance & Standard Deviation
Worked Example 2 (Unequal variances)
Two investment companies A and B each have 100 observations.
Company A: mean = 28 (000 Rs), variance = 25.
Company B: mean = 37 (000 Rs), variance = 36.
Compute the combined standard deviation.
- Compute pooled variance (assuming equal population variance is not justified here, so we use the unbiased combined variance).
- Combined standard deviation
When to use pooled variance
If a preliminary test (e.g., Levene’s test) shows no significant difference in variances, the simpler pooled variance formula can be applied, saving one term in the numerator.
4. Five‑Number Summary & Box Plot
Worked Example 3
Data (in thousands of Rs):
- Order the data
Minimum = 8, Maximum = 70
Median (Q2) – 6th value (since ) → 32
Lower quartile (Q1) – median of lower half (first 5 values) → median of {8,11,15,20,26} = 15
Upper quartile (Q3) – median of upper half (last 5 values) → median of {39,45,52,60,70} = 52
Five‑number summary:
Shape comment: The median (32) is closer to Q1 (15) than to Q3 (52), indicating a slight right‑skew (longer upper tail).
5. Quartiles from Raw Data
Worked Example 4
Data:
Sort the data (n = 25).
Median (13th value) = 30
Lower quartile (Q1) = median of first 12 values = 11
Upper quartile (Q3) = median of last 12 values = 78
6. Coefficient of Variation (CV)
If the mean and CV = 30 %, then
Thus the standard deviation is 6.
7. Combining Regression Coefficients (Brief)
When two independent data sets share the same explanatory variable and we wish to fit a single regression line, we can pool the sums of squares:
The exam rarely asks for full derivation; knowing the principle is enough.
8. Comparison of Methods
| Situation | Use pooled variance? | Reason |
|---|---|---|
| Variances statistically equal (Levene’s p > 0.05) | Yes | Simpler, more precise estimate |
| Variances differ markedly | No | Use unbiased combined variance formula |
| Want to compare dispersion across groups | No | Keep separate SDs or CVs |
9. Applications in Business
| Business need | Combined statistic used | How it helps |
|---|---|---|
| Aggregating daily sales from multiple stores | Combined mean & SD | Forecast total revenue, set inventory levels |
| Evaluating loan portfolio risk across regions | Pooled variance | Estimate overall risk, price interest rates |
| Monitoring network traffic from several towers | Combined CV | Detect abnormal spikes relative to average usage |
| Merging customer satisfaction scores from two product lines | Five‑number summary & box plot | Visualise overall satisfaction spread and outliers |
10. In the real world
eSewa – The digital wallet aggregates transaction amounts from dozens of merchants each day. By calculating the combined mean transaction value, eSewa can set minimum top‑up thresholds and predict cash‑out requirements.
Daraz order queue – Daraz monitors order processing times for two major warehouses. Using the combined standard deviation, the logistics team estimates the overall variability in delivery time, allowing them to allocate extra staff during peak periods.
Ncell network usage – Ncell collects data‑usage (GB) from urban and rural towers. The coefficient of variation of the combined data tells the network engineers whether usage is uniformly spread (low CV) or highly variable (high CV), guiding capacity upgrades.
Worked example tie‑in: The height‑comparison problem (Indian vs. British) mirrors how a multinational retailer might compare average basket sizes across two countries before deciding on a unified pricing strategy.
11. Step‑by‑Step Procedure (Mermaid)
flowchart TD A["Start"] --> B["Compute Combined Mean"] B --> C["Compute Combined Variance"] C --> D["Compute Coefficient of Variation"] D --> E["Create Five‑Number Summary"] E --> F["Plot Box Plot"] F --> G["End"]Step‑by‑Step Procedure for Combined Data Analysis
12. Summary of Key Formulas
In the real world
- eSewa aggregates transaction amounts from thousands of merchants daily; it uses the combined mean to set minimum top‑up thresholds and forecast cash‑out requirements.
- Daraz monitors order processing times from two major warehouses; the combined standard deviation helps the logistics team estimate overall delivery‑time variability and allocate staff during peak periods.
- Ncell collects data‑usage (GB) from urban and rural towers; the coefficient of variation of the combined data indicates whether usage is uniformly spread or highly variable, guiding capacity upgrades.
Exam tip
- Memorise the combined‑mean formula; it appears in every multi‑sample question.
- When the question gives means, SDs, and sample sizes, plug directly into the unbiased combined‑variance formula—don’t try to reconstruct raw data.
- For five‑number summary questions, always sort the data first; then locate positions using for Q1 and Q3 (round to nearest integer).
- Box‑plot drawing: mark min, Q1, median, Q3, max; whiskers extend to the smallest/largest values within 1.5 IQR. Outliers (if any) are plotted individually.
- Time management: allocate ~2 minutes for sorting, ~3 minutes for calculations, and ~1 minute for drawing the box plot.
Based on the TU BBM syllabus for Business Statistics (STT201), unit 13.
Discussion
Loading…