Business StatisticsUnit 612 min read
Sampling & Sampling Distributions: Methods, Bias, Errors & Central Limit Theorem
Unit 6 of Business Statistics covers sampling techniques (probability vs. non-probability), sampling distributions, the Central Limit Theorem, and how to calculate sampling errors—critical for real-world data collection in IT management, finance, and market research.
TAKEAWAYS:
- Sampling vs. Census: Sampling is faster, cheaper, and practical for large populations (e.g., NTC surveys), while census is exhaustive but costly (e.g., NEPSE’s annual reports).
- Probability Sampling: Methods like simple random, stratified, and cluster sampling reduce bias; non-probability methods (e.g., convenience sampling) risk skewing results.
- Sampling Distribution: The distribution of a statistic (e.g., sample mean) across repeated samples follows a normal distribution (CLT), even if the population isn’t normal.
- Central Limit Theorem (CLT): For large samples (n ≥ 30), the sampling distribution of the mean is normal, regardless of the population distribution. This justifies using z-tests and confidence intervals.
- Sampling Error: The difference between a sample statistic (e.g., sample mean) and the population parameter. Reduce it by increasing sample size or using better sampling methods.
- Real-World Tie: Khalti’s fraud detection uses stratified sampling to analyze transaction patterns across regions, while Pathao’s driver allocation relies on cluster sampling for efficient route planning.
1. Why Sample? Sampling vs. Census
Definitions
- Population: The entire group being studied (e.g., all Nepalese IT students, all Daraz customers).
- Sample: A subset of the population used to estimate population parameters.
- Census: Collecting data from every member of the population (e.g., Nepal’s national census).
Advantages of Sampling
graph TD
A["Sampling"] --> B["Faster"]
A --> C["Cheaper"]
A --> D["Practical for large populations"]
A --> E["Less invasive"]
A --> F["Allows statistical inference"]Disadvantages of Sampling
- Sampling Error: The difference between the sample statistic and the true population parameter.
- Non-response Bias: If sampled individuals refuse to participate (e.g., low response rates in NTC’s customer satisfaction surveys).
- Selection Bias: When the sample isn’t representative (e.g., surveying only Kathmandu residents for national trends).
When to Use Census?
- Small populations (e.g., employees in a single company).
- Critical decisions (e.g., NEPSE’s annual audits).
Visual: Two Venn diagrams showing how populations are divided into strata (homogeneous groups) vs. clusters (heterogeneous groups).
2. Types of Sampling Methods
A. Probability Sampling (Unbiased)
Every member has a known chance of being selected. Four key methods:
| Method | How It Works | Example in Nepal | When to Use |
|---|---|---|---|
| Simple Random | Every member has equal probability (e.g., lottery). | Selecting 500 households from Pokhara’s 20,000 for an NTC survey. | Homogeneous populations. |
| Stratified | Population divided into strata (e.g., age, income), then random samples from each. | Surveying IT students from TU, PU, and KU separately for a tech skills report. | Heterogeneous groups (e.g., urban vs. rural). |
| Cluster | Population divided into clusters (e.g., wards), then entire clusters sampled. | Selecting 5 wards from Kathmandu’s 33 for a traffic congestion study. | Geographically spread populations. |
| Systematic | Select every k-th member (e.g., every 10th customer at Daraz). | Sampling every 50th transaction in Khalti’s database to detect fraud. | Ordered populations (e.g., customer IDs). |
B. Non-Probability Sampling (Biased but Practical)
Used when probability sampling is infeasible. Three common methods:
| Method | How It Works | Example in Nepal | Risk of Bias |
|---|---|---|---|
| Convenience | Sample whoever is easiest to reach. | Surveying TU students in Dhulikhel for a national education report. | Overrepresents accessible groups. |
| Purposive | Select members based on specific criteria. | Interviewing top 100 NEPSE traders for market analysis. | Excludes non-experts. |
| Snowball | Existing subjects recruit others (e.g., referrals). | Studying underground economy via contacts of initial informants. | Limited to networks of initial subjects. |
WORKED EXAMPLE 1: Stratified Sampling for Ncell’s Customer Satisfaction Ncell wants to survey 500 customers across 3 regions (East, West, Central) with populations of 2M, 1.5M, and 2.5M respectively. Design a stratified sample.
Steps:
- Calculate stratum sizes:
- East:
- West:
- Central:
- Randomly select 167 East, 125 West, and 208 Central customers using simple random sampling within each stratum.
Why Stratified?
- Ensures representation from all regions.
- Reduces sampling error compared to simple random sampling.
3. Sampling Distribution and the Central Limit Theorem (CLT)
Key Concepts
- Sampling Distribution: The distribution of a statistic (e.g., sample mean, ) across all possible samples of size n.
- CLT: For large n (≥30), the sampling distribution of is approximately normal, regardless of the population distribution.
Visualizing the CLT
Key Observations:
- The sampling distribution is normal, even if the population is skewed (e.g., income data).
- The mean of the sampling distribution () equals the population mean ().
- The standard deviation of the sampling distribution ( or SE) is: where = population standard deviation.
WORKED EXAMPLE 2: CLT in Action – Daraz’s Order Processing Daraz processes orders with a population mean delivery time days and days. If a sample of n=40 orders is taken:
- What is the probability that the sample mean delivery time is >5.5 days?
- How does the sampling distribution change if n=100?
Solution:
- Calculate :
- Standardize the value:
- Find P(Z > 1.56) from z-table: 0.0586 (5.86%).
- For n=100: The distribution becomes narrower, reducing sampling error.
Real-World Tie: Daraz uses the CLT to predict delivery delays by analyzing sample means from different regions. A sample mean >5.5 days triggers alerts for logistics teams.
4. Sampling Error and How to Reduce It
Types of Sampling Error
- Random Error: Due to chance (e.g., selecting a non-representative sample).
- Systematic Error: Due to flawed methods (e.g., biased sampling frame).
How to Reduce Sampling Error
| Method | Effect |
|---|---|
| Increase sample size | Reduces (e.g., from n=30 to n=100). |
| Use stratified sampling | Ensures representation of subgroups (e.g., urban vs. rural in NTC surveys). |
| Avoid non-response bias | Follow-ups for missing data (e.g., Khalti’s SMS reminders). |
| Pilot testing | Check for systematic errors before full survey. |
WORKED EXAMPLE 3: Sample Size Calculation for NEPSE NEPSE wants to estimate the average stock price with a margin of error (E) of Rs. 500 and 95% confidence. The population standard deviation () is Rs. 2,000. Calculate the required sample size.
Formula: where for 95% confidence.
Calculation:
Interpretation: NEPSE needs 62 stocks in the sample to ensure the estimate is within Rs. 500 of the true mean with 95% confidence.
5. Common Pitfalls in Sampling
graph TD
A["Poor Sampling Frame"] --> B["Excludes key groups"]
A --> C["Biased results"]
D["Small Sample Size"] --> E["High sampling error"]
D --> F["Unreliable estimates"]
G["Non-Response Bias"] --> H["Overrepresents respondents"]
G --> I["Underrepresents hard-to-reach groups"]Example:
- NTC’s 2022 survey used convenience sampling (only landline users), leading to underrepresentation of mobile-only customers.
In the Real World
Khalti’s Fraud Detection
- Idea Used: Stratified sampling by transaction type (e.g., peer-to-peer, merchant payments) and region.
- How: Khalti divides transactions into strata (e.g., high-value vs. low-value) and samples proportionally to detect anomalies. For example, if 10% of high-value transactions are flagged in Kathmandu, they investigate further.
Pathao’s Driver Allocation
- Idea Used: Cluster sampling by city zones (e.g., Thapathali, Lakshmi Marg).
- How: Pathao clusters drivers by location and samples entire clusters during peak hours to optimize ride matching. This reduces the need for real-time GPS data for every driver.
NTC’s Customer Satisfaction Surveys
- Idea Used: Systematic sampling of phone numbers from a random dialing frame.
- Challenge: Non-response bias (e.g., elderly users may not answer). NTC mitigates this by offering incentives (e.g., free internet minutes).
Daraz’s Inventory Management
- Idea Used: Central Limit Theorem to predict demand.
- How: Daraz analyzes sample means of daily orders across product categories. For example, if the sample mean for laptops is 500 units/day (±50), they stock accordingly, reducing over/under-stocking.
NEPSE’s Index Calculation
- Idea Used: Random sampling of listed companies to calculate the NEPSE index.
- How: NEPSE randomly selects 50–100 stocks from its 200+ listings to compute the index, ensuring representativeness.
Exam Tip
What Examiners Look For
Distinguish between sampling methods:
- Probability vs. non-probability: Always state whether the method is random or biased.
- Stratified vs. cluster: Explain how the population is divided and why (e.g., "stratified by income to ensure representation").
CLT applications:
- When to use: Always check if n ≥ 30 before applying the CLT.
- Formulas: Know and .
- Real-world link: Connect to confidence intervals (e.g., "Since n=100 > 30, we can use the CLT to construct a 95% CI for the mean").
Sampling error calculations:
- Margin of error (E): .
- Sample size (n): Derive from .
Common mistakes to avoid:
- Ignoring non-response bias: Always mention how it affects results.
- Assuming normality without CLT: Only use z-tests if n ≥ 30 or the population is normal.
- Confusing stratified and cluster sampling: Stratified = homogeneous groups; cluster = heterogeneous groups.
High-Score Strategies
Draw diagrams: Sketch sampling distributions (e.g., population vs. sampling distribution) for CLT questions.
Link to real data: Even in theoretical questions, relate to Nepalese examples (e.g., "Like NTC’s surveys, stratified sampling ensures...").
Show calculations step-by-step: Examiners reward clarity. For example:
"Given , , n=25, find P(). Step 1: . Step 2: . Step 3: P(Z > 1) = 0.1587."
Interpret results: Always explain what the probability/sampling error means in context (e.g., "There’s a 15.87% chance the sample mean exceeds 52, indicating potential delays in Pathao’s delivery times").
Based on the TU BITM syllabus for Business Statistics (STT201), unit 6.
Discussion
Loading…