Business StatisticsUnit 612 min read
Sampling Methods, Distributions & Errors: Types, Bias, and Central Limit Theorem
Unit 6 of Business Statistics explores how to select representative samples from populations, the mathematical properties of sampling distributions, and how to avoid common pitfalls like bias and non-response errors—essential for market research, quality control, and financial forecasting.
TAKEAWAYS:
- Sampling vs. Census: Learn why sampling is preferred in practice (cost, time, feasibility) and when a census is unavoidable (e.g., voter registration).
- Probability vs. Non-Probability Sampling: Master the trade-offs between randomness (accuracy) and convenience (speed) in real-world data collection.
- Sampling Distributions: Understand how sample statistics (mean, proportion) vary around the population parameter and why the Central Limit Theorem (CLT) is the foundation of inferential statistics.
- Bias and Errors: Identify and quantify systematic errors (selection bias, non-response bias) and random errors (sampling error) to improve data quality.
- Sample Size Calculation: Apply formulas to determine the minimum sample size needed for a given margin of error and confidence level (critical for surveys and polls).
- Real-World Applications: See how companies like Nepal Rastra Bank (NRB) use stratified sampling to audit bank transactions, or how Daraz optimizes inventory using cluster sampling in warehouses.
1. Why Sample? Sampling vs. Census
Definitions
- Population (N): The entire group of individuals or items about which information is desired (e.g., all voters in Nepal, all Daraz customers in 2024).
- Sample (n): A subset of the population selected for analysis (e.g., 1,000 Daraz customers surveyed for satisfaction).
- Census: Collecting data from every member of the population (e.g., Nepal’s population census every 10 years).
Why Not Always Use a Census?
| Factor | Census | Sampling |
|---|---|---|
| Cost | High (e.g., NTC surveying all users) | Low (e.g., Pathao surveying 5% riders) |
| Time | Months/years (e.g., NEPSE auditing all listed companies) | Days/weeks (e.g., Daraz analyzing 10% orders) |
| Feasibility | Impossible for large populations (e.g., all WhatsApp users) | Practical for any population size |
| Accuracy | 100% precise but may have errors (e.g., miscounting in a census) | Approximate but statistically reliable if designed well |
When to Use a Census?
- Small, well-defined populations (e.g., employees in a single office, students in one class).
- Example: Nepal’s National ID card verification (2023) used a census for all 30 million registered citizens to ensure accuracy in voter lists.
2. Types of Sampling Methods
A. Probability Sampling (Random Selection)
Every member has a known chance of being selected. Reduces bias but requires more effort.
mindmap
root((Probability Sampling))
Random Sampling["Simple Random: Every nth item (e.g., lottery)"]
Stratified["Divide population into subgroups (strata), then sample each (e.g., age groups in a survey)"]
Cluster["Divide into clusters (e.g., districts), then sample entire clusters (e.g., NTC testing 5 districts)"]
Systematic["Select every kth item from a list (e.g., auditing every 100th transaction at Ncell)"]Worked Example 1: Simple Random Sampling
Scenario: A bank wants to survey 100 customers out of 10,000 to estimate average loan repayment time. Method: Assign each customer a number (1–10,000), then use a random number generator to pick 100 unique numbers. Visual: Advantages:
- Unbiased (everyone has equal chance).
- Easy to analyze statistically. Disadvantages:
- May miss subgroups (e.g., rural vs. urban customers).
- Expensive for large populations.
Worked Example 2: Stratified Sampling
Scenario: Nepal Rastra Bank (NRB) wants to audit bank transactions but must ensure rural and urban banks are represented proportionally. Population: 50 rural banks, 50 urban banks (total 100). Sample: 20 banks total. Method:
- Divide into strata: Rural (50%), Urban (50%).
- Sample 10 rural banks and 10 urban banks (proportional allocation). Visual:
B. Non-Probability Sampling (Convenience)
Used when random sampling is impractical. Higher risk of bias but faster/cheaper.
mindmap
root((Non-Probability Sampling))
Convenience["Easiest to reach (e.g., students outside TU campus)"]
Judgmental["Expert selects samples (e.g., Daraz hiring focus groups)"]
Snowball["Existing subjects recruit others (e.g., rare disease studies)"]
Quota["Fill predefined quotas (e.g., 50% male, 50% female)"]Worked Example 3: Quota Sampling
Scenario: A market research firm wants to survey 200 people in Kathmandu about eSewa usage, with quotas:
- 50% male, 50% female.
- 30% under 30, 70% over 30. Method:
- Interview people until quotas are filled (e.g., stop after 50 females under 30). Visual: Advantages:
- Quick and cost-effective.
- Ensures representation of key groups. Disadvantages:
- Bias if quotas are poorly defined (e.g., overrepresenting TU students).
3. Sampling Errors and Biases
A. Sampling Error (Random)
- Occurs due to chance variation in selecting a sample.
- Reduced by: Increasing sample size.
- Example: If you sample 50 Daraz customers and get an average order value of ₹5,000, but the true average is ₹5,200, the difference (₹200) is sampling error.
B. Non-Sampling Errors (Systematic)
More dangerous than random errors because they skew results permanently.
| Error Type | Definition | Example |
|---|---|---|
| Selection Bias | Sample not representative of population | Surveying only TU students for national eSewa usage (misses rural users). |
| Non-Response Bias | People who refuse to participate differ from those who do | Online polls where only tech-savvy people respond. |
| Measurement Bias | Flawed data collection tools | Asking "Do you use Khalti?" but excluding illiterate respondents. |
| Survivorship Bias | Only "successful" cases are sampled | Studying only Daraz sellers who succeeded (ignoring those who failed). |
Worked Example 4: Non-Response Bias
Scenario: A company sends a satisfaction survey to 1,000 customers but only 200 respond. Problem: The 800 who didn’t respond might be angry customers (lower satisfaction scores). Solution: Use follow-ups or incentives (e.g., discounts) to increase response rate.
4. Sampling Distributions and the Central Limit Theorem (CLT)
Key Idea
- If you take many random samples from a population and calculate their means, those means will form a normal distribution (bell curve), even if the original population is not normal.
- This is the Central Limit Theorem (CLT).
Visualizing the CLT
Parameters:
- Population Mean (μ): 50
- Population Std (σ): 10
- Sample Size (n): 30
- Std of Sampling Distribution (σₓ̄):
Why CLT Matters
- Allows us to use normal distribution to estimate population parameters from samples.
- Justifies confidence intervals and hypothesis testing.
Worked Example 5: CLT in Action
Scenario: Nepal Electricity Authority (NEA) wants to estimate the average monthly electricity bill (μ) for households in Kathmandu. They take 50 random samples of 30 households each. Population: Known to be right-skewed (most bills are low, but a few are very high). Sampling Distribution:
- Mean of sample means = μ (population mean).
- Std of sample means = . Conclusion: Even though individual bills are skewed, the average bill across samples will be normally distributed.
5. Sample Size Determination
Formula
The minimum sample size () for a given margin of error () and confidence level (z-score) is: Where:
- : Confidence level (1.96 for 95% confidence).
- : Population standard deviation (or use if unknown).
- : Margin of error (e.g., ±5%).
Worked Example 6: Sample Size for a Poll
Scenario: A political party wants to estimate support for a candidate with a 95% confidence level and ±3% margin of error. Assumption: From past data, (standard deviation of support percentages). Calculation: Interpretation: The party needs to survey at least 1,067 voters to ensure their estimate is within ±3% of the true value.
Adjustments for Finite Populations
If sampling without replacement (e.g., from a small population like 1,000 customers), use the finite population correction factor (FPC): Where = population size. Example: For , : Result: Only 533 samples are needed (instead of 1,067) because the population is small.
6. Real-World Applications
A. Nepal Rastra Bank (NRB) and Financial Audits
- Method: Stratified sampling to audit bank transactions.
- Strata: Rural banks, urban banks, foreign exchange dealers.
- Sample: 10% of transactions in each stratum.
- Why? Ensures no group is underrepresented, and detects fraud patterns.
B. Daraz and Inventory Management
- Method: Cluster sampling for warehouse stock checks.
- Clusters: Warehouses in Kathmandu, Pokhara, Biratnagar.
- Sample: Check 30% of items in 2 randomly selected warehouses.
- Why? Reduces cost while ensuring inventory accuracy across regions.
C. Pathao and Rider Satisfaction Surveys
- Method: Systematic sampling (every 50th rider).
- Problem: Non-response bias (only happy riders reply).
- Solution: Offer incentives (₹50 credit) to increase response rate.
D. NTC and Customer Satisfaction
- Method: Quota sampling to ensure representation by age, gender, and region.
- Example Quotas:
- 40% female, 60% male.
- 30% under 30, 70% over 30.
- Outcome: More reliable than convenience sampling (e.g., only surveying call center users).
Exam Tip
- Distinguish between sampling methods: Always explain why a method is chosen (e.g., stratified for heterogeneous populations, cluster for geographic efficiency).
- CLT is key: Expect questions on how sample means form a normal distribution, even if the population is skewed.
- Sample size formulas: Memorize the formula and know when to apply the finite population correction.
- Bias vs. error:
- Bias = Systematic (e.g., selection bias).
- Error = Random (e.g., sampling error).
- Real-world ties: Exams often ask how companies (e.g., NRB, Daraz) apply sampling. Link theory to practice!
- Graphs: Be ready to sketch:
- Sampling distributions (normal curve).
- Stratified vs. cluster sampling diagrams.
- Venn diagrams for bias sources.
Practice Questions
- NRB Audit: If NRB wants to audit 5% of all bank transactions in Nepal (10 million transactions), how would you design a stratified sample? What biases might arise if they use convenience sampling instead?
- Daraz Inventory: Daraz has 10 warehouses. To estimate average stock levels, they sample 2 warehouses entirely. Is this cluster sampling? What’s the advantage over simple random sampling?
- CLT Application: A population has a uniform distribution (not normal). If you take samples of size 30, will the sampling distribution of the mean be normal? Why?
- Sample Size: Calculate the sample size needed to estimate the average monthly salary in Kathmandu with a 99% confidence level and ±₹500 margin of error, given . Assume the population is 1 million.
Based on the TU BIM syllabus for Business Statistics (STT201), unit 6.
Discussion
Loading…