STT201 Business Statistics

Business StatisticsUnit 612 min read

Sampling & Sampling Distributions: Methods, Bias, Errors & Central Limit Theorem

Unit 6 of Business Statistics covers sampling techniques (probability vs. non-probability), sampling distributions, the Central Limit Theorem, and how to calculate sampling errors—critical for real-world data collection in IT management, finance, and market research.

TAKEAWAYS:

  • Sampling vs. Census: Sampling is faster, cheaper, and practical for large populations (e.g., NTC surveys), while census is exhaustive but costly (e.g., NEPSE’s annual reports).
  • Probability Sampling: Methods like simple random, stratified, and cluster sampling reduce bias; non-probability methods (e.g., convenience sampling) risk skewing results.
  • Sampling Distribution: The distribution of a statistic (e.g., sample mean) across repeated samples follows a normal distribution (CLT), even if the population isn’t normal.
  • Central Limit Theorem (CLT): For large samples (n ≥ 30), the sampling distribution of the mean is normal, regardless of the population distribution. This justifies using z-tests and confidence intervals.
  • Sampling Error: The difference between a sample statistic (e.g., sample mean) and the population parameter. Reduce it by increasing sample size or using better sampling methods.
  • Real-World Tie: Khalti’s fraud detection uses stratified sampling to analyze transaction patterns across regions, while Pathao’s driver allocation relies on cluster sampling for efficient route planning.

1. Why Sample? Sampling vs. Census

Definitions

  • Population: The entire group being studied (e.g., all Nepalese IT students, all Daraz customers).
  • Sample: A subset of the population used to estimate population parameters.
  • Census: Collecting data from every member of the population (e.g., Nepal’s national census).

Advantages of Sampling

graph TD
    A["Sampling"] --> B["Faster"]
    A --> C["Cheaper"]
    A --> D["Practical for large populations"]
    A --> E["Less invasive"]
    A --> F["Allows statistical inference"]

Disadvantages of Sampling

  • Sampling Error: The difference between the sample statistic and the true population parameter.
  • Non-response Bias: If sampled individuals refuse to participate (e.g., low response rates in NTC’s customer satisfaction surveys).
  • Selection Bias: When the sample isn’t representative (e.g., surveying only Kathmandu residents for national trends).

When to Use Census?

  • Small populations (e.g., employees in a single company).
  • Critical decisions (e.g., NEPSE’s annual audits).

Visual: Two Venn diagrams showing how populations are divided into strata (homogeneous groups) vs. clusters (heterogeneous groups).


2. Types of Sampling Methods

A. Probability Sampling (Unbiased)

Every member has a known chance of being selected. Four key methods:

Method How It Works Example in Nepal When to Use
Simple Random Every member has equal probability (e.g., lottery). Selecting 500 households from Pokhara’s 20,000 for an NTC survey. Homogeneous populations.
Stratified Population divided into strata (e.g., age, income), then random samples from each. Surveying IT students from TU, PU, and KU separately for a tech skills report. Heterogeneous groups (e.g., urban vs. rural).
Cluster Population divided into clusters (e.g., wards), then entire clusters sampled. Selecting 5 wards from Kathmandu’s 33 for a traffic congestion study. Geographically spread populations.
Systematic Select every k-th member (e.g., every 10th customer at Daraz). Sampling every 50th transaction in Khalti’s database to detect fraud. Ordered populations (e.g., customer IDs).

B. Non-Probability Sampling (Biased but Practical)

Used when probability sampling is infeasible. Three common methods:

Method How It Works Example in Nepal Risk of Bias
Convenience Sample whoever is easiest to reach. Surveying TU students in Dhulikhel for a national education report. Overrepresents accessible groups.
Purposive Select members based on specific criteria. Interviewing top 100 NEPSE traders for market analysis. Excludes non-experts.
Snowball Existing subjects recruit others (e.g., referrals). Studying underground economy via contacts of initial informants. Limited to networks of initial subjects.

WORKED EXAMPLE 1: Stratified Sampling for Ncell’s Customer Satisfaction Ncell wants to survey 500 customers across 3 regions (East, West, Central) with populations of 2M, 1.5M, and 2.5M respectively. Design a stratified sample.

Steps:

  1. Calculate stratum sizes:
    • East:
    • West:
    • Central:
  2. Randomly select 167 East, 125 West, and 208 Central customers using simple random sampling within each stratum.

Why Stratified?

  • Ensures representation from all regions.
  • Reduces sampling error compared to simple random sampling.

3. Sampling Distribution and the Central Limit Theorem (CLT)

Key Concepts

  • Sampling Distribution: The distribution of a statistic (e.g., sample mean, ) across all possible samples of size n.
  • CLT: For large n (≥30), the sampling distribution of is approximately normal, regardless of the population distribution.

Visualizing the CLT

Key Observations:

  1. The sampling distribution is normal, even if the population is skewed (e.g., income data).
  2. The mean of the sampling distribution () equals the population mean ().
  3. The standard deviation of the sampling distribution ( or SE) is: where = population standard deviation.

WORKED EXAMPLE 2: CLT in Action – Daraz’s Order Processing Daraz processes orders with a population mean delivery time days and days. If a sample of n=40 orders is taken:

  1. What is the probability that the sample mean delivery time is >5.5 days?
  2. How does the sampling distribution change if n=100?

Solution:

  1. Calculate :
  2. Standardize the value:
  3. Find P(Z > 1.56) from z-table: 0.0586 (5.86%).
  4. For n=100: The distribution becomes narrower, reducing sampling error.

Real-World Tie: Daraz uses the CLT to predict delivery delays by analyzing sample means from different regions. A sample mean >5.5 days triggers alerts for logistics teams.


4. Sampling Error and How to Reduce It

Types of Sampling Error

  1. Random Error: Due to chance (e.g., selecting a non-representative sample).
  2. Systematic Error: Due to flawed methods (e.g., biased sampling frame).

How to Reduce Sampling Error

Method Effect
Increase sample size Reduces (e.g., from n=30 to n=100).
Use stratified sampling Ensures representation of subgroups (e.g., urban vs. rural in NTC surveys).
Avoid non-response bias Follow-ups for missing data (e.g., Khalti’s SMS reminders).
Pilot testing Check for systematic errors before full survey.

WORKED EXAMPLE 3: Sample Size Calculation for NEPSE NEPSE wants to estimate the average stock price with a margin of error (E) of Rs. 500 and 95% confidence. The population standard deviation () is Rs. 2,000. Calculate the required sample size.

Formula: where for 95% confidence.

Calculation:

Interpretation: NEPSE needs 62 stocks in the sample to ensure the estimate is within Rs. 500 of the true mean with 95% confidence.


5. Common Pitfalls in Sampling

graph TD
    A["Poor Sampling Frame"] --> B["Excludes key groups"]
    A --> C["Biased results"]
    D["Small Sample Size"] --> E["High sampling error"]
    D --> F["Unreliable estimates"]
    G["Non-Response Bias"] --> H["Overrepresents respondents"]
    G --> I["Underrepresents hard-to-reach groups"]

Example:

  • NTC’s 2022 survey used convenience sampling (only landline users), leading to underrepresentation of mobile-only customers.

In the Real World

  1. Khalti’s Fraud Detection

    • Idea Used: Stratified sampling by transaction type (e.g., peer-to-peer, merchant payments) and region.
    • How: Khalti divides transactions into strata (e.g., high-value vs. low-value) and samples proportionally to detect anomalies. For example, if 10% of high-value transactions are flagged in Kathmandu, they investigate further.
  2. Pathao’s Driver Allocation

    • Idea Used: Cluster sampling by city zones (e.g., Thapathali, Lakshmi Marg).
    • How: Pathao clusters drivers by location and samples entire clusters during peak hours to optimize ride matching. This reduces the need for real-time GPS data for every driver.
  3. NTC’s Customer Satisfaction Surveys

    • Idea Used: Systematic sampling of phone numbers from a random dialing frame.
    • Challenge: Non-response bias (e.g., elderly users may not answer). NTC mitigates this by offering incentives (e.g., free internet minutes).
  4. Daraz’s Inventory Management

    • Idea Used: Central Limit Theorem to predict demand.
    • How: Daraz analyzes sample means of daily orders across product categories. For example, if the sample mean for laptops is 500 units/day (±50), they stock accordingly, reducing over/under-stocking.
  5. NEPSE’s Index Calculation

    • Idea Used: Random sampling of listed companies to calculate the NEPSE index.
    • How: NEPSE randomly selects 50–100 stocks from its 200+ listings to compute the index, ensuring representativeness.

Exam Tip

What Examiners Look For

  1. Distinguish between sampling methods:

    • Probability vs. non-probability: Always state whether the method is random or biased.
    • Stratified vs. cluster: Explain how the population is divided and why (e.g., "stratified by income to ensure representation").
  2. CLT applications:

    • When to use: Always check if n ≥ 30 before applying the CLT.
    • Formulas: Know and .
    • Real-world link: Connect to confidence intervals (e.g., "Since n=100 > 30, we can use the CLT to construct a 95% CI for the mean").
  3. Sampling error calculations:

    • Margin of error (E): .
    • Sample size (n): Derive from .
  4. Common mistakes to avoid:

    • Ignoring non-response bias: Always mention how it affects results.
    • Assuming normality without CLT: Only use z-tests if n ≥ 30 or the population is normal.
    • Confusing stratified and cluster sampling: Stratified = homogeneous groups; cluster = heterogeneous groups.

High-Score Strategies

  • Draw diagrams: Sketch sampling distributions (e.g., population vs. sampling distribution) for CLT questions.

  • Link to real data: Even in theoretical questions, relate to Nepalese examples (e.g., "Like NTC’s surveys, stratified sampling ensures...").

  • Show calculations step-by-step: Examiners reward clarity. For example:

    "Given , , n=25, find P(). Step 1: . Step 2: . Step 3: P(Z > 1) = 0.1587."

  • Interpret results: Always explain what the probability/sampling error means in context (e.g., "There’s a 15.87% chance the sample mean exceeds 52, indicating potential delays in Pathao’s delivery times").

Based on the TU BITM syllabus for Business Statistics (STT201), unit 6.

Discussion

Loading…