STT201 Business Statistics

Business StatisticsUnit 48 min read

Sampling Techniques & Distributions: Methods, Bias, Errors & Normal Approximation

Unit 4 of Business Statistics covers sampling techniques (probability vs. non-probability), sampling distributions, central limit theorem, and sampling errors—essential for designing surveys, quality control, and market research in Nepal’s business sector (e.g., NTC’s customer satisfaction polls, Daraz’s product testin

Core Concepts: What is Sampling?

Sampling is the process of selecting a subset (sample) from a larger group (population) to study its characteristics. Why sample?

  • Cost-effective: Testing all 30 million Nepali voters is impossible; a sample of 1,500 suffices for election polls.
  • Time-saving: Measuring every Daraz order’s delivery time would halt operations; sampling tracks trends.
  • Practical: Destroying a batch of Ncell batteries to test quality isn’t feasible—sampling avoids waste.

Key Definitions

Population vs. Sample:

Term Example (Nepal) Symbol Formula
Population All 30M Nepali adults Mean =
Sample 1,500 adults surveyed for NEPSE reports Mean =

Sampling Techniques: How to Select Your Sample

Sampling methods are classified into two broad categories:

1. Probability Sampling (Random Selection)

Every population unit has a known chance of being selected. Reduces bias but requires a sampling frame (e.g., NTC’s road network map).

Types of Probability Sampling

graph TD
    A["Probability Sampling"] --> B["Simple Random"]
    A --> C["Stratified"]
    A --> D["Cluster"]
    A --> E["Systematic"]
    B -->|"Example"| F["Lottery method for NEPSE investor surveys"]
    C -->|"Example"| G["Divide Nepali households by income tiers, then sample each"]
    D -->|"Example"| H["Select 5 districts, then survey all households in those districts"]
    E -->|"Example"| I["Every 10th customer at a Khalti ATM"]

Worked Example: Simple Random Sampling

Scenario: NTC wants to estimate average traffic speed on 500 roads. They randomly select 50 roads using a random number generator.

  • Steps:
    1. Assign numbers 1–500 to all roads.
    2. Use RAND() in Excel to pick 50 unique numbers.
    3. Measure speed on those roads.
  • Visual:

Advantages:

  • Unbiased (every road has equal chance).
  • Easy to analyze statistically.

Disadvantages:

  • Expensive if population is spread out (e.g., rural vs. urban Nepal).
  • May miss subgroups (e.g., only high-speed roads if sample is skewed).

2. Non-Probability Sampling (Judgment-Based)

Used when random sampling is impractical (e.g., rare diseases, expert opinions). Higher bias risk but faster/cheaper.

Types of Non-Probability Sampling

graph TD
    A["Non-Probability Sampling"] --> B["Convenience"]
    A --> C["Purposive"]
    A --> D["Snowball"]
    A --> E["Quota"]
    B -->|"Example"| F["Surveying Pathao riders waiting at a single station"]
    C -->|"Example"| G["Selecting top 10 Daraz sellers for feedback"]
    D -->|"Example"| H["Asking Ncell employees to refer tech-savvy friends"]
    E -->|"Example"| I["Interview 50 people: 20 young, 20 middle-aged, 10 elderly"]

Worked Example: Quota Sampling

Scenario: A bank wants to study loan repayment behavior. They set quotas:

  • 200 loans: 50 from rural areas, 70 from urban, 80 from semi-urban.
  • Selection: Interview the first 200 customers matching these criteria who walk into a branch.

Visual:

Advantages:

  • Quick and cost-effective.
  • Ensures representation of key groups (e.g., age, location).

Disadvantages:

  • Bias: Researchers may unconsciously favor certain groups.
  • Not generalizable: Results may not apply to the entire population.

Sampling Errors: Why Your Sample Might Be Wrong

Even with perfect methods, samples can misrepresent the population due to:

1. Sampling Error (Random Variation)

  • Cause: Natural variability in the sample (e.g., your 50 NTC roads might all be in Kathmandu by chance).
  • Fix: Increase sample size (law of large numbers).
  • Example:
    • Population mean speed (μ): 40 km/h.
    • Sample mean (x̄): 42 km/h (error = ±2 km/h).

2. Non-Sampling Error (Systematic Bias)

  • Cause: Flawed methodology (e.g., surveying only daytime traffic).
  • Types:
    • Selection bias: Excluding certain groups (e.g., only surveying Khalti users online).
    • Response bias: People lying (e.g., overreporting income to banks).
    • Non-response bias: Only 30% of NEPSE investors reply to your survey.

Real-World Impact:

  • 2017 Nepal Election Polls: Some exit polls missed rural trends due to urban sampling bias.
  • Daraz Customer Satisfaction: If you only survey buyers who left reviews (not silent users), you overestimate happiness.

Sampling Distributions: The Bridge to Statistics

A sampling distribution shows how a statistic (e.g., mean, proportion) varies across many samples from the same population.

Key Idea: Central Limit Theorem (CLT)

"No matter the population shape, the sampling distribution of the mean will be normal if the sample size is large enough (n ≥ 30)."

Visual Proof:

Why CLT Matters:

  • Lets us use normal distribution tables for confidence intervals.
  • Justifies why n=1,000 is standard for polls (n ≥ 30 ensures normality).

Sample Size Determination: How Big Should Your Sample Be?

Use this formula for proportions (e.g., % of Nepali adults using Khalti):

  • : Confidence level (1.96 for 95% confidence).
  • : Expected proportion (e.g., 50% if unsure).
  • : Margin of error (e.g., ±5%).

Worked Example: Goal: Estimate % of Pathao riders who’d switch to a competitor, with 95% confidence and ±4% error.

  • (worst-case variability).
  • .
  • → 601 riders.

Visual:


In the Real World

  1. eSewa’s Fraud Detection:

    • Idea: Stratified sampling by transaction size.
    • How: eSewa divides payments into tiers (₹1–10K, 10K–50K, >50K) and samples 10% from each to detect fraud patterns. High-value transactions get higher sampling rates (purposive sampling).
  2. NTC’s Traffic Congestion Study:

    • Idea: Cluster sampling + systematic sampling.
    • How: NTC divides Kathmandu into 50 zones (clusters), then picks every 5th zone (systematic). Within each zone, they measure speed at random intervals (simple random). This balances cost and coverage.
  3. Daraz’s Product Quality Control:

    • Idea: Random sampling with acceptance sampling.
    • How: For every 1,000 orders, Daraz randomly selects 30 items for quality checks. If >2 are defective, the batch is rejected (uses binomial distribution to set thresholds).

Exam Tip: How to Score Full Marks

  1. Define Clearly:

    • Always start with definitions. For example:

      "Stratified sampling is a probability technique where the population is divided into homogeneous subgroups (strata) before random sampling within each stratum."

  2. Compare Methods:

    • Examiners love tables. For example:
      Method Type Bias Risk When to Use
      Simple Random Probability Low Small, homogeneous groups
      Quota Non-prob High Quick market research
      Cluster Probability Moderate Large, spread-out groups
  3. Show Calculations:

    • For sample size, always write the formula and plug in numbers. Example:

      "For a margin of error of ±3% and p=0.5, n = (1.96² × 0.5 × 0.5)/0.03² = 1,068."

  4. Link to Nepal:

    • Use local examples (NTC, Daraz, banks) to explain concepts. Example:

      "If Ncell wanted to test battery life, they’d use systematic sampling by selecting every 100th battery from the production line to ensure even coverage."

  5. Avoid Common Mistakes:

    • ❌ "Stratified sampling is the same as cluster sampling." → Wrong! Stratified divides into groups before sampling; cluster samples whole groups.
    • ❌ "Increasing sample size reduces non-sampling error." → Wrong! Only reduces sampling error.

Final Visual Summary:

flowchart TD
    A["Start"] --> B["Is sampling frame available?"]
    B -->|"Yes"| C["Use Probability Sampling"]
    B -->|"No"| D["Use Non-Probability Sampling"]
    C --> E["Simple/Stratified/Cluster/Systematic"]
    D --> F["Convenience/Purposive/Quota/Snowball"]
    E --> G["Calculate Sample Size\n(n = Z²p(1-p)/e²)"]
    F --> H["Accept Higher Bias"]
    G --> I["Collect Data"]
    H --> I
    I --> J["Analyze\n(CLT applies if n ≥ 30)"]

Based on the TU BBM syllabus for Business Statistics (STT201), unit 4.

Discussion

Loading…