Business StatisticsUnit 814 min read

Sampling Methods, Sampling Distributions & Estimation Techniques

Unit 8 of Business Statistics covers sampling techniques (probability vs. non-probability), sampling distributions, point and interval estimation, and confidence intervals—essential for making data-driven business decisions with limited resources.

TAKEAWAYS:

  • Sampling saves time/money: Use representative samples (e.g., 500 Nepali voters) instead of studying entire populations (millions).
  • Probability sampling is unbiased: Methods like stratified sampling (dividing Daraz customers by region) ensure fair representation.
  • Sampling distributions explain variability: The Central Limit Theorem shows why sample means cluster around the true mean (e.g., Ncell’s average call duration).
  • Confidence intervals quantify uncertainty: A 95% CI for NEPSE’s stock returns means we’re 95% sure the true return lies within ±X%.
  • Estimation types differ: Point estimates (e.g., "average Khalti transaction = Rs. 2,500") vs. interval estimates (e.g., "Rs. 2,400–2,600").
  • Sample size matters: Larger samples (e.g., 1,000 vs. 100 Pathao riders) reduce margin of error but cost more.

1. Why Sample? The Basics

Definition: Sampling is the process of selecting a subset (sample) from a population to estimate characteristics (e.g., average income, customer satisfaction) without studying everyone. Critical for businesses like eSewa (surveying 1,000 users instead of 10 million) or NTC (testing 500 internet speeds across Nepal).

Key Terms:

Term Definition Example
Population Entire group of interest (e.g., all Nepali smartphone users). All 30 million mobile subscribers in Nepal.
Sample Subset of the population (e.g., 500 users surveyed). 500 Pathao riders in Kathmandu.
Parameter Numerical description of a population (e.g., mean income = Rs. 45,000). Average monthly Khalti transaction in Nepal.
Statistic Numerical description of a sample (e.g., sample mean income = Rs. 44,000). Average Daraz order value from 200 customers.

Why Sample?

  • Cost-effective: Surveying 1,000 Daraz customers costs less than studying all 5 million.
  • Time-saving: NEPSE can’t wait months to analyze every stock; samples give quick insights.
  • Practical: Impossible to test every Ncell SIM card’s battery life.

2. Sampling Methods: Probability vs. Non-Probability

A. Probability Sampling (Unbiased)

Every population member has a known chance of being selected. Ensures representativeness.

graph TD
    A["Probability Sampling"] --> B["Simple Random"]
    A --> C["Stratified"]
    A --> D["Cluster"]
    A --> E["Systematic"]
    B -->|"Example:"| F["Picking 100 NEPSE investors randomly from a list."]
    C -->|"Example:"| G["Divide Nepali voters by region (Hills, Terai, Mountains), then sample each."]
    D -->|"Example:"| H["Select 5 Kathmandu wards, then survey all businesses in those wards."]
    E -->|"Example:"| I["Every 10th Khalti user in a transaction log."]

1. Simple Random Sampling (SRS)

  • Every member has an equal chance of selection (e.g., lottery).
  • Pros: Simple, unbiased.
  • Cons: Expensive (e.g., mailing surveys to randomly selected Nepali households).
  • Example: To estimate average monthly electricity bill in Nepal:
    • Population: 8 million NTC customers.
    • Sample: 1,000 customers selected via random numbers.
    • Worked Example: Suppose bills (in Rs.) for 5 randomly selected customers: 1,200; 1,800; 900; 2,500; 1,500. Sample mean = . If the population mean (parameter) is Rs. 1,600, our estimate is close!

2. Stratified Sampling

  • Divide population into homogeneous subgroups (strata), then sample from each.
  • Pros: More precise (e.g., comparing rural vs. urban Khalti usage).
  • Cons: Complex design.
  • Example: Ncell wants to study call duration across Nepal.
    • Strata: Urban (Kathmandu, Pokhara), Semi-urban (Biratnagar, Dharan), Rural (Doti, Achham).
    • Sample: 300 urban, 200 semi-urban, 100 rural users.
    • Worked Example:
      Stratum Sample Size Avg. Call Duration (mins)
      Urban 300 12.5
      Semi-urban 200 8.0
      Rural 100 5.0
      Weighted mean = mins.

3. Cluster Sampling

  • Divide population into heterogeneous clusters, randomly select clusters, then survey all in them.
  • Pros: Cost-effective (e.g., surveying all schools in 10 districts instead of all 77).
  • Cons: Less precise.
  • Example: Daraz wants to study return rates across Nepal.
    • Clusters: 7 provinces.
    • Sample: Randomly pick 3 provinces (e.g., Gandaki, Lumbini, Karnali), then survey all orders in those provinces.

4. Systematic Sampling

  • Select every k-th member from a list (e.g., every 10th Khalti transaction).
  • Pros: Simple, uniform coverage.
  • Cons: Risk of periodicity bias (e.g., if transactions repeat every 10th entry).
  • Example: To study WhatsApp usage, select every 50th user from a list of 5 million Nepali WhatsApp accounts.

B. Non-Probability Sampling (Biased but Practical)

Used when probability sampling is impossible or too costly. Results cannot generalize to the population.

graph TD
    A["Non-Probability Sampling"] --> B["Convenience"]
    A --> C["Judgmental/Purposive"]
    A --> D["Snowball"]
    A --> E["Quota"]
    B -->|"Example:"| F["Surveying first 100 Daraz customers who visit a pop-up stall."]
    C -->|"Example:"| G["Selecting top 50 NEPSE traders based on experience."]
    D -->|"Example:"| H["Asking Khalti users to refer friends for a survey."]
    E -->|"Example:"| I["Interviewing 50 men and 50 women in Kathmandu (fixed quotas)."]

1. Convenience Sampling

  • Select easiest-to-reach members (e.g., students in a class).
  • Pros: Cheap, fast.
  • Cons: Highly biased (e.g., surveying only TU students won’t represent all Nepali students).
  • Example: A startup tests its app by giving free trials to first 200 people who download it (not representative of all potential users).

2. Judgmental/Purposive Sampling

  • Expert selects "typical" members (e.g., top 10 NEPSE analysts).
  • Pros: Targets specific insights.
  • Cons: Subjective, not generalizable.
  • Example: NTC hires consultants to study internet outages by selecting known trouble spots (e.g., Bhaktapur, Lalitpur).

3. Snowball Sampling

  • Initial respondents refer others (used for rare populations).
  • Pros: Useful for hard-to-reach groups (e.g., underground economy).
  • Cons: Bias accumulates.
  • Example: A study on Nepali freelancers starts with 10 known freelancers, who each refer 2 more.

4. Quota Sampling

  • Fix quotas (e.g., 30% female, 70% male) but select conveniently.
  • Pros: Ensures representation on key variables.
  • Cons: Still biased.
  • Example: A market research firm interviews 50 men and 50 women in Thamel, but only those walking past their office (not representative of all Nepalis).

3. Sampling Distributions & the Central Limit Theorem (CLT)

Definition: A sampling distribution shows how a statistic (e.g., sample mean) varies across repeated samples.

-3-2-1123-10000-8000-6000-4000-2000200040006000800010000xyNormal Distribution (Population)Sampling Distribution (n=30)μx̄
Sampling distribution of the mean (CLT: as n increases, x̄ approaches N(μ, σ/√n)).

Key Idea:

  • Even if the population distribution is skewed, the sampling distribution of the mean becomes normal as sample size increases (Central Limit Theorem).
  • Why it matters: Lets us use normal distribution tables for confidence intervals.

Worked Example: CLT in Action

  • Population: Call durations (mins) of Ncell users (skewed right: most calls are short, few are very long). Data: 3, 5, 7, 8, 10, 12, 15, 20, 30.
  • Sample: Take 100 samples of size n=5, calculate each sample’s mean.
  • Result: The distribution of these 100 means is normal, centered at the population mean (μ = 10).

4. Estimation: Point vs. Interval

012345678910Point Estimate (x̄)Lower BoundUpper Bound
Confidence interval for population mean (e.g., 95% CI: [5, 7]).

A. Point Estimation

  • Single value estimate (e.g., "average Khalti transaction = Rs. 2,500").
  • Unbiased estimator: On average, equals the parameter (e.g., sample mean = population mean).
  • Example: From a sample of 200 Daraz orders, the sample mean order value is Rs. 3,200. This is our point estimate for the population mean.

B. Interval Estimation (Confidence Intervals)

  • Range where the true parameter likely lies (e.g., "95% confident NEPSE’s return is between 2% and 5%").
  • Formula:
    • = sample mean
    • = z-score (1.96 for 95% CI)
    • = population standard deviation (or sample s if σ unknown)
    • = sample size

Worked Example: Confidence Interval for NEPSE Returns

  • Sample: 50 NEPSE stocks, average return last year = 3.5%, s = 1.2%.
  • 95% CI: CI: [3.16%, 3.84%]. Interpretation: We’re 95% confident the true population mean return lies between 3.16% and 3.84%.

5. Determining Sample Size

Goal: Balance precision (small margin of error) and cost (larger samples = more expensive). Formula:

  • = margin of error (e.g., ±2%)
  • = 1.96 (for 95% CI)
  • = population standard deviation (or pilot study estimate)

Worked Example: Sample Size for Khalti Transactions

  • Desired margin of error (E): ±Rs. 100.
  • Pilot study: Sample s = Rs. 500.
  • Calculation: Answer: Need a sample of 97 transactions for a 95% CI with ±Rs. 100 error.

In the Real World

  1. eSewa’s User Surveys

    • Method: Stratified sampling by region (Hills, Terai, Mountains) to estimate average transaction value.
    • Why? Ensures rural/urban differences are captured (e.g., Terai users may pay more for electricity bills).
  2. Ncell’s Network Quality Tests

    • Method: Cluster sampling—select 10 districts, then test all towers in those districts for call drop rates.
    • Real Impact: Helps Ncell identify regional outages (e.g., high drops in Doti vs. Kathmandu).
  3. Daraz’s Return Rate Analysis

    • Method: Systematic sampling—every 50th order is reviewed to estimate national return rate (3–5%).
    • Business Use: Reduces logistics costs by identifying high-return product categories.
  4. NEPSE’s Stock Index Calculation

    • Method: Probability sampling of 10–15 "representative" stocks (e.g., NMB, NBL, CG) to compute the NEPSE index.
    • Why? Impossible to track all 150+ listed companies daily.
  5. Pathao’s Driver Satisfaction Survey

    • Method: Snowball sampling—happy drivers refer friends, creating a biased but engaged sample for feedback.
    • Limitation: Overrepresents drivers who love Pathao (ignores unhappy ones).

Exam Tip

What Examiners Look For

  1. Distinguish probability vs. non-probability sampling:

    • Probability = random selection, generalizable (e.g., SRS, stratified).
    • Non-probability = convenient/judgmental, not generalizable (e.g., quota, snowball).
    • Common mistake: Calling convenience sampling "random."
  2. Apply the Central Limit Theorem (CLT):

    • Key phrase: "For large n, the sampling distribution of the mean is normal, regardless of population shape."
    • Exam question: Given a skewed population, why can we use z-tests for sample means? Answer: Because of CLT—sample means approximate a normal distribution.
  3. Calculate confidence intervals:

    • Formula: .
    • Interpretation: "We are 95% confident the true mean lies between X and Y."
    • Trap: Forgetting to use s (sample SD) if σ is unknown.
  4. Sample size questions:

    • Given: Desired margin of error (E), standard deviation (σ).
    • Find: n using .
    • Example: "How many NTC customers must be surveyed to estimate mean bill within ±Rs. 50 with 95% confidence? Assume σ = Rs. 200."
  5. Real-world connections:

    • eSewa: Stratified sampling by region.
    • Ncell: Cluster sampling by district.
    • NEPSE: Probability sampling of index stocks.
    • Daraz: Systematic sampling for return rates.

How to Score Full Marks

  • Diagrams: Draw sampling distribution curves (show normal shape for sample means).
  • Tables: Compare sampling methods (e.g., pros/cons of SRS vs. stratified).
  • Worked examples: Always label variables (e.g., , , ).
  • Interpretation: Explain why a method is chosen (e.g., "Stratified sampling ensures rural/urban differences are captured in Khalti’s user base.").

Practice Questions (Exam-Style)

  1. Short Answer:

    • Define stratified sampling and give a Nepali business example.
    • What is the Central Limit Theorem, and why is it useful for estimation?
  2. Calculation:

    • A sample of 100 NEPSE stocks has a mean return of 4% and s = 1.5%. Calculate a 90% CI for the true mean return. (Hint: z = 1.645 for 90% CI.)
  3. Application:

    • NTC wants to estimate the average internet speed in Nepal. They have a budget to survey 500 users. Which sampling method would you recommend, and why? Justify with a real-world advantage.
  4. Critical Thinking:

    • A market researcher uses convenience sampling to survey shoppers in Thamel about their spending habits. Can the results be generalized to all Nepalis? Why or why not? Suggest an alternative method.

Based on the PU BBA (PU) syllabus for Business Statistics, unit 8.

Discussion

Loading…