Elective Research Fundamentals

Research FundamentalsUnit 712 min read

Sampling Methods, Techniques & Applications

Unit 7 of Research Fundamentals explains how to select representative samples from populations, covering probability vs. non-probability techniques, sampling errors, and real-world applications in engineering research.

What is Sampling?

Sampling is the process of selecting a subset (sample) from a larger group (population) to study, analyze, or make inferences about the entire group. Proper sampling ensures representativeness, generalizability, and cost-efficiency in research.

Why Sample?

  • Population too large: Studying every voter in Nepal (30+ million) is impractical.
  • Resource constraints: Time, budget, and manpower limit full population studies.
  • Precision vs. cost: A well-chosen sample can yield results as accurate as a census at a fraction of the cost.

Population vs. Sample

Definition: Entire group to be studied (e.g., all PU ComputeExample: All 50,000+ students enrolled in PUPopulationDefinition: Subset of the population selected for studyExample: 500 randomly chosen students from 10 campusesSampleResearch Universe
Hierarchical relationship between population and sample

Types of Sampling Techniques

Sampling methods are classified into probability (random) and non-probability (non-random) techniques.

023466992Simple Random85Systematic78Stratified92Cluster88Convenience65Purposive72Snowball55Quota70
Relative frequency of sampling method usage in Nepalese research studies (2020-2024)

1. Probability Sampling

Every member of the population has a known chance of being selected. Ensures statistical validity but may be costly/time-consuming.

A. Simple Random Sampling

  • Definition: Every individual has an equal probability of selection (e.g., lottery method).
  • How it works:
    1. Define the population (e.g., all Daraz customers in Kathmandu).
    2. Assign a unique number to each member.
    3. Use a random number generator to select samples.
  • Example:
    • Nepal Electricity Authority (NEA) surveys 500 households out of 5 million to estimate energy consumption.
    • Worked Example: If NEA wants to sample 1,000 households from 100,000, it assigns numbers 1–100,000 and picks every 100th number (100, 200, 300...) using a random start.

B. Systematic Sampling

  • Definition: Selects every k-th individual from a list (e.g., every 50th student in a register).
  • Steps:
    1. Determine sampling interval: .
    2. Randomly select a starting point between 1 and k.
    3. Select every k-th individual thereafter.
  • Example:
    • Pathao surveys every 200th ride out of 50,000 daily rides to assess driver satisfaction.
    • Worked Example: For a population of 5,000 students and a sample of 200, . If the random start is 7, the sample includes students numbered: 7, 32, 57, 82, ...

C. Stratified Sampling

  • Definition: Population divided into homogeneous subgroups (strata), then samples are taken from each stratum.
  • When to use: When subgroups vary significantly (e.g., age, income, region).
  • Example:
    • Nepal Rastra Bank (NRB) studies inflation by sampling:
      • Urban households (stratum 1: 30% sample)
      • Rural households (stratum 2: 50% sample)
      • Semi-urban (stratum 3: 20% sample)
    • Worked Example: A study on smartphone usage among PU students divides students into:
      Stratum Size (N) Sample Size (n) Sampling Method
      1st Year 10,000 200 Simple Random
      2nd Year 12,000 240 Systematic (k=50)
      3rd Year 8,000 160 Stratified Random
      4th Year 5,000 100 Convenience (if limited)

D. Cluster Sampling

  • Definition: Population divided into heterogeneous clusters, then entire clusters are randomly selected.
  • When to use: When population is geographically dispersed (e.g., nationwide surveys).
  • Example:
    • NTC (Nepal Telecom) surveys internet usage by randomly selecting 10 districts out of 77, then surveying all households in those districts.
    • Worked Example: A study on electric vehicle adoption in Nepal:
      1. Divide Nepal into 7 provinces.
      2. Randomly select 3 provinces (e.g., Province 3, 5, 7).
      3. Survey all vehicle owners in those provinces.

2. Non-Probability Sampling

Selection is not random; used when probability methods are impractical. Less generalizable but faster/cheaper.

A. Convenience Sampling

  • Definition: Samples are chosen based on accessibility (e.g., students in a class).
  • Example:
    • A PU professor surveys 50 students in their class about online learning tools.
    • Risk: May not represent all PU students (e.g., only tech-savvy students respond).

B. Purposive Sampling

  • Definition: Researcher intentionally selects individuals with specific traits.
  • Example:
    • Studying cybersecurity threats in Nepal, a researcher targets IT professionals at Ncell or NTC.
    • Worked Example: A study on elderly health in Kathmandu targets only residents of senior citizen homes.

C. Snowball Sampling

  • Definition: Initial samples refer others with similar traits (used for hard-to-reach populations).
  • Example:
    • Studying undocumented migrant workers in Nepal: Start with 5 workers, who refer 10 more, and so on.
    • Risk: Sample may be homogeneous (e.g., only friends of initial respondents).

D. Quota Sampling

  • Definition: Population divided into strata, then non-randomly fill quotas per stratum.
  • Example:
    • A market research firm for Daraz sets quotas:
      • 40% urban, 30% semi-urban, 30% rural.
      • Interviewers stop once quotas are met (no random selection within strata).

Sampling Errors & Biases

Sampling Error (30%)Non-Sampling Error (25%)Measurement Bias (20%)Coverage Error (25%)
Distribution of error types in a sample of 50 published Nepali studies

1. Sampling Errors

Occur when the sample does not represent the population due to random variation.

  • Types:
    • Random error: Natural variation (e.g., a sample of 500 may overrepresent young voters).
    • Systematic error: Flawed method (e.g., surveying only daytime shoppers at Daraz).

2. Non-Sampling Errors

Due to data collection issues (not sample selection).

  • Examples:
    • Response bias: People lie or avoid sensitive questions (e.g., income in surveys).
    • Non-response bias: Only certain groups respond (e.g., wealthy users of eSewa vs. poor).
    • Measurement error: Poorly worded questions (e.g., "Do you use WhatsApp daily?" may exclude light users).

How to Choose the Right Sampling Method?


In the Real World

  1. eSewa & Khalti (Digital Payments)

    • Method: Stratified sampling by income groups.
    • How: eSewa surveys users segmented into:
      • Low-income (<Rs. 20,000/month)
      • Middle-income (Rs. 20,000–100,000)
      • High-income (>Rs. 100,000)
    • Why: Ensures feedback from all user segments, not just frequent high-value users.
  2. Pathao (Ride-Hailing App)

    • Method: Systematic sampling of ride data.
    • How: Pathao logs every 500th ride in Kathmandu to analyze peak hours and driver earnings.
    • Real Example:
      • If Pathao has 200,000 rides/day, it samples 400 rides (every 500th ride) to estimate average fare (Rs. 120 vs. actual Rs. 118).
  3. Nepal Stock Exchange (NEPSE)

    • Method: Cluster sampling of listed companies.
    • How: NEPSE divides companies into sectors (banking, hydropower, FMCG) and randomly selects 20% of companies in each sector for regulatory audits.
    • Why: Ensures all sectors are represented without surveying all 250+ listed companies.

Worked Example: Traffic Congestion Study in Kathmandu

Problem: Estimate average daily traffic delay in Kathmandu’s Thapathali–Kageshwori route. Population: All vehicles (cars, buses, motorcycles) on the route. Sample Size: 500 vehicles (due to budget constraints). Method: Stratified Systematic Sampling

Steps:

  1. Divide into strata (vehicle types):
    • Cars (40% of traffic)
    • Buses (20%)
    • Motorcycles (30%)
    • Others (10%)
  2. Calculate sample per stratum:
    • Cars:
    • Buses:
    • Motorcycles:
    • Others:
  3. Systematic selection:
    • For cars: List all cars passing a checkpoint, pick every 50th car (if 10,000 cars pass daily, sample 200).
    • Repeat for other strata.
  4. Measure delay for each sampled vehicle using GPS timestamps.

Result: Average delay = 18 minutes (vs. census estimate of 17.5 minutes).


Advantages & Disadvantages of Sampling Methods

Method Advantages Disadvantages Best For
Simple Random Unbiased, easy to analyze Time-consuming, may miss subgroups Small, homogeneous populations
Systematic Simple, evenly spaced Periodicity bias (e.g., surveying every 100th student in a repeating pattern) Ordered populations (e.g., phone books)
Stratified Represents all subgroups Complex, requires population data Heterogeneous populations (e.g., income groups)
Cluster Cost-effective for large areas Less precise, clusters may be homogeneous Geographically dispersed populations (e.g., nationwide surveys)
Convenience Fast, cheap High bias, not generalizable Preliminary studies, pilot tests
Purposive Targets specific traits Subjective, researcher bias Rare populations (e.g., experts)
Snowball Accesses hidden populations Non-random, may be homogeneous Undocumented groups (e.g., migrants)
Quota Quick, controls subgroups Non-random, interviewer bias Market research (e.g., Daraz surveys)

Sampling Size Determination

The sample size (n) depends on:

  1. Population size (N): Larger populations need larger samples.
  2. Confidence level: Higher confidence (e.g., 95%) requires larger n.
  3. Margin of error (e): Smaller error needs larger n.
  4. Population variability (σ): More variation → larger n.

Formula:

For large populations (N > 10,000):

  • : Z-score (1.96 for 95% confidence)
  • : Standard deviation (use pilot data or 0.5 if unknown)
  • : Margin of error (e.g., 5% = 0.05)

Example:

  • Study: PU student satisfaction with online exams.
  • Population (N): 50,000 students.
  • Confidence level: 95% ()
  • Margin of error (e): 5% (0.05)
  • Variability (σ): Assume 0.5 (moderate variation).

Conclusion: Survey 385 students for reliable results.


Exam Tip

What Examiners Look For

  1. Definitions: Clearly distinguish between probability vs. non-probability sampling.
  2. Applications: Link methods to real-world scenarios (e.g., "How would NTC use cluster sampling?").
  3. Worked Examples: Show step-by-step calculations for sample size or stratified sampling.
  4. Bias Awareness: Identify sampling errors in flawed studies (e.g., "Why is convenience sampling risky?").
  5. Visuals: Draw comparison tables or flowcharts to explain methods.

Common Pitfalls

  • Confusing stratified vs. cluster sampling:
    • Stratified: Subgroups are homogeneous (e.g., age groups).
    • Cluster: Subgroups are heterogeneous (e.g., districts).
  • Ignoring non-response bias: Always mention how to minimize it (e.g., follow-ups, incentives).
  • Overlooking sample size: Examiners may ask you to calculate n given parameters.

Model Answer Structure

For a 5-mark question on sampling methods:

  1. Define the method (1 mark).
  2. Explain how it works (1 mark).
  3. Give a real-world example (e.g., "Nepal Rastra Bank uses stratified sampling...") (1 mark).
  4. Discuss advantages/disadvantages (1 mark).
  5. Compare with another method (1 mark).

Based on the PU BE Computer (PU) syllabus for Research Fundamentals, unit 7.

Discussion

Loading…