Statistics IUnit 214 min read
Sampling Techniques: Probability vs. Non-Probability Methods & Real-World Applications
Unit 2 of Statistics I covers the core concepts of sampling techniques, distinguishing between probability and non-probability sampling methods, their applications, advantages, and limitations, with real-world examples from Nepalese and global industries.
TAKEAWAYS:
- Sampling is the process of selecting a subset of a population to estimate characteristics of the whole, reducing cost and effort while maintaining accuracy.
- Probability sampling ensures every member has a known chance of selection, guaranteeing representativeness (e.g., simple random, stratified, systematic, cluster sampling).
- Non-probability sampling relies on subjective judgment or convenience, often faster but prone to bias (e.g., quota, purposive, snowball sampling).
- Real-world applications include market research (e.g., Daraz customer surveys), public health studies (e.g., NTC’s traffic congestion analysis), and financial audits (e.g., banks sampling transaction records).
- Key trade-offs: Probability sampling is rigorous but resource-intensive; non-probability sampling is flexible but may lack generalizability.
- Exam focus: Define methods, compare advantages/disadvantages, and apply to scenarios (e.g., calculating sample sizes, identifying biases in given studies).
1. Introduction to Sampling Techniques
Sampling is the foundation of statistical inference. Instead of studying an entire population (e.g., all voters in Nepal, all Daraz customers, or all Ncell subscribers), we analyze a representative subset to draw conclusions about the whole. Poor sampling leads to biased or unreliable results, while good sampling ensures validity and generalizability.
Why Sample?
- Cost-effective: Surveying 1,000 Nepali households is cheaper than surveying all 29 million.
- Time-saving: Analyzing 100 NEPSE stock trades is faster than tracking every transaction.
- Practicality: Destroying a sample (e.g., testing battery life) is acceptable if the sample represents the population.
2. Probability Sampling: Fair and Representative
Probability sampling methods ensure every member of the population has a known, non-zero chance of being selected. This reduces bias and allows for statistical inference (e.g., calculating margins of error).
Types of Probability Sampling
graph TD
A["Probability Sampling"] --> B["Simple Random Sampling"]
A --> C["Stratified Sampling"]
A --> D["Systematic Sampling"]
A --> E["Cluster Sampling"]A. Simple Random Sampling (SRS)
- Definition: Every possible sample of size n has an equal chance of being selected.
- How it works:
- Assign a unique number to each population member (e.g., list all Ncell subscribers).
- Use a random number generator to select n members.
- Example: Selecting 500 voters from Kathmandu’s 2 million registered voters using a random digit table.
- Advantages:
- Unbiased, representative.
- Easy to analyze statistically.
- Disadvantages:
- Expensive and time-consuming for large populations.
- May not guarantee proportional representation of subgroups (e.g., rural vs. urban voters).
B. Stratified Sampling
- Definition: Divide the population into homogeneous subgroups (strata) and randomly sample from each stratum.
- How it works:
- Identify strata (e.g., age groups: 18–30, 31–50, 50+ for a Khalti user survey).
- Randomly sample from each stratum proportionally.
- Example: A bank auditing loan defaults might stratify by loan amount (small, medium, large) and sample 10% from each.
- Advantages:
- Ensures representation of all subgroups.
- More precise estimates for strata.
- Disadvantages:
- Requires prior knowledge of strata.
- More complex than SRS.
C. Systematic Sampling
- Definition: Select every k-th member from a list after a random start.
- How it works:
- Calculate k = population size / sample size (e.g., k = 1000 if sampling 100 from 10,000).
- Randomly select a start (1–1000), then pick every 1000th member thereafter.
- Example: Inspecting every 100th bag of rice from a Daraz warehouse shipment.
- Advantages:
- Simple and uniform coverage.
- Suitable for ordered populations (e.g., production lines).
- Disadvantages:
- Risk of periodic bias if the population has hidden patterns (e.g., defects every 50th item).
D. Cluster Sampling
- Definition: Divide the population into heterogeneous clusters, randomly select clusters, and sample all members within them.
- How it works:
- Group population into clusters (e.g., schools in Nepal, city blocks in Kathmandu).
- Randomly select clusters (e.g., 5 schools out of 100).
- Survey all students in selected schools.
- Example: NTC studying traffic congestion might cluster by district (e.g., select 3 districts out of 77) and survey all households in those districts.
- Advantages:
- Cost-effective for large, spread-out populations.
- No need for a complete population list.
- Disadvantages:
- Less precise than SRS if clusters are homogeneous.
3. Non-Probability Sampling: Convenient but Biased
Non-probability sampling methods rely on availability or judgment, not randomness. Results cannot be generalized to the population but are useful for exploratory or qualitative studies.
Types of Non-Probability Sampling
graph TD
A["Non-Probability Sampling"] --> B["Convenience Sampling"]
A --> C["Quota Sampling"]
A --> D["Purposive Sampling"]
A --> E["Snowball Sampling"]A. Convenience Sampling
- Definition: Select the easiest-to-reach members (e.g., students in a TU classroom).
- Example: A Pathao driver surveying passengers at a single pickup point in Thapathali.
- Advantages:
- Fast and cheap.
- Disadvantages:
- Highly biased (e.g., only tech-savvy users if sampling at a cybercafé).
B. Quota Sampling
- Definition: Divide the population into strata and fill quotas based on characteristics (e.g., 50% male, 50% female).
- Example: A Daraz market research team might set quotas for age groups (18–25, 26–40, 40+) and stop sampling once quotas are met.
- Advantages:
- Ensures representation of key groups.
- Flexible and faster than stratified sampling.
- Disadvantages:
- Samplers may cherry-pick easy-to-reach members, introducing bias.
C. Purposive Sampling
- Definition: Select members based on specific criteria (e.g., expert opinions).
- Example: Interviewing 10 top NEPSE analysts to predict stock trends.
- Advantages:
- Targets specific information needs.
- Disadvantages:
- Not representative; results may not apply broadly.
D. Snowball Sampling
- Definition: Start with a few members, then ask them to refer others with similar traits.
- Example: Studying rare diseases by asking diagnosed patients to refer others.
- Advantages:
- Useful for hard-to-reach populations (e.g., underground markets).
- Disadvantages:
- Risk of homogeneity bias (all samples may share traits).
4. Probability vs. Non-Probability Sampling: Comparison
| Feature | Probability Sampling | Non-Probability Sampling |
|---|---|---|
| Selection Method | Random, known chance for all members | Convenience, judgment, or criteria-based |
| Representativeness | High (generalizable) | Low (biased) |
| Cost | High (time/resources) | Low |
| Statistical Inference | Possible (margins of error calculable) | Not possible |
| Examples | Simple random, stratified, systematic, cluster | Convenience, quota, purposive, snowball |
| Use Cases | Government surveys, clinical trials | Pilot studies, exploratory research |
5. Real-World Applications in Nepal and Globally
A. eSewa and Khalti: Customer Satisfaction Surveys
- Method: Stratified random sampling (strata = transaction frequency: low, medium, high).
- How it works:
- Divide users into strata based on transaction volume.
- Randomly sample 500 users from each stratum.
- Analyze feedback to improve UX.
- Why it matters: Ensures feedback from all user segments, not just heavy users.
B. Daraz: Inventory Management
- Method: Systematic sampling for quality checks.
- How it works:
- Daraz receives 10,000 electronics from a supplier.
- Every 100th item is inspected for defects (sample size = 100).
- Why it matters: Reduces costs while maintaining product standards.
C. NTC: Traffic Congestion Study
- Method: Cluster sampling (clusters = districts).
- How it works:
- Nepal has 77 districts; NTC selects 10 districts randomly.
- Surveys all households in these districts about traffic issues.
- Why it matters: Provides a national estimate of congestion without surveying every household.
D. Ncell: Customer Churn Prediction
- Method: Purposive sampling (targeting high-value customers).
- How it works:
- Identify customers who switched to competitors in the last 6 months.
- Interview them to find patterns (e.g., poor customer service).
- Why it matters: Helps design retention strategies for at-risk users.
E. Banks: Loan Default Analysis
- Method: Stratified sampling (strata = loan amounts: small, medium, large).
- Example:
- Population: 10,000 loans.
- Strata:
- Small loans (<5 lakhs): 60% of population.
- Medium loans (5–20 lakhs): 30%.
- Large loans (>20 lakhs): 10%.
- Sample: 600 small, 300 medium, 100 large loans.
- Why it matters: Ensures default rates are analyzed proportionally across loan sizes.
6. Worked Examples
Example 1: Simple Random Sampling for a TU Survey
Scenario: TU wants to survey 200 students about hostel facilities. There are 10,000 students. Steps:
- Assign numbers 1–10,000 to all students.
- Use a random number generator to select 200 unique numbers.
- Survey the corresponding students.
Example 2: Stratified Sampling for a Khalti User Study
Scenario: Khalti wants to survey 500 users. Breakdown by age:
- 18–30: 60%
- 31–50: 30%
- 50+: 10%
Steps:
- Calculate sample sizes:
- 18–30: 500 × 0.6 = 300 users.
- 31–50: 500 × 0.3 = 150 users.
- 50+: 500 × 0.1 = 50 users.
- Randomly select 300 users from the 18–30 age group, etc.
Example 3: Systematic Sampling for a Daraz Quality Check
Scenario: Daraz receives 5,000 mobile phones. Sample 100 for quality checks. Steps:
- Calculate k = 5,000 / 100 = 50.
- Randomly select a start between 1–50 (e.g., 12).
- Inspect items 12, 62, 112, ..., 4962.
Example 4: Quota Sampling for a Pathao Driver Survey
Scenario: Pathao wants to survey 200 drivers. Quotas:
- Bike drivers: 60%
- Car drivers: 30%
- Auto drivers: 10%
Steps:
- Set quotas:
- Bike: 120 drivers.
- Car: 60 drivers.
- Auto: 20 drivers.
- Stop sampling once quotas are filled (e.g., survey drivers at pickup points until quotas are met).
7. Common Pitfalls and How to Avoid Them
- Undercoverage: Missing a subgroup (e.g., ignoring rural users in an urban survey).
- Fix: Use stratified sampling to include all groups.
- Non-response Bias: Surveyed members refuse to participate.
- Fix: Follow up with incentives (e.g., Khalti cashback for survey completion).
- Sampling Frame Errors: Using an incomplete or outdated list (e.g., voter lists with outdated addresses).
- Fix: Update the sampling frame regularly (e.g., NTC’s traffic surveys use GPS data).
- Interviewer Bias: Influencing responses (e.g., leading questions).
- Fix: Train interviewers to ask neutral questions.
8. Exam Tip: How to Score Full Marks
Definitions:
- Always define sampling methods clearly. For example:
"Stratified sampling is a probability technique where the population is divided into homogeneous subgroups (strata), and samples are randomly selected from each stratum proportionally."
- Always define sampling methods clearly. For example:
Comparison Tables:
- Examiners love tables. Use them to compare probability vs. non-probability methods (as shown above).
Real-World Applications:
- Tie examples to Nepalese contexts (e.g., NTC traffic studies, bank audits, Daraz quality checks). Always name the company/product and explain the method used.
Worked Examples:
- Show all steps with clear calculations. For systematic sampling, always state:
- Population size (N).
- Sample size (n).
- Sampling interval (k = N/n).
- Random start point.
- Show all steps with clear calculations. For systematic sampling, always state:
Avoid Common Mistakes:
- ❌ Saying "random sampling" is the only method (there are 8!).
- ❌ Confusing stratified and cluster sampling (strata = homogeneous groups; clusters = heterogeneous groups).
- ❌ Forgetting to mention probability of selection in probability sampling.
Diagrams:
- Draw timelines for systematic sampling, Venn diagrams for strata, or pie charts for quota distributions. Label every part.
9. Practice Questions (Exam-Style)
- Define stratified sampling and explain how a bank might use it to audit loan defaults.
- A company wants to sample 200 employees from 2,000. Describe two probability and two non-probability methods, stating their advantages and disadvantages.
- NEPSE wants to survey 500 investors. The population is divided into:
- Small investors (<1 crore): 70%
- Medium investors (1–10 crore): 20%
- Large investors (>10 crore): 10% How would you design a stratified sample? Show calculations.
- Explain why convenience sampling might lead to biased results in a Pathao driver satisfaction survey.
- Draw a timeline to represent systematic sampling for inspecting 1,000 Daraz orders with a sample size of 50.
10. Key Formulas to Remember
| Method | Formula/Calculation |
|---|---|
| Simple Random Sampling | k = N/n (sampling interval) |
| Stratified Sampling | n_h = (n/N) × N_h (sample per stratum) |
| Systematic Sampling | Start = random(1, k); select x, x+k, x+2k, ... |
| Sampling Error | SE = σ / √n (for SRS, where σ = population SD) |
11. Summary Checklist for Exams
Before submitting your answer, ensure you’ve covered: ✅ Definitions of all sampling methods. ✅ Advantages/disadvantages of each method. ✅ At least one real-world Nepalese example (e.g., NTC, Daraz, Khalti). ✅ Visuals (timelines, pie charts, Venn diagrams). ✅ Calculations (if numerical questions are asked). ✅ Comparison between probability and non-probability methods.
Based on the TU BSc CSIT syllabus for Statistics I (STA169), unit 2.
Discussion
Loading…