STA169 Statistics I

Statistics IUnit 88 min read

Data Collection: Primary vs. Secondary Data – Sources, Uses & Analysis

Unit 8 of Statistics I explains the difference between primary and secondary data, their sources, advantages/disadvantages, and how they are used in real-world studies (e.g., eSewa transactions, NTC traffic analysis). Includes worked examples, comparison tables, and exam-focused tips.

TAKEAWAYS:

  • Primary data is collected firsthand for a specific purpose (e.g., surveys, experiments), while secondary data is pre-existing (e.g., government records, company reports).
  • Sources of primary data: Direct observation, surveys, experiments, interviews.
  • Sources of secondary data: Government publications, academic journals, business reports, online databases (e.g., NEPSE stock data, NTC traffic reports).
  • Advantages of primary data: High accuracy, relevance, and control over collection methods.
  • Advantages of secondary data: Cost-effective, time-saving, and broad coverage.
  • Disadvantages: Primary data is expensive/time-consuming; secondary data may lack specificity or be outdated.

1. Definitions: Primary vs. Secondary Data

Primary Data

  • Definition: Data collected directly by the researcher for a specific study or purpose.
    • Example: A survey conducted by Nepal Electricity Authority (NEA) to measure household electricity consumption in Kathmandu.
    • Key Feature: Original, raw, and tailored to the research question.

Secondary Data

  • Definition: Data already collected by someone else for a different purpose, repurposed for analysis.
    • Example: Using NTC’s annual traffic reports to analyze road congestion patterns in Pokhara.
    • Key Feature: Pre-existing, often publicly available (e.g., World Bank datasets, company annual reports).

2. Sources of Primary and Secondary Data

Primary Data Sources

mindmap
  root((Primary Data Sources))
    Direct Observation
    Surveys
      Questionnaires
      Interviews
    Experiments
      Controlled Tests
      Field Trials
    Case Studies
    Focus Groups

Secondary Data Sources

mindmap
  root((Secondary Data Sources))
    Government Publications
      NTC Reports
      NEPSE Stock Data
      CBS (Central Bureau of Statistics)
    Academic Journals
    Business Reports
      Company Annual Reports (e.g., Ncell, Daraz)
      Market Research Firms
    Online Databases
      World Bank Open Data
      Google Trends
      eSewa Transaction Logs
    Internal Records
      Bank Transaction Histories
      Hospital Patient Records

3. Comparison Table: Primary vs. Secondary Data

Feature Primary Data Secondary Data
Collection Collected by researcher for specific use Collected by others for different purposes
Cost High (time, labor, resources) Low (often free or inexpensive)
Relevance Highly tailored to research question May not fully match research needs
Accuracy High (controlled collection) Depends on source reliability
Time Time-consuming Quick to access
Examples (Nepal) Survey on student satisfaction at TU NTC’s traffic accident data (2023)

4. Worked Example: Choosing Data for a Study

Scenario: A researcher wants to analyze electricity consumption patterns in Nepal to predict demand for NEA’s new solar projects.

Option 1: Primary Data (Survey)

  • Method: Conduct a household survey in 500 homes across Kathmandu, Pokhara, and Biratnagar.
  • Data Collected:
    • Daily electricity usage (kWh)
    • Peak usage times
    • Preferences for solar vs. grid power
  • Advantages:
    • Customized for solar demand prediction.
    • Can include behavioral insights (e.g., "Do households reduce usage during high tariff hours?").
  • Disadvantages:
    • Expensive (~Rs. 500,000 for fieldwork).
    • Time-consuming (3–6 months to collect).

Option 2: Secondary Data (NEA Reports)

  • Source: NEA’s annual electricity consumption reports (2018–2023).
  • Data Available:
    • Monthly consumption per district (kWh).
    • Peak demand hours.
    • Solar panel installations (2020–2023).
  • Advantages:
    • Free and readily available.
    • Covers entire Nepal (not just 500 households).
  • Disadvantages:
    • No behavioral data (e.g., why households switch to solar).
    • May lack granularity (e.g., no breakdown by income group).

Decision:

  • Use both for a robust analysis:
    • Primary data for behavioral insights (survey).
    • Secondary data for national trends (NEA reports).

5. Real-World Applications

Example 1: eSewa Transaction Analysis

  • Primary Data: eSewa collects real-time transaction logs (e.g., mobile recharge, bill payments) to analyze user behavior.
    • Use: Detect fraud patterns or optimize payment processing.
  • Secondary Data: eSewa uses Nepal Rastra Bank’s inflation reports to adjust transaction fees.
    • Use: Ensure fee structures remain competitive.

Example 2: Daraz Logistics Optimization

  • Primary Data: Daraz tracks delivery times via GPS in its warehouses (e.g., Kathmandu, Lalitpur).
    • Use: Identify bottlenecks (e.g., traffic delays in Thapathali).
  • Secondary Data: Uses NTC’s road construction schedules to predict delays.
    • Use: Adjust delivery timelines dynamically.

Example 3: NTC Traffic Management

  • Primary Data: NTC installs traffic cameras at busy intersections (e.g., Kalanki, New Baneshwor) to count vehicles.
    • Use: Design signal timings to reduce congestion.
  • Secondary Data: Uses police accident reports to prioritize road safety measures.
    • Use: Install speed bumps where accidents are frequent.

6. When to Use Each Type?


7. Exam-Style Problem: Identifying Data Types

Question: A student wants to study the impact of study hours on exam scores in TU’s B.Sc. CSIT program.

  • Primary Data Sources:
    • Conduct a survey asking students: "How many hours do you study per week?"
    • Collect exam score records from the TU exam office.
  • Secondary Data Sources:
    • Use past TU result analyses published in journals.
    • Refer to NBE’s (National Board of Examination) historical pass rates.

Answer:

  • Primary: Survey + exam records (collected firsthand).
  • Secondary: TU/NBE reports (pre-existing data).

8. Common Pitfalls

  1. Assuming Secondary Data is Always Accurate:

    • Example: Using old NTC traffic data (2015) to plan 2024 road expansions → outdated.
    • Fix: Cross-validate with primary sources (e.g., recent surveys).
  2. Overlooking Bias in Primary Data:

    • Example: Surveying only TU students for a national study on internet usage → biased sample.
    • Fix: Use random sampling across regions.
  3. Ignoring Data Limitations:

    • Example: Using Khalti transaction logs to estimate national poverty → misses cash-based economies.
    • Fix: Combine with CBS household surveys.

9. Visual: Data Collection Methods in Nepal

Source: Adapted from TU’s 2023 Statistics Research Trends Report.


10. Exam Tip

  1. Define Clearly:

    • Start answers with:

      "Primary data is original data collected for a specific research purpose, while secondary data is pre-existing data repurposed for analysis."

  2. Use Real Examples:

    • Link to Nepali contexts (e.g., NTC, NEPSE, eSewa) to score marks for applicability.
  3. Compare Advantages/Disadvantages:

    • Exams often ask for a table or bullet points comparing the two. Always include:
      • Cost, time, relevance, and reliability.
  4. Watch for Tricks:

    • Questions may ask: "Which data type would you use for X study?"
    • Answer: Justify why primary or secondary is better (e.g., "Primary is needed for behavioral insights").
  5. Practice Calculations:

    • Some questions mix data types with percentiles or moments (e.g., "Compute Q1 using primary survey data").
    • Always show steps for partial credit.

11. Quick Revision Checklist

  • Can you list 3 primary and 3 secondary data sources for a Nepali study?
  • Do you know when to prefer primary over secondary (and vice versa)?
  • Can you critique a study’s data collection method (e.g., "Why is this survey biased?").
  • Are you ready to apply this to exam questions (e.g., "How would you collect data for a Daraz delivery study?").

Final Note:

*"In exams, always justify your choice of data type. For example, if asked about studying student stress levels, explain why a survey (primary) is better than using old TU alumni reports (secondary)—because stress levels change yearly!"*

Based on the TU BSc CSIT syllabus for Statistics I (STA169), unit 8.

Discussion

Loading…