STA154 Basic Statistics

Basic StatisticsUnit 913 min read

Secondary Data & Statistical Applications: Sources, Uses & Analysis

Unit 9 of Basic Statistics explores secondary data—its definition, sources (especially in Nepal), advantages/disadvantages, and practical applications in business, government, and research. Learn how to analyze pre-collected data, interpret statistical reports, and apply findings to real-world problems like NEPSE stock

TAKEAWAYS:

  • Understand secondary data (pre-existing data) vs. primary data, including its sources (government, NGOs, private sectors) with a focus on Nepal’s context (NPC, CBS, NEPSE, etc.).
  • Master advantages (cost-effective, time-saving) and limitations (outdated, bias) of secondary data, with comparisons to primary data.
  • Learn how to analyze secondary data using statistical tools (trends, correlations) and interpret reports from organizations like NTC, NEPSE, or World Bank.
  • Apply secondary data to real-world problems: e.g., predicting Daraz sales trends using NPC consumer reports or analyzing Ncell’s network performance via telecom regulatory data.
  • Differentiate between absolute (raw numbers) and relative (percentages, ratios) statistical measures and when to use each.
  • Solve exam-style problems: fitting regression lines to secondary datasets (e.g., blood pressure vs. age) or calculating correlation coefficients from pre-collected data (e.g., student preferences for DELL vs. HP).


What is Secondary Data?

Secondary data refers to pre-existing data collected by someone else for a different purpose but reused for current analysis. Unlike primary data (collected firsthand by researchers), secondary data is already available—saving time and money.

Key Characteristics

mindmap
  root((Secondary Data))
    Definition
    Sources
    Advantages
    Limitations
    Applications
    Analysis Techniques

Sources of Secondary Data (With Nepal Focus)

Secondary data comes from diverse sources. Below is a classified table with Nepal-specific examples:

UABCBS Nepal reports, Nepal Budget documentsWorld Bank data, UNICEF Nepal statisticsGovernment sources, International sources, Academic sources
Overlap between national (CBS) and international (World Bank) secondary data sources
Source Type Examples (Nepal) Example Use Case
Government Central Bureau of Statistics (CBS), NPC Analyzing GDP growth trends for economic reports.
NGOs/International UNICEF, World Bank, ADB Studying child malnutrition rates in Nepal.
Private Sector NEPSE (stock market), NTC (telecom), banks Predicting stock prices or network traffic.
Media & Surveys Kantipur, Republica, Nepal Rastra Bank reports Assessing public opinion on political policies.
Academic/Research Tribhuvan University (TU) publications Reviewing past BIT student performance data.

Advantages and Limitations of Secondary Data

Advantages

pie
  title Advantages of Secondary Data
  "Cost-effective" : 35
  "Time-saving" : 30
  "Wide coverage" : 20
  "Reusable" : 15
  • Cost-effective: No need for expensive data collection (e.g., NTC’s network data is free for researchers).
  • Time-saving: Ready-to-use datasets (e.g., CBS’s census data).
  • Wide coverage: Large-scale data (e.g., NPC’s economic reports).
  • Reusable: Can be analyzed for multiple purposes (e.g., NEPSE data for stock predictions).

Limitations

Outdated (30%) (30%)Bias (25%) (25%)Incomplete (20%) (20%)Mismatched purpose (15%) (15%)Quality issues (10%) (10%)
Limitations of secondary data with percentage breakdown (Nepal context)
  • Outdated: Data may not reflect current trends (e.g., a 2011 census for 2024 analysis).
  • Bias: Collected for a different purpose (e.g., NEPSE reports may favor certain stocks).
  • Incomplete: Missing variables (e.g., CBS data may lack regional breakdowns).
  • Quality issues: Errors or inconsistencies (e.g., NTC’s network data may have gaps).

Comparison Table: Primary vs. Secondary Data

Feature Primary Data Secondary Data
Collection Collected by researcher Already exists
Cost High (surveys, experiments) Low (free/cheap)
Relevance Tailored to research needs May not fit perfectly
Time Time-consuming Instant access
Example (Nepal) Surveying BIT students’ laptop preferences Using CBS’s education reports

How to Analyze Secondary Data?

Secondary data is analyzed using statistical tools like:

  1. Descriptive Statistics: Mean, median, mode (e.g., average NEPSE stock price).
  2. Trend Analysis: Time-series data (e.g., NTC’s internet usage growth).
  3. Correlation & Regression: Relationships between variables (e.g., age vs. blood pressure).
  4. Index Numbers: Comparing trends (e.g., inflation rate over years).

Worked Example 1: Fitting a Regression Line to Secondary Data

Problem: The following table shows age (X) and blood pressure (Y) data from a hospital’s secondary records. Fit a regression line to estimate blood pressure for a 40-year-old.

12345622.533.544.55yRegression line: y = 0.5x + 2Nepal’s GDP growth vs. education spending (2015-2019)
Regression analysis using Nepal’s secondary economic data (hypothetical)
Age (X) 5 6 4 2 7 2 3 6 3 4
BP (Y) 147 125 160 118 149 128 110 150 130 140

Solution: We use the least squares method to find the regression equation: where:

Step 1: Calculate Sums

Step 2: Compute Slope (b)

Step 3: Compute Intercept (a)

Final Regression Equation:

Prediction for X = 40: Correction: Regression should only be used within the range of data (here, X = 2 to 7). For X = 40, we’d need more data or a different model.

Visualization:

Real-World Tie-In: Nepal’s NTC uses secondary data (past network traffic) to predict future demand and plan infrastructure upgrades. Similarly, NEPSE analysts use historical stock prices to forecast trends.


Worked Example 2: Pearson’s Rank Correlation (Secondary Data)

Problem: The following table shows 10 students’ preferences for DELL and HP computers (data from a secondary survey). Calculate Pearson’s rank correlation coefficient (r).

Student DELL (X) HP (Y)
1 5 10
2 2 5
3 9 8
4 8 1
5 1 10
6 3 4
7 4 6
8 6 2
9 7 9
10 10 3

Solution: Pearson’s rank correlation formula: where .

Step 1: Rank X and Y

Student X Rank Y Rank d = X-Y d²
1 8 10 -2 4
2 2 5 -3 9
3 9 9 0 0
4 7 1 6 36
5 1 10 -9 81
6 3 4 -1 1
7 4 7 -3 9
8 6 2 4 16
9 5 8 -3 9
10 10 3 7 49
Sum 214

Step 2: Compute r

Interpretation:

  • indicates a weak negative correlation between DELL and HP preferences.
  • This suggests students who prefer DELL slightly dislike HP, but the relationship is not strong.

Visualization:

Real-World Tie-In: Daraz (Nepal’s Amazon) uses secondary data on customer preferences (e.g., DELL vs. HP sales trends) to decide inventory levels. A negative correlation might signal that promoting one brand could boost the other’s sales (complementary demand).


Absolute vs. Relative Statistical Measures

Measure Type Definition Example (Nepal) When to Use
Absolute Raw, unadjusted numbers NEPSE’s total stock volume: 500 million Comparing exact quantities.
Relative Adjusted (percentages, ratios) Inflation rate: 5% (vs. last year) Comparing proportions or trends.

Worked Example:

  • Absolute: Nepal’s total internet users = 25 million (NTC data).
  • Relative: Penetration rate = (25M / 30M population) × 100 = 83.3%.

Why It Matters:

  • NTC reports absolute numbers (e.g., 10M 4G users) but analyzes relative growth (e.g., 20% increase YoY).
  • NEPSE uses relative measures (e.g., stock price change %) to compare performance.

## In the Real World

  1. NEPSE (Nepal Stock Exchange)

    • Idea Used: Secondary data analysis (historical stock prices, trading volumes).
    • How: Analysts use past data to predict trends (e.g., regression models for stock prices). For example, if NEPSE’s NIBL Bank stock has historically risen with GDP growth, investors use CBS’s GDP data (secondary) to make decisions.
  2. NTC (Nepal Telecom Authority)

    • Idea Used: Trend analysis and correlation.
    • How: NTC tracks network traffic data (secondary) to predict demand. For instance, if internet usage correlates with smartphone sales (data from NPC), they plan infrastructure upgrades accordingly.
  3. Daraz (Nepal’s E-Commerce Giant)

    • Idea Used: Regression analysis on secondary sales data.
    • How: Daraz uses past sales trends (e.g., laptop sales vs. student enrollment data from TU) to forecast demand. For example:
      • If BIT student enrollment (secondary data from TU) increases by 10%, Daraz expects a 5% rise in laptop sales (derived from regression analysis).
  4. Ncell (Telecom Provider)

    • Idea Used: Index numbers (relative measures).
    • How: Ncell compares current 4G coverage (e.g., 80%) to last year’s (70%) to report a 14% improvement—a relative measure used in marketing.
  5. World Bank Reports (Used by Nepalese Policymakers)

    • Idea Used: Secondary data for policy decisions.
    • How: Nepal’s Ministry of Finance uses World Bank’s poverty data (secondary) to design subsidies. For example, if 20% of Nepal’s population lives below $2/day (World Bank data), they allocate budgets for food security programs.

## Exam Tip

  1. Understand the Difference:

    • Secondary data = pre-existing (e.g., CBS, NEPSE).
    • Primary data = newly collected (e.g., your own survey).
    • Exam Question: "Define secondary data and list 3 Nepalese sources." → Answer with CBS, NEPSE, NTC.
  2. Regression and Correlation Are Key:

    • Always check the range when using regression (e.g., don’t predict blood pressure for age 100 if data is only for 2–70).
    • Pearson’s r is for linear relationships; if data is ranked, use Spearman’s rank correlation.
  3. Absolute vs. Relative:

    • Absolute = raw numbers (e.g., "1000 students").
    • Relative = percentages/ratios (e.g., "50% prefer DELL").
    • Exam Tip: If asked to "compare," use relative measures (e.g., "HP’s market share grew from 30% to 40%").
  4. Real-World Applications:

    • NEPSE/NTC/Daraz use secondary data for predictions.
    • Government reports (CBS, NPC) are primary sources for secondary analysis.
    • Always tie examples to Nepal (e.g., "NTC’s network data can be used to predict...").
  5. Common Mistakes to Avoid:

    • Extrapolating regression beyond the data range.
    • Ignoring units (e.g., mixing years and months in time-series data).
    • Assuming causation from correlation (e.g., "More ice cream sales → more drowning" doesn’t mean ice cream causes drowning).

Final Note: Secondary data is everywhere in Nepal—from NEPSE’s stock trends to NTC’s network reports. Mastering its analysis will help you ace exams and solve real-world problems like a pro! 🚀

Based on the TU BIT syllabus for Basic Statistics (STA154), unit 9.

Discussion

Loading…