Basic StatisticsUnit 913 min read
Secondary Data & Statistical Applications: Sources, Uses & Analysis
Unit 9 of Basic Statistics explores secondary data—its definition, sources (especially in Nepal), advantages/disadvantages, and practical applications in business, government, and research. Learn how to analyze pre-collected data, interpret statistical reports, and apply findings to real-world problems like NEPSE stock
TAKEAWAYS:
- Understand secondary data (pre-existing data) vs. primary data, including its sources (government, NGOs, private sectors) with a focus on Nepal’s context (NPC, CBS, NEPSE, etc.).
- Master advantages (cost-effective, time-saving) and limitations (outdated, bias) of secondary data, with comparisons to primary data.
- Learn how to analyze secondary data using statistical tools (trends, correlations) and interpret reports from organizations like NTC, NEPSE, or World Bank.
- Apply secondary data to real-world problems: e.g., predicting Daraz sales trends using NPC consumer reports or analyzing Ncell’s network performance via telecom regulatory data.
- Differentiate between absolute (raw numbers) and relative (percentages, ratios) statistical measures and when to use each.
- Solve exam-style problems: fitting regression lines to secondary datasets (e.g., blood pressure vs. age) or calculating correlation coefficients from pre-collected data (e.g., student preferences for DELL vs. HP).
What is Secondary Data?
Secondary data refers to pre-existing data collected by someone else for a different purpose but reused for current analysis. Unlike primary data (collected firsthand by researchers), secondary data is already available—saving time and money.
Key Characteristics
mindmap
root((Secondary Data))
Definition
Sources
Advantages
Limitations
Applications
Analysis TechniquesSources of Secondary Data (With Nepal Focus)
Secondary data comes from diverse sources. Below is a classified table with Nepal-specific examples:
| Source Type | Examples (Nepal) | Example Use Case |
|---|---|---|
| Government | Central Bureau of Statistics (CBS), NPC | Analyzing GDP growth trends for economic reports. |
| NGOs/International | UNICEF, World Bank, ADB | Studying child malnutrition rates in Nepal. |
| Private Sector | NEPSE (stock market), NTC (telecom), banks | Predicting stock prices or network traffic. |
| Media & Surveys | Kantipur, Republica, Nepal Rastra Bank reports | Assessing public opinion on political policies. |
| Academic/Research | Tribhuvan University (TU) publications | Reviewing past BIT student performance data. |
Advantages and Limitations of Secondary Data
Advantages
pie title Advantages of Secondary Data "Cost-effective" : 35 "Time-saving" : 30 "Wide coverage" : 20 "Reusable" : 15
- Cost-effective: No need for expensive data collection (e.g., NTC’s network data is free for researchers).
- Time-saving: Ready-to-use datasets (e.g., CBS’s census data).
- Wide coverage: Large-scale data (e.g., NPC’s economic reports).
- Reusable: Can be analyzed for multiple purposes (e.g., NEPSE data for stock predictions).
Limitations
- Outdated: Data may not reflect current trends (e.g., a 2011 census for 2024 analysis).
- Bias: Collected for a different purpose (e.g., NEPSE reports may favor certain stocks).
- Incomplete: Missing variables (e.g., CBS data may lack regional breakdowns).
- Quality issues: Errors or inconsistencies (e.g., NTC’s network data may have gaps).
Comparison Table: Primary vs. Secondary Data
| Feature | Primary Data | Secondary Data |
|---|---|---|
| Collection | Collected by researcher | Already exists |
| Cost | High (surveys, experiments) | Low (free/cheap) |
| Relevance | Tailored to research needs | May not fit perfectly |
| Time | Time-consuming | Instant access |
| Example (Nepal) | Surveying BIT students’ laptop preferences | Using CBS’s education reports |
How to Analyze Secondary Data?
Secondary data is analyzed using statistical tools like:
- Descriptive Statistics: Mean, median, mode (e.g., average NEPSE stock price).
- Trend Analysis: Time-series data (e.g., NTC’s internet usage growth).
- Correlation & Regression: Relationships between variables (e.g., age vs. blood pressure).
- Index Numbers: Comparing trends (e.g., inflation rate over years).
Worked Example 1: Fitting a Regression Line to Secondary Data
Problem: The following table shows age (X) and blood pressure (Y) data from a hospital’s secondary records. Fit a regression line to estimate blood pressure for a 40-year-old.
| Age (X) | 5 | 6 | 4 | 2 | 7 | 2 | 3 | 6 | 3 | 4 |
|---|---|---|---|---|---|---|---|---|---|---|
| BP (Y) | 147 | 125 | 160 | 118 | 149 | 128 | 110 | 150 | 130 | 140 |
Solution: We use the least squares method to find the regression equation: where:
Step 1: Calculate Sums
Step 2: Compute Slope (b)
Step 3: Compute Intercept (a)
Final Regression Equation:
Prediction for X = 40: Correction: Regression should only be used within the range of data (here, X = 2 to 7). For X = 40, we’d need more data or a different model.
Visualization:
Real-World Tie-In: Nepal’s NTC uses secondary data (past network traffic) to predict future demand and plan infrastructure upgrades. Similarly, NEPSE analysts use historical stock prices to forecast trends.
Worked Example 2: Pearson’s Rank Correlation (Secondary Data)
Problem: The following table shows 10 students’ preferences for DELL and HP computers (data from a secondary survey). Calculate Pearson’s rank correlation coefficient (r).
| Student | DELL (X) | HP (Y) |
|---|---|---|
| 1 | 5 | 10 |
| 2 | 2 | 5 |
| 3 | 9 | 8 |
| 4 | 8 | 1 |
| 5 | 1 | 10 |
| 6 | 3 | 4 |
| 7 | 4 | 6 |
| 8 | 6 | 2 |
| 9 | 7 | 9 |
| 10 | 10 | 3 |
Solution: Pearson’s rank correlation formula: where .
Step 1: Rank X and Y
| Student | X Rank | Y Rank | d = X-Y | d² |
|---|---|---|---|---|
| 1 | 8 | 10 | -2 | 4 |
| 2 | 2 | 5 | -3 | 9 |
| 3 | 9 | 9 | 0 | 0 |
| 4 | 7 | 1 | 6 | 36 |
| 5 | 1 | 10 | -9 | 81 |
| 6 | 3 | 4 | -1 | 1 |
| 7 | 4 | 7 | -3 | 9 |
| 8 | 6 | 2 | 4 | 16 |
| 9 | 5 | 8 | -3 | 9 |
| 10 | 10 | 3 | 7 | 49 |
| Sum | 214 |
Step 2: Compute r
Interpretation:
- indicates a weak negative correlation between DELL and HP preferences.
- This suggests students who prefer DELL slightly dislike HP, but the relationship is not strong.
Visualization:
Real-World Tie-In: Daraz (Nepal’s Amazon) uses secondary data on customer preferences (e.g., DELL vs. HP sales trends) to decide inventory levels. A negative correlation might signal that promoting one brand could boost the other’s sales (complementary demand).
Absolute vs. Relative Statistical Measures
| Measure Type | Definition | Example (Nepal) | When to Use |
|---|---|---|---|
| Absolute | Raw, unadjusted numbers | NEPSE’s total stock volume: 500 million | Comparing exact quantities. |
| Relative | Adjusted (percentages, ratios) | Inflation rate: 5% (vs. last year) | Comparing proportions or trends. |
Worked Example:
- Absolute: Nepal’s total internet users = 25 million (NTC data).
- Relative: Penetration rate = (25M / 30M population) × 100 = 83.3%.
Why It Matters:
- NTC reports absolute numbers (e.g., 10M 4G users) but analyzes relative growth (e.g., 20% increase YoY).
- NEPSE uses relative measures (e.g., stock price change %) to compare performance.
## In the Real World
NEPSE (Nepal Stock Exchange)
- Idea Used: Secondary data analysis (historical stock prices, trading volumes).
- How: Analysts use past data to predict trends (e.g., regression models for stock prices). For example, if NEPSE’s NIBL Bank stock has historically risen with GDP growth, investors use CBS’s GDP data (secondary) to make decisions.
NTC (Nepal Telecom Authority)
- Idea Used: Trend analysis and correlation.
- How: NTC tracks network traffic data (secondary) to predict demand. For instance, if internet usage correlates with smartphone sales (data from NPC), they plan infrastructure upgrades accordingly.
Daraz (Nepal’s E-Commerce Giant)
- Idea Used: Regression analysis on secondary sales data.
- How: Daraz uses past sales trends (e.g., laptop sales vs. student enrollment data from TU) to forecast demand. For example:
- If BIT student enrollment (secondary data from TU) increases by 10%, Daraz expects a 5% rise in laptop sales (derived from regression analysis).
Ncell (Telecom Provider)
- Idea Used: Index numbers (relative measures).
- How: Ncell compares current 4G coverage (e.g., 80%) to last year’s (70%) to report a 14% improvement—a relative measure used in marketing.
World Bank Reports (Used by Nepalese Policymakers)
- Idea Used: Secondary data for policy decisions.
- How: Nepal’s Ministry of Finance uses World Bank’s poverty data (secondary) to design subsidies. For example, if 20% of Nepal’s population lives below $2/day (World Bank data), they allocate budgets for food security programs.
## Exam Tip
Understand the Difference:
- Secondary data = pre-existing (e.g., CBS, NEPSE).
- Primary data = newly collected (e.g., your own survey).
- Exam Question: "Define secondary data and list 3 Nepalese sources." → Answer with CBS, NEPSE, NTC.
Regression and Correlation Are Key:
- Always check the range when using regression (e.g., don’t predict blood pressure for age 100 if data is only for 2–70).
- Pearson’s r is for linear relationships; if data is ranked, use Spearman’s rank correlation.
Absolute vs. Relative:
- Absolute = raw numbers (e.g., "1000 students").
- Relative = percentages/ratios (e.g., "50% prefer DELL").
- Exam Tip: If asked to "compare," use relative measures (e.g., "HP’s market share grew from 30% to 40%").
Real-World Applications:
- NEPSE/NTC/Daraz use secondary data for predictions.
- Government reports (CBS, NPC) are primary sources for secondary analysis.
- Always tie examples to Nepal (e.g., "NTC’s network data can be used to predict...").
Common Mistakes to Avoid:
- Extrapolating regression beyond the data range.
- Ignoring units (e.g., mixing years and months in time-series data).
- Assuming causation from correlation (e.g., "More ice cream sales → more drowning" doesn’t mean ice cream causes drowning).
Final Note: Secondary data is everywhere in Nepal—from NEPSE’s stock trends to NTC’s network reports. Mastering its analysis will help you ace exams and solve real-world problems like a pro! 🚀
Based on the TU BIT syllabus for Basic Statistics (STA154), unit 9.
Discussion
Loading…