Basic StatisticsUnit 210 min read
Data Collection & Presentation: Methods, Tables & Visuals
Unit 2 of Basic Statistics: Learn how to design surveys, classify data, create frequency tables, and visualize data with bar charts, histograms, and pie charts—with real-world examples from eSewa, Daraz, and NTC.
TAKEAWAYS
- Data collection methods include primary (surveys, experiments) and secondary (published reports, government records) sources, each with trade-offs.
- Classification organizes raw data into meaningful groups (e.g., age brackets, income ranges) before analysis.
- Frequency tables summarize raw data into counts (frequency) and relative frequencies (percentages) for easier interpretation.
- Graphical representation (bar charts, histograms, pie charts) reveals patterns like trends, peaks, and distributions at a glance.
- Grouped data (intervals or ranges) is used when raw data is too large or continuous (e.g., monthly sales, exam scores).
- Real-world tools like eSewa’s transaction logs (frequency tables) and Daraz’s order queues (histograms) rely on these techniques.
1. Introduction to Data Collection
Data is the raw material of statistics. Before analysis, it must be collected systematically. There are two broad categories:
1.1 Primary vs. Secondary Data
| Aspect | Primary Data | Secondary Data |
|---|---|---|
| Source | Collected firsthand by the researcher. | Obtained from existing sources. |
| Reliability | High (direct observation). | Depends on source (may be outdated). |
| Cost | High (time, effort, resources). | Low (often free or cheap). |
| Flexibility | Customizable for specific needs. | Limited to pre-defined categories. |
| Example (Nepal) | A survey on Kathmandu traffic congestion conducted by NTC. | NTC’s published annual traffic accident reports. |
1.2 Methods of Primary Data Collection
Primary data is gathered through:
- Surveys/Questionnaires: Structured questions to collect opinions or facts. Example: A BIT student survey on preferred programming languages (Python, Java, C++).
- Observation: Directly recording behaviors (e.g., NTC monitoring traffic flow).
- Experiments: Controlled tests (e.g., testing a new eSewa payment feature).
- Interviews: One-on-one or group discussions (e.g., Daraz customer feedback).
Mermaid Diagram: Survey Design Process
flowchart TD
A["Define Research Objective"] --> B["Design Questionnaire"]
B --> C["Pilot Test"]
C --> D["Collect Data"]
D --> E["Analyze & Interpret"]1.3 Methods of Secondary Data Collection
Secondary data comes from:
- Published reports (e.g., NEPSE’s stock market data).
- Government databases (e.g., Nepal Rastra Bank’s inflation reports).
- Company records (e.g., Khalti’s transaction logs).
- Internet sources (e.g., World Bank data on Nepal’s GDP).
Advantages:
- Saves time and cost.
- Provides historical context (e.g., comparing Ncell’s revenue over 5 years).
Disadvantages:
- May lack detail or relevance.
- Risk of bias (e.g., NEPSE reports may favor certain stocks).
2. Classification of Data
Raw data is disorganized and overwhelming. Classification groups it into meaningful categories.
2.1 Types of Classification
| Type | Description | Example |
|---|---|---|
| Qualitative | Categorical (non-numeric). | Colors of Daraz packages (red, blue, white). |
| Quantitative | Numeric (discrete or continuous). | Number of Pathao rides per day (120, 150). |
| Temporal | Organized by time (daily, monthly). | NTC’s traffic data by hour of the day. |
| Spatial | Organized by location. | Sales of eSewa in Kathmandu vs. Pokhara. |
2.2 Steps for Classification
- Identify variables: What data is being collected? (e.g., age, income, product preference).
- Define categories: Group similar data (e.g., age groups: 18-25, 26-35).
- Assign frequencies: Count how many fall into each category.
Example: Classifying 10 students’ computer brand preferences (DELL or HP).
| Brand | Frequency |
|---|---|
| DELL | 5 |
| HP | 5 |
3. Frequency Distribution Tables
Frequency tables organize raw data into counts and percentages.
3.1 Simple Frequency Table
For discrete data (countable, distinct values):
| Value (X) | Frequency (f) | Relative Frequency (%) |
|---|---|---|
| 1 | 2 | 20% |
| 2 | 3 | 30% |
| 3 | 5 | 50% |
Mermaid Diagram: Frequency Table Construction
flowchart TD
A["Raw Data"] --> B["Count Frequencies"]
B --> C["Calculate Relative Frequencies"]
C --> D["Create Table"]3.2 Grouped Frequency Table
For continuous data (e.g., heights, exam scores), we use intervals (classes):
| Class Interval | Frequency (f) | Midpoint | Relative Frequency (%) |
|---|---|---|---|
| 10-20 | 5 | 15 | 25% |
| 20-30 | 8 | 25 | 40% |
| 30-40 | 7 | 35 | 35% |
Example: Grouped data for experience (X) vs. performance (Y) of computer operators (from past exam).
Assume the following raw data for Experience (years):
1, 6, 1, 2, 1, 8, 4, 3, 1, 0, 5, 1, 2, 4, 3, 1, 0, 5, 1, 2
Step 1: Define intervals (e.g., 0-2, 2-4, 4-6, 6-8). Step 2: Count frequencies.
| Experience (X) | Frequency (f) |
|---|---|
| 0-2 | 8 |
| 2-4 | 6 |
| 4-6 | 3 |
| 6-8 | 3 |
4. Presentation of Data: Graphical Methods
Visuals make data interpretable at a glance. Common graphs:
4.1 Bar Chart
- Used for categorical data (e.g., brand preferences).
- Example: Computer brand preferences (DELL vs. HP).
| DELL | HP |
|------|----|
| 5 | 5 |
Visual: Two bars of equal height (5 units each) for DELL and HP.
4.2 Histogram
- Used for continuous grouped data (e.g., experience vs. performance).
- Example: Experience of computer operators (from grouped table above).
[0-2] [2-4] [4-6] [6-8]
|-----| |-----| |-----| |-----|
| | | | | | | |
Visual: Bars with heights proportional to frequencies (8, 6, 3, 3).
4.3 Pie Chart
- Shows proportions of a whole (e.g., market share).
- Example: Distribution of eSewa transactions by payment method (credit card, UPI, bank transfer).
[Credit Card: 30%]
[UPI: 40%]
[Bank Transfer: 30%]
Visual: A circle divided into 3 slices (30%, 40%, 30%).
4.4 Line Graph
- Used for trends over time (e.g., Ncell’s monthly revenue).
- Example: Monthly sales of Daraz (in Rs. lakhs).
Month: Jan | Feb | Mar | Apr | May
Sales: 50 | 60 | 70 | 80 | 90
Visual: A line connecting points (50, 60, 70, 80, 90) with upward trend.
5. Real-World Applications
In the Real World
eSewa:
- Idea: Frequency tables track transaction types (e.g., bill payments, peer-to-peer).
- How: eSewa’s backend uses grouped frequency tables to analyze peak transaction hours (e.g., 6-8 PM).
Daraz:
- Idea: Histograms visualize order volumes by day/week.
- How: Daraz’s logistics team uses histograms to predict delivery bottlenecks (e.g., high orders on weekends).
NTC:
- Idea: Line graphs monitor traffic accidents over time.
- How: NTC plots monthly accident data to identify high-risk periods (e.g., monsoon season).
Worked Example: Daraz Order Queue
Scenario: Daraz receives 100 orders in a day. The number of orders per hour is:
12, 15, 8, 20, 10, 18, 14, 22, 9, 16
Task: Present this data in a frequency table and histogram.
Step 1: Create a frequency table.
| Orders per Hour (X) | Frequency (f) |
|---|---|
| 8-12 | 3 (12, 10, 8) |
| 12-16 | 4 (12, 15, 14, 16) |
| 16-20 | 3 (18, 20, 16) |
| 20-24 | 2 (20, 22) |
Step 2: Plot a histogram.
[8-12] [12-16] [16-20] [20-24]
|-----| |-----| |-----| |-----|
| | | | | | | |
Visual: Bars with heights 3, 4, 3, 2.
Interpretation:
- Peak hours: 12-16 orders/hour (highest frequency).
- Lowest activity: 20-24 orders/hour.
6. Exam Tip
- Focus on:
- Defining primary vs. secondary data clearly.
- Constructing frequency tables (both simple and grouped).
- Drawing bar charts, histograms, and pie charts accurately (label axes, use proper scales).
- Understanding real-world applications (e.g., NTC traffic data, Daraz sales trends).
- Common mistakes:
- Forgetting to label axes in graphs.
- Incorrectly grouping continuous data (e.g., overlapping intervals).
- Misinterpreting relative frequencies (e.g., confusing percentages with counts).
- Past exam pattern:
- 30%: Definitions (primary/secondary data, frequency tables).
- 40%: Practical (construct tables/graphs from raw data).
- 30%: Application (link to real-world scenarios like NEPSE or NTC).
Final Note: Always show your work—examiners reward clear, step-by-step solutions. For graphs, sketch neatly and label axes properly. Use real-world examples (e.g., eSewa, Daraz) to explain your reasoning.
Based on the TU BIT syllabus for Basic Statistics (STA154), unit 2.
Discussion
Loading…