StatisticsUnit 115 min read

Statistics Basics: Data Types, Sources & Classification

Unit 1 of Statistics introduces core concepts like data types (primary/secondary), sources, classification methods (qualitative/quantitative), and their applications in tourism and business decision-making.

TAKEAWAYS:

  • Data classification is the foundation of statistical analysis, dividing raw information into meaningful categories (qualitative/quantitative, discrete/continuous).
  • Primary vs. secondary data determines the reliability and cost of your analysis—primary data is original but expensive; secondary data is cheaper but may lack relevance.
  • Tourism applications use data classification to analyze visitor demographics, seasonal trends, and revenue patterns (e.g., NTC classifying passenger traffic by season).
  • Business decisions rely on classified data for inventory management (Daraz’s product categorization), marketing segmentation (Pathao’s rider demographics), and financial forecasting (bank loan approvals).
  • Visual tools like frequency distributions and ogives transform raw data into actionable insights (e.g., NEPSE’s stock trend analysis).
  • Exam focus: Expect questions on distinguishing data types, classifying real-world datasets (e.g., hotel occupancy rates), and interpreting visual representations (ogives, bar charts).

1. Introduction to Statistics: Why It Matters

Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data to make informed decisions. In Travel and Tourism Management, statistics helps:

  • Predict tourist arrivals (e.g., NTC’s annual passenger data).
  • Optimize pricing (e.g., Daraz’s dynamic discounts based on demand).
  • Assess customer satisfaction (e.g., Pathao’s rider feedback analysis).

Key Definitions

  • Population: The entire group being studied (e.g., all hotels in Kathmandu).
  • Sample: A representative subset (e.g., 20 hotels surveyed for cleanliness ratings).
  • Parameter: A fixed value describing the population (e.g., mean income of all Nepali tourists).
  • Statistic: A variable describing the sample (e.g., mean income of surveyed tourists).

2. Types of Data: Qualitative vs. Quantitative

Data is classified based on measurement scale and nature:

UABDiscrete (e.g., number of students in a class: 25, 26, 27, .Continuous (e.g., height of students in cm, temperature in °Quantitative Data
Quantitative Data: Discrete (countable) vs. Continuous (measurable)
UABDiscrete (e.g., number of tourists at Tribhuvan Airport: 50,Continuous (e.g., luggage weight in kg, temperature in °C), Quantitative Data
Quantitative Data Subtypes: Discrete vs. Continuous
UABNominal (e.g., hotel star ratings: 1★, 2★, 3★, 4★, 5★), eSewOrdinal (e.g., customer satisfaction: Poor < Fair < Good < EQualitative Data
Qualitative Data Subtypes: Nominal (no order) vs. Ordinal (ordered categories)

A. Qualitative (Categorical) Data

Descriptive data divided into categories (no numerical value).

  • Nominal: No order (e.g., hotel star ratings: 1★, 2★, 3★).
  • Ordinal: Ordered categories (e.g., customer satisfaction: Poor, Average, Good, Excellent).

Real-World Example:

  • eSewa classifies transactions by service type (electricity, water, telecom) — this is nominal data.
  • Pathao categorizes riders by vehicle type (bike, car, auto) — ordinal if ordered by capacity.

B. Quantitative Data

Numerical data that can be measured.

  • Discrete: Countable values (e.g., number of tourists arriving daily at Tribhuvan Airport).
  • Continuous: Measurable values (e.g., weight of luggage in kg, temperature in °C).

Worked Example 1: Classifying Tourism Data Classify the following data collected by NTC:

  1. Number of domestic flights per day.
  2. Passenger complaints categorized as "delay," "lost baggage," or "rude staff."
  3. Average waiting time at immigration (in minutes).
  4. Hotel ratings (1 to 5 stars).

Solution:

Data Type Subtype
Number of domestic flights Quantitative Discrete
Passenger complaints Qualitative Nominal
Average waiting time Quantitative Continuous
Hotel ratings (1-5 stars) Qualitative Ordinal

3. Sources of Data: Primary vs. Secondary

Data can be collected from two primary sources:

A. Primary Data

Collected firsthand for a specific purpose. Advantages:

  • Highly relevant to the study.
  • Up-to-date and accurate. Disadvantages:
  • Time-consuming and expensive.
  • Requires skilled personnel.

Methods of Collection: Real-World Example:

  • NTC conducts surveys to collect data on tourist satisfaction at airports (primary data).
  • Daraz uses customer feedback forms to gather opinions on product quality.

B. Secondary Data

Collected from existing sources for other purposes. Advantages:

  • Saves time and money.
  • Wider scope of data. Disadvantages:
  • May not fit the current study’s needs.
  • Risk of outdated or biased data.

Sources: Worked Example 2: Identifying Data Sources A tourism student wants to analyze backpacker trends in Nepal. Identify whether the following sources are primary or secondary:

  1. Interviews with 50 backpackers at Thamel.
  2. Data from the Nepal Tourism Board’s 2023 report.
  3. Observations of backpacker behavior at Pokhara Lake.
  4. A study on "Backpacking in Southeast Asia" published in 2020.

Solution:

Source Type Reasoning
Interviews with 50 backpackers Primary Collected directly for this specific study.
NTB’s 2023 report Secondary Existing data collected for general tourism statistics.
Observations at Pokhara Lake Primary Firsthand collection of behavioral data.
2020 study on backpacking trends Secondary Data collected for a different purpose (published research).

4. Data Classification: Organizing Raw Data

Raw data is unorganized and meaningless. Classification involves:

  1. Grouping data into categories.
  2. Arranging data in a meaningful order.
  3. Summarizing data for easier analysis.
Primary Data (35%)Secondary Data (65%)
Typical Data Sources in Nepali Studies (Example: 35% Primary, 65% Secondary)
00.751.52.25350-60260-70370-80180-901Frequency (Number of Days)
Frequency Distribution of Tourist Arrival Times at Pokhara Airport (Example)

A. Frequency Distribution

A table showing how often each value occurs. Example: Number of tourists visiting Pokhara per day for a week.

Number of Tourists Frequency (f) Relative Frequency (f/N)
50-60 2 2/7 ≈ 0.286
60-70 3 3/7 ≈ 0.429
70-80 1 1/7 ≈ 0.143
80-90 1 1/7 ≈ 0.143
Total (N) 7 1.000

Worked Example 3: Creating a Frequency Table The following data shows the number of days tourists stayed in Nepal (sample of 20 tourists): 5, 7, 3, 8, 6, 4, 9, 5, 7, 6, 8, 4, 5, 6, 7, 9, 10, 5, 6, 8

Solution:

  1. Determine classes: Use a range of 2-3 days (e.g., 3-5, 6-8, etc.).
  2. Tally frequencies:
Days Stayed Tally Frequency (f)
3-5
6-8
9-11
Total 20

B. Ogives (Cumulative Frequency Graphs)

Ogives help find median, quartiles, and percentiles visually.

Types:

  1. Less Than Ogive: Shows cumulative frequency below a class.
  2. More Than Ogive: Shows cumulative frequency above a class.

Worked Example 4: Drawing Ogives Using the height data from past exam questions:

Height (cm) Persons (f) Cumulative Frequency (≤) Cumulative Frequency (>)
62-63 2 2 20
63-64 6 8 18
64-65 14 22 12
65-66 16 38 6
66-67 8 46 2
67-68 3 49 1
68-69 1 50 0

Solution: Finding the Median:

  • Total frequency (N) = 50.
  • Median position = .
  • On the less than ogive, locate 25 on the y-axis and drop to the x-axis → Median height ≈ 64.5 cm.

5. Applications in Tourism and Business

A. Real-World Example 1: NTC’s Passenger Traffic Analysis

Problem: NTC wants to classify passenger traffic by season to optimize flight schedules. Data Classification:

  • Qualitative: Passenger type (domestic/international).
  • Quantitative: Number of passengers per month (discrete).
  • Primary Data: Monthly surveys at airports.
  • Secondary Data: Historical flight records.

Visualization: Insight: Peak in June (monsoon season) → NTC can increase flights during this period.

B. Real-World Example 2: Daraz’s Inventory Management

Problem: Daraz needs to classify products to manage stock efficiently. Data Classification:

  • Qualitative: Product category (electronics, fashion, groceries).
  • Quantitative: Sales volume per product (discrete).
  • Primary Data: Daily sales tracking.
  • Secondary Data: Market trend reports.

Frequency Table:

Product Category Sales Volume (units/day) Frequency
Electronics 500-1000 15
Fashion 300-800 20
Groceries 200-600 10

Insight: Fashion has the highest variability → Daraz can focus on dynamic pricing for this category.

**C. Real-World Example 3: Bank Loan Approvals (Nepal)

Problem: A bank classifies loan applicants to assess risk. Data Classification:

  • Qualitative: Employment type (salaried, self-employed, business).
  • Quantitative: Income (continuous), loan amount (discrete).
  • Primary Data: Applicant interviews.
  • Secondary Data: Credit bureau reports.

Decision Tree:

flowchart TD
    A["Loan Applicant"] --> B["Income < Rs. 500,000?"]
    B -->|"Yes"| C["Self-Employed?"]
    C -->|"Yes"| D["Reject (High Risk)"]
    C -->|"No"| E["Approve (Low Risk)"]
    B -->|"No"| F["Business Owner?"]
    F -->|"Yes"| G["Approve with Collateral"]
    F -->|"No"| H["Approve"]

Insight: Salaried applicants with stable income get faster approvals.


6. Common Mistakes to Avoid

  1. Misclassifying data: Treating ordinal data as interval (e.g., assuming hotel ratings are numerically equal).
  2. Ignoring data sources: Using outdated secondary data (e.g., 2020 tourism trends for 2024 decisions).
  3. Incorrect frequency tables: Forgetting to include all classes or miscounting frequencies.
  4. Ogive errors: Plotting cumulative frequency against the wrong axis (always x = class boundary, y = cumulative frequency).

Exam Tip: How to Score Full Marks

  1. Definitions: Always define terms clearly (e.g., "Primary data is collected directly by the researcher for a specific purpose").
  2. Examples: Use real-world tourism/business examples (NTC, Daraz, banks) to illustrate concepts.
  3. Visuals: For ogives/frequency tables, label axes clearly and show calculations (e.g., median position).
  4. Comparisons: Use tables to distinguish between primary/secondary data or qualitative/quantitative types.
  5. Practical Application: In case studies, link data classification to decision-making (e.g., "This classification helps NTC allocate resources efficiently").

Sample Exam Question: "Distinguish between primary and secondary data with examples from the tourism industry. How would a hotel manager use each type to improve operations?"

Model Answer: Primary data is collected firsthand (e.g., a hotel conducting guest satisfaction surveys to measure service quality). Secondary data is existing data (e.g., using Nepal Tourism Board reports to analyze competitor occupancy rates).

  • Primary use: The manager can adjust staffing based on real-time feedback.
  • Secondary use: The manager can compare performance against industry benchmarks.

Key Formulas to Remember

Concept Formula
Median position (ungrouped) or
Median position (grouped)
Relative Frequency

Summary Table: Data Classification

Aspect Qualitative Quantitative
Definition Categorical, non-numerical Numerical, measurable
Subtypes Nominal, Ordinal Discrete, Continuous
Example (Tourism) Hotel star ratings (1-5★) Number of tourists per day
Data Source Surveys, observations Counts, measurements
Analysis Tool Bar charts, pie charts Histograms, line graphs

Final Checklist Before Exam

✅ Can you distinguish primary vs. secondary data with examples? ✅ Can you classify tourism/business data into qualitative/quantitative? ✅ Can you construct a frequency table and draw ogives? ✅ Can you explain how NTC/Daraz use data classification in real decisions? ✅ Do you know how to find the median from an ogive?

Based on the TU BTTM syllabus for Statistics (STT301), unit 1.

Discussion

Loading…