STA169 Statistics I

Statistics IUnit 110 min read

Measurement Scales & Data Types: Classification, Examples & Applications

Unit 1 of Statistics I covers the four measurement scales (nominal, ordinal, interval, ratio) and four data types (qualitative/quantitative, discrete/continuous), with real-world applications in Nepalese tech companies, worked examples, and exam-focused visuals.

TAKEAWAYS:

  • Measurement scales (nominal, ordinal, interval, ratio) determine how data can be analyzed—nominal (labels only) to ratio (true zero + arithmetic).
  • Data types split into qualitative (descriptive) and quantitative (numerical), with subcategories discrete (countable) and continuous (measurable).
  • eSewa uses ordinal scales for user ratings (1–5 stars) and ratio for transaction amounts (Rs. 0–∞).
  • Khalti applies interval scales for time (e.g., 12:00 PM vs. 12:30 PM) and discrete data for transaction counts.
  • Worked examples show how to classify real datasets (e.g., Daraz order quantities vs. customer feedback).
  • Exam tip: Always label scales/data types in answers and justify choices (e.g., "Temperature in °C is interval because it has equal intervals but no true zero").

1. Measurement Scales: The Four Levels of Data

Measurement scales define how data is categorized, ordered, or quantified. They determine which statistical operations are valid. Below is a hierarchy of scales from least to most informative:

graph LR
    A["Nominal"] --> B["Ordinal"]
    B --> C["Interval"]
    C --> D["Ratio"]
    A -->|"Labels only"| E["Categories"]
    B -->|"Order + labels"| F["Rankings"]
    C -->|"Equal intervals"| G["Temperature (°C), IQ"]
    D -->|"True zero + arithmetic"| H["Weight (kg), Age (years)"]

Key Definitions & Examples

Scale Definition Examples (Nepal/Global) Allowed Operations Not Allowed
Nominal Labels/categories with no order. Gender (Male/Female), Blood group (A+, B-), eSewa user IDs, Daraz product categories. Counting, Mode. Mean, Median, Subtraction.
Ordinal Categories with meaningful order but no equal intervals. Customer ratings (1–5 stars), Traffic light colors (Red < Green < Yellow), Khalti transaction priority (Low/Medium/High). Mode, Median, Rank comparisons. Mean, Subtraction (e.g., "5 stars" – "3 stars" ≠ 2).
Interval Equal intervals but no true zero (arbitrary zero). Temperature (°C or °F), IQ scores, Year (e.g., 2000 vs. 2023). Mean, Median, Subtraction (e.g., 30°C – 20°C = 10). Multiplication/division (e.g., "20°C is not twice 10°C").
Ratio True zero + equal intervals + arithmetic valid. Weight (kg), Height (m), Age (years), Ncell data usage (MB), Daraz order quantities. Mean, Median, Mode, All arithmetic. None.
010203040Nominal10Ordinal25Interval40Ratio25Number of Examples in Nepalese Context
Frequency of each measurement scale in real-world Nepali data (e.g., eSewa, Khalti)

A pyramid showing nominal (bottom) to ratio (top) with icons for each scale type.

Worked Example 1: Classifying Data from eSewa

Dataset: eSewa user feedback survey responses:

  • Q1: "How satisfied are you with our service?" (Options: Very Dissatisfied, Dissatisfied, Neutral, Satisfied, Very Satisfied).
  • Q2: "How much did you pay for your last transaction?" (Rs. 500, Rs. 1200, Rs. 3500, etc.).
  • Q3: "What is your age group?" (18–25, 26–35, 36–45, 45+).

Solution:

  1. Q1: Ordinal (ordered categories but no equal intervals between "Dissatisfied" and "Neutral").
  2. Q2: Ratio (true zero, arithmetic valid: Rs. 1200 is 2.4× Rs. 500).
  3. Q3: Ordinal (age groups are ordered but intervals are unequal).

Interpretation: The mode (most frequent) is "Satisfied," but we cannot calculate the mean satisfaction score because the scale is ordinal.


2. Data Types: Qualitative vs. Quantitative

Data is further classified into qualitative (descriptive) and quantitative (numeric), with subcategories:

graph TD
    A["Data Types"] --> B["Qualitative"]
    A --> C["Quantitative"]
    B --> D["Nominal/Ordinal"]
    C --> E["Discrete"]
    C --> F["Continuous"]
    D -->|"Text/Labels"| G["Gender, Blood Group"]
    E -->|"Countable"| H["Number of Daraz orders, Ncell calls"]
    F -->|"Measurable"| I["Height, Temperature, Traffic speed"]

Definitions & Examples

Type Subtype Definition Examples (Nepal/Global) Statistical Tools
Qualitative Nominal Non-numeric labels. eSewa user IDs, Khalti payment methods (Debit/Credit), Blood groups. Frequency tables, Mode.
Ordinal Ordered categories. Customer reviews (1–5 stars), Traffic signals (Red/Yellow/Green), Education level (Primary/Secondary/University). Median, Percentiles.
Quantitative Discrete Countable, finite values. Number of Pathao rides, Ncell SMS sent, Daraz orders placed, NEPSE stock trades. Mean, Variance, Binomial distribution.
Continuous Measurable, infinite values within a range. Height (165.5 cm), Temperature (28.3°C), Traffic speed (65.2 km/h), NTC electricity usage (kWh). Mean, Standard deviation, Normal distribution.

A Venn diagram showing qualitative (left) and quantitative (right) with subtypes and icons (e.g., a bar chart for discrete, a line graph for continuous).

Worked Example 2: Classifying Data from NTC

Dataset: NTC’s monthly electricity consumption (in kWh) for 10 households:

  • Household A: 250, 260, 275, 280, 290
  • Household B: Low, Medium, High, Low, Medium
  • Household C: "High", "Very High", "High", "Medium", "Low"

Solution:

  1. Household A: Continuous quantitative (measurable, infinite possible values between 250 and 290).
  2. Household B: Ordinal qualitative (ordered categories but no numeric values).
  3. Household C: Nominal qualitative (labels only, no order implied by "High" vs. "Very High").

Key Insight: Only Household A can have its mean consumption calculated (271 kWh). Households B and C require non-parametric methods.


3. Real-World Applications in Nepalese Tech

Example 1: eSewa (Ordinal + Ratio Scales)

  • Ordinal: User ratings (1–5 stars) for service quality.
    • Why? "5 stars" > "3 stars," but the difference between "5" and "4" isn’t numerically meaningful.
  • Ratio: Transaction amounts (Rs. 500, Rs. 2000).
    • Why? Rs. 0 is valid (no transaction), and Rs. 2000 is 4× Rs. 500.

Example 2: Khalti (Interval + Discrete Data)

  • Interval: Time of transactions (e.g., 10:30 AM vs. 11:00 AM).
    • Why? 30-minute intervals are equal, but "10:30 AM" isn’t twice "5:15 AM."
  • Discrete: Number of transactions per user (0, 1, 2, ...).
    • Why? You can’t have a fraction of a transaction.

Example 3: Daraz (Nominal + Ratio Data)

  • Nominal: Product categories (Electronics, Grocery, Fashion).
    • Why? No order or arithmetic applies.
  • Ratio: Order quantities (1 kg, 5 kg, 10 kg of rice).
    • Why? 10 kg is 10× 1 kg, and 0 kg is valid (no order).

4. Common Mistakes & How to Avoid Them

Mistake Why It’s Wrong Correct Approach
Treating ordinal as interval. Assuming "5 stars" – "3 stars" = 2 units of satisfaction. Use median or percentiles, not mean.
Using mean on nominal data. Averaging blood groups (A+, B–) is meaningless. Report frequencies (e.g., "60% are A+").
Confusing discrete and continuous. Saying "height is discrete" because it’s measured in cm. Height is continuous (165.5 cm, 165.51 cm, etc.).
Ignoring true zero in ratio data. Calculating mean temperature in °C as if it were ratio. Convert to Kelvin (true zero) for ratio operations.

A flowchart showing "Is the data nominal/ordinal/interval/ratio?" with arrows to correct operations.


5. Exam Tip: How to Score Full Marks

  1. Always justify your classification:

    • ❌ "This is ratio data."
    • ✅ "This is ratio data because it has a true zero (e.g., 0 kg of rice) and equal intervals (1 kg increments), allowing arithmetic operations like calculating the mean order quantity."
  2. Use real-world examples:

    • For interval data, cite temperature or IQ scores.
    • For discrete data, use counts (e.g., "number of Ncell calls").
  3. Visual aids in answers:

    • Draw a table for classification (like the one above).
    • Sketch a bar chart for qualitative data or a histogram for quantitative data.
  4. Watch out for trick questions:

    • Age groups (e.g., 18–25, 26–35) are ordinal, not interval, because the intervals (7 years vs. 9 years) are unequal.
    • Years (e.g., 2000 vs. 2023) are interval, not ratio, because there’s no true zero (year 0 doesn’t mean "no time").
  5. Practice past exam questions:

    • For Q1, D7, P58 (percentiles/deciles), always convert grouped data to cumulative frequencies first.
    • For moments about an arbitrary point, use the formula: where is the arbitrary point (e.g., 4 in the past exam).

A table with formulas for mean, median, mode per scale type (e.g., "Mean: Only for interval/ratio").

Based on the TU BSc CSIT syllabus for Statistics I (STA169), unit 1.

Discussion

Loading…