Research MethodologyUnit 611 min read
Measurement & Scaling: Types, Techniques & Applications
Unit 6 of Research Methodology explores how to quantify variables in research, covering measurement scales (nominal, ordinal, interval, ratio), scaling techniques (Likert, semantic differential, etc.), and their applications in data analysis. Learn how to classify data correctly, choose appropriate scales, and avoid co
TAKEAWAYS:
- Measurement scales determine how data is analyzed—nominal (categories), ordinal (ranked), interval (equal intervals), and ratio (absolute zero).
- Likert scales (agree/disagree) and semantic differential scales (bipolar adjectives) are common scaling techniques for surveys.
- Reliability (consistency) and validity (accuracy) are critical in measurement tools like questionnaires.
- Real-world applications include customer satisfaction surveys (e.g., Daraz, Pathao), financial risk assessment (banks), and policy evaluation (NTC, NEPSE).
- Misclassifying scales (e.g., treating ordinal as interval) leads to invalid statistical tests—a common exam mistake.
- Pilot testing ensures scales work before full deployment (e.g., eSewa’s user feedback forms).
1. What is Measurement in Research?
Measurement is the process of assigning numbers or labels to variables to describe and analyze phenomena systematically. Without proper measurement, research data becomes meaningless. For example:
- Nominal: Gender (Male/Female) → Categories only.
- Ordinal: Customer satisfaction (Poor, Fair, Good, Excellent) → Ranked but no equal intervals.
- Interval: Temperature in °C → Equal intervals but no true zero.
- Ratio: Income in NPR → True zero (NPR 0 = no income).
Why does this matter?
- Statistical tests (e.g., mean, standard deviation, t-tests) require specific scales.
- You cannot calculate a mean for nominal data (e.g., "red," "blue").
- You can rank ordinal data but not perform advanced math.
2. Levels of Measurement (Scales of Measurement)
The four scales differ in precision, mathematical operations allowed, and statistical tests applicable. Use this table to decide which scale fits your data:
| Scale Type | Example | Mathematical Operations Allowed | Statistical Tests | Real-World Use Case |
|---|---|---|---|---|
| Nominal | Gender (Male/Female) | Counting frequencies | Chi-square, Mode | Blood type (A, B, AB, O) in medical studies |
| Ordinal | Education level (Primary, Secondary, Bachelor’s) | Ranking (>, <) | Spearman’s rank, Mann-Whitney U | Customer reviews (1–5 stars) on Daraz |
| Interval | IQ scores (80, 100, 120) | Addition/subtraction, mean, median | Pearson correlation, t-tests | Temperature (°C), SAT scores |
| Ratio | Height (150 cm, 160 cm) | Multiplication/division, ratios | ANOVA, regression, geometric mean | Income (NPR 50,000 vs. 100,000), weight |
3. Common Scaling Techniques
Scaling converts qualitative data into quantitative form for analysis. Two widely used methods:
A. Likert Scale
- Measures attitudes or opinions on a symmetric agree-disagree spectrum.
- Example items:
- "I am satisfied with Pathao’s delivery service."
- 1 (Strongly Disagree) → 5 (Strongly Agree)
- "I am satisfied with Pathao’s delivery service."
- Strengths:
- Simple to design and administer.
- Captures degree of agreement (not just yes/no).
- Weaknesses:
- Ordinal data (cannot assume equal intervals between points).
- Central tendency bias (respondents may cluster around "neutral").
Worked Example: Pathao Driver Satisfaction Survey Suppose Pathao wants to measure driver satisfaction with a 5-point Likert scale:
- Question: "How satisfied are you with your earnings as a Pathao driver?"
- 1 (Very Dissatisfied) → 5 (Very Satisfied)
- Analysis:
- If 60% of drivers select 4 or 5, Pathao can infer high satisfaction.
- But: You cannot say "Satisfaction increased by 2 points" because the scale is ordinal.
B. Semantic Differential Scale
- Uses bipolar adjectives (e.g., "Good-Bad," "Fast-Slow") to measure perceptions.
- Example:
- "How would you describe Daraz’s customer service?"
- Poor 1 2 3 4 5 Excellent
- "How would you describe Daraz’s customer service?"
- Strengths:
- Captures nuanced opinions (e.g., "slightly fast" vs. "very slow").
- Useful for brand perception (e.g., Ncell vs. NTC).
- Weaknesses:
- Subjective interpretation (what "3" means varies by respondent).
- Requires careful wording to avoid bias.
Worked Example: NTC vs. Ncell Customer Perception A study uses a semantic differential scale to compare:
- NTC: "Reliable" (1–7) vs. "Unreliable"
- Ncell: "Fast" (1–7) vs. "Slow" Findings:
- If Ncell scores 6.2 for "Fast" and NTC scores 4.5, Ncell is perceived as faster.
- But: The scale is interval, so you can calculate a mean difference.
4. Measurement Errors and How to Avoid Them
Even the best scales can produce invalid or unreliable data. Common errors:
| Error Type | Cause | Example | Solution |
|---|---|---|---|
| Reliability | Inconsistent results | A Likert scale gives different results on retest | Pilot test the scale before full survey. |
| Validity | Measures the wrong thing | Asking "How often do you use WhatsApp?" to measure "digital literacy" | Face validity: Ensure questions align with research goals. |
| Bias | Leading or ambiguous questions | "Don’t you agree Daraz’s prices are too high?" | Use neutral wording (e.g., "What do you think of Daraz’s prices?"). |
| Social Desirability | Respondents lie to appear "better" | Overestimating income in a bank survey | Use anonymous surveys or indirect questions. |
5. Choosing the Right Scale for Your Research
Step-by-Step Decision Guide:
- Define your variable:
- Is it a category (nominal), rank (ordinal), interval, or ratio?
- Decide on data collection method:
- Surveys: Likert or semantic differential scales.
- Experiments: Ratio scales (e.g., time taken, weight).
- Pilot test:
- Give the scale to a small group (e.g., 10–20 people) and check:
- Are questions clear?
- Are responses consistent?
- Give the scale to a small group (e.g., 10–20 people) and check:
- Analyze statistically:
- Nominal → Frequencies, Chi-square.
- Ordinal → Spearman’s rank.
- Interval/Ratio → t-tests, ANOVA, regression.
Worked Example: eSewa User Feedback eSewa wants to measure user satisfaction with a new feature.
- Option 1: Nominal ("Yes/No" for "Did you use the feature?")
- Problem: Cannot measure degree of satisfaction.
- Option 2: Likert scale (1–5 for "How satisfied are you?")
- Better: Captures variation in satisfaction.
- Option 3: Semantic differential ("Easy-Hard," "Useful-Useless")
- Best for nuanced feedback.
6. Real-World Applications in Nepal
A. E-Commerce: Daraz Customer Reviews
- Scale Used: Ordinal (1–5 stars) for product ratings.
- How It Works:
- Customers rate products on satisfaction, quality, delivery speed.
- Daraz uses mean ratings to rank products (interval-like analysis, though technically ordinal).
- Why It Matters:
- Helps Daraz improve low-rated items (e.g., faster delivery for 2-star products).
- But: If Daraz treats 5-star = "twice as good" as 2.5-star, it’s incorrect (ordinal data cannot be multiplied).
B. Banking: Loan Risk Assessment
- Scale Used: Ratio (credit score, income in NPR) + Ordinal (repayment history: Poor/Fair/Good/Excellent).
- How It Works:
- Banks like NMB or Global IME use:
- Income (ratio): Higher income = lower risk.
- Credit history (ordinal): Late payments reduce score.
- Statistical Test: Logistic regression (predicts loan default probability).
- Banks like NMB or Global IME use:
C. Traffic Management: Kathmandu’s Road Congestion
- Scale Used: Interval (traffic density per km) + Ordinal (congestion level: Low/Medium/High).
- How It Works:
- NTC measures:
- Vehicle count per hour (ratio) → High = congestion.
- Driver frustration (Likert scale: 1–5) → Correlates with accidents.
- Policy Decision: If 70% of drivers rate congestion as "5" (Extreme), NTC may introduce smart traffic lights.
- NTC measures:
7. Common Mistakes to Avoid in Exams
- Treating Ordinal as Interval:
- ❌ "The mean satisfaction score is 3.5" (if scale is Likert).
- ✅ "Most respondents rated satisfaction as 4 or 5."
- Ignoring Pilot Testing:
- ❌ Using a poorly worded Likert scale in the final survey.
- ✅ Test with 10 people first, refine questions.
- Misclassifying Variables:
- ❌ "Temperature in °C is ratio" (it’s interval—no true zero).
- ✅ "Income in NPR is ratio" (true zero exists).
- Overlooking Validity:
- ❌ Asking "How often do you exercise?" to measure "health awareness."
- ✅ Use a multidimensional scale (e.g., Likert for exercise + diet).
Exam Tip
How This Unit is Tested:
- Short Questions (5–10 marks):
- Define nominal vs. ordinal scales.
- Give an example of a Likert scale and its limitations.
- Long Questions (20–30 marks):
- Design a measurement tool for a given scenario (e.g., "Measure student satisfaction with TU’s online classes").
- Steps:
- Choose scale type (e.g., Likert for satisfaction).
- Draft 5–10 questions with clear options.
- Justify why this scale is appropriate.
- Steps:
- Critique a given scale (e.g., "This survey uses a 1–10 scale for pain—is this valid? Why not 1–5?").
- Design a measurement tool for a given scenario (e.g., "Measure student satisfaction with TU’s online classes").
- Data Interpretation (10–15 marks):
- Given a table of Likert responses, calculate:
- Mode (most frequent response).
- Median (middle value).
- Interpret trends (e.g., "70% rated service as 4 or 5 → high satisfaction").
- Given a table of Likert responses, calculate:
Key Formula to Remember:
- Cronbach’s Alpha (α) → Measures internal consistency of a scale (e.g., all Likert items should correlate).
- α ≥ 0.7 = acceptable reliability.
Summary Checklist
Before finalizing your measurement tool, ask:
- Is my scale the right type (nominal/ordinal/interval/ratio) for my data?
- Are my questions clear and unbiased?
- Have I pilot-tested the scale?
- What statistical tests will I use based on the scale?
- Does the scale measure what I intend (validity)?
Final Visual Recap:
Based on the TU BIT syllabus for Research Methodology (RSM354), unit 6.
Discussion
Loading…