Elective Research Fundamentals

Research FundamentalsUnit 817 min read

Data Analysis: Methods, Tools & Real-World Applications

Unit 8 of Research Fundamentals explores quantitative vs. qualitative analysis, statistical techniques, data visualization, and software tools (SPSS, R, Excel) with real-world examples from Nepalese tech companies (e.g., eSewa fraud detection, Ncell customer churn analysis) and global platforms (Google search ranking).

TAKEAWAYS:

  • Quantitative analysis uses statistics (descriptive/inferential) to test hypotheses, while qualitative analysis interprets themes from text/audio (e.g., Ncell customer complaints).
  • Data cleaning (handling missing values, outliers) is critical—eSewa’s fraud detection relies on removing duplicate transactions before analysis.
  • Visualization tools (charts, heatmaps) reveal patterns: Pathao’s ride-demand heatmap shows peak hours in Kathmandu.
  • Software matters: SPSS for surveys, R for predictive modeling (e.g., NTC’s network failure prediction), Excel for quick summaries.
  • Ethics in analysis: Avoid cherry-picking data (e.g., Daraz hiding low ratings) or overfitting models to bias results.
  • Worked example: Calculate NEPSE’s 5-year stock return using mean/median to compare volatility vs. stability.

1. Types of Data Analysis: Quantitative vs. Qualitative

Data analysis splits into two broad methods, each suited to different research goals. The choice depends on your research question, data type, and objective.

00.30.60.91.2Pre-OTP Fraud Incidents1.2Post-OTP Fraud Incidents0.5
eSewa fraud reduction: Mean incidents per user (N=500) before/after OTP implementation.

Quantitative Analysis

Definition: Uses numerical data and statistical methods to measure, quantify, and analyze patterns or relationships. It answers "how much?" or "how often?".

Key Techniques:

  • Descriptive statistics: Summarize data (mean, median, mode, standard deviation).
  • Inferential statistics: Test hypotheses (t-tests, ANOVA, regression).
  • Multivariate analysis: Examines multiple variables (factor analysis, cluster analysis).

When to Use:

  • Testing hypotheses (e.g., "Does eSewa’s new OTP system reduce fraud by 20%?").
  • Measuring trends (e.g., "How does Ncell’s 4G speed vary by district?").
  • Predictive modeling (e.g., "Can Daraz’s algorithm predict customer churn?").

Example Workflow:

  1. Collect data: Survey 500 eSewa users on fraud incidents (pre/post-OTP).
  2. Clean data: Remove incomplete responses, correct typos in transaction IDs.
  3. Analyze:
    • Descriptive: Mean fraud incidents = 1.2 (pre-OTP) vs. 0.5 (post-OTP).
    • Inferential: Run a paired t-test to confirm the reduction is statistically significant (p < 0.05).
  4. Visualize: Bar chart of fraud rates by month.
graph TD
    A["Quantitative Analysis"] --> B["Descriptive Stats\n(Mean, Median, SD)"]
    A --> C["Inferential Stats\n(t-test, ANOVA, Regression)"]
    A --> D["Multivariate Stats\n(Cluster, Factor Analysis)"]
    B --> E["Summarize Data"]
    C --> F["Test Hypotheses"]
    D --> G["Find Patterns in\nMultiple Variables"]

Qualitative Analysis

Definition: Interprets non-numerical data (text, images, audio) to uncover themes, motivations, or experiences. Answers "why?" or "how?".

Key Techniques:

  • Thematic analysis: Identify recurring themes (e.g., "Why do Pathao drivers quit?").
  • Content analysis: Code text for keywords (e.g., "How often does NTC blame ‘technical issues’ for outages?").
  • Grounded theory: Develop theories from data (e.g., "What drives small businesses to use Khalti over eSewa?").

When to Use:

  • Exploring user experiences (e.g., "How do Daraz sellers feel about late payments?").
  • Understanding cultural contexts (e.g., "Why do Nepali students prefer WhatsApp over email?").
  • Pilot studies before quantitative research.

Example Workflow:

  1. Collect data: Interview 10 Ncell customers about their complaints.
  2. Transcribe: Convert audio to text.
  3. Code: Label phrases like "network drops" or "slow speed" as "Technical Issues."
  4. Analyze themes: 70% of complaints = "Signal Problems"; 20% = "Billing Errors."
  5. Report: "Ncell’s primary issue is weak 4G coverage in hilly areas."

Comparison Table: Quantitative vs. Qualitative

Aspect Quantitative Analysis Qualitative Analysis
Data Type Numerical (surveys, experiments, logs) Text, audio, video (interviews, focus groups)
Research Question "How many?", "How much?" "Why?", "How?", "What’s the experience?"
Tools SPSS, R, Excel, Python (Pandas, SciPy) NVivo, ATLAS.ti, Excel (for coding)
Sample Size Large (300+ for reliability) Small (10–50 for depth)
Flexibility Rigid (predefined variables) Flexible (emergent themes)
Example in Nepal Analyzing NEPSE stock trends with moving averages Studying Khalti’s user feedback for UX improvements

2. Data Cleaning: The Hidden Step Before Analysis

Why it matters: Dirty data leads to wrong conclusions. For example:

  • eSewa’s fraud detection fails if duplicate transactions aren’t removed.
  • Ncell’s network reports show false "outages" if GPS coordinates are mislabeled.

Common Data Issues & Fixes

Problem Example in Nepalese Context Solution
Missing values 20% of Daraz order forms left "delivery address" blank Impute (fill) with mode/most common address or flag as "incomplete."
Outliers A single NEPSE stock price of Rs. 50,000 (vs. usual Rs. 1,000–5,000) Check for errors; if valid, analyze separately.
Inconsistent formats Dates written as "2023/05/15" and "15-05-2023" Standardize to YYYY-MM-DD for sorting.
Duplicate entries Same Khalti transaction ID appears twice Use SQL DISTINCT or Python drop_duplicates().
Categorical errors "Male" coded as "1" in one survey, "M" in another Recode all as "1" for consistency.

Worked Example: Cleaning NTC’s Network Outage Data Raw Data:

Date District Outage Duration (mins) Cause
2023-10-01 Kathmandu 120 "Technical Issue"
2023-10-01 Kathmandu 120 "Technical Issue"
2023-10-02 Lalitpur 45 "Tree Fall"
2023-10-02 Lalitpur NULL "Tree Fall"

Steps:

  1. Remove duplicates: Keep one row for 2023-10-01, Kathmandu.
  2. Impute missing values: Replace NULL with mean duration (e.g., 60 mins).
  3. Standardize "Cause": Combine "Technical Issue" and "Hardware Failure" into "Infrastructure."
  4. Verify: Check if outages in hilly districts (e.g., Dhading) are underreported.

Tool Tip: Use Excel’s Remove Duplicates or Python’s pandas:

import pandas as pd
df = pd.read_csv("ntc_outages.csv")
df_clean = df.drop_duplicates().fillna(df.mean())

3. Statistical Techniques: From Descriptive to Predictive

graph TD
    A["Descriptive Stats"] --> B["Mean/Median"]
    A --> C["Standard Deviation"]
    A --> D["Visualization"]
    B --> E["NEPSE Stock Return: Rs. 4,500"]
    C --> F["Volatility: Rs. 800 SD"]
    D --> G["Line Chart: 5-Year Trend"]
Workflow for analyzing NEPSE stock returns using descriptive statistics.

A. Descriptive Statistics: Summarizing Data

Purpose: Simplify large datasets into measures of central tendency and dispersion.

Statistic Formula When to Use Example
Mean Symmetrical data (e.g., NEPSE daily returns) Mean return = 0.5% (but hide volatility!)
Median Middle value Skewed data (e.g., Khalti transaction amounts) Median = Rs. 2,000 (mean = Rs. 5,000 due to outliers)
Mode Most frequent value Categorical data (e.g., most common Ncell complaint) Mode = "Signal drops" (appears 40% of feedback)
Standard Deviation Measure spread (e.g., Daraz delivery delays) SD = 15 mins → Most orders arrive within 30 mins.

Visual: Always pair numbers with charts. For NEPSE stock:


B. Inferential Statistics: Testing Hypotheses

Goal: Generalize findings from a sample to a population.

Test When to Use Example in Nepal Interpretation
t-test Compare means of two groups (independent or paired) Does eSewa’s new OTP reduce fraud vs. old PIN? p < 0.05 → OTP is significantly better.
ANOVA Compare means of >2 groups Do Ncell, NTC, and SmartCell have different outage rates? p < 0.05 → At least one differs.
Chi-square Test relationship between categorical variables Are Khalti users more likely to be under 30? p < 0.05 → Yes, 70% of users are <30.
Regression Predict one variable from others How does NEPSE index predict bank loan defaults? R² = 0.6 → 60% of defaults explained by NEPSE.

Worked Example: eSewa Fraud Reduction Hypothesis: "The new OTP system reduces fraud by 20%."

  1. Data: 300 users pre-OTP (mean fraud = 1.2 incidents), 300 post-OTP (mean = 0.5).
  2. Test: Paired t-test (same users before/after).
  3. Result: t = 4.2, p = 0.0001 → Reject null hypothesis (fraud reduced significantly).
  4. Conclusion: OTP works! But qualitative follow-up needed: "Why do users still get scammed?"

4. Data Visualization: Making Data Speak

Rule: "A picture is worth 1,000 data points." Poor visuals mislead—like Daraz hiding low ratings in tiny text.

When to Use Which Chart

Chart Type Best For Nepalese Example Avoid If...
Bar Chart Compare categories (e.g., fraud by bank) Ncell vs. NTC vs. SmartCell outage rates Data is continuous (use histogram instead).
Line Graph Trends over time (e.g., NEPSE index) Monthly Khalti transactions (2020–2023) Categories aren’t ordered (use bar chart).
Pie Chart Parts of a whole (but rarely useful) Reasons for Pathao driver quits (30% low pay, 20% traffic) >5 categories (hard to read).
Histogram Distribution of continuous data Distribution of Daraz order delays Data is categorical (use bar chart).
Scatter Plot Relationships between two variables NEPSE index vs. bank loan defaults No clear pattern (add regression line).
Heatmap Intensity in 2D (e.g., time + location) Pathao ride demand by hour + district Data isn’t spatial/temporal.

Example: NEPSE Stock Volatility


Interpretation:

  • Upper Band: Stock is overvalued (sell signal).
  • Lower Band: Stock is undervalued (buy signal).
  • Narrow Bands: Low volatility (e.g., NEPSE in 2021).

5. Software Tools for Data Analysis

Tool Best For Nepalese Use Case Learning Curve
Excel Quick summaries, basic charts Bank loan repayment tracking Low
SPSS Statistical tests (t-tests, ANOVA) Ncell customer satisfaction surveys Medium
R Advanced stats, predictive modeling NTC network failure prediction High
Python (Pandas, SciPy) Large datasets, automation Daraz’s recommendation algorithm Medium-High
NVivo Qualitative analysis (themes, coding) Khalti user feedback analysis High

Worked Example: Python for NEPSE Analysis

import pandas as pd
import matplotlib.pyplot as plt

# Load NEPSE data
data = pd.read_csv("nepse_index.csv")
data['Date'] = pd.to_datetime(data['Date'])

# Calculate moving average (50 days)
data['MA50'] = data['Close'].rolling(50).mean()

# Plot
plt.plot(data['Date'], data['Close'], label='NEPSE Index')
plt.plot(data['Date'], data['MA50'], label='50-Day MA', color='red')
plt.title("NEPSE Index with Moving Average")
plt.legend()
plt.show()

Output:



In the Real World

  1. eSewa’s Fraud Detection

    • Idea: Anomaly detection (statistical outliers).
    • How: Uses z-scores to flag transactions where amount > mean + 3*SD. Example: A Rs. 50,000 transfer (mean = Rs. 2,000) triggers a review.
    • Impact: Reduced fraud by 15% in 2022.
  2. Pathao’s Ride Demand Heatmap

    • Idea: Geospatial visualization (heatmaps).
    • How: Overlays ride requests per km² with traffic data to predict surge pricing zones. Example: Heatmap shows high demand in Thapathali (2 PM–5 PM).
    • Impact: Drivers earn 30% more in peak hours.
  3. Ncell’s Customer Churn Prediction

    • Idea: Logistic regression (predictive modeling).
    • How: Trains a model on call duration, complaints, and usage to predict which users will switch. Example: Users with >3 complaints/month have 80% churn risk.
    • Impact: Targeted discounts reduced churn by 12%.
  4. Daraz’s Order Fulfillment Delays

    • Idea: Queueing theory (probability of delays).
    • How: Models order arrival rate (λ) vs. warehouse processing rate (μ). Example: If λ = 100 orders/hour and μ = 80, 20% of orders face delays.
    • Impact: Added 2 warehouses in Kathmandu to balance λ and μ.
  5. NEPSE’s Index Calculation

    • Idea: Weighted average (statistical aggregation).
    • How: Combines stock prices of top 10 companies with market caps as weights. Example: If NMB Bank (20% weight) rises 5%, it boosts NEPSE by 1%.

Exam Tip

  1. Know the Difference: Always distinguish between quantitative (numbers, stats) and qualitative (text, themes). Examiners love questions like:

    • "When would you use a t-test vs. thematic analysis?"
    • Answer: "A t-test compares means (quantitative), while thematic analysis finds patterns in interview transcripts (qualitative)."
  2. Data Cleaning is Critical: Expect 3–5 marks on cleaning steps. For example:

    • Question: "How would you handle missing values in a survey on Khalti usage?"
    • Answer:
      • Missing = "No" (e.g., "Do you use Khalti for bills?" → Missing = "No").
      • Impute mean for numerical data (e.g., "How often?").
      • Exclude if >20% missing (but justify!).
  3. Visuals = Marks: If asked to "analyze NEPSE data", always sketch a chart (even on paper). Example:

    • Question: "Describe the trend in NEPSE from 2020–2023."
    • Answer: *"A line graph shows a 2020 dip (-15%) due to COVID, followed by a 2021–2022 recovery (+30%), and 2023 volatility (±10%)."*
  4. Software Shortcuts: Memorize these SPSS/R/Python commands for quick answers:

    • SPSS: Analyze > Compare Means > Independent-Samples T Test.
    • R: t.test(group1 ~ group2, data = df).
    • Python: from scipy import stats; stats.ttest_ind(group1, group2).
  5. Ethics Trap: Examiners test unethical data practices. Example:

    • Question: "Is it ethical to exclude outliers in Ncell’s outage data?"
    • Answer: "No—unless outliers are errors. If they’re valid (e.g., 2021 Kathmandu blackout), exclude them only if justified (e.g., ‘extreme event’). Otherwise, analyze separately."
  6. Worked Example Formula: For any numerical question, follow:

    1. State the test (e.g., "We’ll use a chi-square test.").
    2. Show the formula (even if not calculated).
    3. Interpret the result (e.g., "p < 0.05 → relationship is significant.").

Final Visual Summary:

mindmap
  root((Data Analysis))
    Quantitative
      Descriptive Stats
        Mean/Median/Mode
        Standard Deviation
      Inferential Stats
        t-test
        ANOVA
        Regression
    Qualitative
      Thematic Analysis
      Content Analysis
    Data Cleaning
      Missing Values
      Outliers
      Duplicates
    Visualization
      Bar Charts
      Line Graphs
      Heatmaps
    Tools
      SPSS
      R
      Python

In the real world

  • eSewa uses quantitative analysis to detect fraud by flagging transactions with outliers (e.g., Rs. 50,000 transfers to unknown accounts) and missing values (incomplete KYC data).
  • NTC applies qualitative analysis to customer complaints (e.g., coding themes like "signal drops" vs. "billing errors") to prioritize network upgrades in Kathmandu’s hilly districts.
  • Pathao leverages predictive modeling (quantitative) to forecast driver shortages during Dashain by analyzing historical ride-demand heatmaps (visualization) from the previous year.

Based on the PU BE Computer (PU) syllabus for Research Fundamentals, unit 8.

Discussion

Loading…