Research FundamentalsUnit 817 min read
Data Analysis: Methods, Tools & Real-World Applications
Unit 8 of Research Fundamentals explores quantitative vs. qualitative analysis, statistical techniques, data visualization, and software tools (SPSS, R, Excel) with real-world examples from Nepalese tech companies (e.g., eSewa fraud detection, Ncell customer churn analysis) and global platforms (Google search ranking).
TAKEAWAYS:
- Quantitative analysis uses statistics (descriptive/inferential) to test hypotheses, while qualitative analysis interprets themes from text/audio (e.g., Ncell customer complaints).
- Data cleaning (handling missing values, outliers) is critical—eSewa’s fraud detection relies on removing duplicate transactions before analysis.
- Visualization tools (charts, heatmaps) reveal patterns: Pathao’s ride-demand heatmap shows peak hours in Kathmandu.
- Software matters: SPSS for surveys, R for predictive modeling (e.g., NTC’s network failure prediction), Excel for quick summaries.
- Ethics in analysis: Avoid cherry-picking data (e.g., Daraz hiding low ratings) or overfitting models to bias results.
- Worked example: Calculate NEPSE’s 5-year stock return using mean/median to compare volatility vs. stability.
1. Types of Data Analysis: Quantitative vs. Qualitative
Data analysis splits into two broad methods, each suited to different research goals. The choice depends on your research question, data type, and objective.
Quantitative Analysis
Definition: Uses numerical data and statistical methods to measure, quantify, and analyze patterns or relationships. It answers "how much?" or "how often?".
Key Techniques:
- Descriptive statistics: Summarize data (mean, median, mode, standard deviation).
- Inferential statistics: Test hypotheses (t-tests, ANOVA, regression).
- Multivariate analysis: Examines multiple variables (factor analysis, cluster analysis).
When to Use:
- Testing hypotheses (e.g., "Does eSewa’s new OTP system reduce fraud by 20%?").
- Measuring trends (e.g., "How does Ncell’s 4G speed vary by district?").
- Predictive modeling (e.g., "Can Daraz’s algorithm predict customer churn?").
Example Workflow:
- Collect data: Survey 500 eSewa users on fraud incidents (pre/post-OTP).
- Clean data: Remove incomplete responses, correct typos in transaction IDs.
- Analyze:
- Descriptive: Mean fraud incidents = 1.2 (pre-OTP) vs. 0.5 (post-OTP).
- Inferential: Run a paired t-test to confirm the reduction is statistically significant (p < 0.05).
- Visualize: Bar chart of fraud rates by month.
graph TD
A["Quantitative Analysis"] --> B["Descriptive Stats\n(Mean, Median, SD)"]
A --> C["Inferential Stats\n(t-test, ANOVA, Regression)"]
A --> D["Multivariate Stats\n(Cluster, Factor Analysis)"]
B --> E["Summarize Data"]
C --> F["Test Hypotheses"]
D --> G["Find Patterns in\nMultiple Variables"]Qualitative Analysis
Definition: Interprets non-numerical data (text, images, audio) to uncover themes, motivations, or experiences. Answers "why?" or "how?".
Key Techniques:
- Thematic analysis: Identify recurring themes (e.g., "Why do Pathao drivers quit?").
- Content analysis: Code text for keywords (e.g., "How often does NTC blame ‘technical issues’ for outages?").
- Grounded theory: Develop theories from data (e.g., "What drives small businesses to use Khalti over eSewa?").
When to Use:
- Exploring user experiences (e.g., "How do Daraz sellers feel about late payments?").
- Understanding cultural contexts (e.g., "Why do Nepali students prefer WhatsApp over email?").
- Pilot studies before quantitative research.
Example Workflow:
- Collect data: Interview 10 Ncell customers about their complaints.
- Transcribe: Convert audio to text.
- Code: Label phrases like "network drops" or "slow speed" as "Technical Issues."
- Analyze themes: 70% of complaints = "Signal Problems"; 20% = "Billing Errors."
- Report: "Ncell’s primary issue is weak 4G coverage in hilly areas."
Comparison Table: Quantitative vs. Qualitative
| Aspect | Quantitative Analysis | Qualitative Analysis |
|---|---|---|
| Data Type | Numerical (surveys, experiments, logs) | Text, audio, video (interviews, focus groups) |
| Research Question | "How many?", "How much?" | "Why?", "How?", "What’s the experience?" |
| Tools | SPSS, R, Excel, Python (Pandas, SciPy) | NVivo, ATLAS.ti, Excel (for coding) |
| Sample Size | Large (300+ for reliability) | Small (10–50 for depth) |
| Flexibility | Rigid (predefined variables) | Flexible (emergent themes) |
| Example in Nepal | Analyzing NEPSE stock trends with moving averages | Studying Khalti’s user feedback for UX improvements |
2. Data Cleaning: The Hidden Step Before Analysis
Why it matters: Dirty data leads to wrong conclusions. For example:
- eSewa’s fraud detection fails if duplicate transactions aren’t removed.
- Ncell’s network reports show false "outages" if GPS coordinates are mislabeled.
Common Data Issues & Fixes
| Problem | Example in Nepalese Context | Solution |
|---|---|---|
| Missing values | 20% of Daraz order forms left "delivery address" blank | Impute (fill) with mode/most common address or flag as "incomplete." |
| Outliers | A single NEPSE stock price of Rs. 50,000 (vs. usual Rs. 1,000–5,000) | Check for errors; if valid, analyze separately. |
| Inconsistent formats | Dates written as "2023/05/15" and "15-05-2023" | Standardize to YYYY-MM-DD for sorting. |
| Duplicate entries | Same Khalti transaction ID appears twice | Use SQL DISTINCT or Python drop_duplicates(). |
| Categorical errors | "Male" coded as "1" in one survey, "M" in another | Recode all as "1" for consistency. |
Worked Example: Cleaning NTC’s Network Outage Data Raw Data:
| Date | District | Outage Duration (mins) | Cause |
|---|---|---|---|
| 2023-10-01 | Kathmandu | 120 | "Technical Issue" |
| 2023-10-01 | Kathmandu | 120 | "Technical Issue" |
| 2023-10-02 | Lalitpur | 45 | "Tree Fall" |
| 2023-10-02 | Lalitpur | NULL | "Tree Fall" |
Steps:
- Remove duplicates: Keep one row for 2023-10-01, Kathmandu.
- Impute missing values: Replace
NULLwith mean duration (e.g., 60 mins). - Standardize "Cause": Combine "Technical Issue" and "Hardware Failure" into "Infrastructure."
- Verify: Check if outages in hilly districts (e.g., Dhading) are underreported.
Tool Tip: Use Excel’s Remove Duplicates or Python’s pandas:
import pandas as pd
df = pd.read_csv("ntc_outages.csv")
df_clean = df.drop_duplicates().fillna(df.mean())
3. Statistical Techniques: From Descriptive to Predictive
graph TD
A["Descriptive Stats"] --> B["Mean/Median"]
A --> C["Standard Deviation"]
A --> D["Visualization"]
B --> E["NEPSE Stock Return: Rs. 4,500"]
C --> F["Volatility: Rs. 800 SD"]
D --> G["Line Chart: 5-Year Trend"]Workflow for analyzing NEPSE stock returns using descriptive statistics.A. Descriptive Statistics: Summarizing Data
Purpose: Simplify large datasets into measures of central tendency and dispersion.
| Statistic | Formula | When to Use | Example |
|---|---|---|---|
| Mean | Symmetrical data (e.g., NEPSE daily returns) | Mean return = 0.5% (but hide volatility!) | |
| Median | Middle value | Skewed data (e.g., Khalti transaction amounts) | Median = Rs. 2,000 (mean = Rs. 5,000 due to outliers) |
| Mode | Most frequent value | Categorical data (e.g., most common Ncell complaint) | Mode = "Signal drops" (appears 40% of feedback) |
| Standard Deviation | Measure spread (e.g., Daraz delivery delays) | SD = 15 mins → Most orders arrive within 30 mins. |
Visual: Always pair numbers with charts. For NEPSE stock:
B. Inferential Statistics: Testing Hypotheses
Goal: Generalize findings from a sample to a population.
| Test | When to Use | Example in Nepal | Interpretation |
|---|---|---|---|
| t-test | Compare means of two groups (independent or paired) | Does eSewa’s new OTP reduce fraud vs. old PIN? | p < 0.05 → OTP is significantly better. |
| ANOVA | Compare means of >2 groups | Do Ncell, NTC, and SmartCell have different outage rates? | p < 0.05 → At least one differs. |
| Chi-square | Test relationship between categorical variables | Are Khalti users more likely to be under 30? | p < 0.05 → Yes, 70% of users are <30. |
| Regression | Predict one variable from others | How does NEPSE index predict bank loan defaults? | R² = 0.6 → 60% of defaults explained by NEPSE. |
Worked Example: eSewa Fraud Reduction Hypothesis: "The new OTP system reduces fraud by 20%."
- Data: 300 users pre-OTP (mean fraud = 1.2 incidents), 300 post-OTP (mean = 0.5).
- Test: Paired t-test (same users before/after).
- Result: t = 4.2, p = 0.0001 → Reject null hypothesis (fraud reduced significantly).
- Conclusion: OTP works! But qualitative follow-up needed: "Why do users still get scammed?"
4. Data Visualization: Making Data Speak
Rule: "A picture is worth 1,000 data points." Poor visuals mislead—like Daraz hiding low ratings in tiny text.
When to Use Which Chart
| Chart Type | Best For | Nepalese Example | Avoid If... |
|---|---|---|---|
| Bar Chart | Compare categories (e.g., fraud by bank) | Ncell vs. NTC vs. SmartCell outage rates | Data is continuous (use histogram instead). |
| Line Graph | Trends over time (e.g., NEPSE index) | Monthly Khalti transactions (2020–2023) | Categories aren’t ordered (use bar chart). |
| Pie Chart | Parts of a whole (but rarely useful) | Reasons for Pathao driver quits (30% low pay, 20% traffic) | >5 categories (hard to read). |
| Histogram | Distribution of continuous data | Distribution of Daraz order delays | Data is categorical (use bar chart). |
| Scatter Plot | Relationships between two variables | NEPSE index vs. bank loan defaults | No clear pattern (add regression line). |
| Heatmap | Intensity in 2D (e.g., time + location) | Pathao ride demand by hour + district | Data isn’t spatial/temporal. |
Example: NEPSE Stock Volatility
Interpretation:
- Upper Band: Stock is overvalued (sell signal).
- Lower Band: Stock is undervalued (buy signal).
- Narrow Bands: Low volatility (e.g., NEPSE in 2021).
5. Software Tools for Data Analysis
| Tool | Best For | Nepalese Use Case | Learning Curve |
|---|---|---|---|
| Excel | Quick summaries, basic charts | Bank loan repayment tracking | Low |
| SPSS | Statistical tests (t-tests, ANOVA) | Ncell customer satisfaction surveys | Medium |
| R | Advanced stats, predictive modeling | NTC network failure prediction | High |
| Python (Pandas, SciPy) | Large datasets, automation | Daraz’s recommendation algorithm | Medium-High |
| NVivo | Qualitative analysis (themes, coding) | Khalti user feedback analysis | High |
Worked Example: Python for NEPSE Analysis
import pandas as pd
import matplotlib.pyplot as plt
# Load NEPSE data
data = pd.read_csv("nepse_index.csv")
data['Date'] = pd.to_datetime(data['Date'])
# Calculate moving average (50 days)
data['MA50'] = data['Close'].rolling(50).mean()
# Plot
plt.plot(data['Date'], data['Close'], label='NEPSE Index')
plt.plot(data['Date'], data['MA50'], label='50-Day MA', color='red')
plt.title("NEPSE Index with Moving Average")
plt.legend()
plt.show()
Output:
In the Real World
eSewa’s Fraud Detection
- Idea: Anomaly detection (statistical outliers).
- How: Uses z-scores to flag transactions where amount > mean + 3*SD. Example: A Rs. 50,000 transfer (mean = Rs. 2,000) triggers a review.
- Impact: Reduced fraud by 15% in 2022.
Pathao’s Ride Demand Heatmap
- Idea: Geospatial visualization (heatmaps).
- How: Overlays ride requests per km² with traffic data to predict surge pricing zones. Example: Heatmap shows high demand in Thapathali (2 PM–5 PM).
- Impact: Drivers earn 30% more in peak hours.
Ncell’s Customer Churn Prediction
- Idea: Logistic regression (predictive modeling).
- How: Trains a model on call duration, complaints, and usage to predict which users will switch. Example: Users with >3 complaints/month have 80% churn risk.
- Impact: Targeted discounts reduced churn by 12%.
Daraz’s Order Fulfillment Delays
- Idea: Queueing theory (probability of delays).
- How: Models order arrival rate (λ) vs. warehouse processing rate (μ). Example: If λ = 100 orders/hour and μ = 80, 20% of orders face delays.
- Impact: Added 2 warehouses in Kathmandu to balance λ and μ.
NEPSE’s Index Calculation
- Idea: Weighted average (statistical aggregation).
- How: Combines stock prices of top 10 companies with market caps as weights. Example: If NMB Bank (20% weight) rises 5%, it boosts NEPSE by 1%.
Exam Tip
Know the Difference: Always distinguish between quantitative (numbers, stats) and qualitative (text, themes). Examiners love questions like:
- "When would you use a t-test vs. thematic analysis?"
- Answer: "A t-test compares means (quantitative), while thematic analysis finds patterns in interview transcripts (qualitative)."
Data Cleaning is Critical: Expect 3–5 marks on cleaning steps. For example:
- Question: "How would you handle missing values in a survey on Khalti usage?"
- Answer:
- Missing = "No" (e.g., "Do you use Khalti for bills?" → Missing = "No").
- Impute mean for numerical data (e.g., "How often?").
- Exclude if >20% missing (but justify!).
Visuals = Marks: If asked to "analyze NEPSE data", always sketch a chart (even on paper). Example:
- Question: "Describe the trend in NEPSE from 2020–2023."
- Answer: *"A line graph shows a 2020 dip (-15%) due to COVID, followed by a 2021–2022 recovery (+30%), and 2023 volatility (±10%)."*
Software Shortcuts: Memorize these SPSS/R/Python commands for quick answers:
- SPSS:
Analyze > Compare Means > Independent-Samples T Test. - R:
t.test(group1 ~ group2, data = df). - Python:
from scipy import stats; stats.ttest_ind(group1, group2).
- SPSS:
Ethics Trap: Examiners test unethical data practices. Example:
- Question: "Is it ethical to exclude outliers in Ncell’s outage data?"
- Answer: "No—unless outliers are errors. If they’re valid (e.g., 2021 Kathmandu blackout), exclude them only if justified (e.g., ‘extreme event’). Otherwise, analyze separately."
Worked Example Formula: For any numerical question, follow:
- State the test (e.g., "We’ll use a chi-square test.").
- Show the formula (even if not calculated).
- Interpret the result (e.g., "p < 0.05 → relationship is significant.").
Final Visual Summary:
mindmap
root((Data Analysis))
Quantitative
Descriptive Stats
Mean/Median/Mode
Standard Deviation
Inferential Stats
t-test
ANOVA
Regression
Qualitative
Thematic Analysis
Content Analysis
Data Cleaning
Missing Values
Outliers
Duplicates
Visualization
Bar Charts
Line Graphs
Heatmaps
Tools
SPSS
R
PythonIn the real world
- eSewa uses quantitative analysis to detect fraud by flagging transactions with outliers (e.g., Rs. 50,000 transfers to unknown accounts) and missing values (incomplete KYC data).
- NTC applies qualitative analysis to customer complaints (e.g., coding themes like "signal drops" vs. "billing errors") to prioritize network upgrades in Kathmandu’s hilly districts.
- Pathao leverages predictive modeling (quantitative) to forecast driver shortages during Dashain by analyzing historical ride-demand heatmaps (visualization) from the previous year.
Based on the PU BE Computer (PU) syllabus for Research Fundamentals, unit 8.
Discussion
Loading…