Business Research MethodsUnit 811 min read
Data Processing & Analysis: Techniques, Tools & Interpretation
Unit 8 of Business Research Methods explores systematic data processing (coding, cleaning, tabulation) and statistical/qualitative analysis techniques (descriptive, inferential, thematic) used to derive actionable insights from raw research data, with real-world applications in Nepali business contexts.
Core Concepts
What is Data Processing and Analysis?
Data processing transforms raw data into meaningful information through:
- Coding: Assigning numerical/alphabetical values to qualitative responses (e.g., "Strongly Agree" → 5).
- Cleaning: Handling missing values, outliers, and inconsistencies.
- Tabulation: Organizing data into tables for pattern recognition.
graph LR
A["Raw Data"] --> B["Data Processing"]
B --> C["Coding"]
B --> D["Cleaning"]
B --> E["Tabulation"]
E --> F["Analyzed Data"]
F --> G["Insights/Reports"]Why it matters: Without proper processing, even the best-collected data becomes unusable (like a Daraz order system with corrupted customer records).
Data Processing Techniques
1. Coding
Definition: Converting qualitative data (e.g., survey responses) into quantitative form for analysis.
Example:
| Raw Response (Likert Scale) | Coded Value |
|---|---|
| Strongly Disagree | 1 |
| Disagree | 2 |
| Neutral | 3 |
| Agree | 4 |
| Strongly Agree | 5 |
Worked Example (Nepali Context): Problem: A PU student surveys 50 small business owners in Pokhara about their satisfaction with Ncell’s digital payment services (1=Very Dissatisfied to 5=Very Satisfied). Solution:
- Code responses as above.
- Calculate mean satisfaction score: .
- If , conclude "Moderate satisfaction" (use a 3-point scale: <2.5=Low, 2.5–3.5=Moderate, >3.5=High).
Real-World Tie:
- Khalti uses coded transaction data to detect fraud (e.g., "Failed" → 0, "Successful" → 1) and flag outliers for review.
2. Data Cleaning
Common Issues and Fixes:
| Issue | Example | Solution |
|---|---|---|
| Missing Data | "Age: —" in a customer survey | Impute (replace with mean/median) or exclude cases. |
| Outliers | A Daraz order with ₹99,999 value when others are <₹5,000 | Check for data entry errors; cap at 99th percentile. |
| Inconsistencies | "Gender: Male/Female/Other" | Standardize to binary (M/F) or add "Prefer not to say." |
Exam Tip: Always justify your cleaning method (e.g., "Excluded 3% missing responses to avoid bias").
3. Tabulation
Purpose: Organize data into tables to identify patterns, trends, or relationships.
Types of Tables:
- Frequency Distribution: Shows how often each response occurs.
Payment Method Frequency (n) Percentage (%) Khalti 30 60 E-sewa 15 30 Cash on Delivery 5 10 - Cross-Tabulation: Compares two variables (e.g., age group vs. preferred banking app).
Age Group Nabil App Global IME Other Total 18–30 12 28 5 45 31–50 20 10 2 32
Real-World Example:
- NTC uses cross-tabulation to analyze internet usage by region (e.g., "Lalitpur: 80% use mobile data; Kathmandu: 60% use fiber") to allocate infrastructure budgets.
Data Analysis Techniques
1. Descriptive Statistics
Tools: Measures of central tendency, dispersion, and shape.
- Central Tendency:
- Mean (): Average (sensitive to outliers).
- Median: Middle value (robust to outliers).
- Mode: Most frequent value.
- Dispersion:
- Range: Max − Min.
- Standard Deviation (): Average deviation from the mean.
Worked Example (Nepali Context): Scenario: A Chaudhary Group store in Thapathali records daily footfall for 30 days. Data: [250, 300, 280, 270, ..., 320] (sample data). Analysis:
- Mean footfall = customers/day.
- Standard deviation = 20 (shows moderate variability). Insight: "Store sees ~290 customers/day with ±20 variation; stock inventory for 310/day to cover peak demand."
2. Inferential Statistics
Purpose: Make predictions or inferences about a population from a sample. Key Techniques:
- Hypothesis Testing: Compare sample statistics to population parameters (covered in Unit 9).
- Confidence Intervals: Estimate population parameters with a margin of error.
- Example: "95% CI for average Daraz order value = ₹1,200 ± ₹50."
Real-World Example:
- Pathao uses inferential statistics to predict driver demand in Pokhara:
- Sample 20% of rides → estimate average ride duration = 12 minutes (95% CI: 11–13 minutes).
- Adjust surge pricing dynamically based on this interval.
3. Qualitative Data Analysis
Methods:
- Thematic Analysis: Identify recurring themes in open-ended responses.
- Example: Survey question: "What challenges do you face using eSewa?"
- Themes: "Slow transactions," "Lack of awareness," "Technical glitches."
- Example: Survey question: "What challenges do you face using eSewa?"
- Content Analysis: Quantify words/phrases in texts (e.g., count "complaint" mentions in customer reviews).
Visualization:
mindmap
root((Qualitative Analysis))
Thematic Analysis
Step 1: Transcribe Data
Step 2: Code Responses
Step 3: Identify Themes
Step 4: Validate Themes
Content Analysis
Step 1: Define Categories
Step 2: Count Occurrences
Step 3: Calculate FrequenciesCase Study: Nabil Bank Customer Feedback
- Data: 500 open-ended responses to "How can we improve our mobile banking app?"
- Themes:
- Technical Issues (30%): "App crashes frequently."
- User Experience (40%): "Navigation is confusing."
- Features (20%): "Need biometric login."
- Action: Bank prioritized app stability fixes and added a tutorial video.
4. Data Visualization
Purpose: Communicate insights clearly. Tools:
| Chart Type | Best For | Example (Nepali Context) |
|---|---|---|
| Bar Chart | Comparing categories | Sales of Himalayan Java coffee flavors. |
| Line Graph | Trends over time | Monthly NEPSE index movement. |
| Pie Chart | Proportions of a whole | Market share of Nepali banks (2023). |
| Histogram | Distribution of continuous data | Age distribution of Daraz customers. |
Worked Example: Scenario: A PU student analyzes NTC’s customer satisfaction scores (1–5) across 4 regions. Visualization:
Insight: "Kathmandu leads in satisfaction; Dhangadhi needs targeted service improvements."
Software Tools for Data Processing and Analysis
| Tool | Use Case | Nepali Example |
|---|---|---|
| SPSS | Statistical analysis, hypothesis testing | PU research projects on consumer behavior. |
| Excel | Basic tabulation, pivot tables | Nabil Bank’s monthly loan portfolio analysis. |
| R/Python | Advanced analytics, machine learning | Daraz’s demand forecasting models. |
| NVivo | Qualitative data analysis | NTC’s customer complaint thematic analysis. |
| Tableau/Power BI | Interactive dashboards | Khalti’s transaction trend visualizations. |
Common Pitfalls and Best Practices
Pitfalls:
- Ignoring Missing Data: Can bias results (e.g., excluding non-respondents who may be dissatisfied).
- Overlooking Outliers: A single ₹1,000,000 Daraz order can skew average order value.
- Misinterpreting Correlations: "Ice cream sales and drowning incidents rise in summer" ≠ causation.
Best Practices:
- Validate Data: Cross-check with secondary sources (e.g., verify survey responses with NRA’s economic reports).
- Document Steps: Keep a log of coding decisions, cleaning rules, and analysis choices.
- Triangulate: Use multiple methods (e.g., quantitative survey + qualitative interviews) for robustness.
In the Real World
Khalti’s Fraud Detection:
- Idea Used: Outlier detection in transaction data.
- How: Transactions >₹50,000 are flagged for manual review (assuming most users spend <₹10,000/month). Uses z-score to identify deviations from the mean.
Daraz’s Inventory Management:
- Idea Used: Time-series forecasting (descriptive + inferential stats).
- How: Analyzes past 12 months of sales data to predict demand for Diwali season. Uses moving averages to smooth trends and confidence intervals to set safety stock levels.
NTC’s Network Expansion:
- Idea Used: Cross-tabulation and geographic analysis.
- How: Compares internet usage by district (e.g., "Lalitpur: 90% mobile data; Sindhupalchowk: 70% dial-up") to prioritize fiber vs. 4G upgrades. Visualized in choropleth maps.
Exam Tip
What Examiners Look For:
- Structure:
- Clearly label steps (e.g., "Step 1: Coding → Step 2: Cleaning").
- Use headings like "Descriptive Analysis" and "Inferential Analysis."
- Justification:
- Always explain why you chose a method (e.g., "Used median instead of mean because data was skewed").
- Visuals:
- Include one table or chart per question (even if not asked, show you can visualize data).
- Label axes and units (e.g., "Y-axis: Number of Customers (n)").
- Real-World Link:
- Tie answers to Nepali businesses (e.g., "Like Nabil Bank’s loan approval process, we should use a 70% confidence interval for risk assessment").
- Common Mistakes to Avoid:
- Forgetting to define terms (e.g., "Standard deviation is...").
- Ignoring units (e.g., "Mean age = 30" vs. "Mean age = 30 years").
- Overcomplicating (e.g., using regression when a simple mean suffices).
Sample Exam Question and Answer:
Question: "A researcher collects data on customer satisfaction with eSewa’s new USSD service from 100 users in Kathmandu. The responses are on a Likert scale (1–5). Show how you would process and analyze this data to present insights to eSewa’s management."
Model Answer:
Data Processing:
- Coding: Convert Likert responses to numerical values (1–5).
- Cleaning: Exclude 5 missing responses; recode "Don’t know" (3 cases) as neutral (3).
- Tabulation:
Score Frequency Percentage 1 8 8.4% 2 12 12.6% 3 25 26.3% 4 38 40.0% 5 17 18.1%
Descriptive Analysis:
- Mean satisfaction = .
- Standard deviation = 1.1 (moderate spread).
- Insight: "Average satisfaction is ‘Agree’ (3.6/5), but 8% are very dissatisfied (score=1)."
Visualization:
Recommendation:
- Address pain points for the 20% scoring 1–2 (e.g., survey them for common issues).
- Highlight success: 58% scored 4–5 ("Very Satisfied/Satisfied").
Why This Scores Full Marks:
- Shows complete processing steps (coding, cleaning, tabulation).
- Uses descriptive stats with interpretation.
- Includes a clear visualization.
- Ends with actionable insights (not just numbers).
Based on the PU BBA (PU) syllabus for Business Research Methods, unit 8.
Discussion
Loading…