RCH201 Business Research Methods

Business Research MethodsUnit 812 min read

Data Processing & Analysis: Cleaning, Coding, Tabulation, Analysis & Interpretation

Unit 8 of Business Research Methods covers transforming raw data into meaningful insights through cleaning, coding, tabulation, statistical analysis, and interpretation—essential for drawing valid conclusions in business research.

TAKEAWAYS

  • Data processing involves cleaning, coding, and tabulating raw data to ensure accuracy and consistency before analysis.
  • Descriptive statistics (mean, median, mode, standard deviation) summarize data trends, while inferential statistics (hypothesis testing, regression) help generalize findings.
  • SPSS, Excel, and R are common tools for data analysis, each suited for different research scales and complexities.
  • Data visualization (charts, graphs, tables) enhances clarity and aids decision-making in business reports.
  • Interpretation links statistical results to research objectives, ensuring findings are actionable and contextually relevant.

1. Data Processing: From Raw Data to Usable Information

Data processing is the systematic transformation of raw data into a structured format for analysis. It includes three key steps:

1.1 Data Cleaning (Data Sanitization)

  • Definition: Removing or correcting errors, inconsistencies, and missing values in raw data to improve accuracy.
  • Common Issues:
    • Incomplete data (missing responses in surveys).
    • Inconsistent data (different formats for dates, e.g., "2023-10-05" vs. "05/10/2023").
    • Outliers (extreme values that distort analysis, e.g., a customer spending ₹10,000 in a ₹500 average basket).
    • Duplicates (same respondent filling the survey twice).

How to Clean Data?

flowchart TD
    A["Raw Data"] --> B["Check for Missing Values"]
    B --> C{"Missing Values?"}
    C -->|"Yes"| D["Impute or Remove"]
    C -->|"No"| E["Check for Inconsistencies"]
    E --> F["Standardize Formats"]
    F --> G["Remove Outliers"]
    G --> H["Final Clean Dataset"]

Example (Nepali Context): In a Nepal Stock Exchange (NEPSE) survey on investor behavior, suppose 10% of responses have missing ages. You might:

  • Impute missing ages using the median age of the sample.
  • Remove responses if >20% of data is missing (to avoid bias).

1.2 Data Coding

  • Definition: Assigning numerical or categorical values to qualitative responses to enable statistical analysis.
  • Types of Coding:
    Type Example Purpose
    Numerical Age (25, 30, 35) Quantitative analysis
    Categorical Gender (1=Male, 2=Female, 3=Other) Grouping responses
    Ordinal Customer satisfaction (1=Poor, 2=Fair, 3=Good, 4=Excellent) Ranking responses
    Binary "Did you use Daraz in the last month?" (1=Yes, 0=No) Simple yes/no analysis

Example (Pathao Driver Survey): If a question asks: "How often do you use Pathao’s ‘Earn More’ feature?" Responses:

  • Never (1)
  • Rarely (2)
  • Sometimes (3)
  • Often (4)
  • Always (5)

This converts qualitative answers into quantifiable data for analysis.

1.3 Data Tabulation

  • Definition: Organizing cleaned and coded data into tables (frequency distributions, cross-tabulations) for easier interpretation.
  • Types of Tables:
    • Frequency Distribution: Shows how often each response occurs.
    • Cross-Tabulation: Examines relationships between two variables (e.g., age vs. Khalti usage frequency).

Example (NTC Customer Satisfaction Survey):

Service Quality Rating Frequency Percentage
Poor (1) 15 5%
Fair (2) 40 14%
Good (3) 120 42%
Excellent (4) 90 32%
Total 265 100%

2. Data Analysis: Turning Data into Insights

Data analysis involves applying statistical techniques to interpret patterns, trends, and relationships in the data.

2.1 Descriptive Statistics

Summarizes data using measures of central tendency and dispersion.

Measure Formula When to Use
Mean Average performance (e.g., average Daraz order value).
Median Middle value in ordered data Best for skewed data (e.g., income distribution in Kathmandu).
Mode Most frequent value Identifying most common response (e.g., preferred payment method in eSewa).
Standard Deviation Measuring variability (e.g., fluctuation in NEPSE stock prices).

Example (Bank Loan Interest Calculation): Suppose Nabil Bank collects data on loan repayment times (in months): Data: 12, 18, 24, 30, 36

  • Mean = months
  • Median = 24 months (middle value)
  • Mode = None (all values unique)
  • Standard Deviation ≈ 8.4 months (shows repayment times vary significantly).

2.2 Inferential Statistics

Uses sample data to make predictions or inferences about a larger population.

Key Techniques:

  • Hypothesis Testing (Unit 9 covers this in detail, but briefly):
    • Null Hypothesis (H₀): No effect (e.g., "Khalti’s new feature does not increase user retention").
    • Alternative Hypothesis (H₁): There is an effect.
    • p-value: Probability of observing data if H₀ is true. If p < 0.05, reject H₀.
  • Regression Analysis: Predicts relationships (e.g., "How does advertising spend affect Daraz sales?").
  • ANOVA: Compares means across groups (e.g., "Do different age groups respond differently to NTC’s new tariffs?").

Example (WhatsApp Business Usage): A study collects data on small business owners’ WhatsApp Business usage:

  • Independent Variable (X): Hours spent daily on WhatsApp Business.
  • Dependent Variable (Y): Monthly sales increase (%). A linear regression might show: → For every extra hour spent, sales increase by 2.5%.

3. Data Visualization: Making Data Speak

Visuals simplify complex data and highlight trends.

3.1 Common Charts & Graphs

Chart Type Best For Example
Bar Chart Comparing discrete categories (e.g., market share of banks in Nepal). Nabil Bank (30%), Global IME (25%), Standard Chartered (15%).
Pie Chart Showing proportions (e.g., payment methods in eSewa). Credit Card (40%), Mobile Banking (35%), Cash (25%).
Line Graph Trends over time (e.g., NEPSE index from 2019–2023). Monthly stock price movements.
Histogram Distribution of continuous data (e.g., customer ages using Daraz). Age groups: 18–25 (30%), 26–35 (45%), 36+ (25%).
Scatter Plot Relationships between two variables (e.g., ad spend vs. sales). X-axis: Ad spend (₹), Y-axis: Sales (₹).

Example (NTC Revenue Growth):

graph LR
    A["2019"] -->|"₹50B"| B["2020"]
    B -->|"₹60B"| C["2021"]
    C -->|"₹75B"| D["2022"]
    D -->|"₹90B"| E["2023"]

Interpretation: NTC’s revenue grew steadily, with a sharp increase in 2022 due to 5G rollout.

3.2 Tools for Visualization

  • Excel/Google Sheets: Quick charts for basic analysis.
  • SPSS: Advanced statistical visualizations.
  • Tableau/Power BI: Interactive dashboards for business reporting.

4. Data Interpretation: Linking Results to Research Objectives

Interpretation explains why findings matter in the context of the research question.

Steps:

  1. Compare with Hypotheses: Did the data support your initial assumptions?
  2. Contextualize: Relate findings to industry trends (e.g., "Low Khalti usage among rural users aligns with digital divide challenges").
  3. Recommend Actions: Suggest improvements (e.g., "NTC should offer cheaper data plans for low-income users").

Example (Himalayan Java Customer Feedback):

  • Finding: 60% of customers prefer online orders over in-store.
  • Interpretation: Digital adoption is high, but in-store experience may lack appeal due to long queues.
  • Recommendation: Optimize online ordering with faster delivery (e.g., partner with Pathao for last-mile delivery).

In the Real World

  1. eSewa’s Payment Data Analysis

    • Idea Used: Data cleaning and tabulation to identify fraudulent transactions.
    • How? eSewa flags outliers (e.g., a ₹50,000 transfer at 3 AM) for manual review, reducing fraud by 30%.
  2. Daraz’s Customer Segmentation

    • Idea Used: Cross-tabulation and regression analysis to group buyers by demographics and spending habits.
    • How? Daraz uses age vs. purchase frequency to tailor promotions (e.g., discounts for 18–25-year-olds on electronics).
  3. NTC’s Network Performance Monitoring

    • Idea Used: Descriptive statistics (mean download speed, standard deviation) to identify slow zones.
    • How? NTC’s dashboard shows that Kathmandu’s average speed is 40 Mbps (σ=15 Mbps), prompting infrastructure upgrades in low-performing areas.

Exam Tip

  1. Data Cleaning Questions:

    • Expect scenario-based questions (e.g., "How would you handle missing values in a survey on Pathao driver earnings?").
    • Key Points to Mention:
      • Imputation methods (mean/median for numerical, mode for categorical).
      • Removal criteria (e.g., "delete if >30% data missing").
      • Standardization (e.g., converting all dates to YYYY-MM-DD).
  2. Data Analysis Questions:

    • Calculate measures: Always show formulas (e.g., mean, standard deviation).
    • Interpret results: Link statistics to business implications (e.g., "High standard deviation in NEPSE stock prices suggests volatility—investors should diversify").
    • Tool usage: Mention Excel for small datasets, SPSS for large-scale surveys, or R for advanced regression.
  3. Visualization Questions:

    • Choose the right chart: Match the data type (e.g., pie chart for proportions, line graph for trends).
    • Label axes clearly: Avoid vague titles like "Sales Data"—specify units (e.g., "Monthly Sales (₹ in millions)").
    • Highlight trends: Use annotations (e.g., "Peak sales in June due to Dashain").
  4. Interpretation Questions:

    • Structure your answer:
      1. State the finding (e.g., "65% of Nabil Bank customers prefer online banking").
      2. Explain why it matters (e.g., "Indicates high digital adoption but may overlook elderly users").
      3. Recommend solutions (e.g., "Offer multichannel support for non-tech-savvy customers").

Common Pitfalls to Avoid:

  • Ignoring units (e.g., reporting mean age as "25" instead of "25 years").
  • Overcomplicating visuals (stick to 1–2 variables per chart).
  • Misinterpreting correlation as causation (e.g., "More ice cream sales → more drownings" is correlation, not cause).

Case Study: Chaudhary Group’s Supply Chain Optimization

Problem: Chaudhary Group (owners of Daraz, Mega Mart) faced inefficiencies in inventory management due to inconsistent sales data.

Solution:

  1. Data Cleaning:
    • Removed duplicate orders and corrected misspelled product names.
    • Imputed missing stock levels using the previous day’s data.
  2. Data Analysis:
    • Used regression analysis to predict demand based on seasonality (e.g., rice sales spike before Dashain).
    • Calculated standard deviation to identify high-variability products (e.g., electronics).
  3. Visualization:
    • Created a heatmap of sales by region and product category to identify hotspots.
  4. Interpretation:
    • Found that online orders grew 40% YoY, but rural areas lagged due to poor logistics.
    • Recommendation: Expanded delivery partnerships with Pathao in Tier-2 cities.

Result: Reduced overstocking by 20% and improved delivery times by 30%.


Final Checklist for Full Marks: ✅ Definitions: Clearly define terms (e.g., "Data coding is the process of converting qualitative data into quantitative form"). ✅ Examples: Use Nepali companies (NTC, Nabil Bank, Daraz) and global tools (Excel, SPSS). ✅ Visuals: Include at least 3 diagrams (flowchart for cleaning, table for statistics, chart for trends). ✅ Real-World Links: Connect theory to eSewa, Khalti, or NEPSE in examples. ✅ Exam Strategy: Practice calculating mean/median and interpreting p-values under time pressure.

Based on the TU BITM syllabus for Business Research Methods (RCH201), unit 8.

Discussion

Loading…