CMP422 Data Science and Analytics

Data Science and AnalyticsUnit 313 min read

Exploratory Data Analysis: Techniques, Tools & Insights

Unit 3 of Data Science and Analytics explores how to summarize, visualize, and interpret raw data to uncover patterns, anomalies, and trends before formal modeling. This note covers key techniques (univariate/multivariate analysis, statistical summaries, and visualization), tools (Python/R libraries), and real-world ap

Core Concepts

1. What is Exploratory Data Analysis (EDA)?

EDA is the iterative process of examining data to understand its structure, detect patterns, and identify outliers or anomalies before applying advanced modeling techniques. It bridges raw data and statistical modeling by answering:

  • What does the data look like?
  • Are there missing values or errors?
  • What relationships exist between variables?

Why is EDA critical?

  • Avoids blind modeling: Garbage in → garbage out (GIGO). EDA ensures data quality.
  • Guides feature engineering: Helps select relevant variables for modeling.
  • Reveals hidden insights: Often uncovers trends not obvious in raw tables.

2. Key Steps in EDA

flowchart TD
    A["1. Data Understanding"] --> B["2. Data Cleaning"]
    B --> C["3. Univariate Analysis"]
    C --> D["4. Bivariate/Multivariate Analysis"]
    D --> E["5. Hypothesis Generation"]
    E --> F["6. Visualization & Reporting"]

Step 1: Data Understanding

  • Goal: Get familiar with the dataset’s shape, size, and structure.
  • Actions:
    • Read the first/last 5 rows (df.head(), df.tail() in Python).
    • Check data types (df.dtypes), missing values (df.isnull().sum()), and descriptive statistics (df.describe()).
    • Example: For a dataset of Nepal’s traffic accidents (2018–2023), you’d first check:
      • Number of records (rows) vs. features (columns).
      • Data types (e.g., date, int, str).
      • Missing values in injury_severity or weather_conditions.

Step 2: Data Cleaning

  • Goal: Handle missing values, duplicates, and inconsistencies.
  • Common Issues & Fixes:
    Issue Example Solution
    Missing values NaN in age of accident victims Impute (mean/median) or drop rows.
    Duplicates Same accident recorded twice Drop duplicates (df.drop_duplicates()).
    Outliers Age = 150 years Cap at 99th percentile or investigate.
    Incorrect data types date stored as string Convert to datetime (pd.to_datetime()).

Worked Example: Cleaning Nepal Traffic Data

import pandas as pd
# Load data
df = pd.read_csv("nepal_traffic_accidents.csv")
# Check missing values
print(df.isnull().sum())
# Drop rows with missing 'injury_severity'
df_clean = df.dropna(subset=['injury_severity'])
# Convert 'accident_date' to datetime
df_clean['accident_date'] = pd.to_datetime(df_clean['accident_date'])

3. Univariate Analysis

Definition: Examining one variable at a time to summarize its distribution, central tendency, and dispersion.

Key Metrics

Metric Python/R Function Interpretation
Central Tendency mean(), median(), mode() Typical value (e.g., average age of victims).
Dispersion std(), var(), range() Spread of data (e.g., variance in accident counts by district).
Shape skew(), kurtosis() Symmetry (skewness) and tailedness.

Visualizations for Univariate Data

Data Type Best Visualization Example
Numerical (continuous) Histogram, Boxplot Distribution of victim ages.
Numerical (discrete) Bar plot, Pie chart Number of accidents by vehicle type.
Categorical Count plot, Frequency table Accidents by district (Kathmandu vs. Pokhara).

Worked Example: Univariate Analysis of Accident Severity

import seaborn as sns
import matplotlib.pyplot as plt
# Plot distribution of injury severity
sns.countplot(data=df_clean, x='injury_severity')
plt.title("Distribution of Injury Severity in Nepal (2018–2023)")
plt.show()

Output: ![Count plot showing "minor," "moderate," and "severe" injuries with severe being the least frequent.]


4. Bivariate and Multivariate Analysis

Definition:

  • Bivariate: Relationship between two variables (e.g., age vs. injury severity).
  • Multivariate: Relationships among three or more variables (e.g., age, district, and weather).

Key Techniques

Technique Use Case Python/R Tool
Correlation matrix Linear relationships (e.g., temperature vs. accidents) df.corr(), corrplot
Scatter plots Relationship between two numerical variables sns.scatterplot()
Grouped bar plots Comparison across categories (e.g., accidents by district and season) sns.barplot(hue=...)
Heatmaps Correlation patterns in multivariate data sns.heatmap()

Worked Example: Bivariate Analysis

Question: Does accident severity vary by district?

# Cross-tabulation of injury severity by district
cross_tab = pd.crosstab(df_clean['district'], df_clean['injury_severity'])
print(cross_tab)

Output:

District Minor Moderate Severe
Kathmandu 1200 850 300
Pokhara 450 300 120
Lalitpur 600 400 150

Visualization:

sns.boxplot(data=df_clean, x='district', y='injury_severity_score')
plt.title("Injury Severity by District (1=minor, 3=severe)")

Insight: Kathmandu has higher median severity scores, suggesting worse road conditions or higher-speed traffic.


5. Hypothesis Generation

EDA often leads to testable hypotheses for further analysis or modeling. Examples:

  1. Hypothesis: "Accidents increase during monsoon (June–September)."
    • Test: Compare accident counts across seasons using a chi-square test.
  2. Hypothesis: "Older drivers (>50 years) have more severe injuries."
    • Test: Use a t-test or ANOVA to compare injury severity by age group.

In the Real World

  1. eSewa (Nepal)

    • Idea Used: Multivariate analysis of transaction data.
    • How: eSewa analyzes spending patterns (e.g., time of day, location, transaction amount) to detect fraudulent activities. For example, a sudden spike in transactions from a single IP address triggers an alert.
    • EDA Step: They use heatmaps of transaction frequencies by hour/day to identify unusual clusters.
  2. Pathao (Ride-Hailing App)

    • Idea Used: Bivariate analysis of ride demand vs. weather.
    • How: Pathao’s data team plots scatter plots of ride requests against rainfall data to predict surge pricing during monsoons. In Kathmandu, they found a 30% increase in ride demand during heavy rain.
    • Worked Example: If Pathao’s EDA shows:
      • Correlation (r = 0.7) between rain intensity and ride requests,
      • They might automate surge pricing during monsoon seasons.
  3. Nepal Electricity Authority (NEA)

    • Idea Used: Univariate analysis of power outage data.
    • How: NEA uses boxplots of outage durations by district to identify regions with chronic power cuts. For example, EDA might reveal that Dharan has 2x longer outages than Kathmandu, prompting targeted infrastructure upgrades.

Tools for EDA

Tool Language Key Features
Pandas Python Data cleaning, describe(), groupby()
NumPy Python Numerical operations, arrays
Seaborn/Matplotlib Python Advanced visualizations
ggplot2 R Grammar of graphics for plots
Tableau/Power BI GUI Drag-and-drop dashboards

Common Pitfalls in EDA

  1. Ignoring Missing Data: Assuming missing values are random can bias results.
    • Fix: Use df.dropna() or imputation (e.g., SimpleImputer in scikit-learn).
  2. Overlooking Outliers: Outliers can skew metrics like mean/standard deviation.
    • Fix: Use median for skewed data or IQR (Interquartile Range) to cap outliers.
  3. Choosing the Wrong Visualization:
    • Bad: Pie charts for continuous data (e.g., ages of victims).
    • Good: Histograms or boxplots.
  4. Not Documenting Assumptions: EDA findings depend on data quality. Always note:
    • How missing values were handled.
    • Why certain outliers were removed.

Exam Tip

What Examiners Look For

  1. Structured Approach: Start with data understanding, then cleaning, followed by univariate → bivariate → multivariate analysis. Examiners reward a logical flow.
  2. Visual Evidence: Always show plots (even in written exams, describe them clearly). For example:
    • "The histogram of accident counts reveals a right-skewed distribution, with most districts having <50 accidents but a few (e.g., Kathmandu) exceeding 100."
  3. Insights Over Raw Data: Don’t just list statistics—interpret them. Example:
    • ❌ "The mean age is 35."
    • ✅ "The mean victim age of 35 suggests most accidents involve working-age adults, implying high economic impact."
  4. Tool Mention: If using Python/R, name the specific functions (e.g., sns.boxplot(), df.corr()). Examiners test both conceptual and technical knowledge.
  5. Real-World Connection: Relate your EDA to a Nepalese context (e.g., traffic data, e-commerce, healthcare). Example:
    • "Like Daraz’s analysis of customer purchase patterns, our EDA of accident data can help NTA prioritize road safety measures in high-risk districts."

Common Exam Questions

Question Type How to Answer
"Describe the EDA process." Follow the 6-step flowchart above. Include one visualization per step.
"How would you analyze [dataset]?" Start with df.head(), then clean → univariate → bivariate. End with insights.
"What tools would you use?" Name Pandas + Seaborn (Python) or ggplot2 (R). Justify with features.
"Interpret this plot:" Describe shape, outliers, and trends. Example: "The boxplot shows Lalitpur has lower median severity than Kathmandu, suggesting better emergency response."

Practice Problem

Dataset: Nepal’s Nepse stock prices (2023) with columns:

  • date, open, high, low, close, volume, sector.

Task: Perform EDA to answer:

  1. What is the distribution of daily trading volume? (Use a histogram.)
  2. Is there a correlation between stock price and trading volume? (Use a scatter plot + correlation coefficient.)
  3. Which sector has the highest average daily return? (Use grouped bar plots.)

Solution Outline:

  1. Univariate:

    sns.histplot(df['volume'], bins=30, kde=True)
    plt.title("Distribution of Daily Trading Volume (Nepse 2023)")
    
    • Insight: Right-skewed distribution; most days have <500,000 shares traded, but some spike to 2M.
  2. Bivariate:

    sns.scatterplot(data=df, x='close', y='volume', hue='sector')
    plt.title("Stock Price vs. Trading Volume by Sector")
    
    • Insight: Positive correlation (r = 0.45); higher prices attract more volume in sectors like finance and hydro.
  3. Multivariate:

    df.groupby('sector')['close'].mean().sort_values().plot(kind='barh')
    
    • Insight: Hydro sector has the highest average closing price, suggesting strong investor confidence.

Key Formulas to Remember

Concept Formula When to Use
Mean Central tendency for symmetric data.
Standard Deviation Measure of spread.
Correlation (Pearson) Linear relationships between two variables.
Coefficient of Variation Compare dispersion across datasets.

Final Checklist for EDA

Before submitting your analysis (or exam answer), ensure you’ve covered:

  1. Data Understanding: Shape, types, missing values.
  2. Cleaning: Handled duplicates, outliers, and incorrect formats.
  3. Univariate: Summary stats + at least one plot (histogram/boxplot).
  4. Bivariate/Multivariate: At least two relationships explored (e.g., scatter plot + correlation).
  5. Insights: 3–4 actionable takeaways (e.g., "Prioritize road safety in Kathmandu").
  6. Tools: Named Pandas/Seaborn or R/ggplot2 with specific functions.

Based on the PU BE Computer (PU) syllabus for Data Science and Analytics (CMP422), unit 3.

Discussion

Loading…