Data Science and AnalyticsUnit 313 min read
Exploratory Data Analysis: Techniques, Tools & Insights
Unit 3 of Data Science and Analytics explores how to summarize, visualize, and interpret raw data to uncover patterns, anomalies, and trends before formal modeling. This note covers key techniques (univariate/multivariate analysis, statistical summaries, and visualization), tools (Python/R libraries), and real-world ap
Core Concepts
1. What is Exploratory Data Analysis (EDA)?
EDA is the iterative process of examining data to understand its structure, detect patterns, and identify outliers or anomalies before applying advanced modeling techniques. It bridges raw data and statistical modeling by answering:
- What does the data look like?
- Are there missing values or errors?
- What relationships exist between variables?
Why is EDA critical?
- Avoids blind modeling: Garbage in → garbage out (GIGO). EDA ensures data quality.
- Guides feature engineering: Helps select relevant variables for modeling.
- Reveals hidden insights: Often uncovers trends not obvious in raw tables.
2. Key Steps in EDA
flowchart TD
A["1. Data Understanding"] --> B["2. Data Cleaning"]
B --> C["3. Univariate Analysis"]
C --> D["4. Bivariate/Multivariate Analysis"]
D --> E["5. Hypothesis Generation"]
E --> F["6. Visualization & Reporting"]Step 1: Data Understanding
- Goal: Get familiar with the dataset’s shape, size, and structure.
- Actions:
- Read the first/last 5 rows (
df.head(),df.tail()in Python). - Check data types (
df.dtypes), missing values (df.isnull().sum()), and descriptive statistics (df.describe()). - Example: For a dataset of Nepal’s traffic accidents (2018–2023), you’d first check:
- Number of records (rows) vs. features (columns).
- Data types (e.g.,
date,int,str). - Missing values in
injury_severityorweather_conditions.
- Read the first/last 5 rows (
Step 2: Data Cleaning
- Goal: Handle missing values, duplicates, and inconsistencies.
- Common Issues & Fixes:
Issue Example Solution Missing values NaNinageof accident victimsImpute (mean/median) or drop rows. Duplicates Same accident recorded twice Drop duplicates ( df.drop_duplicates()).Outliers Age = 150 years Cap at 99th percentile or investigate. Incorrect data types datestored as stringConvert to datetime(pd.to_datetime()).
Worked Example: Cleaning Nepal Traffic Data
import pandas as pd
# Load data
df = pd.read_csv("nepal_traffic_accidents.csv")
# Check missing values
print(df.isnull().sum())
# Drop rows with missing 'injury_severity'
df_clean = df.dropna(subset=['injury_severity'])
# Convert 'accident_date' to datetime
df_clean['accident_date'] = pd.to_datetime(df_clean['accident_date'])
3. Univariate Analysis
Definition: Examining one variable at a time to summarize its distribution, central tendency, and dispersion.
Key Metrics
| Metric | Python/R Function | Interpretation |
|---|---|---|
| Central Tendency | mean(), median(), mode() |
Typical value (e.g., average age of victims). |
| Dispersion | std(), var(), range() |
Spread of data (e.g., variance in accident counts by district). |
| Shape | skew(), kurtosis() |
Symmetry (skewness) and tailedness. |
Visualizations for Univariate Data
| Data Type | Best Visualization | Example |
|---|---|---|
| Numerical (continuous) | Histogram, Boxplot | Distribution of victim ages. |
| Numerical (discrete) | Bar plot, Pie chart | Number of accidents by vehicle type. |
| Categorical | Count plot, Frequency table | Accidents by district (Kathmandu vs. Pokhara). |
Worked Example: Univariate Analysis of Accident Severity
import seaborn as sns
import matplotlib.pyplot as plt
# Plot distribution of injury severity
sns.countplot(data=df_clean, x='injury_severity')
plt.title("Distribution of Injury Severity in Nepal (2018–2023)")
plt.show()
Output: ![Count plot showing "minor," "moderate," and "severe" injuries with severe being the least frequent.]
4. Bivariate and Multivariate Analysis
Definition:
- Bivariate: Relationship between two variables (e.g., age vs. injury severity).
- Multivariate: Relationships among three or more variables (e.g., age, district, and weather).
Key Techniques
| Technique | Use Case | Python/R Tool |
|---|---|---|
| Correlation matrix | Linear relationships (e.g., temperature vs. accidents) | df.corr(), corrplot |
| Scatter plots | Relationship between two numerical variables | sns.scatterplot() |
| Grouped bar plots | Comparison across categories (e.g., accidents by district and season) | sns.barplot(hue=...) |
| Heatmaps | Correlation patterns in multivariate data | sns.heatmap() |
Worked Example: Bivariate Analysis
Question: Does accident severity vary by district?
# Cross-tabulation of injury severity by district
cross_tab = pd.crosstab(df_clean['district'], df_clean['injury_severity'])
print(cross_tab)
Output:
| District | Minor | Moderate | Severe |
|---|---|---|---|
| Kathmandu | 1200 | 850 | 300 |
| Pokhara | 450 | 300 | 120 |
| Lalitpur | 600 | 400 | 150 |
Visualization:
sns.boxplot(data=df_clean, x='district', y='injury_severity_score')
plt.title("Injury Severity by District (1=minor, 3=severe)")
Insight: Kathmandu has higher median severity scores, suggesting worse road conditions or higher-speed traffic.
5. Hypothesis Generation
EDA often leads to testable hypotheses for further analysis or modeling. Examples:
- Hypothesis: "Accidents increase during monsoon (June–September)."
- Test: Compare accident counts across seasons using a chi-square test.
- Hypothesis: "Older drivers (>50 years) have more severe injuries."
- Test: Use a t-test or ANOVA to compare injury severity by age group.
In the Real World
eSewa (Nepal)
- Idea Used: Multivariate analysis of transaction data.
- How: eSewa analyzes spending patterns (e.g., time of day, location, transaction amount) to detect fraudulent activities. For example, a sudden spike in transactions from a single IP address triggers an alert.
- EDA Step: They use heatmaps of transaction frequencies by hour/day to identify unusual clusters.
Pathao (Ride-Hailing App)
- Idea Used: Bivariate analysis of ride demand vs. weather.
- How: Pathao’s data team plots scatter plots of ride requests against rainfall data to predict surge pricing during monsoons. In Kathmandu, they found a 30% increase in ride demand during heavy rain.
- Worked Example: If Pathao’s EDA shows:
- Correlation (r = 0.7) between rain intensity and ride requests,
- They might automate surge pricing during monsoon seasons.
Nepal Electricity Authority (NEA)
- Idea Used: Univariate analysis of power outage data.
- How: NEA uses boxplots of outage durations by district to identify regions with chronic power cuts. For example, EDA might reveal that Dharan has 2x longer outages than Kathmandu, prompting targeted infrastructure upgrades.
Tools for EDA
| Tool | Language | Key Features |
|---|---|---|
| Pandas | Python | Data cleaning, describe(), groupby() |
| NumPy | Python | Numerical operations, arrays |
| Seaborn/Matplotlib | Python | Advanced visualizations |
| ggplot2 | R | Grammar of graphics for plots |
| Tableau/Power BI | GUI | Drag-and-drop dashboards |
Common Pitfalls in EDA
- Ignoring Missing Data: Assuming missing values are random can bias results.
- Fix: Use
df.dropna()or imputation (e.g.,SimpleImputerin scikit-learn).
- Fix: Use
- Overlooking Outliers: Outliers can skew metrics like mean/standard deviation.
- Fix: Use median for skewed data or IQR (Interquartile Range) to cap outliers.
- Choosing the Wrong Visualization:
- Bad: Pie charts for continuous data (e.g., ages of victims).
- Good: Histograms or boxplots.
- Not Documenting Assumptions: EDA findings depend on data quality. Always note:
- How missing values were handled.
- Why certain outliers were removed.
Exam Tip
What Examiners Look For
- Structured Approach: Start with data understanding, then cleaning, followed by univariate → bivariate → multivariate analysis. Examiners reward a logical flow.
- Visual Evidence: Always show plots (even in written exams, describe them clearly). For example:
- "The histogram of accident counts reveals a right-skewed distribution, with most districts having <50 accidents but a few (e.g., Kathmandu) exceeding 100."
- Insights Over Raw Data: Don’t just list statistics—interpret them. Example:
- ❌ "The mean age is 35."
- ✅ "The mean victim age of 35 suggests most accidents involve working-age adults, implying high economic impact."
- Tool Mention: If using Python/R, name the specific functions (e.g.,
sns.boxplot(),df.corr()). Examiners test both conceptual and technical knowledge. - Real-World Connection: Relate your EDA to a Nepalese context (e.g., traffic data, e-commerce, healthcare). Example:
- "Like Daraz’s analysis of customer purchase patterns, our EDA of accident data can help NTA prioritize road safety measures in high-risk districts."
Common Exam Questions
| Question Type | How to Answer |
|---|---|
| "Describe the EDA process." | Follow the 6-step flowchart above. Include one visualization per step. |
| "How would you analyze [dataset]?" | Start with df.head(), then clean → univariate → bivariate. End with insights. |
| "What tools would you use?" | Name Pandas + Seaborn (Python) or ggplot2 (R). Justify with features. |
| "Interpret this plot:" | Describe shape, outliers, and trends. Example: "The boxplot shows Lalitpur has lower median severity than Kathmandu, suggesting better emergency response." |
Practice Problem
Dataset: Nepal’s Nepse stock prices (2023) with columns:
date,open,high,low,close,volume,sector.
Task: Perform EDA to answer:
- What is the distribution of daily trading volume? (Use a histogram.)
- Is there a correlation between stock price and trading volume? (Use a scatter plot + correlation coefficient.)
- Which sector has the highest average daily return? (Use grouped bar plots.)
Solution Outline:
Univariate:
sns.histplot(df['volume'], bins=30, kde=True) plt.title("Distribution of Daily Trading Volume (Nepse 2023)")- Insight: Right-skewed distribution; most days have <500,000 shares traded, but some spike to 2M.
Bivariate:
sns.scatterplot(data=df, x='close', y='volume', hue='sector') plt.title("Stock Price vs. Trading Volume by Sector")- Insight: Positive correlation (r = 0.45); higher prices attract more volume in sectors like finance and hydro.
Multivariate:
df.groupby('sector')['close'].mean().sort_values().plot(kind='barh')- Insight: Hydro sector has the highest average closing price, suggesting strong investor confidence.
Key Formulas to Remember
| Concept | Formula | When to Use |
|---|---|---|
| Mean | Central tendency for symmetric data. | |
| Standard Deviation | Measure of spread. | |
| Correlation (Pearson) | Linear relationships between two variables. | |
| Coefficient of Variation | Compare dispersion across datasets. |
Final Checklist for EDA
Before submitting your analysis (or exam answer), ensure you’ve covered:
- Data Understanding: Shape, types, missing values.
- Cleaning: Handled duplicates, outliers, and incorrect formats.
- Univariate: Summary stats + at least one plot (histogram/boxplot).
- Bivariate/Multivariate: At least two relationships explored (e.g., scatter plot + correlation).
- Insights: 3–4 actionable takeaways (e.g., "Prioritize road safety in Kathmandu").
- Tools: Named Pandas/Seaborn or R/ggplot2 with specific functions.
Based on the PU BE Computer (PU) syllabus for Data Science and Analytics (CMP422), unit 3.
Discussion
Loading…