Programming with PythonUnit 99 min read
NumPy Arrays, Pandas DataFrames & Data Analysis Workflows
Unit 9 of Programming with Python covers NumPy’s array operations, Pandas’ DataFrame manipulation, and real-world data analysis workflows—from loading datasets to cleaning, aggregating, and visualizing data with Python libraries.
Core Concepts
1. NumPy: The Foundation for Numerical Computing
NumPy (Numerical Python) is a library for working with multi-dimensional arrays and mathematical operations. It provides:
- Efficient array objects (
ndarray) for storing large datasets. - Vectorized operations (applying functions to entire arrays without loops).
- Mathematical functions (linear algebra, Fourier transforms, random number generation).
Why NumPy?
- Speed: Written in C, NumPy operations are 100x faster than Python lists.
- Memory efficiency: Stores data in contiguous blocks.
- Compatibility: Works seamlessly with Pandas, SciPy, and Matplotlib.
Key Data Structures in NumPy
classDiagram
class ndarray {
+shape: tuple
+ndim: int
+dtype: data type
+itemsize: bytes per element
+size: total elements
+methods: reshape(), transpose(), broadcast()
}
class ufunc {
+universal functions (e.g., sin(), add())
}
ndarray --> ufunc : "applies to"Example: Creating and Manipulating Arrays
import numpy as np
# Create a 1D array
arr1 = np.array([1, 2, 3, 4])
print("1D Array:", arr1)
# Create a 2D array (matrix)
arr2 = np.array([[1, 2, 3], [4, 5, 6]])
print("2D Array:\n", arr2)
# Array operations
print("Sum:", arr1 + 5) # Vectorized addition
print("Shape:", arr2.shape) # (2, 3)
print("Transpose:\n", arr2.T)
Trace of Operations:
| Step | Code Executed | Output |
|---|---|---|
arr1 = np.array(...) |
Initialization | [1 2 3 4] |
arr1 + 5 |
Vectorized addition | [6 7 8 9] |
arr2.T |
Transpose | [[1 4], [2 5], [3 6]] |
2. Pandas: Data Analysis with DataFrames
Pandas builds on NumPy to provide high-level data structures for tabular data (like Excel sheets or SQL tables). Its two main objects:
- Series: 1D array-like object (like a column in a table).
- DataFrame: 2D table with labeled rows and columns (most used in data analysis).
Why Pandas?
- Data cleaning: Handle missing values (
NaN), duplicates, and outliers. - Data aggregation: GroupBy, pivot tables, and statistical summaries.
- I/O: Read/write data from/to CSV, Excel, SQL, and APIs.
Key Operations
flowchart TD
A["Load Data"] --> B["Clean Data"]
B --> C["Inspect Data"]
C --> D["Transform Data"]
D --> E["Aggregate Data"]
E --> F["Visualize Data"]Example: Loading and Cleaning a Dataset
import pandas as pd
# Load a CSV file (e.g., NEPSE stock data)
df = pd.read_csv("nepse_data.csv")
print("First 5 rows:\n", df.head())
# Check for missing values
print("\nMissing values:\n", df.isnull().sum())
# Drop rows with missing values
df_clean = df.dropna()
print("\nCleaned Data Shape:", df_clean.shape)
# Add a new column (e.g., moving average)
df_clean['MA_5'] = df_clean['Close'].rolling(window=5).mean()
Trace of Data Cleaning:
| Step | Code Executed | Output |
|---|---|---|
pd.read_csv(...) |
Load data | DataFrame with 1000 rows |
df.isnull().sum() |
Check missing values | Price: 5, Volume: 0 |
df.dropna() |
Remove missing rows | DataFrame with 995 rows |
df_clean['MA_5'] |
Add moving average column | New column with 5-day MA values |
In the Real World
NEPSE Stock Analysis
- Tool: Pandas + NumPy
- Use Case: Calculate daily returns, moving averages, and volatility for stocks like Nabil Bank or NMB.
- Example: A trader uses
df['Return'] = df['Close'].pct_change()to compute percentage changes for decision-making.
Khalti Transaction Fraud Detection
- Tool: Pandas DataFrames + NumPy
- Use Case: Detect anomalies in transaction data (e.g., sudden large amounts) using statistical methods like Z-score.
- Example:
from scipy import stats z_scores = np.abs(stats.zscore(df['Amount'])) df['Fraud'] = z_scores > 3 # Flag outliers
Daraz Customer Segmentation
- Tool: Pandas + NumPy
- Use Case: Group customers by purchase behavior (e.g.,
df.groupby('Customer_ID')['Total_Spent'].sum()) to target marketing. - Example: Identify high-value customers for loyalty programs.
3. Data Analysis Workflow
A typical workflow in Pandas/NumPy:
- Load Data:
pd.read_csv(),np.load(). - Clean Data: Handle
NaN, duplicates, and outliers. - Explore Data:
df.describe(),df.info(), visualizations. - Transform Data: Normalization, feature engineering.
- Aggregate Data:
groupby(),pivot_table(). - Visualize Data: Use Matplotlib/Seaborn.
flowchart TD
A["Load Data: `pd.read_csv()`"] --> B["Clean Data: `dropna()`, `fillna()`"]
B --> C["Explore: `df.describe()`, `df.info()`"]
C --> D["Transform: `rolling()`, `pct_change()`"]
D --> E["Aggregate: `groupby()`, `pivot_table()`"]
E --> F["Visualize: `plot()`, `seaborn`"]Pandas data analysis workflow (NEPSE stock example)Example: Analyzing NTC Internet Usage Data
# Load NTC monthly data (hypothetical)
ntc_data = pd.read_csv("ntc_internet_usage.csv")
# Calculate monthly growth rate
ntc_data['Growth_Rate'] = ntc_data['Users'].pct_change() * 100
# Plot growth trend
import matplotlib.pyplot as plt
ntc_data['Growth_Rate'].plot(title="NTC User Growth Rate (%)")
plt.show()
Visualization Output:
4. Comparing NumPy and Pandas
| Feature | NumPy | Pandas |
|---|---|---|
| Data Structure | ndarray (homogeneous) |
DataFrame (heterogeneous) |
| Missing Values | Not supported | Supported (NaN) |
| Operations | Vectorized math | Data alignment, grouping |
| Use Case | Numerical computing | Data analysis, ETL |
| Performance | Faster for pure math | Slower but flexible |
5. Advanced Operations
a) Merging and Joining DataFrames
# Merge two DataFrames (e.g., customer and order data)
customer_df = pd.DataFrame({'ID': [1, 2], 'Name': ['Alice', 'Bob']})
order_df = pd.DataFrame({'Order_ID': [101, 102], 'Customer_ID': [1, 2], 'Amount': [500, 300]})
merged_df = pd.merge(customer_df, order_df, left_on='ID', right_on='Customer_ID')
print(merged_df)
Output:
ID Name Order_ID Customer_ID Amount
0 1 Alice 101 1 500
1 2 Bob 102 2 300
b) Time Series Analysis
# Convert 'Date' column to datetime
df['Date'] = pd.to_datetime(df['Date'])
# Resample data (e.g., daily to monthly)
monthly_data = df.set_index('Date').resample('M').mean()
c) Handling Categorical Data
# Convert categorical column to numerical
df['Category'] = df['Category'].astype('category')
df['Category_Code'] = df['Category'].cat.codes
Exam Tip
Understand the Difference Between NumPy and Pandas:
- NumPy is for numerical operations (e.g.,
np.sin(),np.dot()). - Pandas is for tabular data (e.g.,
df.groupby(),df.pivot()).
- NumPy is for numerical operations (e.g.,
Practice Common Operations:
- Loading data (
pd.read_csv()). - Cleaning data (
dropna(),fillna()). - Aggregating data (
groupby(),agg()). - Merging data (
pd.merge()).
- Loading data (
Visualize Data:
- Use
df.plot()for quick visualizations. - Know how to interpret plots (e.g., line charts for trends, bar charts for comparisons).
- Use
Real-World Scenarios:
- Be ready to apply concepts to financial data (NEPSE), transaction logs (Khalti), or sensor data (NTC).
- Example question:
"Given a dataset of Daraz orders, write a Pandas script to find the top 5 customers by total spending."
Common Pitfalls:
- Forgetting to handle
NaNvalues (usedropna()orfillna()). - Misaligning indices when merging DataFrames (use
left_on/right_on). - Not converting strings to datetime before time-series operations.
- Forgetting to handle
Key Formula to Remember:
For moving average (MA) of window size n:
where is the price at time .
Based on the TU BIM syllabus for Programming with Python (IT243), unit 9.
Discussion
Loading…