IT243 Programming with Python

Programming with PythonUnit 99 min read

NumPy Arrays, Pandas DataFrames & Data Analysis Workflows

Unit 9 of Programming with Python covers NumPy’s array operations, Pandas’ DataFrame manipulation, and real-world data analysis workflows—from loading datasets to cleaning, aggregating, and visualizing data with Python libraries.

Core Concepts

1. NumPy: The Foundation for Numerical Computing

NumPy (Numerical Python) is a library for working with multi-dimensional arrays and mathematical operations. It provides:

  • Efficient array objects (ndarray) for storing large datasets.
  • Vectorized operations (applying functions to entire arrays without loops).
  • Mathematical functions (linear algebra, Fourier transforms, random number generation).
60718293
Result of vectorized addition: `arr1 + 5` (output: [6 7 8 9])
1,2,304,5,61
2D NumPy array (matrix): `np.array([[1, 2, 3], [4, 5, 6]])` (shape: (2, 3))
10213243
1D NumPy array: `np.array([1, 2, 3, 4])` (shape: (4,), dtype: int64)

Why NumPy?

  • Speed: Written in C, NumPy operations are 100x faster than Python lists.
  • Memory efficiency: Stores data in contiguous blocks.
  • Compatibility: Works seamlessly with Pandas, SciPy, and Matplotlib.

Key Data Structures in NumPy

classDiagram
    class ndarray {
        +shape: tuple
        +ndim: int
        +dtype: data type
        +itemsize: bytes per element
        +size: total elements
        +methods: reshape(), transpose(), broadcast()
    }
    class ufunc {
        +universal functions (e.g., sin(), add())
    }
    ndarray --> ufunc : "applies to"

Example: Creating and Manipulating Arrays

import numpy as np

# Create a 1D array
arr1 = np.array([1, 2, 3, 4])
print("1D Array:", arr1)

# Create a 2D array (matrix)
arr2 = np.array([[1, 2, 3], [4, 5, 6]])
print("2D Array:\n", arr2)

# Array operations
print("Sum:", arr1 + 5)  # Vectorized addition
print("Shape:", arr2.shape)  # (2, 3)
print("Transpose:\n", arr2.T)

Trace of Operations:

Step Code Executed Output
arr1 = np.array(...) Initialization [1 2 3 4]
arr1 + 5 Vectorized addition [6 7 8 9]
arr2.T Transpose [[1 4], [2 5], [3 6]]

2. Pandas: Data Analysis with DataFrames

Pandas builds on NumPy to provide high-level data structures for tabular data (like Excel sheets or SQL tables). Its two main objects:

  • Series: 1D array-like object (like a column in a table).
  • DataFrame: 2D table with labeled rows and columns (most used in data analysis).
0CloseVolume1MA_52—3—
Pandas DataFrame as a hash-table-like structure (simplified): Columns as keys, rows as values.

Why Pandas?

  • Data cleaning: Handle missing values (NaN), duplicates, and outliers.
  • Data aggregation: GroupBy, pivot tables, and statistical summaries.
  • I/O: Read/write data from/to CSV, Excel, SQL, and APIs.

Key Operations

flowchart TD
    A["Load Data"] --> B["Clean Data"]
    B --> C["Inspect Data"]
    C --> D["Transform Data"]
    D --> E["Aggregate Data"]
    E --> F["Visualize Data"]

Example: Loading and Cleaning a Dataset

import pandas as pd

# Load a CSV file (e.g., NEPSE stock data)
df = pd.read_csv("nepse_data.csv")
print("First 5 rows:\n", df.head())

# Check for missing values
print("\nMissing values:\n", df.isnull().sum())

# Drop rows with missing values
df_clean = df.dropna()
print("\nCleaned Data Shape:", df_clean.shape)

# Add a new column (e.g., moving average)
df_clean['MA_5'] = df_clean['Close'].rolling(window=5).mean()

Trace of Data Cleaning:

Step Code Executed Output
pd.read_csv(...) Load data DataFrame with 1000 rows
df.isnull().sum() Check missing values Price: 5, Volume: 0
df.dropna() Remove missing rows DataFrame with 995 rows
df_clean['MA_5'] Add moving average column New column with 5-day MA values

In the Real World

  1. NEPSE Stock Analysis

    • Tool: Pandas + NumPy
    • Use Case: Calculate daily returns, moving averages, and volatility for stocks like Nabil Bank or NMB.
    • Example: A trader uses df['Return'] = df['Close'].pct_change() to compute percentage changes for decision-making.
  2. Khalti Transaction Fraud Detection

    • Tool: Pandas DataFrames + NumPy
    • Use Case: Detect anomalies in transaction data (e.g., sudden large amounts) using statistical methods like Z-score.
    • Example:
      from scipy import stats
      z_scores = np.abs(stats.zscore(df['Amount']))
      df['Fraud'] = z_scores > 3  # Flag outliers
      
  3. Daraz Customer Segmentation

    • Tool: Pandas + NumPy
    • Use Case: Group customers by purchase behavior (e.g., df.groupby('Customer_ID')['Total_Spent'].sum()) to target marketing.
    • Example: Identify high-value customers for loyalty programs.

3. Data Analysis Workflow

A typical workflow in Pandas/NumPy:

  1. Load Data: pd.read_csv(), np.load().
  2. Clean Data: Handle NaN, duplicates, and outliers.
  3. Explore Data: df.describe(), df.info(), visualizations.
  4. Transform Data: Normalization, feature engineering.
  5. Aggregate Data: groupby(), pivot_table().
  6. Visualize Data: Use Matplotlib/Seaborn.
flowchart TD
    A["Load Data: `pd.read_csv()`"] --> B["Clean Data: `dropna()`, `fillna()`"]
    B --> C["Explore: `df.describe()`, `df.info()`"]
    C --> D["Transform: `rolling()`, `pct_change()`"]
    D --> E["Aggregate: `groupby()`, `pivot_table()`"]
    E --> F["Visualize: `plot()`, `seaborn`"]
Pandas data analysis workflow (NEPSE stock example)

Example: Analyzing NTC Internet Usage Data

# Load NTC monthly data (hypothetical)
ntc_data = pd.read_csv("ntc_internet_usage.csv")

# Calculate monthly growth rate
ntc_data['Growth_Rate'] = ntc_data['Users'].pct_change() * 100

# Plot growth trend
import matplotlib.pyplot as plt
ntc_data['Growth_Rate'].plot(title="NTC User Growth Rate (%)")
plt.show()

Visualization Output:



4. Comparing NumPy and Pandas

Feature NumPy Pandas
Data Structure ndarray (homogeneous) DataFrame (heterogeneous)
Missing Values Not supported Supported (NaN)
Operations Vectorized math Data alignment, grouping
Use Case Numerical computing Data analysis, ETL
Performance Faster for pure math Slower but flexible

5. Advanced Operations

a) Merging and Joining DataFrames

# Merge two DataFrames (e.g., customer and order data)
customer_df = pd.DataFrame({'ID': [1, 2], 'Name': ['Alice', 'Bob']})
order_df = pd.DataFrame({'Order_ID': [101, 102], 'Customer_ID': [1, 2], 'Amount': [500, 300]})

merged_df = pd.merge(customer_df, order_df, left_on='ID', right_on='Customer_ID')
print(merged_df)

Output:

   ID   Name  Order_ID  Customer_ID  Amount
0   1  Alice       101             1     500
1   2    Bob       102             2     300

b) Time Series Analysis

# Convert 'Date' column to datetime
df['Date'] = pd.to_datetime(df['Date'])

# Resample data (e.g., daily to monthly)
monthly_data = df.set_index('Date').resample('M').mean()

c) Handling Categorical Data

# Convert categorical column to numerical
df['Category'] = df['Category'].astype('category')
df['Category_Code'] = df['Category'].cat.codes

Exam Tip

  1. Understand the Difference Between NumPy and Pandas:

    • NumPy is for numerical operations (e.g., np.sin(), np.dot()).
    • Pandas is for tabular data (e.g., df.groupby(), df.pivot()).
  2. Practice Common Operations:

    • Loading data (pd.read_csv()).
    • Cleaning data (dropna(), fillna()).
    • Aggregating data (groupby(), agg()).
    • Merging data (pd.merge()).
  3. Visualize Data:

    • Use df.plot() for quick visualizations.
    • Know how to interpret plots (e.g., line charts for trends, bar charts for comparisons).
  4. Real-World Scenarios:

    • Be ready to apply concepts to financial data (NEPSE), transaction logs (Khalti), or sensor data (NTC).
    • Example question:

      "Given a dataset of Daraz orders, write a Pandas script to find the top 5 customers by total spending."

  5. Common Pitfalls:

    • Forgetting to handle NaN values (use dropna() or fillna()).
    • Misaligning indices when merging DataFrames (use left_on/right_on).
    • Not converting strings to datetime before time-series operations.

Key Formula to Remember: For moving average (MA) of window size n: where is the price at time .

Based on the TU BIM syllabus for Programming with Python (IT243), unit 9.

Discussion

Loading…