Machine LearningUnit 212 min read

Data Preprocessing: Cleaning, Transforming & Feature Engineering

Unit 2 of Machine Learning covers essential techniques to prepare raw data for modeling—handling missing values, scaling, encoding, feature selection, and dimensionality reduction—with real-world applications in Nepalese tech (e.g., Khalti fraud detection, Pathao route optimization) and global platforms (Google’s recom

Why Preprocessing Matters

Machine learning models are only as good as the data they’re trained on. Garbage in, garbage out (GIGO) applies here. Real-world datasets from sources like Nepal’s NTC call logs, Daraz customer reviews, or Khalti transaction records are messy: incomplete, inconsistent, or unstructured. Preprocessing transforms this "noisy" data into a clean, structured format that algorithms can learn from efficiently.


1. Data Cleaning: Fixing the Mess

Definition: Removing or correcting errors, inconsistencies, and irrelevant data to improve quality. Key Steps:

  • Handling Missing Values: Data is often incomplete (e.g., 30% of NEPSE stock records might lack closing prices).
  • Removing Duplicates: Identical records (e.g., duplicate Khalti transactions) skew analysis.
  • Outlier Detection: Extreme values (e.g., a Pathao rider’s speed of 200 km/h) may need removal or transformation.
  • Noise Reduction: Smoothing irregularities (e.g., sensor data from Nepal’s air quality monitors).
02200440066008800Complete Records8800Missing Amount1200Missing Timestamp500Duplicate Transactions300Number of Records (Khalti Dataset)
Distribution of data quality issues in Khalti’s 10,000 transactions

How It Works: Missing Data Strategies

Strategy When to Use Example (Nepal Context)
Deletion <5% missing data Drop rows with missing NEPSE dividend yields.
Mean/Median Imputation Numerical data with <30% missingness Fill missing monthly rainfall data in Kathmandu.
Mode Imputation Categorical data (e.g., "city") Replace missing "district" in Daraz delivery records.
Predictive Imputation High missingness (>50%) Use regression to predict missing NTC call durations.

Worked Example: Khalti Fraud Detection Suppose Khalti’s dataset has 10,000 transactions with 1,200 missing amount values (12% missingness).

  • Step 1: Check if missingness is random or correlated with fraud flags.
  • Step 2: Use median imputation (robust to outliers) since transaction amounts are skewed.
    • Median amount = ₹4,500.
    • Replace all missing amount with ₹4,500.
  • Step 3: Flag transactions where amount > 10× median as potential fraud (e.g., ₹50,000).
graph LR
    A["Raw Khalti Data\n(12% missing amounts)"] --> B["Detect Missingness\n(1,200 rows)"]
    B --> C["Choose Strategy:\nMedian Imputation"]
    C --> D["Replace Missing\nwith ₹4,500"]
    D --> E["Flag Outliers\n(amount > ₹45,000)"]
    E --> F["Cleaned Data\nReady for Model"]

2. Data Integration: Combining Datasets

Definition: Merging data from multiple sources (e.g., NTC call logs + Ncell tower locations) to create a unified view. Methods:

  • Horizontal Integration: Stacking rows (e.g., combining monthly sales data from Daraz and Sastodeal).
  • Vertical Integration: Adding columns (e.g., merging NEPSE stock prices with macroeconomic indicators like Nepal’s inflation rate).
  • Temporal Integration: Aligning time-series data (e.g., Pathao’s daily rides with Kathmandu traffic camera feeds).

Real Picture:


3. Data Transformation: Reshaping Features

Definition: Converting data into a format suitable for modeling (e.g., normalizing Khalti transaction timestamps).

5000100001500020000250003000035000400004500050000-4000-2000200040006000800010000xLog-transformed Prices (₹)Original Prices (₹)₹50 → 3.91₹50,000 → 10.82
Log transformation linearizes Daraz product price scale (original vs. log)

A. Normalization vs. Standardization

Technique Formula When to Use Example
Min-Max Bounded ranges (e.g., ages 18–65) Scale Pathao rider ratings (1–5 stars) to [0,1].
Z-Score Gaussian-like data (e.g., NEPSE returns) Standardize Khalti transaction amounts.
Log Transform Skewed data (e.g., house prices in Nepal) Transform Daraz product prices.

Worked Example: Daraz Price Comparison Daraz’s dataset has product prices ranging from ₹50 to ₹50,000. A log transform helps linear models (like regression) perform better:

  • Original prices: [50, 100, 500, 50,000]
  • Log-transformed: [3.91, 4.61, 6.21, 10.82]
  • Now, the scale is linearized for distance-based algorithms (e.g., k-NN).

B. Encoding Categorical Data

Categorical variables (e.g., "district" in Daraz orders) must be converted to numerical form.

Method Use Case Example
Label Encoding Ordinal categories (e.g., "low/medium/high") Convert NTC call quality: "Poor"→0, "Good"→1.
One-Hot Encoding Nominal categories (e.g., "Kathmandu/Lalitpur") Daraz orders: district_Kathmandu=1, district_Lalitpur=0.
Embedding High-cardinality features (e.g., 100+ cities) Use neural networks to encode "district" in Khalti’s fraud model.

Worked Example: NEPSE Sector Classification NEPSE stocks are categorized into sectors (e.g., "Banking," "Hydro"). Use one-hot encoding:

  • Original: sector = ["Banking", "Hydro", "Banking"]
  • Encoded:
    sector_Banking = [1, 0, 1]
    sector_Hydro   = [0, 1, 0]
    

4. Feature Selection: Picking the Best Inputs

Definition: Selecting the most relevant features to reduce noise and improve model performance. Methods:

  • Filter Methods: Use statistical tests (e.g., correlation, chi-square) to rank features.
    • Example: Select top 10 features from 100 in NTC’s call duration dataset based on correlation with "fraud."
  • Wrapper Methods: Exhaustive search (e.g., recursive feature elimination).
    • Example: Train a model on all features, remove the least important, and repeat.
  • Embedded Methods: Built into models (e.g., decision trees’ feature importance).
    • Example: Use a random forest to rank features in Pathao’s route optimization.
classDiagram
  class KhaltiFraudModel {
    +amount: float
    +time: datetime
    +location: str
    +device: str
    +is_fraud: bool
  }
  class FeatureImportance {
    <<enumeration>>
    AMOUNT
    TIME
    LOCATION
    DEVICE
  }
  KhaltiFraudModel --> FeatureImportance : "Top 2 features: "
  note for KhaltiFraudModel "Decision Tree splits on:
  1. amount > ₹50,000 (Gini=0.45)
  2. time after midnight (Gini=0.20)"
Feature importance hierarchy for Khalti’s fraud detection model

Visual: Feature Importance in a Decision Tree

graph TD
    A["Root: Is transaction amount > ₹50,000?"] --> B["Yes\n(Fraud Probability: 0.85)"]
    A --> C["No\n(Fraud Probability: 0.10)"]
    B --> D["Check: Is time after midnight?"]
    C --> E["Check: Is device location Kathmandu?"]

Key Insight: The top feature (amount) splits the data most effectively, reducing impurity (Gini index) the most.


5. Dimensionality Reduction

Definition: Reducing the number of input features while retaining most information. Techniques:

  • PCA (Principal Component Analysis): Linear transformation to uncorrelated features.
    • Example: Reduce 50 features in NEPSE stock data to 10 principal components.
  • t-SNE/UMAP: Non-linear reduction for visualization (e.g., clustering Daraz customers).
  • Feature Extraction: Create new features (e.g., "average transaction amount per user" in Khalti).

Worked Example: NTC Call Data Original features: call_duration, time_of_day, tower_location, user_age (4D).

  • Step 1: Standardize all features (Z-score).
  • Step 2: Compute covariance matrix.
  • Step 3: Eigen decomposition → select top 2 principal components (explaining 85% variance).
  • Result: Reduce 4D → 2D for faster fraud detection.

In the Real World

  1. Khalti’s Fraud Detection

    • Idea Used: Data Cleaning (missing value imputation) + Feature Engineering (transaction velocity features).
    • How: Khalti flags suspicious transactions by:
      • Imputing missing amount with median values.
      • Creating features like "transactions per minute per user" to detect bots.
      • Using log transformation on amounts to handle skewness.
  2. Pathao’s Route Optimization

    • Idea Used: Data Integration (traffic camera + GPS data) + Dimensionality Reduction (PCA for tower locations).
    • How: Pathao merges:
      • Real-time traffic data from Nepal Police cameras.
      • Rider locations from Ncell towers.
      • Applies PCA to reduce 100+ tower coordinates to 3 principal components for faster route calculations.
  3. Daraz’s Recommendation System

    • Idea Used: One-Hot Encoding (for product categories) + Normalization (for price ranges).
    • How: Daraz encodes:
      • Product categories (e.g., "Electronics" → one-hot vector).
      • Normalizes prices using Min-Max scaling to recommend similar-priced items.

6. Handling Imbalanced Data

Problem: Classes are uneven (e.g., 99% legitimate Khalti transactions vs. 1% fraud). Solutions:

  • Resampling: Oversample fraud cases or undersample legitimate ones.
  • Synthetic Data: SMOTE generates synthetic fraud examples.
  • Class Weights: Penalize misclassification of the minority class (e.g., fraud) more heavily.

Worked Example: NEPSE Stock Anomaly Detection

  • Data: 95% normal days, 5% anomalies (e.g., sudden price drops).
  • Solution: Use SMOTE to create synthetic anomaly examples, then train an SVM.

Exam Tip

  1. Know the Trade-offs:

    • Deletion vs. Imputation: Deletion loses data; imputation introduces bias. Justify your choice (e.g., "Deletion is acceptable here because missingness is <5%").
    • Normalization vs. Standardization: Min-Max for bounded ranges; Z-score for Gaussian data.
  2. Practical Scenarios:

    • Always tie examples to Nepal: "In Khalti’s dataset, if 20% of amount is missing, how would you handle it?" → Median imputation.
    • Feature Engineering: "How would you create a feature to detect fake Daraz reviews?" → Review velocity (reviews per hour per user).
  3. Visuals in Exams:

    • Draw a mermaid flowchart for preprocessing steps (e.g., "Raw Data → Clean → Transform → Select Features").
    • Sketch a PCA transformation (arrows showing original → reduced dimensions).
  4. Common Pitfalls:

    • Data Leakage: Never use future data (e.g., don’t normalize using test set stats).
    • Overfitting: Avoid selecting features based on test set performance.
  5. Short-Answer Questions:

    • "What is the difference between normalization and standardization?" → Normalization scales to [0,1]; standardization uses Z-scores (mean=0, std=1).
    • "Why is one-hot encoding better than label encoding for nominal data?" → Label encoding imposes ordinal relationships (e.g., "Kathmandu" > "Lalitpur"), which is meaningless.

Summary Checklist

Before submitting your answer, ensure you’ve covered: ✅ Data Cleaning: Missing values, duplicates, outliers. ✅ Transformation: Normalization, encoding, log transforms. ✅ Feature Selection: Filter/wrapper/embedded methods. ✅ Dimensionality Reduction: PCA, t-SNE. ✅ Real-World Tie-Ins: Khalti, Pathao, Daraz, NEPSE. ✅ Visuals: At least 3 diagrams (e.g., PCA, decision tree feature importance, preprocessing pipeline).

In the real world

  • Khalti: Uses median imputation for missing transaction amounts (12% missingness) and log transformation of amounts to detect fraud patterns in skewed data.
  • Pathao: Applies one-hot encoding for district names (e.g., Kathmandu/Lalitpur) and Z-score standardization of rider ratings (1–5 stars) to train route optimization models.
  • NEPSE: Employs recursive feature elimination to select top 5 macroeconomic indicators (e.g., inflation rate, NTC call volumes) from 20+ features for stock price prediction.

Based on the PU BE Computer (PU) syllabus for Machine Learning (CMP364), unit 2.

Discussion

Loading…