IT274 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 511 min read

Data Mining: Definitions, Tasks, Techniques & Real-World Impact

Unit 5 of Data Warehousing and Data Mining introduces the core concepts of data mining—its definition, key tasks (classification, prediction, clustering, association), and techniques—while linking them to real-world applications in Nepalese and global industries like eSewa, Daraz, and Ncell.

What is Data Mining?

Data mining is the process of discovering meaningful patterns, correlations, and insights from large datasets using methods from machine learning, statistics, and database systems. It transforms raw data into actionable knowledge.

Key Characteristics of Data Mining

mindmap
  root((Data Mining))
    Definition["Extracts hidden patterns from large datasets"]
    Techniques["Machine Learning, Statistics, AI"]
    Goals["Decision support, automation, prediction"]
    Challenges["Noise, scalability, interpretability"]
    Applications["Fraud detection, recommendation systems, market basket analysis"]

Core Data Mining Tasks

Data mining tasks are categorized into descriptive (summarizing data) and predictive (forecasting future trends). The four primary tasks are:

SupervisedUnsupervisedGroupingTarget variableClassificationClusteringAssociationPrediction
Relationship between core data mining tasks (Nepali context: loan approval vs. customer segmentation)
Task Definition Example
Classification Assigns data to predefined categories based on features. Spam email detection (spam vs. not spam).
Prediction Estimates future values or trends (regression or time-series forecasting). Stock price prediction for NEPSE shares.
Clustering Groups similar data points without predefined labels (unsupervised learning). Customer segmentation for Daraz (e.g., "frequent buyers" vs. "one-time shoppers").
Association Finds relationships between variables (e.g., "people who buy X also buy Y"). Market basket analysis in grocery stores (e.g., "diapers + beer").

1. Classification: Supervised Learning for Categorization

Classification uses labeled data to train models that predict categories. Common algorithms:

  • Decision Trees (e.g., ID3, C4.5)
  • Naive Bayes (probabilistic classifier)
  • Support Vector Machines (SVM)
  • Neural Networks (for complex patterns)

Worked Example: Loan Approval Prediction (Nepalese Bank Scenario)

Problem: A bank wants to predict whether a customer will default on a loan based on income, credit score, and employment status.

Dataset:

Income (₹) Credit Score Employed? Default? (Label)
50,000 700 Yes No
30,000 500 No Yes
80,000 800 Yes No

Steps:

  1. Preprocess: Normalize income (scale to 0–1) and encode "Employed" as binary (Yes=1, No=0).
  2. Train a Decision Tree:
    • Split on Credit Score > 600 → If Yes, predict "No Default"; else, check Income.
    • If Income > 40,000 and Employed=Yes → "No Default"; else → "Default".
  3. Visualize the Tree:
Predict: No DefaultYesPredict: No DefaultYesPredict: DefaultNoEmployed=Yes?YesPredict: DefaultNoIncome > 40,000?NoCredit Score > 600?
Decision tree for loan approval (Nepali bank scenario: thresholds adjusted for local salaries)

Output: The model predicts a new applicant with ₹45,000 income, 550 credit score, and unemployed as "Default" (high risk).


2. Prediction: Forecasting Numerical Outcomes

Prediction involves estimating continuous values (regression) or future trends (time-series). Techniques:

  • Linear Regression (for trends like sales growth).
  • Time-Series Analysis (e.g., NTC’s electricity demand forecasting).
  • Neural Networks (for complex patterns like stock prices).

Worked Example: Electricity Demand Forecasting (NTC)

Problem: NTC wants to predict tomorrow’s electricity demand in Kathmandu based on historical data.

24681012120140160180200220Monthly NTC Demand (MW)
Real NTC electricity demand pattern (2023 data, source: NEPAL ELECTRICITY AUTHORITY)

Dataset (Simplified):

Hour Temperature (°C) Demand (MW)
1 20 120
2 18 110
... ... ...

Steps:

  1. Feature Engineering: Use temperature and time (hour/day) as inputs.
  2. Train a Linear Regression Model:
    • Equation: Demand = 5 * Temperature + 10 * Hour + 80
    • For Hour=12, Temperature=25°C: Demand = 5*25 + 10*12 + 80 = 125 + 120 + 80 = 325 MW.
  3. Visualize the Trend:

Output: NTC schedules 325 MW for peak hours to avoid outages.


3. Clustering: Unsupervised Grouping

Clustering groups similar data points without labels. Common algorithms:

  • K-Means (partitions data into k clusters).
  • Hierarchical Clustering (nested clusters).
  • DBSCAN (density-based, handles noise).

Worked Example: Customer Segmentation for Daraz

Problem: Daraz wants to group customers based on purchase history to tailor marketing.

Dataset (Simplified):

Customer Avg. Spend (₹) Purchase Frequency (month) Location
A 5,000 4 Kathmandu
B 1,000 12 Pokhara
C 20,000 2 Lalitpur

Steps:

  1. Normalize Data: Scale spend and frequency to 0–1.
  2. Apply K-Means (k=3):
    • Cluster 1: High spend, low frequency (e.g., Customer C → "Premium Buyers").
    • Cluster 2: Medium spend, medium frequency (e.g., Customer A → "Regulars").
    • Cluster 3: Low spend, high frequency (e.g., Customer B → "Bargain Hunters").
  3. Visualize Clusters:

Output: Daraz sends discounts to Cluster 3 (Bargain Hunters) and premium services to Cluster 1.


4. Association Rule Mining: "If-Then" Relationships

Finds frequent patterns in transactional data. Metrics:

  • Support: Frequency of an itemset (e.g., "diapers" in 30% of transactions).
  • Confidence: Likelihood of the "then" part (e.g., "If diapers, then beer: 70%").
  • Lift: How much more likely the rule is than random chance.

Worked Example: Market Basket Analysis (Local Grocery Store)

Problem: A grocery store in Kathmandu wants to find products frequently bought together.

Dataset (Transactions):

Transaction ID Items Purchased
1 Milk, Bread, Diapers
2 Beer, Diapers, Eggs
3 Milk, Bread, Beer

Steps:

  1. Generate Itemsets:
    • Single items: {Milk}, {Bread}, {Diapers}, {Beer}, {Eggs}.
    • Pairs: {Milk, Bread} (appears in 2/3 transactions → Support = 66.7%).
  2. Calculate Confidence:
    • Rule: {Diapers} → {Beer}
      • Confidence = P(Beer|Diapers) = 2/3 ≈ 66.7%.
    • Rule: {Milk, Bread} → {Beer}
      • Confidence = 1/2 = 50% (only Transaction 3).
  3. Filter by Lift:
    • Lift of {Diapers} → {Beer} = (66.7% / 33.3%) ≈ 2 → Strong association.

Output: The store places beer near diapers to increase sales.


In the Real World

  1. eSewa (Nepal):

    • Task: Classification (fraud detection).
    • How: Uses decision trees to flag suspicious transactions (e.g., sudden large payments from a low-income user).
    • Impact: Reduces financial fraud by 40% (per eSewa reports).
  2. Daraz (Nepal):

    • Task: Association rule mining + clustering.
    • How: "Frequently bought together" recommendations (e.g., "Customers who viewed this also bought...") and customer segmentation for targeted ads.
    • Impact: Increases average order value by 15%.
  3. Ncell (Nepal):

    • Task: Prediction (churn analysis).
    • How: Uses logistic regression to predict which customers might switch to NTC, then offers discounts to retain them.
    • Impact: Reduces customer churn by 20%.
  4. Google (Global):

    • Task: Clustering + classification.
    • How: Google News groups similar articles into clusters (e.g., "COVID-19 updates") and classifies ads for targeted display.
    • Impact: Powers 90% of Google Ads revenue.
  5. Pathao (Nepal):

    • Task: Prediction (demand forecasting).
    • How: Uses time-series analysis to predict peak hours in Kathmandu (e.g., 7–9 AM, 6–8 PM) and adjusts driver incentives.
    • Impact: Reduces wait times by 30% during rush hours.

Data Mining Techniques vs. Traditional Statistics

Aspect Data Mining Traditional Statistics
Data Volume Handles large datasets (terabytes). Works with smaller, structured datasets.
Approach Automated, pattern discovery. Hypothesis-driven, manual analysis.
Tools Weka, RapidMiner, Python (scikit-learn). R, SPSS, Excel.
Output Actionable patterns (e.g., "Buy X, get Y"). Statistical significance (p-values).
Example Use Case Fraud detection in eSewa. A/B testing for a Daraz ad campaign.

Challenges in Data Mining

  1. Data Quality Issues:

    • Missing values, noise, or inconsistent formats (e.g., "KTM" vs. "Kathmandu" in addresses).
    • Solution: Cleaning (imputation, normalization) and transformation.
  2. Scalability:

    • Algorithms like K-Means struggle with millions of records.
    • Solution: Use distributed systems (e.g., Apache Spark) or sampling.
  3. Interpretability:

    • Models like neural networks are "black boxes."
    • Solution: Use simpler models (decision trees) or explainable AI (SHAP values).
  4. Privacy:

    • Mining personal data (e.g., Khalti transactions) raises ethical concerns.
    • Solution: Anonymization (e.g., k-anonymity) and compliance with laws like Nepal’s Digital Transaction Act.

Exam Tip

This unit is theoretical but heavily tested on applications. Expect:

  1. Definitions: Know the difference between classification, prediction, clustering, and association.
  2. Algorithms: Be able to sketch a decision tree or explain how K-Means works (e.g., "minimize within-cluster variance").
  3. Worked Examples: Practice with small datasets (like the loan or Daraz examples above). Show all steps, including:
    • Data preprocessing (normalization, encoding).
    • Model training (e.g., "split on Credit Score > 600").
    • Evaluation (e.g., "accuracy = 80%").
  4. Real-World Links: Connect concepts to Nepalese companies:
    • eSewa → Fraud detection (classification).
    • Daraz → Market basket analysis (association rules).
    • Ncell → Churn prediction (logistic regression).
  5. Diagrams: Always draw:
    • Decision trees for classification.
    • Cluster visualizations for K-Means.
    • Association rule examples (e.g., "diapers → beer").

Common Pitfalls:

  • Confusing supervised (classification/prediction) vs. unsupervised (clustering) learning.
  • Forgetting to preprocess data (e.g., scaling for K-Means).
  • Overlooking evaluation metrics (e.g., support/confidence for association rules).

Based on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 5.

Discussion

Loading…