IT274 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 512 min read

Data Mining: Definitions, Tasks, Techniques & Real-World Impact

Unit 5 of Data Warehousing and Data Mining introduces the core concepts of data mining—its definition, key tasks (classification, prediction, clustering, etc.), techniques (supervised vs. unsupervised learning), and real-world applications in Nepalese and global industries. This note explains how data mining extracts a

What is Data Mining?

Data mining is the process of discovering meaningful patterns, correlations, and insights from large datasets using statistical, machine learning, and database techniques. It bridges the gap between raw data and actionable knowledge, enabling businesses to make data-driven decisions.

Key Characteristics of Data Mining

mindmap
  root((Data Mining))
    Characteristics
      Large Datasets: Handles terabytes/petabytes of data
      Automated Discovery: Uncovers hidden patterns without manual intervention
      Pattern-Oriented: Focuses on trends, clusters, and associations
      Non-Trivial: Extracts knowledge beyond simple queries
      Practical Evaluation: Results must be useful and interpretable

Why is it called "mining"?

  • Just as miners extract valuable minerals from the earth, data miners extract valuable knowledge from vast datasets.

In the Real World

  1. eSewa (Nepal) uses association rule mining to recommend services (e.g., "Users who paid utility bills also booked flights") by analyzing transaction logs. This increases cross-selling by 15%.
  2. Pathao (ride-hailing app) employs clustering algorithms to group high-demand areas in Kathmandu, optimizing driver dispatch and reducing wait times by 20%.
  3. Nepal Rastra Bank (NRB) applies fraud detection models (a classification task) to flag suspicious transactions in real-time, preventing financial crimes worth NPR 500M+ annually.

IMAGE: "data mining process flowchart" | A labeled diagram showing the stages: Data Collection → Data Cleaning → Data Transformation → Mining → Pattern Evaluation → Knowledge Presentation.


Core Tasks in Data Mining

Data mining tasks are categorized into predictive (forecasting future trends) and descriptive (summarizing past data). The syllabus emphasizes four primary tasks:

Classification (35%)Clustering (25%)Association Rule Mining (20%)Prediction (Regression) (20%)
Proportion of real-world data mining tasks (approximate industry distribution).
Task Definition Example in Nepal Techniques Used
Classification Categorizes data into predefined groups based on labeled training data. Predicting loan defaults (yes/no) for Nabil Bank customers. Decision Trees, Naive Bayes, SVM, Neural Nets
Prediction Estimates future values or trends (regression or time-series forecasting). Forecasting NEPSE stock prices for next quarter. Linear Regression, Time-Series Analysis
Clustering Groups unlabeled data into clusters based on similarity. Segmenting Daraz customers into "high-spenders," "bargain hunters," etc. K-Means, Hierarchical Clustering, DBSCAN
Association Finds relationships between variables (e.g., "if-then" rules). "Customers who buy Khalti recharge cards also buy mobile data plans." Apriori, FP-Growth Algorithms

Worked Example: Classification (Loan Approval)

Scenario: A bank (e.g., Global IME Bank) wants to predict whether a customer will default on a loan based on:

  • Age (25, 30, 40, 50)
  • Income (NPR 50K, 100K, 200K)
  • Credit Score (300–850)
  • Label: Default (Yes/No)

Step 1: Data Preprocessing

  • Handle missing values (e.g., impute average income for missing entries).
  • Normalize credit scores to a 0–1 scale.

Step 2: Choose an Algorithm

  • Decision Tree (simple, interpretable) or Random Forest (more accurate).

Step 3: Train the Model

from sklearn.tree import DecisionTreeClassifier
model = DecisionTreeClassifier(criterion='gini', max_depth=3)
model.fit(X_train, y_train)  # X_train: [Age, Income, Credit Score], y_train: [Default]

Step 4: Predict & Evaluate

  • Input: Age=30, Income=100K, Credit Score=650 → Predicted Output: Default = No (92% confidence).
  • Evaluation Metrics: Accuracy = 88%, Precision = 85%, Recall = 80%.

Visualization of the Decision Tree:

Credit Score > 500?Income > 80K?Age > 35?
Decision tree for loan approval (Step 4: Predict & Evaluate). Root splits on Credit Score (threshold 500).

Interpretation: Customers with credit scores ≤ 500 are automatically rejected, while others are evaluated based on income and age.


Supervised vs. Unsupervised Learning

Aspect Supervised Learning Unsupervised Learning
Data Labels Uses labeled data (e.g., "Default: Yes/No"). Uses unlabeled data (no predefined categories).
Goal Predicts or classifies new data. Discovers hidden patterns or groupings.
Examples Spam detection, stock price prediction. Customer segmentation, anomaly detection.
Algorithms Decision Trees, SVM, Neural Networks. K-Means, PCA, Apriori.
Real-World Use Ncell predicts churn (customer attrition). NTC clusters high-traffic areas for network optimization.

IMAGE: "supervised vs unsupervised learning comparison" | A side-by-side table with icons (e.g., a teacher for supervised, a detective for unsupervised).


Data Mining Techniques

1. Classification (Supervised)

  • How it works: The model learns from labeled examples (e.g., "Customer X defaulted because...") and generalizes to new data.
  • Example: Khalti uses classification to detect fraudulent transactions by training on past fraud cases.

Worked Example: Fraud Detection Dataset:

Transaction ID Amount (NPR) Location Time Is Fraud?
T001 50,000 Kathmandu 3 AM Yes
T002 10,000 Pokhara 2 PM No

Algorithm: Logistic Regression (outputs probability of fraud). Decision Rule: Flag transactions with P(fraud) > 0.7.


2. Clustering (Unsupervised)

  • How it works: Groups similar data points without prior labels (e.g., grouping customers by spending habits).
  • Example: Daraz uses clustering to identify "VIP customers" for personalized discounts.

Worked Example: Customer Segmentation Dataset: 100 customers with features: [Age, Income, Purchase Frequency]. Algorithm: K-Means (K=3 clusters). Output:

  • Cluster 1: Young, low-income, frequent small purchases (target with flash sales).
  • Cluster 2: Middle-aged, high-income, occasional large purchases (offer premium services).
  • Cluster 3: Old, medium-income, rare purchases (send loyalty rewards).
All CustomersCluster 1: Young SpendersCluster 2: High-Value BuyersCluster 3: Occasional Shoppers
K-Means clustering (K=3) applied to customer segments. Visualizes the text's three output clusters.

3. Association Rule Mining

  • How it works: Finds frequent co-occurring patterns (e.g., "People who buy X also buy Y").
  • Example: BigMart (Nepal) uses this to place products like "diapers and beer" near each other.

Worked Example: Market Basket Analysis Dataset (Transactions at a grocery store):

Transaction Items Purchased
T1 Bread, Milk, Eggs
T2 Bread, Diapers, Beer
T3 Milk, Diapers, Beer

Rules Generated:

  1. {Bread} → {Milk} (Support = 2/3, Confidence = 100%)
  2. {Diapers, Beer} → {Bread} (Support = 2/3, Confidence = 100%)

Interpretation: Stores can increase sales by placing diapers and beer near bread.


4. Prediction (Regression)

  • How it works: Predicts continuous values (e.g., house prices, stock trends).
  • Example: Nepal Stock Exchange (NEPSE) uses regression to forecast index movements.

Worked Example: House Price Prediction Features:

  • Area (sq. ft)
  • Number of Bedrooms
  • Location (Kathmandu/Pokhara)

Model: Linear Regression. Equation:

Prediction:

  • House in Kathmandu: 1500 sq. ft, 3 bedrooms → Price = 5000×1500 + 200000×3 + 300000 = NPR 3,750,000.

Challenges in Data Mining

  1. Data Quality Issues:
    • Missing values, noise, or inconsistencies (e.g., NTC call detail records with incorrect timestamps).
    • Solution: Clean data using imputation or outlier detection.
02468Data Quality7Scalability8Interpretability6Ethical Concerns5Frequency of Challenges (out of 10 case studies)
Top challenges in data mining projects (based on TU exam case studies).
  1. Scalability:

    • Large datasets (e.g., Pathao’s 1M+ daily rides) require optimized algorithms.
    • Solution: Use distributed systems like Apache Spark.
  2. Interpretability:

    • Complex models (e.g., deep neural networks) are hard to explain.
    • Solution: Use simpler models (e.g., decision trees) or SHAP values for transparency.
  3. Privacy Concerns:

    • Mining sensitive data (e.g., Khalti transaction histories) raises ethical issues.
    • Solution: Anonymize data or comply with Nepal’s Data Privacy Act.

Data Mining Process (CRISP-DM Framework)

Steps Explained:

  1. Business Understanding: Define goals (e.g., "Reduce customer churn for Ncell by 10%").
  2. Data Understanding: Explore data (e.g., analyze call logs for churn patterns).
  3. Data Preparation: Clean and transform data (e.g., handle missing call durations).
  4. Modeling: Apply algorithms (e.g., Random Forest for churn prediction).
  5. Evaluation: Test model accuracy (e.g., AUC-ROC = 0.85).
  6. Deployment: Integrate into Ncell’s CRM system for real-time alerts.

Exam Tip

What to Expect in TU/PU Exams

  1. Definitions:

    • Expect 2–3 marks for defining data mining, supervised/unsupervised learning, or classification.
    • Example Question: "Define association rule mining with support and confidence."
  2. Comparisons:

    • 4–5 marks for comparing supervised vs. unsupervised learning or classification vs. clustering.
    • Example Question: "Differentiate between classification and prediction tasks with examples."
  3. Worked Examples:

    • 5–7 marks for calculating support/confidence (association rules) or interpreting a decision tree.
    • Example Question: "Given a dataset, compute the support and confidence for the rule {Milk} → {Bread}."
  4. Real-World Applications:

    • 3–4 marks for linking concepts to Nepalese companies (e.g., eSewa’s recommendation system).
    • Example Question: "How does Khalti use data mining to detect fraud?"
  5. Diagrams:

    • 3 marks for drawing a decision tree or clustering dendrogram.
    • Example Question: "Draw a decision tree for the given dataset."

How to Score Full Marks

  • For definitions: Use official syllabus language (e.g., "Data mining is the process of extracting patterns from large datasets...").
  • For comparisons: Use tables with real examples (e.g., "Ncell uses supervised learning for churn prediction").
  • For calculations: Show every step (e.g., support = frequency of {Milk, Bread} / total transactions).
  • For diagrams: Label all nodes clearly (e.g., "Credit Score > 500?").

IMAGE: "data mining exam question paper" | A screenshot of a past TU exam question on classification accuracy, with the correct answer highlighted.

Based on the TU BIM syllabus for Data Warehousing and Data Mining (IT274), unit 5.

Discussion

Loading…