Data Warehousing and Data MiningUnit 512 min read
Data Mining: Definitions, Tasks, Techniques & Real-World Impact
Unit 5 of Data Warehousing and Data Mining introduces the core concepts of data mining—its definition, key tasks (classification, prediction, clustering, etc.), techniques (supervised vs. unsupervised learning), and real-world applications in Nepalese and global industries. This note explains how data mining extracts a
What is Data Mining?
Data mining is the process of discovering meaningful patterns, correlations, and insights from large datasets using statistical, machine learning, and database techniques. It bridges the gap between raw data and actionable knowledge, enabling businesses to make data-driven decisions.
Key Characteristics of Data Mining
mindmap
root((Data Mining))
Characteristics
Large Datasets: Handles terabytes/petabytes of data
Automated Discovery: Uncovers hidden patterns without manual intervention
Pattern-Oriented: Focuses on trends, clusters, and associations
Non-Trivial: Extracts knowledge beyond simple queries
Practical Evaluation: Results must be useful and interpretableWhy is it called "mining"?
- Just as miners extract valuable minerals from the earth, data miners extract valuable knowledge from vast datasets.
In the Real World
- eSewa (Nepal) uses association rule mining to recommend services (e.g., "Users who paid utility bills also booked flights") by analyzing transaction logs. This increases cross-selling by 15%.
- Pathao (ride-hailing app) employs clustering algorithms to group high-demand areas in Kathmandu, optimizing driver dispatch and reducing wait times by 20%.
- Nepal Rastra Bank (NRB) applies fraud detection models (a classification task) to flag suspicious transactions in real-time, preventing financial crimes worth NPR 500M+ annually.
IMAGE: "data mining process flowchart" | A labeled diagram showing the stages: Data Collection → Data Cleaning → Data Transformation → Mining → Pattern Evaluation → Knowledge Presentation.
Core Tasks in Data Mining
Data mining tasks are categorized into predictive (forecasting future trends) and descriptive (summarizing past data). The syllabus emphasizes four primary tasks:
| Task | Definition | Example in Nepal | Techniques Used |
|---|---|---|---|
| Classification | Categorizes data into predefined groups based on labeled training data. | Predicting loan defaults (yes/no) for Nabil Bank customers. | Decision Trees, Naive Bayes, SVM, Neural Nets |
| Prediction | Estimates future values or trends (regression or time-series forecasting). | Forecasting NEPSE stock prices for next quarter. | Linear Regression, Time-Series Analysis |
| Clustering | Groups unlabeled data into clusters based on similarity. | Segmenting Daraz customers into "high-spenders," "bargain hunters," etc. | K-Means, Hierarchical Clustering, DBSCAN |
| Association | Finds relationships between variables (e.g., "if-then" rules). | "Customers who buy Khalti recharge cards also buy mobile data plans." | Apriori, FP-Growth Algorithms |
Worked Example: Classification (Loan Approval)
Scenario: A bank (e.g., Global IME Bank) wants to predict whether a customer will default on a loan based on:
- Age (25, 30, 40, 50)
- Income (NPR 50K, 100K, 200K)
- Credit Score (300–850)
- Label: Default (Yes/No)
Step 1: Data Preprocessing
- Handle missing values (e.g., impute average income for missing entries).
- Normalize credit scores to a 0–1 scale.
Step 2: Choose an Algorithm
- Decision Tree (simple, interpretable) or Random Forest (more accurate).
Step 3: Train the Model
from sklearn.tree import DecisionTreeClassifier
model = DecisionTreeClassifier(criterion='gini', max_depth=3)
model.fit(X_train, y_train) # X_train: [Age, Income, Credit Score], y_train: [Default]
Step 4: Predict & Evaluate
- Input: Age=30, Income=100K, Credit Score=650 → Predicted Output: Default = No (92% confidence).
- Evaluation Metrics: Accuracy = 88%, Precision = 85%, Recall = 80%.
Visualization of the Decision Tree:
Interpretation: Customers with credit scores ≤ 500 are automatically rejected, while others are evaluated based on income and age.
Supervised vs. Unsupervised Learning
| Aspect | Supervised Learning | Unsupervised Learning |
|---|---|---|
| Data Labels | Uses labeled data (e.g., "Default: Yes/No"). | Uses unlabeled data (no predefined categories). |
| Goal | Predicts or classifies new data. | Discovers hidden patterns or groupings. |
| Examples | Spam detection, stock price prediction. | Customer segmentation, anomaly detection. |
| Algorithms | Decision Trees, SVM, Neural Networks. | K-Means, PCA, Apriori. |
| Real-World Use | Ncell predicts churn (customer attrition). | NTC clusters high-traffic areas for network optimization. |
IMAGE: "supervised vs unsupervised learning comparison" | A side-by-side table with icons (e.g., a teacher for supervised, a detective for unsupervised).
Data Mining Techniques
1. Classification (Supervised)
- How it works: The model learns from labeled examples (e.g., "Customer X defaulted because...") and generalizes to new data.
- Example: Khalti uses classification to detect fraudulent transactions by training on past fraud cases.
Worked Example: Fraud Detection Dataset:
| Transaction ID | Amount (NPR) | Location | Time | Is Fraud? |
|---|---|---|---|---|
| T001 | 50,000 | Kathmandu | 3 AM | Yes |
| T002 | 10,000 | Pokhara | 2 PM | No |
Algorithm: Logistic Regression (outputs probability of fraud). Decision Rule: Flag transactions with P(fraud) > 0.7.
2. Clustering (Unsupervised)
- How it works: Groups similar data points without prior labels (e.g., grouping customers by spending habits).
- Example: Daraz uses clustering to identify "VIP customers" for personalized discounts.
Worked Example: Customer Segmentation Dataset: 100 customers with features: [Age, Income, Purchase Frequency]. Algorithm: K-Means (K=3 clusters). Output:
- Cluster 1: Young, low-income, frequent small purchases (target with flash sales).
- Cluster 2: Middle-aged, high-income, occasional large purchases (offer premium services).
- Cluster 3: Old, medium-income, rare purchases (send loyalty rewards).
3. Association Rule Mining
- How it works: Finds frequent co-occurring patterns (e.g., "People who buy X also buy Y").
- Example: BigMart (Nepal) uses this to place products like "diapers and beer" near each other.
Worked Example: Market Basket Analysis Dataset (Transactions at a grocery store):
| Transaction | Items Purchased |
|---|---|
| T1 | Bread, Milk, Eggs |
| T2 | Bread, Diapers, Beer |
| T3 | Milk, Diapers, Beer |
Rules Generated:
- {Bread} → {Milk} (Support = 2/3, Confidence = 100%)
- {Diapers, Beer} → {Bread} (Support = 2/3, Confidence = 100%)
Interpretation: Stores can increase sales by placing diapers and beer near bread.
4. Prediction (Regression)
- How it works: Predicts continuous values (e.g., house prices, stock trends).
- Example: Nepal Stock Exchange (NEPSE) uses regression to forecast index movements.
Worked Example: House Price Prediction Features:
- Area (sq. ft)
- Number of Bedrooms
- Location (Kathmandu/Pokhara)
Model: Linear Regression. Equation:
Prediction:
- House in Kathmandu: 1500 sq. ft, 3 bedrooms → Price = 5000×1500 + 200000×3 + 300000 = NPR 3,750,000.
Challenges in Data Mining
- Data Quality Issues:
- Missing values, noise, or inconsistencies (e.g., NTC call detail records with incorrect timestamps).
- Solution: Clean data using imputation or outlier detection.
Scalability:
- Large datasets (e.g., Pathao’s 1M+ daily rides) require optimized algorithms.
- Solution: Use distributed systems like Apache Spark.
Interpretability:
- Complex models (e.g., deep neural networks) are hard to explain.
- Solution: Use simpler models (e.g., decision trees) or SHAP values for transparency.
Privacy Concerns:
- Mining sensitive data (e.g., Khalti transaction histories) raises ethical issues.
- Solution: Anonymize data or comply with Nepal’s Data Privacy Act.
Data Mining Process (CRISP-DM Framework)
Steps Explained:
- Business Understanding: Define goals (e.g., "Reduce customer churn for Ncell by 10%").
- Data Understanding: Explore data (e.g., analyze call logs for churn patterns).
- Data Preparation: Clean and transform data (e.g., handle missing call durations).
- Modeling: Apply algorithms (e.g., Random Forest for churn prediction).
- Evaluation: Test model accuracy (e.g., AUC-ROC = 0.85).
- Deployment: Integrate into Ncell’s CRM system for real-time alerts.
Exam Tip
What to Expect in TU/PU Exams
Definitions:
- Expect 2–3 marks for defining data mining, supervised/unsupervised learning, or classification.
- Example Question: "Define association rule mining with support and confidence."
Comparisons:
- 4–5 marks for comparing supervised vs. unsupervised learning or classification vs. clustering.
- Example Question: "Differentiate between classification and prediction tasks with examples."
Worked Examples:
- 5–7 marks for calculating support/confidence (association rules) or interpreting a decision tree.
- Example Question: "Given a dataset, compute the support and confidence for the rule {Milk} → {Bread}."
Real-World Applications:
- 3–4 marks for linking concepts to Nepalese companies (e.g., eSewa’s recommendation system).
- Example Question: "How does Khalti use data mining to detect fraud?"
Diagrams:
- 3 marks for drawing a decision tree or clustering dendrogram.
- Example Question: "Draw a decision tree for the given dataset."
How to Score Full Marks
- For definitions: Use official syllabus language (e.g., "Data mining is the process of extracting patterns from large datasets...").
- For comparisons: Use tables with real examples (e.g., "Ncell uses supervised learning for churn prediction").
- For calculations: Show every step (e.g., support = frequency of {Milk, Bread} / total transactions).
- For diagrams: Label all nodes clearly (e.g., "Credit Score > 500?").
IMAGE: "data mining exam question paper" | A screenshot of a past TU exam question on classification accuracy, with the correct answer highlighted.
Based on the TU BIM syllabus for Data Warehousing and Data Mining (IT274), unit 5.
Discussion
Loading…