Data Warehousing and Data MiningUnit 511 min read
Data Mining: Definitions, Tasks, Techniques & Real-World Impact
Unit 5 of Data Warehousing and Data Mining introduces the core concepts of data mining—its definition, key tasks (classification, prediction, clustering, association), and techniques—while linking them to real-world applications in Nepalese and global industries like eSewa, Daraz, and Ncell.
What is Data Mining?
Data mining is the process of discovering meaningful patterns, correlations, and insights from large datasets using methods from machine learning, statistics, and database systems. It transforms raw data into actionable knowledge.
Key Characteristics of Data Mining
mindmap
root((Data Mining))
Definition["Extracts hidden patterns from large datasets"]
Techniques["Machine Learning, Statistics, AI"]
Goals["Decision support, automation, prediction"]
Challenges["Noise, scalability, interpretability"]
Applications["Fraud detection, recommendation systems, market basket analysis"]Core Data Mining Tasks
Data mining tasks are categorized into descriptive (summarizing data) and predictive (forecasting future trends). The four primary tasks are:
| Task | Definition | Example |
|---|---|---|
| Classification | Assigns data to predefined categories based on features. | Spam email detection (spam vs. not spam). |
| Prediction | Estimates future values or trends (regression or time-series forecasting). | Stock price prediction for NEPSE shares. |
| Clustering | Groups similar data points without predefined labels (unsupervised learning). | Customer segmentation for Daraz (e.g., "frequent buyers" vs. "one-time shoppers"). |
| Association | Finds relationships between variables (e.g., "people who buy X also buy Y"). | Market basket analysis in grocery stores (e.g., "diapers + beer"). |
1. Classification: Supervised Learning for Categorization
Classification uses labeled data to train models that predict categories. Common algorithms:
- Decision Trees (e.g., ID3, C4.5)
- Naive Bayes (probabilistic classifier)
- Support Vector Machines (SVM)
- Neural Networks (for complex patterns)
Worked Example: Loan Approval Prediction (Nepalese Bank Scenario)
Problem: A bank wants to predict whether a customer will default on a loan based on income, credit score, and employment status.
Dataset:
| Income (₹) | Credit Score | Employed? | Default? (Label) |
|---|---|---|---|
| 50,000 | 700 | Yes | No |
| 30,000 | 500 | No | Yes |
| 80,000 | 800 | Yes | No |
Steps:
- Preprocess: Normalize income (scale to 0–1) and encode "Employed" as binary (Yes=1, No=0).
- Train a Decision Tree:
- Split on Credit Score > 600 → If Yes, predict "No Default"; else, check Income.
- If Income > 40,000 and Employed=Yes → "No Default"; else → "Default".
- Visualize the Tree:
Output: The model predicts a new applicant with ₹45,000 income, 550 credit score, and unemployed as "Default" (high risk).
2. Prediction: Forecasting Numerical Outcomes
Prediction involves estimating continuous values (regression) or future trends (time-series). Techniques:
- Linear Regression (for trends like sales growth).
- Time-Series Analysis (e.g., NTC’s electricity demand forecasting).
- Neural Networks (for complex patterns like stock prices).
Worked Example: Electricity Demand Forecasting (NTC)
Problem: NTC wants to predict tomorrow’s electricity demand in Kathmandu based on historical data.
Dataset (Simplified):
| Hour | Temperature (°C) | Demand (MW) |
|---|---|---|
| 1 | 20 | 120 |
| 2 | 18 | 110 |
| ... | ... | ... |
Steps:
- Feature Engineering: Use temperature and time (hour/day) as inputs.
- Train a Linear Regression Model:
- Equation:
Demand = 5 * Temperature + 10 * Hour + 80 - For Hour=12, Temperature=25°C:
Demand = 5*25 + 10*12 + 80 = 125 + 120 + 80 = 325 MW.
- Equation:
- Visualize the Trend:
Output: NTC schedules 325 MW for peak hours to avoid outages.
3. Clustering: Unsupervised Grouping
Clustering groups similar data points without labels. Common algorithms:
- K-Means (partitions data into k clusters).
- Hierarchical Clustering (nested clusters).
- DBSCAN (density-based, handles noise).
Worked Example: Customer Segmentation for Daraz
Problem: Daraz wants to group customers based on purchase history to tailor marketing.
Dataset (Simplified):
| Customer | Avg. Spend (₹) | Purchase Frequency (month) | Location |
|---|---|---|---|
| A | 5,000 | 4 | Kathmandu |
| B | 1,000 | 12 | Pokhara |
| C | 20,000 | 2 | Lalitpur |
Steps:
- Normalize Data: Scale spend and frequency to 0–1.
- Apply K-Means (k=3):
- Cluster 1: High spend, low frequency (e.g., Customer C → "Premium Buyers").
- Cluster 2: Medium spend, medium frequency (e.g., Customer A → "Regulars").
- Cluster 3: Low spend, high frequency (e.g., Customer B → "Bargain Hunters").
- Visualize Clusters:
Output: Daraz sends discounts to Cluster 3 (Bargain Hunters) and premium services to Cluster 1.
4. Association Rule Mining: "If-Then" Relationships
Finds frequent patterns in transactional data. Metrics:
- Support: Frequency of an itemset (e.g., "diapers" in 30% of transactions).
- Confidence: Likelihood of the "then" part (e.g., "If diapers, then beer: 70%").
- Lift: How much more likely the rule is than random chance.
Worked Example: Market Basket Analysis (Local Grocery Store)
Problem: A grocery store in Kathmandu wants to find products frequently bought together.
Dataset (Transactions):
| Transaction ID | Items Purchased |
|---|---|
| 1 | Milk, Bread, Diapers |
| 2 | Beer, Diapers, Eggs |
| 3 | Milk, Bread, Beer |
Steps:
- Generate Itemsets:
- Single items: {Milk}, {Bread}, {Diapers}, {Beer}, {Eggs}.
- Pairs: {Milk, Bread} (appears in 2/3 transactions → Support = 66.7%).
- Calculate Confidence:
- Rule:
{Diapers} → {Beer}- Confidence = P(Beer|Diapers) = 2/3 ≈ 66.7%.
- Rule:
{Milk, Bread} → {Beer}- Confidence = 1/2 = 50% (only Transaction 3).
- Rule:
- Filter by Lift:
- Lift of
{Diapers} → {Beer}= (66.7% / 33.3%) ≈ 2 → Strong association.
- Lift of
Output: The store places beer near diapers to increase sales.
In the Real World
eSewa (Nepal):
- Task: Classification (fraud detection).
- How: Uses decision trees to flag suspicious transactions (e.g., sudden large payments from a low-income user).
- Impact: Reduces financial fraud by 40% (per eSewa reports).
Daraz (Nepal):
- Task: Association rule mining + clustering.
- How: "Frequently bought together" recommendations (e.g., "Customers who viewed this also bought...") and customer segmentation for targeted ads.
- Impact: Increases average order value by 15%.
Ncell (Nepal):
- Task: Prediction (churn analysis).
- How: Uses logistic regression to predict which customers might switch to NTC, then offers discounts to retain them.
- Impact: Reduces customer churn by 20%.
Google (Global):
- Task: Clustering + classification.
- How: Google News groups similar articles into clusters (e.g., "COVID-19 updates") and classifies ads for targeted display.
- Impact: Powers 90% of Google Ads revenue.
Pathao (Nepal):
- Task: Prediction (demand forecasting).
- How: Uses time-series analysis to predict peak hours in Kathmandu (e.g., 7–9 AM, 6–8 PM) and adjusts driver incentives.
- Impact: Reduces wait times by 30% during rush hours.
Data Mining Techniques vs. Traditional Statistics
| Aspect | Data Mining | Traditional Statistics |
|---|---|---|
| Data Volume | Handles large datasets (terabytes). | Works with smaller, structured datasets. |
| Approach | Automated, pattern discovery. | Hypothesis-driven, manual analysis. |
| Tools | Weka, RapidMiner, Python (scikit-learn). | R, SPSS, Excel. |
| Output | Actionable patterns (e.g., "Buy X, get Y"). | Statistical significance (p-values). |
| Example Use Case | Fraud detection in eSewa. | A/B testing for a Daraz ad campaign. |
Challenges in Data Mining
Data Quality Issues:
- Missing values, noise, or inconsistent formats (e.g., "KTM" vs. "Kathmandu" in addresses).
- Solution: Cleaning (imputation, normalization) and transformation.
Scalability:
- Algorithms like K-Means struggle with millions of records.
- Solution: Use distributed systems (e.g., Apache Spark) or sampling.
Interpretability:
- Models like neural networks are "black boxes."
- Solution: Use simpler models (decision trees) or explainable AI (SHAP values).
Privacy:
- Mining personal data (e.g., Khalti transactions) raises ethical concerns.
- Solution: Anonymization (e.g., k-anonymity) and compliance with laws like Nepal’s Digital Transaction Act.
Exam Tip
This unit is theoretical but heavily tested on applications. Expect:
- Definitions: Know the difference between classification, prediction, clustering, and association.
- Algorithms: Be able to sketch a decision tree or explain how K-Means works (e.g., "minimize within-cluster variance").
- Worked Examples: Practice with small datasets (like the loan or Daraz examples above). Show all steps, including:
- Data preprocessing (normalization, encoding).
- Model training (e.g., "split on Credit Score > 600").
- Evaluation (e.g., "accuracy = 80%").
- Real-World Links: Connect concepts to Nepalese companies:
- eSewa → Fraud detection (classification).
- Daraz → Market basket analysis (association rules).
- Ncell → Churn prediction (logistic regression).
- Diagrams: Always draw:
- Decision trees for classification.
- Cluster visualizations for K-Means.
- Association rule examples (e.g., "diapers → beer").
Common Pitfalls:
- Confusing supervised (classification/prediction) vs. unsupervised (clustering) learning.
- Forgetting to preprocess data (e.g., scaling for K-Means).
- Overlooking evaluation metrics (e.g., support/confidence for association rules).
Based on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 5.
Discussion
Loading…