Data Warehousing and Data MiningUnit 214 min read
Data Mining: Definitions, Tasks, Techniques & Real-World Impact
Unit 2 of Data Warehousing and Data Mining introduces the core concepts of data mining—its definition, key tasks (classification, clustering, association, prediction), techniques (supervised vs. unsupervised), and real-world applications in Nepalese and global industries. This note explains how data mining extracts hid
What is Data Mining?
Data mining is the process of discovering meaningful patterns, correlations, and insights from large datasets using statistical, machine learning, and database techniques. It bridges the gap between raw data and actionable knowledge.
Key Characteristics of Data Mining
mindmap
root((Data Mining))
Definition: "Extracting patterns from data"
Goals: ["Prediction", "Description", "Discovery"]
Techniques: ["Machine Learning", "Statistics", "Database Systems"]
Applications: ["Business", "Healthcare", "Finance", "Social Networks"]
Challenges: ["Data Quality", "Scalability", "Privacy"]Why is it called "mining"? Just like mining gold from the earth, data mining extracts valuable knowledge from vast amounts of raw data (e.g., transaction records, social media posts, sensor logs).
Core Tasks in Data Mining
Data mining performs five primary tasks, each serving different analytical needs:
| Task | Definition | Example in Nepal | Example Worldwide |
|---|---|---|---|
| Classification | Assigns data to predefined categories based on labeled training data. | Predicting loan defaults for Nabil Bank customers using past data. | Google’s spam email classifier (labels emails as spam/non-spam). |
| Prediction | Estimates future values or trends (a type of classification/regression). | Forecasting NEPSE stock prices using historical trends. | Netflix’s movie recommendation system (predicts user ratings). |
| Clustering | Groups similar data points without predefined labels (unsupervised learning). | Segmenting Daraz customers into groups (e.g., high-spenders, occasional buyers). | Facebook’s friend suggestion system (groups users by interests). |
| Association | Finds relationships between variables (e.g., "people who buy X also buy Y"). | Khalti’s fraud detection: If a transaction > Rs. 50,000 is followed by a refund, flag it. | Amazon’s "Frequently bought together" suggestions. |
| Anomaly Detection | Identifies rare or unusual patterns (e.g., fraud, errors). | Detecting fake Ncell SIM registrations using unusual call patterns. | Credit card fraud detection (e.g., sudden large transactions in a new location). |
Supervised vs. Unsupervised Learning
Data mining techniques are broadly classified into two types based on whether the data is labeled or not.
1. Supervised Learning (Labeled Data)
- Uses pre-labeled data to train models.
- Goal: Predict outcomes for new data.
- Examples:
- Classification: Spam detection (labeled as "spam" or "not spam").
- Regression: Predicting house prices based on size/location.
flowchart LR
A["Training Data (Labeled)"] --> B["Model Training"]
B --> C["Prediction on New Data"]
C --> D["Output: Class/Value"]Real-World Example:
- eSewa’s Bill Payment Prediction:
- Task: Predict if a user will pay their electricity bill on time.
- Data: Past payment history, user location, bill amount.
- Model: Logistic regression (classifies as "paid" or "unpaid").
- Outcome: eSewa sends reminders only to high-risk users, reducing defaults.
2. Unsupervised Learning (Unlabeled Data)
- Works with unlabeled data to find hidden patterns.
- Goal: Discover structures or groupings.
- Examples:
- Clustering: Grouping customers by purchasing behavior.
- Association: Market basket analysis (e.g., "beer and diapers" correlation).
flowchart LR
A["Unlabeled Data"] --> B["Pattern Discovery"]
B --> C["Clusters/Associations"]
C --> D["Output: Groups/Relationships"]Real-World Example:
- Pathao’s Driver Routing Optimization:
- Task: Group drivers by efficiency (e.g., fast pickups, low cancellation rates).
- Technique: K-means clustering (groups drivers into clusters based on trip data).
- Outcome: Pathao assigns high-demand routes to the most efficient drivers, reducing wait times.
Data Mining Techniques & Algorithms
Different algorithms solve specific data mining tasks. Below are the most common ones:
| Technique | Algorithm Examples | When to Use | Nepalese Use Case |
|---|---|---|---|
| Classification | Decision Trees, SVM, Naive Bayes | Predicting categories (e.g., loan approval, disease diagnosis). | Nepal Rastra Bank’s credit scoring for SME loans. |
| Clustering | K-means, Hierarchical, DBSCAN | Grouping similar data (e.g., customer segmentation). | Daraz’s dynamic pricing: Groups products by demand elasticity. |
| Association | Apriori, FP-Growth | Finding co-occurring items (e.g., market basket analysis). | Foodland’s "Buy X, Get Y Free" promotions. |
| Regression | Linear Regression, Neural Networks | Predicting continuous values (e.g., sales, stock prices). | NTC’s network traffic forecasting for bandwidth allocation. |
| Anomaly Detection | Isolation Forest, Autoencoders | Detecting outliers (e.g., fraud, network intrusions). | Ncell’s SIM box detection (unusual call patterns). |
The Data Mining Process (CRISP-DM Model)
The Cross-Industry Standard Process for Data Mining (CRISP-DM) is a structured approach to data mining projects:
flowchart TD
A["Business Understanding"] --> B["Data Understanding"]
B --> C["Data Preparation"]
C --> D["Modeling"]
D --> E["Evaluation"]
E --> F["Deployment"]
F -->|"Feedback Loop"| AStep-by-Step Breakdown
Business Understanding
- Define objectives (e.g., "Reduce customer churn for Ncell by 15%").
- Identify success metrics (e.g., retention rate).
Data Understanding
- Collect and explore data (e.g., Ncell’s call logs, customer demographics).
- Check for missing values, outliers.
Data Preparation (Most Time-Consuming!)
- Clean data (handle missing values, remove duplicates).
- Transform data (normalize, aggregate).
- Example: Convert Khalti transaction timestamps into daily/weekly patterns.
Modeling
- Apply algorithms (e.g., decision trees for churn prediction).
- Tune hyperparameters (e.g., number of clusters in K-means).
Evaluation
- Test model accuracy (e.g., 90% precision in fraud detection).
- Compare with baseline (e.g., random guessing vs. actual model).
Deployment
- Integrate model into business workflow (e.g., Pathao’s dynamic pricing API).
- Monitor performance over time.
In the Real World
Data mining is everywhere—from Nepalese fintech apps to global tech giants. Here’s how companies use it:
1. eSewa & Khalti: Fraud Detection
- Idea Used: Anomaly Detection + Classification
- How?
- Khalti uses Isolation Forest to detect unusual transactions (e.g., a Rs. 10,000 transfer at 3 AM from a new device).
- eSewa applies decision trees to flag high-risk bill payments (e.g., users with a history of late payments).
- Impact: Reduces fraudulent transactions by 30% (Khalti’s internal report, 2023).
2. Daraz: Personalized Recommendations
- Idea Used: Collaborative Filtering (Clustering + Association)
- How?
- Daraz groups customers using K-means clustering (e.g., "budget shoppers," "luxury buyers").
- Uses Apriori algorithm to suggest products (e.g., "Customers who bought Samsung phones also bought earphones").
- Impact: 20% increase in average order value (Daraz Nepal, 2022).
3. Ncell: Network Optimization
- Idea Used: Regression + Time-Series Forecasting
- How?
- Ncell uses ARIMA models to predict peak call hours in Kathmandu.
- Decision trees optimize tower placements in low-coverage areas (e.g., rural Chitwan).
- Impact: 15% reduction in dropped calls during festivals (Dashain, Tihar).
4. Nepal Rastra Bank (NRB): Economic Forecasting
- Idea Used: Regression Analysis
- How?
- NRB uses linear regression to predict inflation based on past GDP growth, import/export data.
- Clustering identifies economic regions (e.g., Kathmandu Valley vs. Far-Western provinces).
- Impact: Helps set interest rates and monetary policies.
Worked Example: Predicting Loan Defaults (Nabil Bank)
Let’s apply classification to predict if a loan applicant will default.
Step 1: Data Collection
| Feature | Description |
|---|---|
Income (Rs.) |
Annual salary |
LoanAmount (Rs.) |
Amount requested |
CreditScore (0-850) |
Past credit history |
EmploymentYears |
Years at current job |
Default (0/1) |
Target variable (1 = defaulted) |
Sample Data (5 Applicants):
| Income (Rs.) | LoanAmount | CreditScore | EmploymentYears | Default |
|---|---|---|---|---|
| 800,000 | 200,000 | 720 | 5 | 0 |
| 500,000 | 150,000 | 650 | 2 | 1 |
| 1,200,000 | 300,000 | 780 | 8 | 0 |
| 450,000 | 250,000 | 580 | 1 | 1 |
| 900,000 | 100,000 | 750 | 3 | 0 |
Step 2: Choose a Model
We’ll use a Decision Tree (simple and interpretable).
graph TD
A["Income ≤ 600,000?"] -->|"Yes"| B["Default = 1"]
A -->|"No"| C["CreditScore ≤ 700?"]
C -->|"Yes"| D["Default = 1"]
C -->|"No"| E["Default = 0"]Step 3: Predict a New Applicant
New Applicant:
- Income: Rs. 550,000
- LoanAmount: Rs. 180,000
- CreditScore: 680
- EmploymentYears: 1
Prediction:
- Income ≤ 600,000? → Yes → Default = 1 (High risk!).
- Nabil Bank’s Action: Reject loan or offer at a higher interest rate.
Why?
- Low income + low credit score → historically high default rate in the training data.
Advantages and Disadvantages of Data Mining
| Advantages | Disadvantages |
|---|---|
| Reveals hidden patterns in data. | Requires large datasets (garbage in → garbage out). |
| Automates decision-making (e.g., fraud detection). | Privacy risks (e.g., misuse of personal data). |
| Improves business efficiency (e.g., targeted marketing). | High computational cost for complex models. |
| Predicts future trends (e.g., stock prices). | Overfitting (model works only on training data). |
Exam Tip
What to Expect in TU/PU Exams
Definitions:
- Be ready to define data mining, supervised vs. unsupervised learning, and CRISP-DM.
- Example question: "Differentiate between classification and clustering with an example each."
Algorithm Applications:
- Expect questions like: "Which data mining technique would you use to group similar customers in an e-commerce site? Justify your choice." Answer: K-means clustering (unsupervised, groups unlabeled data).
Real-World Scenarios:
- Questions may ask: "How does Khalti use data mining to detect fraud? Explain with one algorithm." Answer: Isolation Forest (anomaly detection) flags transactions deviating from normal patterns.
Process Diagrams:
- Draw the CRISP-DM flowchart or a decision tree from scratch.
- Example: "Illustrate the steps of the data mining process using a flowchart."
Numerical Problems:
- Simple classification/prediction examples (like the Nabil Bank loan default above).
- Example: "Given a dataset of mobile users, predict churn using a decision tree. Show the tree structure."
Common Mistakes to Avoid
- Confusing supervised vs. unsupervised: Remember supervised = labeled data, unsupervised = unlabeled.
- Ignoring data preprocessing: Real-world data is messy! Always mention cleaning/transformation.
- Overcomplicating answers: Exams test concepts, not coding. Use simple examples (like the loan table above).
Key Formulas to Remember
| Concept | Formula |
|---|---|
| Accuracy (Classification) | |
| Precision | |
| Recall (Sensitivity) | |
| Euclidean Distance (Clustering) |
Where TP = True Positive, TN = True Negative, FP = False Positive, FN = False Negative
Summary Checklist
Before the exam, ensure you can: ✅ Define data mining and its five core tasks. ✅ Differentiate supervised vs. unsupervised learning with examples. ✅ Explain the CRISP-DM process and draw its flowchart. ✅ Name three algorithms for classification/clustering and their uses. ✅ Solve a simple classification problem (like the loan example). ✅ Relate data mining to Nepalese companies (eSewa, Khalti, Ncell, etc.).
Based on the TU BIT syllabus for Data Warehousing and Data Mining, unit 2.
Discussion
Loading…