Data Warehousing and Data MiningUnit 1018 min read
Complex Data Mining & Real-World Applications
Unit 10 of Data Warehousing and Data Mining explores mining unstructured data (text, images, video), spatial/temporal patterns, and real-world applications like fraud detection, recommendation systems, and social network analysis—with case studies from Nepalese and global tech.
TAKEAWAYS:
- Complex data types (text, multimedia, spatial/temporal) require specialized mining techniques beyond structured data.
- Text mining extracts insights from unstructured data (e.g., customer reviews, news) using NLP and topic modeling.
- Spatial data mining uncovers patterns in geographic data (e.g., traffic congestion, retail locations) via clustering and association rules.
- Temporal data mining analyzes time-series trends (e.g., stock prices, weather forecasts) using sequence prediction and anomaly detection.
- Real-world applications include fraud detection (banks), recommendation systems (YouTube, Daraz), and social network analysis (Pathao driver routing).
- Challenges include scalability, noise, and interpretability—addressed via hybrid models and visualization tools.
1. Introduction to Complex Data Mining
Complex data refers to unstructured or semi-structured data that lacks a fixed schema, including:
- Text data (emails, social media, documents).
- Multimedia data (images, audio, video).
- Spatial data (GIS, maps, satellite imagery).
- Temporal data (time-series, logs, sensor data).
- Graph/social network data (user interactions, fraud networks).
Unlike structured data (tables in databases), complex data requires domain-specific preprocessing (e.g., NLP for text, image segmentation for visuals) before mining.
Why Mine Complex Data?
- Hidden patterns: Customer sentiment in reviews, fraudulent transactions in bank logs.
- Automation: Chatbots (Khalti customer support), autonomous vehicles (Pathao’s route optimization).
- Decision support: Disease outbreak prediction (NTC’s network traffic), stock market trends (NEPSE).
2. Text Mining: Extracting Insights from Unstructured Data
Text mining combines Natural Language Processing (NLP) and data mining to analyze text for patterns.
Key Techniques
| Technique | Purpose | Example Tools/Libraries |
|---|---|---|
| Tokenization | Split text into words/phrases. | NLTK, spaCy |
| Stopword Removal | Filter common words (e.g., "the", "is"). | Python’s nltk.corpus |
| Stemming/Lemmatization | Reduce words to root form (e.g., "running" → "run"). | Porter Stemmer, WordNet |
| TF-IDF | Weight words by importance. | Scikit-learn |
| Topic Modeling | Group related words into themes. | Latent Dirichlet Allocation (LDA) |
| Sentiment Analysis | Classify text as positive/negative. | VADER, TextBlob |
Worked Example: Analyzing Customer Reviews for Daraz
Scenario: Daraz wants to identify common complaints from product reviews to improve inventory. Steps:
- Preprocess:
- Tokenize:
"Product arrived late, poor packaging"→["Product", "arrived", "late", "poor", "packaging"]. - Remove stopwords:
["arrived", "late", "poor", "packaging"]. - Lemmatize:
["arrive", "late", "poor", "package"].
- Tokenize:
- TF-IDF:
- Calculate weights: "late" appears frequently but isn’t unique; "package" is rare → higher weight.
- Topic Modeling (LDA):
- Group words into topics:
- Topic 1:
{"delay", "ship", "late"} - Topic 2:
{"damage", "package", "poor"}
- Topic 1:
- Group words into topics:
- Insight: Daraz can prioritize faster shipping and better packaging.
Real-World Tie-In:
- Khalti uses sentiment analysis to detect fraudulent transaction descriptions (e.g., "urgent payment" flags for review).
- Ncell analyzes customer service call transcripts to predict network outage causes.
flowchart LR
A["Raw Text Data\n(e.g., Daraz reviews)"] --> B["Preprocessing\n(Tokenization, Cleaning)"]
B --> C["Feature Extraction\n(TF-IDF, Word Embeddings)"]
C --> D["Modeling\n(LDA, SVM for Sentiment)"]
D --> E["Insights\n(Topics: Shipping Delays, Product Quality)"]
E --> F["Action\n(Daraz improves logistics)"]3. Multimedia Data Mining: Images, Audio, and Video
Multimedia data is high-dimensional (e.g., a 1000×1000 pixel image = 1M features). Techniques reduce this complexity.
Key Techniques
| Data Type | Preprocessing | Mining Technique | Example Application |
|---|---|---|---|
| Images | Edge detection, segmentation | Clustering (K-means), CNN | Medical imaging (tumor detection) |
| Audio | Spectrogram conversion | Anomaly detection, classification | Voice assistants (Siri, Google) |
| Video | Frame extraction, object tracking | Sequence mining, activity recognition | Surveillance (NTC traffic monitoring) |
Worked Example: Detecting Traffic Violations in Kathmandu
Scenario: NTC uses CCTV footage to flag jaywalking or speeding. Steps:
- Preprocess:
- Extract frames from video → convert to grayscale.
- Apply edge detection (Canny filter) to highlight moving objects.
- Object Detection:
- Use a pre-trained CNN (e.g., YOLO) to classify objects as "pedestrian" or "vehicle."
- Rule-Based Mining:
- If a pedestrian crosses outside zebra lines → violation alert.
- If a vehicle exceeds speed limit (detected via license plate + GPS) → fine.
- Output: Automated ticketing system integrated with Ncell’s digital payment.
Real-World Tie-In:
- Pathao uses image mining to detect driver behavior (e.g., sudden braking) via phone cameras.
- Google Photos clusters images by faces/places using deep learning.
graph TD
A["Video Input\n(NTC CCTV)"] --> B["Frame Extraction"]
B --> C["Edge Detection\n(Canny Filter)"]
C --> D["Object Detection\n(CNN: YOLO)"]
D --> E["Rule Engine\n(Jaywalking? Speeding?)"]
E --> F["Alert System\n(Ncell Payment Link)"]4. Spatial Data Mining: Geographic Patterns
Spatial data mining analyzes geographic or location-based data to find patterns like:
- Hotspots (e.g., crime, retail sales).
- Clustering (e.g., similar neighborhoods).
- Route optimization (e.g., Pathao driver paths).
Key Techniques
| Technique | Purpose | Example |
|---|---|---|
| Spatial Clustering | Group nearby points (e.g., ATMs). | DBSCAN, K-means |
| Spatial Association | Find co-located items (e.g., cafes + bookstores). | Apriori algorithm adapted for GIS |
| Geographic Outliers | Detect unusual locations (e.g., fraudulent transactions). | LOF (Local Outlier Factor) |
Worked Example: Optimizing Daraz Delivery Routes
Scenario: Daraz wants to reduce delivery costs by clustering orders geographically. Steps:
- Data Collection:
- Orders:
[(Kathmandu-12, "Laptop"), (Lalitpur-3, "Phone"), (Bhaktapur-5, "Books")]. - Locations mapped to coordinates:
(KTM: (27.7172, 85.3240), LLP: (27.6833, 85.3197), BKT: (27.6883, 85.4103)).
- Orders:
- Clustering (DBSCAN):
- Set
ε = 5 km(max distance between points to be clustered). - Result:
Cluster 1 = {Kathmandu, Lalitpur}(close),Cluster 2 = {Bhaktapur}(isolated).
- Set
- Route Optimization:
- Assign a single delivery agent to Cluster 1 (KTM → LLP → KTM).
- Assign another to Bhaktapur.
- Savings: Reduces fuel costs by 30% vs. individual deliveries.
Real-World Tie-In:
- Pathao uses spatial clustering to group nearby passengers for drivers.
- NTC analyzes mobile tower data to predict network congestion zones.
5. Temporal Data Mining: Time-Series Analysis
Temporal data (e.g., stock prices, weather) is analyzed for trends, seasonality, and anomalies.
Key Techniques
| Technique | Purpose | Example |
|---|---|---|
| Time-Series Forecasting | Predict future values (e.g., NEPSE stock). | ARIMA, LSTM (deep learning) |
| Anomaly Detection | Flag unusual patterns (e.g., fraud). | Isolation Forest, Autoencoders |
| Sequence Mining | Find frequent patterns (e.g., user behavior). | PrefixSpan algorithm |
Worked Example: Predicting NEPSE Stock Trends
Scenario: An investor wants to predict NEPSE’s next 5 days using past 6 months of data. Steps:
- Data Preparation:
- Time-series:
[Day1: 1200, Day2: 1210, ..., Day180: 1350]. - Features:
Price,Volume,Moving Average (7-day).
- Time-series:
- Model Selection:
- Use ARIMA (AutoRegressive Integrated Moving Average):
- AR (AutoRegressive): Uses past values to predict next.
- I (Integrated): Makes data stationary (removes trends).
- MA (Moving Average): Accounts for error smoothing.
- Use ARIMA (AutoRegressive Integrated Moving Average):
- Training:
- Fit ARIMA(2,1,2) to historical data.
- Parameters:
p=2(lag observations),d=1(differencing),q=2(error terms).
- Prediction:
- Input: Last 7 days’ prices
[1340, 1345, 1350, 1348, 1352, 1355, 1360]. - Output: Predicted prices for next 5 days:
Day181: 1362 Day182: 1365 Day183: 1368 Day184: 1370 Day185: 1372
- Input: Last 7 days’ prices
- Validation:
- Compare with actual data (if available) to check RMSE (Root Mean Squared Error).
Real-World Tie-In:
- Ncell uses time-series forecasting to predict network traffic spikes (e.g., during festivals).
- Google Trends detects viral topics by analyzing search query patterns over time.
6. Graph and Social Network Mining
Graphs represent entities (nodes) and relationships (edges). Used for:
- Fraud detection (e.g., money laundering networks).
- Recommendation systems (e.g., YouTube "Watch Next").
- Community detection (e.g., Pathao driver groups).
Key Techniques
| Technique | Purpose | Example |
|---|---|---|
| Community Detection | Group nodes with dense connections. | Louvain, Girvan-Newman |
| Centrality Measures | Find influential nodes (e.g., key users). | PageRank, Degree Centrality |
| Link Prediction | Predict missing edges (e.g., new friendships). | Common Neighbors algorithm |
Worked Example: Detecting Fraudulent Transactions in Khalti
Scenario: Khalti flags suspicious transaction patterns using graph mining. Steps:
- Graph Construction:
- Nodes: Users (
U1,U2, ...), Transactions (T1,T2). - Edges:
U1 → T1 → U2(money flow).
- Nodes: Users (
- Anomaly Detection:
- High-degree nodes: Users with many transactions in short time → suspicious.
- Short paths: If
U1 → U2 → U3with small amounts but frequent → money laundering.
- PageRank:
- Assign scores to users based on "importance" in the graph.
- High-score users with sudden large transactions → red flag.
- Action:
- Freeze accounts of
U3(score = 0.95) with transaction pattern:U3 → T100 (₹500) → U4 U3 → T101 (₹600) → U5 ...
- Freeze accounts of
Real-World Tie-In:
- Facebook uses graph mining to detect fake accounts (e.g., nodes with no friends but many posts).
- Pathao analyzes driver-passenger graphs to detect fake reviews (e.g., clustered negative ratings from the same user).
7. Challenges and Solutions in Complex Data Mining
| Challenge | Cause | Solution |
|---|---|---|
| High Dimensionality | Too many features (e.g., pixels). | Dimensionality reduction (PCA, t-SNE). |
| Noise and Missing Data | Incomplete or erroneous data. | Imputation, outlier removal. |
| Scalability | Large datasets (e.g., YouTube videos). | Distributed computing (Spark, Hadoop). |
| Interpretability | Black-box models (e.g., deep learning). | SHAP values, LIME for explanations. |
| Real-Time Processing | Streaming data (e.g., NTC network logs). | Online learning algorithms. |
8. Applications in Nepal and Globally
| Application | Example (Nepal) | Example (Global) |
|---|---|---|
| Fraud Detection | Khalti, Nabil Bank | PayPal, Mastercard |
| Recommendation Systems | Daraz, Hamrobazaar | Amazon, Netflix |
| Traffic Management | NTC, Kathmandu Metro (future) | Google Maps, Uber |
| Healthcare | Hospital patient record analysis | IBM Watson for Genomics |
| Social Media Analysis | Facebook Nepal, Twitter trends | Cambridge Analytica (controversial) |
## In the Real World
Daraz’s Recommendation Engine:
- Uses collaborative filtering (a graph mining technique) to suggest products based on similar users’ purchases.
- Example: If User A buys a laptop and User B buys the same laptop + mouse, Daraz recommends the mouse to User A.
Pathao’s Driver Routing:
- Employs spatial clustering to group nearby passengers for a single driver, reducing empty trips.
- Example: In Thapathali, Pathao clusters 5 passengers within 1 km and assigns them to one driver.
Ncell’s Network Optimization:
- Applies time-series forecasting to predict data usage spikes (e.g., during IPL matches) and pre-allocate resources.
- Example: On match days, Ncell increases tower capacity in Lalitpur by 30% based on historical trends.
Khalti’s Fraud Detection:
- Builds a transaction graph where nodes are users/transactions and edges represent money flow.
- Example: If User X sends ₹500 to User Y, who immediately sends ₹490 to User Z (withdrawal fee), Khalti flags it as potential fraud.
NEPSE’s Stock Prediction:
- Financial analysts use ARIMA models to forecast market trends.
- Example: During Dashain, NEPSE often sees a 2–3% rise; ARIMA captures this seasonality.
## Exam Tip
This unit is conceptual but application-heavy. Expect:
Short Definitions (2 marks):
- "Define spatial clustering in data mining."
- "What is an ARIMA model?"
- Answer: "ARIMA is a time-series forecasting technique combining autoregression (AR), differencing (I), and moving averages (MA) to predict future values."
Scenario-Based Questions (5–10 marks):
- "Daraz wants to reduce delivery costs. Explain how spatial clustering can help with a step-by-step example."
- Structure:
- Problem: High delivery costs due to scattered orders.
- Solution: Use DBSCAN to cluster orders by location.
- Steps: Convert addresses to GPS → cluster → optimize routes.
- Outcome: 20–30% cost reduction.
Comparison Tables (3–5 marks):
- "Compare text mining and multimedia mining techniques."
- Use a table like the one above (Text vs. Image/Audio preprocessing).
Diagram-Based Questions (4–6 marks):
- "Draw a graph showing how Khalti detects fraud using transaction networks."
- Must include:
- Nodes (users, transactions).
- Edges (money flow).
- Anomaly indicators (high-degree nodes, short paths).
Case Study Analysis (10 marks):
- "NTC wants to predict network congestion. Suggest a temporal data mining technique and justify your choice."
- Answer:
- Technique: ARIMA or LSTM (for deep patterns).
- Why:
- ARIMA is interpretable and works well with linear trends.
- LSTM captures non-linear patterns (e.g., sudden spikes).
- Example: Predict traffic during Dashain using past 5 years’ data.
## Key Formulas to Memorize
| Technique | Formula/Concept |
|---|---|
| TF-IDF | |
| ARIMA (p,d,q) | |
| PageRank | |
| DBSCAN (ε, minPts) | Points in same neighborhood (distance < ε) with ≥ minPts neighbors form a cluster. |
## Summary Checklist for Revision
Before the exam, ensure you can:
- Explain the difference between structured and unstructured data with examples.
- Describe preprocessing steps for text, images, and time-series data.
- Draw a graph for:
- DBSCAN clustering (spatial data).
- Transaction fraud network (graph mining).
- ARIMA time-series forecast.
- Apply techniques to real scenarios:
- Use LDA for Daraz reviews.
- Optimize Pathao routes with DBSCAN.
- Predict NEPSE trends with ARIMA.
- Discuss challenges (e.g., scalability) and solutions (e.g., Spark).
- Name 3 Nepalese and 3 global applications of complex data mining.
Based on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 10.
Discussion
Loading…