Data Warehousing and Data MiningUnit 911 min read
Spatial, Multimedia, Text & Web Data Mining
Unit 9 of Data Warehousing and Data Mining explores techniques for extracting insights from unstructured data—geospatial patterns (GPS, maps), multimedia (images, audio), text (documents, social media), and web data (logs, hyperlinks)—using feature extraction, clustering, classification, and specialized algorithms like
Key Concepts and Techniques
1. Spatial Data Mining
Spatial data refers to geographic or location-based data (e.g., coordinates, shapes, or trajectories). Mining this data involves discovering patterns like clusters of similar locations, spatial associations, or outliers.
How It Works
- Data Representation:
- Points (e.g., GPS coordinates of a taxi’s location).
- Lines (e.g., roads, flight paths).
- Polygons (e.g., city boundaries, land parcels).
- Key Algorithms:
- DBSCAN (Density-Based Spatial Clustering): Groups nearby points into clusters without predefined k.
- Spatial Association Rules: Finds co-located patterns (e.g., "If a customer buys coffee near a park, they also buy pastries").
- Trajectory Mining: Analyzes movement patterns (e.g., "Most Pathao riders take Route X to reach Thapathali").
Worked Example: Traffic Hotspots in Kathmandu
Suppose we have GPS data from 10,000 taxis over a month. We use DBSCAN to find clusters of congestion:
- Define ε (max distance between points to be clustered) = 500 meters.
- Define minPts (minimum points to form a cluster) = 20.
- Run DBSCAN on the dataset. The algorithm outputs 5 clusters:
- Cluster 1: Kathmandu Durbar Square (high density, 450 points).
- Cluster 2: Balkhu (moderate density, 180 points).
- Cluster 3: Outskirts (low density, 30 points) → Outlier (no congestion).
- Insight: NTC can prioritize signal upgrades at Durbar Square and Balkhu.
Applications in Nepal
- Pathao/Daraz: Predict delivery delays by mining driver trajectories.
- NTC: Optimize traffic light timings using congestion clusters.
- Agriculture: Identify crop disease hotspots via satellite imagery.
2. Multimedia Data Mining
Multimedia data includes images, audio, and video. Mining involves extracting features (e.g., edges, colors, frequencies) and applying ML to classify or cluster them.
Feature Extraction Techniques
| Data Type | Features Extracted | Example Algorithm |
|---|---|---|
| Images | Edges, textures, color histograms | SIFT (Scale-Invariant Feature Transform) |
| Audio | MFCC (Mel-Frequency Cepstral Coefficients) | K-means clustering |
| Video | Motion vectors, object tracking | Optical flow + SVM classification |
Worked Example: Detecting Fake News Images on Social Media
- Input: 500 images claimed to be from a protest in Lalitpur.
- Feature Extraction:
- Use SIFT to detect keypoints (corners, blobs).
- Compare keypoints against a database of known protest images.
- Classification:
- Train a Random Forest model to flag images with <70% keypoint matches as "likely fake."
- Output: 80 images are flagged; manual review confirms 75 are AI-generated.
Real-World Use Cases
- Google Photos: Auto-tags images using CNN (Convolutional Neural Networks).
- WhatsApp: Detects spam images via hash matching.
- Nepal Police: Uses facial recognition (from CCTV) to identify missing persons.
3. Text Data Mining
Text mining involves processing unstructured text (e.g., tweets, news articles) to extract trends, sentiments, or topics.
Key Steps
- Preprocessing:
- Tokenization (split text into words).
- Stopword removal (e.g., "the," "is").
- Stemming/Lemmatization (reduce words to root form: "running" → "run").
- Feature Representation:
- Bag-of-Words (BoW): Count word frequencies.
- TF-IDF: Weigh words by importance (e.g., "earthquake" is more important in Nepalese news post-2015).
- Mining Techniques:
- Topic Modeling (LDA): Discovers themes in documents (e.g., "NEPSE stock trends," "monsoon delays").
- Sentiment Analysis: Classifies text as positive/negative (e.g., "Khalti’s new feature" → 70% positive tweets).
Worked Example: Analyzing NEPSE Stock Discussions
- Data: 5,000 tweets about NEPSE stocks in 2023.
- Preprocessing:
- Remove hashtags (#NEPSE), mentions (@khalti), and stopwords.
- Lemmatize: "investing" → "invest."
- TF-IDF:
- "Banks" appears in 80% of tweets → TF-IDF score = 0.9.
- "Bitcoin" appears in 5% → TF-IDF score = 0.2.
- Topic Modeling (LDA):
- Identifies 3 topics:
- Topic 1: Bank stocks (words: "NMB," "Global IME," "dividend").
- Topic 2: Inflation fears (words: "rupee," "import," "devaluation").
- Topic 3: Tech stocks (words: "Ncell," "NTC," "5G").
- Identifies 3 topics:
- Insight: Investors are most concerned about banks; NTC’s 5G plans are a minor discussion point.
Applications in Nepal
- eSewa: Analyzes customer complaints to improve service (e.g., "delayed payments" → prioritize backend fixes).
- Kathmandu Post: Uses NLP to auto-categorize news into "Politics," "Economy," "Sports."
- Pathao Drivers: Sentiment analysis of ride reviews to identify unsafe routes.
4. Web Data Mining
Web data includes HTML content, hyperlinks, and usage logs. Techniques focus on:
- Web Structure Mining: Analyzing link patterns (e.g., PageRank).
- Web Content Mining: Extracting text/images from web pages.
- Web Usage Mining: Studying user behavior (e.g., "Most Daraz users add items to cart but don’t checkout").
Key Algorithms
| Technique | Purpose | Example |
|---|---|---|
| PageRank | Rank web pages by importance | Google’s search algorithm |
| Association Rules | Find co-occurring items on a page | "Users who view X also view Y" |
| Session Analysis | Group user visits into patterns | "70% of Daraz users abandon cart at checkout" |
Worked Example: Daraz’s "Frequently Bought Together"
- Data: 1 million user sessions on Daraz.
- Preprocessing:
- Extract items added to cart in the same session.
- Filter sessions with ≥3 items.
- Apriori Algorithm:
- Find frequent itemsets with support ≥ 5% (e.g., {Phone, Charger} appears in 6% of sessions).
- Generate rules: {Phone} → {Charger} with confidence = 80%.
- Output: Daraz displays "Frequently Bought Together" for phones and chargers, increasing sales by 12%.
Real-World Examples
- Google: Uses web usage mining to personalize search results.
- Facebook: Mines user likes/shares to suggest ads (e.g., "Because you liked X, we’re showing Y").
- NTC Website: Analyzes user clicks to improve navigation (e.g., "Most users struggle to find bill payment").
In the Real World
Pathao’s Route Optimization
- Idea Used: Spatial trajectory mining + DBSCAN clustering.
- How: Pathao analyzes driver GPS data to identify congested routes (e.g., Thapathali to Kantipath) and suggests alternative paths during peak hours. This reduces delivery times by 15%.
Khalti’s Fraud Detection
- Idea Used: Text mining (NLP) + anomaly detection.
- How: Khalti’s chatbot flags suspicious transactions by analyzing message patterns (e.g., "urgent payment" + "new bank account" → red flag). In 2023, this blocked 2,000+ fraudulent transfers.
Nepal Police’s Missing Persons Database
- Idea Used: Image mining (facial recognition) + spatial queries.
- How: The police use OpenCV to match CCTV footage against a database of missing persons. In 2022, this helped recover 150+ missing children by cross-referencing facial features with reported cases.
Exam Tip
This unit is conceptual but applied. Expect:
Short Definitions (2 marks):
- Define DBSCAN, TF-IDF, or PageRank in 1–2 sentences.
- Example: "DBSCAN is a density-based clustering algorithm that groups spatial points within a radius ε and requires at least minPts neighbors."
Algorithm Steps (4–6 marks):
- Describe Apriori or LDA in 4–5 bullet points.
- Example for Apriori:
- Input: Transaction database.
- Step 1: Find frequent 1-itemsets (support ≥ min_sup).
- Step 2: Generate candidate 2-itemsets.
- Step 3: Prune candidates with support < min_sup.
- Output: Association rules (e.g., {Bread} → {Butter}).
Worked Examples (6–8 marks):
- Given a small dataset (e.g., 5 GPS points or 3 tweets), perform:
- Clustering (DBSCAN).
- TF-IDF calculation.
- Apriori rule generation.
- Show all steps (e.g., support/confidence calculations).
- Given a small dataset (e.g., 5 GPS points or 3 tweets), perform:
Comparisons (3–4 marks):
- Compare BoW vs. TF-IDF or PageRank vs. HITS.
- Example table:
Criteria BoW TF-IDF Weighting Uniform (all words = 1) Weighs by rarity (IDF) Use Case Simple keyword matching Advanced NLP (e.g., spam detection)
Applications (2–3 marks):
- Link techniques to real-world scenarios (e.g., "How would you use DBSCAN for traffic management in Kathmandu?").
Avoid: Memorizing code. Focus on concepts, steps, and interpretations. For example, in DBSCAN, always explain ε and minPts in the context of the problem (e.g., "ε = 500m captures nearby traffic hotspots").
Based on the TU BIT syllabus for Data Warehousing and Data Mining, unit 9.
Discussion
Loading…