Elective Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 911 min read

Spatial, Multimedia, Text & Web Data Mining

Unit 9 of Data Warehousing and Data Mining explores techniques for extracting insights from unstructured data—geospatial patterns (GPS, maps), multimedia (images, audio), text (documents, social media), and web data (logs, hyperlinks)—using feature extraction, clustering, classification, and specialized algorithms like

Key Concepts and Techniques

1. Spatial Data Mining

Spatial data refers to geographic or location-based data (e.g., coordinates, shapes, or trajectories). Mining this data involves discovering patterns like clusters of similar locations, spatial associations, or outliers.

GPS PointsDBSCAN Cluster 1DBSCAN Cluster 2Outliers
Visual representation of DBSCAN clustering parameters (ε and minPts)

How It Works

  • Data Representation:
    • Points (e.g., GPS coordinates of a taxi’s location).
    • Lines (e.g., roads, flight paths).
    • Polygons (e.g., city boundaries, land parcels).
  • Key Algorithms:
    • DBSCAN (Density-Based Spatial Clustering): Groups nearby points into clusters without predefined k.
    • Spatial Association Rules: Finds co-located patterns (e.g., "If a customer buys coffee near a park, they also buy pastries").
    • Trajectory Mining: Analyzes movement patterns (e.g., "Most Pathao riders take Route X to reach Thapathali").

Worked Example: Traffic Hotspots in Kathmandu

Suppose we have GPS data from 10,000 taxis over a month. We use DBSCAN to find clusters of congestion:

  1. Define ε (max distance between points to be clustered) = 500 meters.
  2. Define minPts (minimum points to form a cluster) = 20.
  3. Run DBSCAN on the dataset. The algorithm outputs 5 clusters:
    • Cluster 1: Kathmandu Durbar Square (high density, 450 points).
    • Cluster 2: Balkhu (moderate density, 180 points).
    • Cluster 3: Outskirts (low density, 30 points) → Outlier (no congestion).
  4. Insight: NTC can prioritize signal upgrades at Durbar Square and Balkhu.
Raw GPS DataDBSCAN (ε=500m, minPts=20)Cluster 1: Kathmandu Durbar Square (450 pts)Cluster 2: Balkhu (180 pts)Outliers (30 pts)NTC Signal Upgrades
DBSCAN clustering workflow for traffic hotspot detection (Kathmandu example)

Applications in Nepal

  • Pathao/Daraz: Predict delivery delays by mining driver trajectories.
  • NTC: Optimize traffic light timings using congestion clusters.
  • Agriculture: Identify crop disease hotspots via satellite imagery.

2. Multimedia Data Mining

Multimedia data includes images, audio, and video. Mining involves extracting features (e.g., edges, colors, frequencies) and applying ML to classify or cluster them.

Feature Extraction Techniques

Data Type Features Extracted Example Algorithm
Images Edges, textures, color histograms SIFT (Scale-Invariant Feature Transform)
Audio MFCC (Mel-Frequency Cepstral Coefficients) K-means clustering
Video Motion vectors, object tracking Optical flow + SVM classification

Worked Example: Detecting Fake News Images on Social Media

  1. Input: 500 images claimed to be from a protest in Lalitpur.
  2. Feature Extraction:
    • Use SIFT to detect keypoints (corners, blobs).
    • Compare keypoints against a database of known protest images.
  3. Classification:
    • Train a Random Forest model to flag images with <70% keypoint matches as "likely fake."
  4. Output: 80 images are flagged; manual review confirms 75 are AI-generated.

Real-World Use Cases

  • Google Photos: Auto-tags images using CNN (Convolutional Neural Networks).
  • WhatsApp: Detects spam images via hash matching.
  • Nepal Police: Uses facial recognition (from CCTV) to identify missing persons.

3. Text Data Mining

Text mining involves processing unstructured text (e.g., tweets, news articles) to extract trends, sentiments, or topics.

011.2522.533.7545Topic 1: Banks45Topic 2: Inflation30Topic 3: Tech25Percentage of Discussion Volume
Topic distribution in NEPSE stock discussions (sample data)

Key Steps

  1. Preprocessing:
    • Tokenization (split text into words).
    • Stopword removal (e.g., "the," "is").
    • Stemming/Lemmatization (reduce words to root form: "running" → "run").
  2. Feature Representation:
    • Bag-of-Words (BoW): Count word frequencies.
    • TF-IDF: Weigh words by importance (e.g., "earthquake" is more important in Nepalese news post-2015).
  3. Mining Techniques:
    • Topic Modeling (LDA): Discovers themes in documents (e.g., "NEPSE stock trends," "monsoon delays").
    • Sentiment Analysis: Classifies text as positive/negative (e.g., "Khalti’s new feature" → 70% positive tweets).

Worked Example: Analyzing NEPSE Stock Discussions

  1. Data: 5,000 tweets about NEPSE stocks in 2023.
  2. Preprocessing:
    • Remove hashtags (#NEPSE), mentions (@khalti), and stopwords.
    • Lemmatize: "investing" → "invest."
  3. TF-IDF:
    • "Banks" appears in 80% of tweets → TF-IDF score = 0.9.
    • "Bitcoin" appears in 5% → TF-IDF score = 0.2.
  4. Topic Modeling (LDA):
    • Identifies 3 topics:
      • Topic 1: Bank stocks (words: "NMB," "Global IME," "dividend").
      • Topic 2: Inflation fears (words: "rupee," "import," "devaluation").
      • Topic 3: Tech stocks (words: "Ncell," "NTC," "5G").
  5. Insight: Investors are most concerned about banks; NTC’s 5G plans are a minor discussion point.
5,000 TweetsPreprocessing (Tokenize, Lemmatize)TF-IDF WeightingLDA (3 Topics)Topic 1: Banks (NMB, Global IME, dividend)Topic 2: Inflation (rupee, import, devaluation)Topic 3: Tech (Ncell, NTC, 5G)
Topic modeling pipeline for NEPSE stock discussion analysis

Applications in Nepal

  • eSewa: Analyzes customer complaints to improve service (e.g., "delayed payments" → prioritize backend fixes).
  • Kathmandu Post: Uses NLP to auto-categorize news into "Politics," "Economy," "Sports."
  • Pathao Drivers: Sentiment analysis of ride reviews to identify unsafe routes.

4. Web Data Mining

Web data includes HTML content, hyperlinks, and usage logs. Techniques focus on:

  • Web Structure Mining: Analyzing link patterns (e.g., PageRank).
  • Web Content Mining: Extracting text/images from web pages.
  • Web Usage Mining: Studying user behavior (e.g., "Most Daraz users add items to cart but don’t checkout").

Key Algorithms

Technique Purpose Example
PageRank Rank web pages by importance Google’s search algorithm
Association Rules Find co-occurring items on a page "Users who view X also view Y"
Session Analysis Group user visits into patterns "70% of Daraz users abandon cart at checkout"

Worked Example: Daraz’s "Frequently Bought Together"

  1. Data: 1 million user sessions on Daraz.
  2. Preprocessing:
    • Extract items added to cart in the same session.
    • Filter sessions with ≥3 items.
  3. Apriori Algorithm:
    • Find frequent itemsets with support ≥ 5% (e.g., {Phone, Charger} appears in 6% of sessions).
    • Generate rules: {Phone} → {Charger} with confidence = 80%.
  4. Output: Daraz displays "Frequently Bought Together" for phones and chargers, increasing sales by 12%.
1M User SessionsFilter (≥3 items)Apriori (support=5%, confidence=80%)Frequent Itemset: {Phone, Charger} (6%)Association Rule: Phone → Charger (confidence=80%)Daraz Product Page
Apriori algorithm workflow for Daraz's 'Frequently Bought Together' feature

Real-World Examples

  • Google: Uses web usage mining to personalize search results.
  • Facebook: Mines user likes/shares to suggest ads (e.g., "Because you liked X, we’re showing Y").
  • NTC Website: Analyzes user clicks to improve navigation (e.g., "Most users struggle to find bill payment").

In the Real World

  1. Pathao’s Route Optimization

    • Idea Used: Spatial trajectory mining + DBSCAN clustering.
    • How: Pathao analyzes driver GPS data to identify congested routes (e.g., Thapathali to Kantipath) and suggests alternative paths during peak hours. This reduces delivery times by 15%.
  2. Khalti’s Fraud Detection

    • Idea Used: Text mining (NLP) + anomaly detection.
    • How: Khalti’s chatbot flags suspicious transactions by analyzing message patterns (e.g., "urgent payment" + "new bank account" → red flag). In 2023, this blocked 2,000+ fraudulent transfers.
  3. Nepal Police’s Missing Persons Database

    • Idea Used: Image mining (facial recognition) + spatial queries.
    • How: The police use OpenCV to match CCTV footage against a database of missing persons. In 2022, this helped recover 150+ missing children by cross-referencing facial features with reported cases.

Exam Tip

This unit is conceptual but applied. Expect:

  1. Short Definitions (2 marks):

    • Define DBSCAN, TF-IDF, or PageRank in 1–2 sentences.
    • Example: "DBSCAN is a density-based clustering algorithm that groups spatial points within a radius ε and requires at least minPts neighbors."
  2. Algorithm Steps (4–6 marks):

    • Describe Apriori or LDA in 4–5 bullet points.
    • Example for Apriori:
      • Input: Transaction database.
      • Step 1: Find frequent 1-itemsets (support ≥ min_sup).
      • Step 2: Generate candidate 2-itemsets.
      • Step 3: Prune candidates with support < min_sup.
      • Output: Association rules (e.g., {Bread} → {Butter}).
  3. Worked Examples (6–8 marks):

    • Given a small dataset (e.g., 5 GPS points or 3 tweets), perform:
      • Clustering (DBSCAN).
      • TF-IDF calculation.
      • Apriori rule generation.
    • Show all steps (e.g., support/confidence calculations).
  4. Comparisons (3–4 marks):

    • Compare BoW vs. TF-IDF or PageRank vs. HITS.
    • Example table:
      Criteria BoW TF-IDF
      Weighting Uniform (all words = 1) Weighs by rarity (IDF)
      Use Case Simple keyword matching Advanced NLP (e.g., spam detection)
  5. Applications (2–3 marks):

    • Link techniques to real-world scenarios (e.g., "How would you use DBSCAN for traffic management in Kathmandu?").

Avoid: Memorizing code. Focus on concepts, steps, and interpretations. For example, in DBSCAN, always explain ε and minPts in the context of the problem (e.g., "ε = 500m captures nearby traffic hotspots").

Based on the TU BIT syllabus for Data Warehousing and Data Mining, unit 9.

Discussion

Loading…