IT274 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 1018 min read

Complex Data Mining & Real-World Applications

Unit 10 of Data Warehousing and Data Mining explores mining unstructured data (text, images, video), spatial/temporal patterns, and real-world applications like fraud detection, recommendation systems, and social network analysis—with case studies from Nepalese and global tech.

TAKEAWAYS:

  • Complex data types (text, multimedia, spatial/temporal) require specialized mining techniques beyond structured data.
  • Text mining extracts insights from unstructured data (e.g., customer reviews, news) using NLP and topic modeling.
  • Spatial data mining uncovers patterns in geographic data (e.g., traffic congestion, retail locations) via clustering and association rules.
  • Temporal data mining analyzes time-series trends (e.g., stock prices, weather forecasts) using sequence prediction and anomaly detection.
  • Real-world applications include fraud detection (banks), recommendation systems (YouTube, Daraz), and social network analysis (Pathao driver routing).
  • Challenges include scalability, noise, and interpretability—addressed via hybrid models and visualization tools.

1. Introduction to Complex Data Mining

Complex data refers to unstructured or semi-structured data that lacks a fixed schema, including:

  • Text data (emails, social media, documents).
  • Multimedia data (images, audio, video).
  • Spatial data (GIS, maps, satellite imagery).
  • Temporal data (time-series, logs, sensor data).
  • Graph/social network data (user interactions, fraud networks).

Unlike structured data (tables in databases), complex data requires domain-specific preprocessing (e.g., NLP for text, image segmentation for visuals) before mining.

Why Mine Complex Data?

  • Hidden patterns: Customer sentiment in reviews, fraudulent transactions in bank logs.
  • Automation: Chatbots (Khalti customer support), autonomous vehicles (Pathao’s route optimization).
  • Decision support: Disease outbreak prediction (NTC’s network traffic), stock market trends (NEPSE).

2. Text Mining: Extracting Insights from Unstructured Data

Text mining combines Natural Language Processing (NLP) and data mining to analyze text for patterns.

Key Techniques

Technique Purpose Example Tools/Libraries
Tokenization Split text into words/phrases. NLTK, spaCy
Stopword Removal Filter common words (e.g., "the", "is"). Python’s nltk.corpus
Stemming/Lemmatization Reduce words to root form (e.g., "running" → "run"). Porter Stemmer, WordNet
TF-IDF Weight words by importance. Scikit-learn
Topic Modeling Group related words into themes. Latent Dirichlet Allocation (LDA)
Sentiment Analysis Classify text as positive/negative. VADER, TextBlob
531Driver ADriver BPassenger XPassenger Y
Pathao driver-passenger interaction graph for fake review detection
011.2522.533.7545Shipping Delays45Product Quality30Customer Service25Percentage of Reviews
Top 3 topics extracted from Daraz reviews using LDA

Worked Example: Analyzing Customer Reviews for Daraz

Scenario: Daraz wants to identify common complaints from product reviews to improve inventory. Steps:

  1. Preprocess:
    • Tokenize: "Product arrived late, poor packaging" → ["Product", "arrived", "late", "poor", "packaging"].
    • Remove stopwords: ["arrived", "late", "poor", "packaging"].
    • Lemmatize: ["arrive", "late", "poor", "package"].
  2. TF-IDF:
    • Calculate weights: "late" appears frequently but isn’t unique; "package" is rare → higher weight.
  3. Topic Modeling (LDA):
    • Group words into topics:
      • Topic 1: {"delay", "ship", "late"}
      • Topic 2: {"damage", "package", "poor"}
  4. Insight: Daraz can prioritize faster shipping and better packaging.

Real-World Tie-In:

  • Khalti uses sentiment analysis to detect fraudulent transaction descriptions (e.g., "urgent payment" flags for review).
  • Ncell analyzes customer service call transcripts to predict network outage causes.

flowchart LR
    A["Raw Text Data\n(e.g., Daraz reviews)"] --> B["Preprocessing\n(Tokenization, Cleaning)"]
    B --> C["Feature Extraction\n(TF-IDF, Word Embeddings)"]
    C --> D["Modeling\n(LDA, SVM for Sentiment)"]
    D --> E["Insights\n(Topics: Shipping Delays, Product Quality)"]
    E --> F["Action\n(Daraz improves logistics)"]

3. Multimedia Data Mining: Images, Audio, and Video

Multimedia data is high-dimensional (e.g., a 1000×1000 pixel image = 1M features). Techniques reduce this complexity.

Key Techniques

Data Type Preprocessing Mining Technique Example Application
Images Edge detection, segmentation Clustering (K-means), CNN Medical imaging (tumor detection)
Audio Spectrogram conversion Anomaly detection, classification Voice assistants (Siri, Google)
Video Frame extraction, object tracking Sequence mining, activity recognition Surveillance (NTC traffic monitoring)

Worked Example: Detecting Traffic Violations in Kathmandu

Scenario: NTC uses CCTV footage to flag jaywalking or speeding. Steps:

  1. Preprocess:
    • Extract frames from video → convert to grayscale.
    • Apply edge detection (Canny filter) to highlight moving objects.
  2. Object Detection:
    • Use a pre-trained CNN (e.g., YOLO) to classify objects as "pedestrian" or "vehicle."
  3. Rule-Based Mining:
    • If a pedestrian crosses outside zebra lines → violation alert.
    • If a vehicle exceeds speed limit (detected via license plate + GPS) → fine.
  4. Output: Automated ticketing system integrated with Ncell’s digital payment.

Real-World Tie-In:

  • Pathao uses image mining to detect driver behavior (e.g., sudden braking) via phone cameras.
  • Google Photos clusters images by faces/places using deep learning.

graph TD
    A["Video Input\n(NTC CCTV)"] --> B["Frame Extraction"]
    B --> C["Edge Detection\n(Canny Filter)"]
    C --> D["Object Detection\n(CNN: YOLO)"]
    D --> E["Rule Engine\n(Jaywalking? Speeding?)"]
    E --> F["Alert System\n(Ncell Payment Link)"]

4. Spatial Data Mining: Geographic Patterns

Spatial data mining analyzes geographic or location-based data to find patterns like:

  • Hotspots (e.g., crime, retail sales).
  • Clustering (e.g., similar neighborhoods).
  • Route optimization (e.g., Pathao driver paths).

Key Techniques

Technique Purpose Example
Spatial Clustering Group nearby points (e.g., ATMs). DBSCAN, K-means
Spatial Association Find co-located items (e.g., cafes + bookstores). Apriori algorithm adapted for GIS
Geographic Outliers Detect unusual locations (e.g., fraudulent transactions). LOF (Local Outlier Factor)

Worked Example: Optimizing Daraz Delivery Routes

Scenario: Daraz wants to reduce delivery costs by clustering orders geographically. Steps:

  1. Data Collection:
    • Orders: [(Kathmandu-12, "Laptop"), (Lalitpur-3, "Phone"), (Bhaktapur-5, "Books")].
    • Locations mapped to coordinates: (KTM: (27.7172, 85.3240), LLP: (27.6833, 85.3197), BKT: (27.6883, 85.4103)).
  2. Clustering (DBSCAN):
    • Set ε = 5 km (max distance between points to be clustered).
    • Result: Cluster 1 = {Kathmandu, Lalitpur} (close), Cluster 2 = {Bhaktapur} (isolated).
  3. Route Optimization:
    • Assign a single delivery agent to Cluster 1 (KTM → LLP → KTM).
    • Assign another to Bhaktapur.
  4. Savings: Reduces fuel costs by 30% vs. individual deliveries.

Real-World Tie-In:

  • Pathao uses spatial clustering to group nearby passengers for drivers.
  • NTC analyzes mobile tower data to predict network congestion zones.

331010KathmanduLalitpurBhaktapur
DBSCAN clustering (ε=5km) showing optimized route between Kathmandu and Lalitpur

5. Temporal Data Mining: Time-Series Analysis

Temporal data (e.g., stock prices, weather) is analyzed for trends, seasonality, and anomalies.

Key Techniques

Technique Purpose Example
Time-Series Forecasting Predict future values (e.g., NEPSE stock). ARIMA, LSTM (deep learning)
Anomaly Detection Flag unusual patterns (e.g., fraud). Isolation Forest, Autoencoders
Sequence Mining Find frequent patterns (e.g., user behavior). PrefixSpan algorithm

Scenario: An investor wants to predict NEPSE’s next 5 days using past 6 months of data. Steps:

  1. Data Preparation:
    • Time-series: [Day1: 1200, Day2: 1210, ..., Day180: 1350].
    • Features: Price, Volume, Moving Average (7-day).
  2. Model Selection:
    • Use ARIMA (AutoRegressive Integrated Moving Average):
      • AR (AutoRegressive): Uses past values to predict next.
      • I (Integrated): Makes data stationary (removes trends).
      • MA (Moving Average): Accounts for error smoothing.
  3. Training:
    • Fit ARIMA(2,1,2) to historical data.
    • Parameters:
      • p=2 (lag observations), d=1 (differencing), q=2 (error terms).
  4. Prediction:
    • Input: Last 7 days’ prices [1340, 1345, 1350, 1348, 1352, 1355, 1360].
    • Output: Predicted prices for next 5 days:
      Day181: 1362
      Day182: 1365
      Day183: 1368
      Day184: 1370
      Day185: 1372
      
  5. Validation:
    • Compare with actual data (if available) to check RMSE (Root Mean Squared Error).

Real-World Tie-In:

  • Ncell uses time-series forecasting to predict network traffic spikes (e.g., during festivals).
  • Google Trends detects viral topics by analyzing search query patterns over time.

2468101213001320134013601380yHistorical NEPSE Data (6 months)ARIMA(2,1,2) Prediction
NEPSE stock trend prediction (80% training, 20% forecast)

6. Graph and Social Network Mining

Graphs represent entities (nodes) and relationships (edges). Used for:

  • Fraud detection (e.g., money laundering networks).
  • Recommendation systems (e.g., YouTube "Watch Next").
  • Community detection (e.g., Pathao driver groups).

Key Techniques

Technique Purpose Example
Community Detection Group nodes with dense connections. Louvain, Girvan-Newman
Centrality Measures Find influential nodes (e.g., key users). PageRank, Degree Centrality
Link Prediction Predict missing edges (e.g., new friendships). Common Neighbors algorithm

Worked Example: Detecting Fraudulent Transactions in Khalti

Scenario: Khalti flags suspicious transaction patterns using graph mining. Steps:

  1. Graph Construction:
    • Nodes: Users (U1, U2, ...), Transactions (T1, T2).
    • Edges: U1 → T1 → U2 (money flow).
  2. Anomaly Detection:
    • High-degree nodes: Users with many transactions in short time → suspicious.
    • Short paths: If U1 → U2 → U3 with small amounts but frequent → money laundering.
  3. PageRank:
    • Assign scores to users based on "importance" in the graph.
    • High-score users with sudden large transactions → red flag.
  4. Action:
    • Freeze accounts of U3 (score = 0.95) with transaction pattern:
      U3 → T100 (₹500) → U4
      U3 → T101 (₹600) → U5
      ...
      

Real-World Tie-In:

  • Facebook uses graph mining to detect fake accounts (e.g., nodes with no friends but many posts).
  • Pathao analyzes driver-passenger graphs to detect fake reviews (e.g., clustered negative ratings from the same user).

100100500500U1U2U3T1T2
Khalti transaction graph showing fraudulent flow (U3: Score = 0.95)

7. Challenges and Solutions in Complex Data Mining

Challenge Cause Solution
High Dimensionality Too many features (e.g., pixels). Dimensionality reduction (PCA, t-SNE).
Noise and Missing Data Incomplete or erroneous data. Imputation, outlier removal.
Scalability Large datasets (e.g., YouTube videos). Distributed computing (Spark, Hadoop).
Interpretability Black-box models (e.g., deep learning). SHAP values, LIME for explanations.
Real-Time Processing Streaming data (e.g., NTC network logs). Online learning algorithms.

8. Applications in Nepal and Globally

Application Example (Nepal) Example (Global)
Fraud Detection Khalti, Nabil Bank PayPal, Mastercard
Recommendation Systems Daraz, Hamrobazaar Amazon, Netflix
Traffic Management NTC, Kathmandu Metro (future) Google Maps, Uber
Healthcare Hospital patient record analysis IBM Watson for Genomics
Social Media Analysis Facebook Nepal, Twitter trends Cambridge Analytica (controversial)

## In the Real World

  1. Daraz’s Recommendation Engine:

    • Uses collaborative filtering (a graph mining technique) to suggest products based on similar users’ purchases.
    • Example: If User A buys a laptop and User B buys the same laptop + mouse, Daraz recommends the mouse to User A.
  2. Pathao’s Driver Routing:

    • Employs spatial clustering to group nearby passengers for a single driver, reducing empty trips.
    • Example: In Thapathali, Pathao clusters 5 passengers within 1 km and assigns them to one driver.
  3. Ncell’s Network Optimization:

    • Applies time-series forecasting to predict data usage spikes (e.g., during IPL matches) and pre-allocate resources.
    • Example: On match days, Ncell increases tower capacity in Lalitpur by 30% based on historical trends.
  4. Khalti’s Fraud Detection:

    • Builds a transaction graph where nodes are users/transactions and edges represent money flow.
    • Example: If User X sends ₹500 to User Y, who immediately sends ₹490 to User Z (withdrawal fee), Khalti flags it as potential fraud.
  5. NEPSE’s Stock Prediction:

    • Financial analysts use ARIMA models to forecast market trends.
    • Example: During Dashain, NEPSE often sees a 2–3% rise; ARIMA captures this seasonality.

## Exam Tip

This unit is conceptual but application-heavy. Expect:

  1. Short Definitions (2 marks):

    • "Define spatial clustering in data mining."
    • "What is an ARIMA model?"
    • Answer: "ARIMA is a time-series forecasting technique combining autoregression (AR), differencing (I), and moving averages (MA) to predict future values."
  2. Scenario-Based Questions (5–10 marks):

    • "Daraz wants to reduce delivery costs. Explain how spatial clustering can help with a step-by-step example."
    • Structure:
      1. Problem: High delivery costs due to scattered orders.
      2. Solution: Use DBSCAN to cluster orders by location.
      3. Steps: Convert addresses to GPS → cluster → optimize routes.
      4. Outcome: 20–30% cost reduction.
  3. Comparison Tables (3–5 marks):

    • "Compare text mining and multimedia mining techniques."
    • Use a table like the one above (Text vs. Image/Audio preprocessing).
  4. Diagram-Based Questions (4–6 marks):

    • "Draw a graph showing how Khalti detects fraud using transaction networks."
    • Must include:
      • Nodes (users, transactions).
      • Edges (money flow).
      • Anomaly indicators (high-degree nodes, short paths).
  5. Case Study Analysis (10 marks):

    • "NTC wants to predict network congestion. Suggest a temporal data mining technique and justify your choice."
    • Answer:
      • Technique: ARIMA or LSTM (for deep patterns).
      • Why:
        • ARIMA is interpretable and works well with linear trends.
        • LSTM captures non-linear patterns (e.g., sudden spikes).
      • Example: Predict traffic during Dashain using past 5 years’ data.

## Key Formulas to Memorize

Technique Formula/Concept
TF-IDF
ARIMA (p,d,q)
PageRank
DBSCAN (ε, minPts) Points in same neighborhood (distance < ε) with ≥ minPts neighbors form a cluster.

## Summary Checklist for Revision

Before the exam, ensure you can:

  1. Explain the difference between structured and unstructured data with examples.
  2. Describe preprocessing steps for text, images, and time-series data.
  3. Draw a graph for:
    • DBSCAN clustering (spatial data).
    • Transaction fraud network (graph mining).
    • ARIMA time-series forecast.
  4. Apply techniques to real scenarios:
    • Use LDA for Daraz reviews.
    • Optimize Pathao routes with DBSCAN.
    • Predict NEPSE trends with ARIMA.
  5. Discuss challenges (e.g., scalability) and solutions (e.g., Spark).
  6. Name 3 Nepalese and 3 global applications of complex data mining.

Based on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 10.

Discussion

Loading…