IT274 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 1012 min read

Complex Data Mining & Real-World Applications

Unit 10 of Data Warehousing and Data Mining explores advanced techniques for mining unstructured data (text, multimedia, streams, graphs), real-time analytics, and practical applications in business, healthcare, and social networks—with Nepalese and global case studies.

TAKEAWAYS:

  • Unstructured data (text, images, audio) requires specialized mining techniques like NLP, image processing, and graph algorithms.
  • Stream mining processes real-time data (e.g., stock prices, social media) using sliding windows and incremental learning.
  • Graph mining uncovers hidden relationships in networks (e.g., fraud detection in banks, social connections in Facebook).
  • Multimedia mining extracts insights from images (e.g., object recognition in Daraz), audio (e.g., speech-to-text in eSewa), and video (e.g., YouTube recommendations).
  • Applications include recommendation systems (Amazon, Netflix), fraud detection (Ncell, banks), and predictive maintenance (NTC power grids).
  • Challenges include scalability, privacy (GDPR), and interpretability of complex models.

1. Mining Unstructured Data

Unstructured data (80% of global data) lacks a predefined schema. Mining it requires domain-specific techniques:

A. Text Mining

Definition: Extracts meaningful patterns from text (emails, reviews, news) using NLP, sentiment analysis, and topic modeling. Key Techniques:

  • Tokenization: Splitting text into words/phrases (e.g., "Nepal earthquake 2015" → ["Nepal", "earthquake", "2015"]).
  • Sentiment Analysis: Classifies text as positive/negative/neutral (e.g., Daraz product reviews).
  • Topic Modeling: Identifies themes (e.g., "NEPSE stock trends" vs. "Kathmandu traffic").

Worked Example: Sentiment Analysis on eSewa Complaints

  1. Input: 1000 eSewa customer complaints (e.g., "My transaction failed twice!").
  2. Preprocessing:
    • Remove stopwords ("my", "the").
    • Stemming: "failed" → "fail".
  3. Model: Train a Naive Bayes classifier on labeled data (positive/negative).
  4. Output:
    Complaint Sentiment Probability
    "Money deducted but not credited" Negative 0.92
    "Service was fast and helpful" Positive 0.88

Real-World Tie-In:

  • WhatsApp Business: Uses text mining to categorize customer queries (e.g., "order status" vs. "complaint") and route them to the right agent.
  • Nepali News Aggregators: Topic modeling groups articles by "politics", "sports", or "technology" for personalized feeds.

B. Image and Video Mining

Definition: Extracts features from pixels (e.g., object detection, facial recognition) using CNNs (Convolutional Neural Networks). Key Techniques:

  • Edge Detection: Highlights boundaries (e.g., identifying potholes in Kathmandu roads).
  • Object Recognition: Classifies items (e.g., Daraz’s "shoes" vs. "electronics").
  • Facial Emotion Recognition: Used in security systems (e.g., airport surveillance).

Visual: CNN Layers for Object Detection

graph LR
    A["Input Image\n(3x3 pixels)"] --> B["Convolution\n(Feature maps)"]
    B --> C["Pooling\n(Dimensionality reduction)"]
    C --> D["Fully Connected\n(Classification: Cat/Dog)"]
    D --> E["Output:\n'Dog' (92% confidence)"]

Worked Example: Daraz’s Product Tagging

  1. Input: Image of a product (e.g., a phone).
  2. CNN Steps:
    • Convolution Layer 1: Detects edges (phone outline).
    • Pooling Layer: Reduces image size while keeping key features.
    • Fully Connected Layer: Outputs probabilities for categories (e.g., "smartphone": 0.95, "charger": 0.03).
  3. Output: Tags the product as "Smartphone" + suggests related items (cases, chargers).

Real-World Tie-In:

  • Pathao’s Driver Verification: Uses facial recognition to match driver IDs with real-time camera feeds.
  • YouTube: Uses video mining to detect inappropriate content (e.g., copyrighted music) via audio fingerprinting.

C. Audio Mining

Definition: Analyzes sound waves for patterns (e.g., speech recognition, music genre classification). Key Techniques:

  • MFCC (Mel-Frequency Cepstral Coefficients): Converts audio into numerical features.
  • Speech-to-Text: Converts spoken Nepali (e.g., in eSewa IVR systems) into text.
  • Music Genre Classification: Identifies "rock" vs. "folk" from audio clips.

Worked Example: eSewa IVR System

  1. Input: Customer says, "Merobat 5000 rupees ko transaction garna chai."
  2. Processing:
    • MFCC extracts features from the audio.
    • A trained RNN (Recurrent Neural Network) converts speech to text: "Transfer 5000 rupees".
  3. Output: System processes the transaction or asks for confirmation.

Real-World Tie-In:

  • Google Assistant: Uses audio mining to transcribe commands (e.g., "Call my mother") in real time.
  • NTC Call Centers: Analyzes customer calls to detect frustration (e.g., long wait times) and reroute agents.

2. Stream Mining (Real-Time Data)

Definition: Processes continuous, high-velocity data streams (e.g., stock prices, social media, sensor data) with limited storage. Key Techniques:

  • Sliding Window: Analyzes fixed-time chunks (e.g., last 5 minutes of Twitter data).
  • Incremental Learning: Updates models without reprocessing all data.
  • Concept Drift Detection: Identifies changes in data patterns (e.g., sudden spike in NEPSE shares).

Visual: Sliding Window for Stock Price Prediction

graph LR
    A["Time Series Data\n(Stock Prices: 2023-01-01 to 2023-12-31)"] --> B["Sliding Window\n(Last 30 days)"]
    B --> C["Feature Extraction\n(Mean, Volatility)"]
    C --> D["Predictive Model\n(LSTM Neural Network)"]
    D --> E["Output:\n'Buy' (78% confidence)"]

Worked Example: NEPSE Real-Time Analytics

  1. Input: Live NEPSE stock data (e.g., 1000 shares/sec).
  2. Sliding Window: Analyzes the last 1-hour window.
  3. Features Extracted:
    • Moving average (7-day).
    • Volatility (standard deviation).
  4. Model: LSTM predicts next 5-minute trend ("Bullish" or "Bearish").
  5. Output: Traders receive alerts via apps like Merostock.

Real-World Tie-In:

  • Ncell Network Monitoring: Stream mining detects unusual call patterns (e.g., sudden drop in signal strength in Bhaktapur) and reroutes traffic.
  • Twitter Trends: Platforms like TweetDeck use stream mining to show real-time hashtag popularity (e.g., #NepalElection2022).

3. Graph Mining

Definition: Extracts patterns from interconnected data (nodes = entities, edges = relationships). Key Techniques:

  • Community Detection: Groups nodes with dense connections (e.g., friend circles in Facebook).
  • Centrality Measures: Identifies influential nodes (e.g., top Daraz sellers).
  • Link Prediction: Suggests new connections (e.g., "People who bought X also bought Y").

Visual: Fraud Detection in Bank Transactions (Graph)

graph TD
    A["Alice\n(Account 123)"] -->|"Transfers 1000"| B["Bob\n(Account 456)"]
    B -->|"Transfers 1000"| C["Charlie\n(Account 789)"]
    C -->|"Transfers 1000"| D["Dave\n(Suspicious Account)"]
    D -->|"Linked to 50+ fake accounts"| E["Fraud Alert"]

Worked Example: Ncell Fraud Detection

  1. Input: Call detail records (CDRs) as a graph (nodes = phone numbers, edges = calls).
  2. Algorithm: Detects communities where:
    • Nodes have unusually high call volumes.
    • Edges show rapid money transfers (via USSD codes).
  3. Output: Flags accounts like 98XXXX1234 as high-risk for SIM swapping.

Real-World Tie-In:

  • Facebook Friend Suggestions: Uses graph mining to recommend connections based on mutual friends.
  • LinkedIn Recruiter: Predicts job candidates by analyzing professional networks (e.g., "Works with X at Y Company").

4. Multimedia Mining Applications

Application Data Type Technique Used Nepalese Example
Recommendation Systems Text + Ratings Collaborative Filtering + NLP Daraz’s "Frequently Bought Together"
Traffic Prediction Video + GPS Object Tracking + Time Series Kathmandu Traffic Police’s AI cameras
Medical Diagnosis X-rays + Reports CNN + Rule-Based Systems Patan Hospital’s tumor detection AI
Sentiment Analysis Social Media NLP + Machine Learning eSewa’s customer feedback dashboard

Visual: Daraz Recommendation Pipeline

flowchart LR
    A["User Browses\n(Shoes Category)"] --> B["Clickstream Data\n(Logged)"]
    B --> C["Collaborative Filtering\n('Users like X also liked Y')"]
    C --> D["NLP on Reviews\n('Comfortable', 'Durable')"]
    D --> E["Hybrid Model\n(Ranks Recommendations)"]
    E --> F["Display:\n'Suggested Products'"]

5. Challenges in Complex Data Mining

Challenge Cause Solution
Scalability Massive datasets (e.g., YouTube) Distributed systems (Apache Spark)
Privacy GDPR, data leaks Federated learning, anonymization
Interpretability Black-box models (e.g., deep CNNs) Explainable AI (SHAP values, LIME)
Real-Time Processing High velocity (e.g., stock trades) Edge computing, stream mining

Real-World Example: GDPR Compliance in Khalti

  • Challenge: Storing customer transaction data while complying with Nepal’s privacy laws.
  • Solution:
    • Differential Privacy: Adds "noise" to data to prevent re-identification.
    • Federated Learning: Trains models on-device (e.g., phones) without centralizing data.

  1. Federated Learning: Trains models across decentralized devices (e.g., Ncell’s AI on user phones).
  2. Explainable AI (XAI): Makes complex models transparent (e.g., why a bank denied a loan).
  3. Quantum Data Mining: Uses quantum computing for faster pattern recognition (future trend).

Visual: Federated Learning Workflow

sequenceDiagram
    participant User as User Device
    participant Model as Global Model
    participant Server as Central Server
    User->>Model: Train locally (data never leaves phone)
    Model->>Server: Send model updates only
    Server->>Model: Aggregate updates
    Server->>User: Improved global model

In the Real World

  1. eSewa’s Fraud Detection

    • Idea Used: Graph mining + stream mining.
    • How: Analyzes transaction graphs to detect money laundering (e.g., rapid transfers between linked accounts). Uses real-time alerts to block suspicious transactions within seconds.
  2. Daraz’s Visual Search

    • Idea Used: Image mining (CNNs).
    • How: Users upload a photo of a product (e.g., a shoe), and Daraz’s AI matches it to inventory using feature extraction, reducing search time from minutes to seconds.
  3. NTC’s Power Outage Prediction

    • Idea Used: Time-series stream mining + geospatial data.
    • How: Sensors across Nepal feed real-time power usage data. Stream mining detects patterns (e.g., sudden drops in Dhulikhel) and predicts outages, allowing NTC to reroute power or dispatch teams proactively.
  4. Pathao’s Dynamic Pricing

    • Idea Used: Real-time stream mining + demand forecasting.
    • How: Analyzes live ride requests, traffic data, and driver availability to adjust fares dynamically (e.g., surge pricing during Dashain in Lalitpur).
  5. Nepali News Aggregators (e.g., Onlinekhabar)

    • Idea Used: Topic modeling + sentiment analysis.
    • How: Groups articles by themes (e.g., "NEPSE crash") and analyzes sentiment to prioritize headlines (e.g., "Market panic" vs. "Stable growth").

Exam Tip

  1. Define Clearly: Start answers with precise definitions (e.g., "Graph mining is the process of discovering patterns in interconnected data represented as nodes and edges").
  2. Compare Techniques: Use tables to contrast methods (e.g., batch vs. stream mining).
  3. Nepalese Context: Always relate to local examples (e.g., "Ncell could use graph mining to detect SIM cloning").
  4. Visuals: Sketch diagrams for:
    • CNN layers (for image mining).
    • Sliding windows (for stream mining).
    • Graphs (for fraud detection).
  5. Challenges: Exams often ask for limitations (e.g., "Why is real-time text mining hard?" → Latency, language complexity).
  6. Applications: Link to real companies (e.g., "Daraz uses CNNs for product tagging").

Common Exam Questions:

  • "How would you detect fraud in Khalti transactions using graph mining?" → Explain community detection + anomaly scoring.
  • "Compare batch processing vs. stream mining." → Use a table with speed, use cases, and tools.
  • "Describe how YouTube recommends videos." → Collaborative filtering + content-based filtering (watch history + video metadata).

Based on the TU BIM syllabus for Data Warehousing and Data Mining (IT274), unit 10.

Discussion

Loading…