CSC410 Data Warehousing and Data Mining

Data Warehousing and Data MiningUnit 311 min read

Data Characterization, Discrimination & Summarization: Methods, Techniques & Applications

Unit 3 of Data Warehousing and Data Mining covers core techniques for extracting meaningful insights from raw data: data characterization (describing data distributions), discrimination (comparing subsets), and summarization (condensing patterns). Learn attribute-oriented induction, data generalization, and visualizati

Key Concepts & Definitions

1. Data Characterization

Definition: The process of describing data using statistical or visual summaries to reveal inherent patterns, trends, or distributions. It answers: "What is the data like?"

062.5125187.5250Kathmandu250Pokhara180Biratnagar120Dharan90Average Monthly Power Consumption (kWh/household)
NEA district-wise power consumption (example from note)

Visual: Data characterization often uses histograms, box plots, or attribute profiles (e.g., min/max/avg values per category).

Example: Suppose Nepal Electricity Authority (NEA) wants to characterize power consumption data across districts. They might compute:

  • Average consumption per household (e.g., 250 kWh/month in Kathmandu vs. 120 kWh in remote areas).
  • Peak usage hours (e.g., 6–9 PM in urban areas).
  • Outliers (e.g., districts with >50% consumption spikes during festivals).

2. Data Discrimination

Definition: Comparing subsets of data to highlight differences. It answers: "How does this group differ from others?" Used in market segmentation, fraud detection, or medical diagnostics.

[object Object][object Object][object Object]Defaulted UsersRepaying Users
Attribute comparison (Khalti loan data from note)

Example: Compare Khalti users who default on loans vs. those who repay:

Attribute Defaulted Users Repaying Users
Avg. Loan Amount NPR 50,000 NPR 25,000
Credit Score <600 >700
Loan Duration 12+ months 3–6 months

Visual: Use parallel coordinates or diverging bar charts to show differences.


3. Data Summarization

Definition: Condensing data into compact forms (e.g., rules, hierarchies, or visualizations) while preserving key insights. Methods include:

  • Attribute-oriented induction (generalizing data).
  • Data cube aggregation (OLAP-style summaries).
  • Automatic summarization (e.g., "Top 5 products bought together on Daraz").

Example: Summarize Pathao driver earnings by city and time:

Rule: IF (City = Kathmandu) AND (Hour = 18–22) THEN Avg_Earnings = NPR 2,500
Support: 60% of rides
Confidence: 85%

Core Techniques

A. Attribute-Oriented Induction (AOI)

How it works:

  1. Generalize specific attribute values (e.g., Age=25 → Age=20–30).
  2. Roll-up data to higher-level concepts (e.g., Product=Samsung Galaxy S23 → Product=Smartphone).
  3. Generate rules (e.g., "Young users prefer smartphones").
graph TD
    A["Price: 450"] -->|"Generalize"| B["400-500"]
    C["Price: 320"] -->|"Generalize"| D["300-400"]
    E["Price: 180"] -->|"Generalize"| F["150-200"]
    B --> G["NABL"]
    D --> H["NMB"]
    F --> I["CMC"]
AOI generalization hierarchy for NEPSE stock prices (example from note)

Steps:

  1. Select attributes to generalize (e.g., Age, Income).
  2. Define generalization hierarchies:
    • Age: 20–25 → 20–30 → 20–40 → Adult
    • Income: <50k → 50k–100k → High
  3. Apply AOI to create a generalized table.

Worked Example: Generalize the following NEPSE stock data for a brokerage app:

Stock Price (NPR) Volume (shares) Sector
NABL 450 12,000 Banking
NMB 320 8,500 Banking
CMC 180 25,000 Cement

Generalized Table (AOI):

Stock Price Range Volume Range Sector
NABL 400–500 10k–20k Banking
NMB 300–400 5k–15k Banking
CMC 150–200 20k–30k Heavy Industry

Visual: Generalization hierarchy for Price:

graph TD
    A["Price"] --> B["400-500"]
    A --> C["300-400"]
    A --> D["150-200"]
    B --> E["450"]
    C --> F["320"]
    D --> G["180"]

B. Data Characterization Methods

Method Description Example
Statistical Mean, median, standard deviation, quartiles. Avg. order value on Daraz: NPR 3,200.
Visual Histograms, scatter plots, heatmaps. Heatmap of NTC internet usage by district.
Rule-Based IF-THEN rules (e.g., "IF rain >50mm THEN traffic delays in Kathmandu"). Pathao’s surge pricing rules.
Cluster-Based Group similar data points (e.g., k-means). Segmenting Ncell customers by usage patterns.

C. Data Discrimination Methods

Method Use Case Example
Attribute Filtering Compare groups by attribute values. Compare Khalti loan defaults by credit score.
Distance-Based Measure dissimilarity (e.g., Euclidean). Identify outliers in NEA power consumption data.
Rule Discovery Find discriminative rules. "Users who buy laptops also buy chargers (confidence: 90%)."

Worked Example: Discriminate between successful and failed NEPSE IPOs using AOI. Original Data:

Company Sector Issue Size (Cr) Subscription % Success
X IT 10 95% Yes
Y Cement 15 60% No
Z Banking 20 85% Yes

Generalized Rules:

  1. IF (Sector = IT OR Banking) AND (Subscription % >80%) THEN Success = Yes (Support: 2/3).
  2. IF (Sector = Cement) THEN Success = No (Support: 1/3).

D. Data Summarization Techniques

  1. Data Cubes (OLAP):

    • Pre-computed aggregations (e.g., "Total sales by region, product, and quarter").
    • Example: Daraz’s sales cube:
      SELECT Region, Product_Category, SUM(Revenue)
      FROM Sales
      GROUP BY Region, Product_Category
      
    • Visual: OLAP cube structure:
  2. Automatic Summarization:

    • Text: "Top 3 complaints in eSewa: Failed transactions, long queues, app crashes."
    • Numerical: "90% of Pathao rides in Pokhara are <5 km."
  3. Visual Summaries:

    • Small multiples: Compare NTC vs. Ncell call drop rates by district.
    • Treemaps: Show NEPSE sector-wise market cap.

In the Real World

  1. Daraz (Nepal):

    • Data Characterization: Analyzes customer purchase histories to identify trends (e.g., "Electronics sales spike during Dashain").
    • Discrimination: Compares urban vs. rural buyer behavior (e.g., rural users prefer bulk purchases).
    • Summarization: Generates rules like: "IF (Cart > NPR 10,000) AND (Location = Kathmandu) THEN Offer 10% discount".
  2. Khalti (Digital Payments):

    • Discrimination: Flags transactions where Amount > NPR 50,000 AND Location = Remote as high-risk (fraud pattern).
    • Summarization: "Top 5 merchant categories by transaction volume: Food, Transport, Retail."
  3. Nepal Electricity Authority (NEA):

    • Characterization: Uses time-series data to predict peak demand (e.g., "Demand rises 30% during winter evenings").
    • Visualization: Heatmaps of power outages by district to prioritize repairs.
  4. Pathao (Ride-Hailing):

    • Clustering: Groups drivers by earnings to identify "high-performing" zones (e.g., Thapathali vs. Bhaktapur).
    • Rule Mining: "IF (Hour = 22–24) AND (Location = Lakeside) THEN Surge Pricing = 1.8x".

Worked Example: AOI for Ncell Customer Churn

Problem: Ncell wants to summarize why customers churn (cancel service). Use AOI to generalize the following data:

Customer ID Age Plan Type Usage (GB) Churned
C1 25 Prepaid 12 Yes
C2 40 Postpaid 8 No
C3 30 Prepaid 5 Yes
C4 22 Prepaid 15 Yes

Steps:

  1. Define hierarchies:
    • Age: 20–25 → 20–30 → 20–40 → Adult
    • Usage: 0–5 → 5–10 → 10–15 → High
  2. Generalize:
    Age Group Plan Type Usage Range Churn Rate
    20–30 Prepaid High 100% (3/3)
    30–40 Postpaid Medium 0% (1/1)

Rule: "IF (Age = 20–30) AND (Plan = Prepaid) AND (Usage = High) THEN Churn = Yes" (Support: 3/4).

Visual: Churn risk by plan type:


Common Pitfalls & Exam Tips

❌ Mistakes to Avoid

  1. Over-generalization: Turning Age=25 into Age=18–60 loses useful detail.
  2. Ignoring hierarchies: AOI requires predefined taxonomies (e.g., product categories).
  3. Confusing characterization vs. discrimination:
    • Characterization = "What is the data?" (e.g., "Avg. order value = NPR 3,200").
    • Discrimination = "How do groups differ?" (e.g., "Urban users spend 2x more").

✅ Exam Tip: Structured Approach

For AOI questions (e.g., "Generalize this dataset"):

  1. Step 1: Identify attributes to generalize (e.g., Age, Income).
  2. Step 2: Define hierarchies (show in a tree diagram).
  3. Step 3: Apply generalization (show before/after tables).
  4. Step 4: Derive rules (e.g., "IF [generalized condition] THEN [outcome]").

Example Question: "Apply AOI to the following data to generalize Salary and Department hierarchies." Your Answer:

  1. Hierarchies:
    graph TD
      A["Salary"] --> B["<50k"] --> C["<30k"]
      A --> D["50k–100k"]
      A --> E[">100k"]
      F["Department"] --> G["IT"] --> H["Software"]
      F --> I["Finance"]
  2. Generalized Table:
    Salary Range Department Avg. Performance
    <50k IT Medium
    50k–100k Finance High

📌 Exam Tip: Real-World Mapping

  • Daraz: Use data cubes for sales analysis (e.g., "Revenue by region and season").
  • Khalti: Apply discrimination to detect fraud (e.g., "Unusual transaction patterns").
  • NEPSE: Use AOI to summarize stock trends (e.g., "Banking stocks outperform in Q4").

Final Note: This unit is highly visual—always include diagrams for hierarchies, rules, or comparisons. For exam questions, trace steps clearly (e.g., show AOI generalization tables) and tie answers to real systems (e.g., "This rule could be used by Daraz’s recommendation engine").

Based on the TU BSc CSIT syllabus for Data Warehousing and Data Mining (CSC410), unit 3.

Discussion

Loading…