Data Warehousing and Data MiningUnit 311 min read
Data Characterization, Discrimination & Summarization: Methods, Techniques & Applications
Unit 3 of Data Warehousing and Data Mining covers core techniques for extracting meaningful insights from raw data: data characterization (describing data distributions), discrimination (comparing subsets), and summarization (condensing patterns). Learn attribute-oriented induction, data generalization, and visualizati
Key Concepts & Definitions
1. Data Characterization
Definition: The process of describing data using statistical or visual summaries to reveal inherent patterns, trends, or distributions. It answers: "What is the data like?"
Visual: Data characterization often uses histograms, box plots, or attribute profiles (e.g., min/max/avg values per category).
Example: Suppose Nepal Electricity Authority (NEA) wants to characterize power consumption data across districts. They might compute:
- Average consumption per household (e.g., 250 kWh/month in Kathmandu vs. 120 kWh in remote areas).
- Peak usage hours (e.g., 6–9 PM in urban areas).
- Outliers (e.g., districts with >50% consumption spikes during festivals).
2. Data Discrimination
Definition: Comparing subsets of data to highlight differences. It answers: "How does this group differ from others?" Used in market segmentation, fraud detection, or medical diagnostics.
Example: Compare Khalti users who default on loans vs. those who repay:
| Attribute | Defaulted Users | Repaying Users |
|---|---|---|
| Avg. Loan Amount | NPR 50,000 | NPR 25,000 |
| Credit Score | <600 | >700 |
| Loan Duration | 12+ months | 3–6 months |
Visual: Use parallel coordinates or diverging bar charts to show differences.
3. Data Summarization
Definition: Condensing data into compact forms (e.g., rules, hierarchies, or visualizations) while preserving key insights. Methods include:
- Attribute-oriented induction (generalizing data).
- Data cube aggregation (OLAP-style summaries).
- Automatic summarization (e.g., "Top 5 products bought together on Daraz").
Example: Summarize Pathao driver earnings by city and time:
Rule: IF (City = Kathmandu) AND (Hour = 18–22) THEN Avg_Earnings = NPR 2,500
Support: 60% of rides
Confidence: 85%
Core Techniques
A. Attribute-Oriented Induction (AOI)
How it works:
- Generalize specific attribute values (e.g.,
Age=25→Age=20–30). - Roll-up data to higher-level concepts (e.g.,
Product=Samsung Galaxy S23→Product=Smartphone). - Generate rules (e.g., "Young users prefer smartphones").
graph TD
A["Price: 450"] -->|"Generalize"| B["400-500"]
C["Price: 320"] -->|"Generalize"| D["300-400"]
E["Price: 180"] -->|"Generalize"| F["150-200"]
B --> G["NABL"]
D --> H["NMB"]
F --> I["CMC"]AOI generalization hierarchy for NEPSE stock prices (example from note)Steps:
- Select attributes to generalize (e.g.,
Age,Income). - Define generalization hierarchies:
- Age: 20–25 → 20–30 → 20–40 → Adult
- Income: <50k → 50k–100k → High
- Apply AOI to create a generalized table.
Worked Example: Generalize the following NEPSE stock data for a brokerage app:
| Stock | Price (NPR) | Volume (shares) | Sector |
|---|---|---|---|
| NABL | 450 | 12,000 | Banking |
| NMB | 320 | 8,500 | Banking |
| CMC | 180 | 25,000 | Cement |
Generalized Table (AOI):
| Stock | Price Range | Volume Range | Sector |
|---|---|---|---|
| NABL | 400–500 | 10k–20k | Banking |
| NMB | 300–400 | 5k–15k | Banking |
| CMC | 150–200 | 20k–30k | Heavy Industry |
Visual: Generalization hierarchy for Price:
graph TD
A["Price"] --> B["400-500"]
A --> C["300-400"]
A --> D["150-200"]
B --> E["450"]
C --> F["320"]
D --> G["180"]B. Data Characterization Methods
| Method | Description | Example |
|---|---|---|
| Statistical | Mean, median, standard deviation, quartiles. | Avg. order value on Daraz: NPR 3,200. |
| Visual | Histograms, scatter plots, heatmaps. | Heatmap of NTC internet usage by district. |
| Rule-Based | IF-THEN rules (e.g., "IF rain >50mm THEN traffic delays in Kathmandu"). | Pathao’s surge pricing rules. |
| Cluster-Based | Group similar data points (e.g., k-means). | Segmenting Ncell customers by usage patterns. |
C. Data Discrimination Methods
| Method | Use Case | Example |
|---|---|---|
| Attribute Filtering | Compare groups by attribute values. | Compare Khalti loan defaults by credit score. |
| Distance-Based | Measure dissimilarity (e.g., Euclidean). | Identify outliers in NEA power consumption data. |
| Rule Discovery | Find discriminative rules. | "Users who buy laptops also buy chargers (confidence: 90%)." |
Worked Example: Discriminate between successful and failed NEPSE IPOs using AOI. Original Data:
| Company | Sector | Issue Size (Cr) | Subscription % | Success |
|---|---|---|---|---|
| X | IT | 10 | 95% | Yes |
| Y | Cement | 15 | 60% | No |
| Z | Banking | 20 | 85% | Yes |
Generalized Rules:
- IF (Sector = IT OR Banking) AND (Subscription % >80%) THEN Success = Yes (Support: 2/3).
- IF (Sector = Cement) THEN Success = No (Support: 1/3).
D. Data Summarization Techniques
Data Cubes (OLAP):
- Pre-computed aggregations (e.g., "Total sales by region, product, and quarter").
- Example: Daraz’s sales cube:
SELECT Region, Product_Category, SUM(Revenue) FROM Sales GROUP BY Region, Product_Category - Visual: OLAP cube structure:
Automatic Summarization:
- Text: "Top 3 complaints in eSewa: Failed transactions, long queues, app crashes."
- Numerical: "90% of Pathao rides in Pokhara are <5 km."
Visual Summaries:
- Small multiples: Compare NTC vs. Ncell call drop rates by district.
- Treemaps: Show NEPSE sector-wise market cap.
In the Real World
Daraz (Nepal):
- Data Characterization: Analyzes customer purchase histories to identify trends (e.g., "Electronics sales spike during Dashain").
- Discrimination: Compares urban vs. rural buyer behavior (e.g., rural users prefer bulk purchases).
- Summarization: Generates rules like:
"IF (Cart > NPR 10,000) AND (Location = Kathmandu) THEN Offer 10% discount".
Khalti (Digital Payments):
- Discrimination: Flags transactions where
Amount > NPR 50,000ANDLocation = Remoteas high-risk (fraud pattern). - Summarization: "Top 5 merchant categories by transaction volume: Food, Transport, Retail."
- Discrimination: Flags transactions where
Nepal Electricity Authority (NEA):
- Characterization: Uses time-series data to predict peak demand (e.g., "Demand rises 30% during winter evenings").
- Visualization: Heatmaps of power outages by district to prioritize repairs.
Pathao (Ride-Hailing):
- Clustering: Groups drivers by earnings to identify "high-performing" zones (e.g., Thapathali vs. Bhaktapur).
- Rule Mining:
"IF (Hour = 22–24) AND (Location = Lakeside) THEN Surge Pricing = 1.8x".
Worked Example: AOI for Ncell Customer Churn
Problem: Ncell wants to summarize why customers churn (cancel service). Use AOI to generalize the following data:
| Customer ID | Age | Plan Type | Usage (GB) | Churned |
|---|---|---|---|---|
| C1 | 25 | Prepaid | 12 | Yes |
| C2 | 40 | Postpaid | 8 | No |
| C3 | 30 | Prepaid | 5 | Yes |
| C4 | 22 | Prepaid | 15 | Yes |
Steps:
- Define hierarchies:
- Age: 20–25 → 20–30 → 20–40 → Adult
- Usage: 0–5 → 5–10 → 10–15 → High
- Generalize:
Age Group Plan Type Usage Range Churn Rate 20–30 Prepaid High 100% (3/3) 30–40 Postpaid Medium 0% (1/1)
Rule: "IF (Age = 20–30) AND (Plan = Prepaid) AND (Usage = High) THEN Churn = Yes" (Support: 3/4).
Visual: Churn risk by plan type:
Common Pitfalls & Exam Tips
❌ Mistakes to Avoid
- Over-generalization: Turning
Age=25intoAge=18–60loses useful detail. - Ignoring hierarchies: AOI requires predefined taxonomies (e.g., product categories).
- Confusing characterization vs. discrimination:
- Characterization = "What is the data?" (e.g., "Avg. order value = NPR 3,200").
- Discrimination = "How do groups differ?" (e.g., "Urban users spend 2x more").
✅ Exam Tip: Structured Approach
For AOI questions (e.g., "Generalize this dataset"):
- Step 1: Identify attributes to generalize (e.g.,
Age,Income). - Step 2: Define hierarchies (show in a tree diagram).
- Step 3: Apply generalization (show before/after tables).
- Step 4: Derive rules (e.g., "IF [generalized condition] THEN [outcome]").
Example Question:
"Apply AOI to the following data to generalize Salary and Department hierarchies."
Your Answer:
- Hierarchies:
graph TD A["Salary"] --> B["<50k"] --> C["<30k"] A --> D["50k–100k"] A --> E[">100k"] F["Department"] --> G["IT"] --> H["Software"] F --> I["Finance"]
- Generalized Table:
Salary Range Department Avg. Performance <50k IT Medium 50k–100k Finance High
📌 Exam Tip: Real-World Mapping
- Daraz: Use data cubes for sales analysis (e.g., "Revenue by region and season").
- Khalti: Apply discrimination to detect fraud (e.g., "Unusual transaction patterns").
- NEPSE: Use AOI to summarize stock trends (e.g., "Banking stocks outperform in Q4").
Final Note: This unit is highly visual—always include diagrams for hierarchies, rules, or comparisons. For exam questions, trace steps clearly (e.g., show AOI generalization tables) and tie answers to real systems (e.g., "This rule could be used by Daraz’s recommendation engine").
Based on the TU BSc CSIT syllabus for Data Warehousing and Data Mining (CSC410), unit 3.
Discussion
Loading…