Data Analysis and VisualizationUnit 710 min read
Text Data Visualization: Techniques, Hierarchies & Real-World Applications
Unit 7 of Data Analysis and Visualization explores how to represent, analyze, and visualize unstructured text data (e.g., tweets, documents, code) using techniques like word clouds, topic modeling, and hierarchical visualizations, with practical examples from Nepali platforms like eSewa and NEPSE.
TAKEAWAYS:
- Text data visualization converts unstructured text into structured visual insights (e.g., word frequencies, sentiment trends) using techniques like word clouds, parallel tag clouds, and topic modeling.
- Hierarchical text visualization (e.g., dendrograms, tree maps) reveals relationships in nested data (e.g., email threads, code repositories).
- Separate, Order, Align (SOA) principles optimize readability in text visualizations (e.g., grouping related terms, sorting by frequency).
- Real-world applications include eSewa’s customer feedback analysis (word clouds for complaint trends) and NEPSE’s stock discussion sentiment tracking (parallel tag clouds).
- Color and typography must align with cognitive load principles (e.g., high-contrast colors for critical terms, font size for hierarchy).
- Exam focus: Compare techniques (e.g., word clouds vs. parallel tag clouds), explain Detroit dataset visualizations, and justify tool choices (e.g., Tableau for interactive text analytics).
1. Text Data: From Raw to Visualizable
Text data is unstructured (no predefined schema) but holds hidden patterns. To visualize it, we first represent it in structured forms:
Levels of Text Representation
Text data can be processed at different granularities:
| Level | Example | Visualization Use Case |
|---|---|---|
| Character-level | Emojis in tweets (😢, 😊) |
Bar charts of emoji frequency in customer reviews. |
| Word-level | "Nepal", "earthquake", "2015" | Word clouds, tag clouds. |
| Sentence/Paragraph | "The earthquake destroyed 50% of Kathmandu." | Parallel tag clouds for sentiment analysis. |
| Document-level | Entire news articles or eSewa complaints | Topic modeling (e.g., clustering similar complaints). |
| Hierarchical | Email threads, code repositories | Tree maps or dendrograms. |
2. Core Techniques for Text Visualization
A. Word Clouds: Frequency as Size
- How it works: Words are sized proportionally to their frequency in a corpus.
- Example: Visualizing eSewa customer complaints (2023 data).
- Input: 10,000 complaints scraped from eSewa’s feedback portal.
- Preprocessing: Remove stopwords (e.g., "the", "and"), stem words ("complaint" → "complain").
- Output: Word cloud where "delay" (500 occurrences) is largest, followed by "payment" (300) and "customer" (200).
graph LR
A["Raw Text\n(eSewa complaints)"] --> B["Preprocess\n(stopword removal, stemming)"]
B --> C["Count Word Frequencies"]
C --> D["Scale Font Sizes\n(by frequency)"]
D --> E["Render Word Cloud\n(see IMAGE above)"]Advantages:
- Simple, intuitive for quick insights.
- Works for single-document or corpus-level analysis.
Disadvantages:
- No context: "Bank" could mean NMB or Global IME.
- Overlap: Small words get lost.
- No hierarchy: Cannot show relationships between terms.
B. Parallel Tag Clouds: Comparative Analysis
- How it works: Multiple word clouds aligned for comparison (e.g., before/after an event).
- Example: NEPSE stock discussions (2022 vs. 2023).
- 2022: "growth", "dividend", "tech" (top terms).
- 2023: "inflation", "interest rate", "recession" (new terms emerge).
- Tool: Use Tableau or Python’s
wordcloudlibrary withmatplotlib.
When to use:
- Comparing two time periods (e.g., Pathao driver complaints pre- vs. post-lockdown).
- Analyzing multiple categories (e.g., Daraz product reviews by star rating).
C. Topic Modeling: Thematic Clusters
- How it works: Algorithms (e.g., Latent Dirichlet Allocation (LDA)) group words into topics.
- Example: Detroit Dataset (car manufacturing complaints).
- Input: 5,000 customer reviews from Ford/Mahindra.
- Output: 5 topics:
- "Engine failure" (words: oil, leak, overheating)
- "Customer service" (words: rude, delay, manager)
- "Safety issues" (words: brake, seatbelt, crash)
- Visualization: Radial dendrogram (shows topic similarity).
graph TD
A["Detroit Dataset\n(Customer Reviews)"] --> B["Preprocess\n(TF-IDF weighting)"]
B --> C["Apply LDA\n(5 topics)"]
C --> D["Visualize as\nRadial Dendrogram"]
D --> E["IMAGE: radial dendrogram of Detroit topics"]Advantages:
- Reveals latent themes (e.g., "safety" vs. "service").
- Scalable for large corpora (e.g., Kathmandu Post archives).
Disadvantages:
- Requires domain knowledge to interpret topics.
- Black-box nature: Hard to debug topic assignments.
D. Hierarchical Text Visualization
For nested text data (e.g., email threads, code repositories), use:
- Tree Maps: Rectangles sized by frequency.
- Example: GitHub repository structure (files as leaves, folders as branches).
- Dendrograms: Hierarchical clustering of documents.
- Example: Ncell customer support tickets clustered by issue type.
3. Separate, Order, Align (SOA) Principles
To make text visualizations readable, apply:
- Separate: Group related terms (e.g., enclose "bank" terms in a box).
- Order: Sort by frequency, alphabetically, or semantically.
- Align: Use grids or baselines for consistency.
Example: Word Cloud for Daraz Product Reviews
- Bad: Random placement of "fast", "slow", "delivery".
- Good: Group "delivery" terms together, sort by complaint severity.
graph LR
A["Raw Reviews\n(Daraz)"] --> B["Separate\n(by category: delivery, quality)"]
B --> C["Order\n(highest complaints first)"]
C --> D["Align\n(grids for categories)"]
D --> E["Final Visualization\n(see IMAGE below)"]4. Real-World Applications in Nepal
| Platform | Text Data | Visualization Technique | Insight Gained |
|---|---|---|---|
| eSewa | Customer complaints | Word clouds + parallel tag clouds | Identified "payment failure" as top issue. |
| NEPSE | Stock discussion forums | Topic modeling | Tracked "inflation" as emerging concern. |
| Pathao | Driver feedback | Tree maps (by city) | Kathmandu drivers complained more about traffic. |
| Ncell | Customer support tickets | Dendrograms | Clustered "billing" and "network" issues. |
| Khalti | Transaction logs | Parallel tag clouds (success vs. failure) | "Bank transfer" failures spiked in 2023. |
5. The Detroit Dataset: A Case Study
Dataset: A collection of car manufacturing customer complaints (1990s–2000s) from Detroit, USA. Key Fields:
complaint_id,customer_name,vehicle_model,issue_description,resolution_status.
Visualization Approach:
- Word Cloud: For all complaints → "engine", "brake", "leak" dominate.
- Parallel Tag Clouds: Compare complaints by
vehicle_model(e.g., Ford vs. Chevrolet). - Topic Modeling: Identify 3 topics:
- Mechanical failures (engine, transmission).
- Customer service (rude, delay).
- Safety (airbag, seatbelt).
Why This Matters for Exams:
- Detroit dataset is a classic example of hierarchical text visualization.
- Expected answer: Combine word clouds (frequency) + topic modeling (themes).
6. Tools for Text Visualization
| Tool | Best For | Example Use Case |
|---|---|---|
| Tableau | Interactive dashboards | eSewa complaint trends over time. |
Python (wordcloud) |
Static word clouds | Quick analysis of NEPSE discussions. |
| Gephi | Network graphs (e.g., co-occurrence) | Mapping terms in Kathmandu Post articles. |
| VOSviewer | Bibliometric maps | Analyzing research papers on "AI in Nepal". |
7. Exam Tip: How to Score Full Marks
For Detroit Dataset (2+3 marks):
- Explain: It’s a car complaint corpus with fields like
issue_description. - Visualization: Use word clouds for frequency + topic modeling for themes.
- Example: "A word cloud would show 'engine' as largest, while topic modeling reveals 'mechanical' vs. 'service' clusters."
- Explain: It’s a car complaint corpus with fields like
For SOA Principles (6+4 marks):
- Separate: "Group 'bank' terms in a box in a Khalti transaction word cloud."
- Order: "Sort Daraz reviews by star rating (1-star first)."
- Align: "Use a grid to align 'delivery' and 'quality' categories."
For Hierarchical Data (1+4 marks):
- Techniques:
- Tree maps: For file structures (e.g., GitHub repos).
- Dendrograms: For document clustering (e.g., Ncell tickets).
- Example: "A tree map of a Python project shows
src/as largest, withmain.pyas a leaf."
- Techniques:
For Importance of Colors (3+2 marks):
- Do: Use high contrast for critical terms (e.g., red for "failure" in Daraz reviews).
- Avoid: Pastel colors (hard to read for colorblind users).
- Example: "In a NEPSE discussion word cloud, 'inflation' in red stands out against 'growth' in green."
8. Worked Example: Analyzing Pathao Driver Complaints
Scenario: Pathao wants to visualize driver complaints in Kathmandu vs. Pokhara. Steps:
- Data: 2,000 complaints (2023), with fields:
driver_id,city,complaint_text,rating. - Preprocess: Remove stopwords, stem "traffic" → "trafic".
- Visualize:
- Word Cloud: Combine both cities → "trafic", "passenger", "app".
- Parallel Tag Clouds: Split by city.
- Kathmandu: "trafic", "delay", "police".
- Pokhara: "route", "navigat", "mountain".
- Insight: Traffic is a Kathmandu-specific issue; Pokhara drivers complain about navigation.
Based on the TU BCA syllabus for Data Analysis and Visualization (CACS455), unit 7.
Discussion
Loading…