CACS455 Data Analysis and Visualization

Data Analysis and VisualizationUnit 710 min read

Text Data Visualization: Techniques, Hierarchies & Real-World Applications

Unit 7 of Data Analysis and Visualization explores how to represent, analyze, and visualize unstructured text data (e.g., tweets, documents, code) using techniques like word clouds, topic modeling, and hierarchical visualizations, with practical examples from Nepali platforms like eSewa and NEPSE.

TAKEAWAYS:

  • Text data visualization converts unstructured text into structured visual insights (e.g., word frequencies, sentiment trends) using techniques like word clouds, parallel tag clouds, and topic modeling.
  • Hierarchical text visualization (e.g., dendrograms, tree maps) reveals relationships in nested data (e.g., email threads, code repositories).
  • Separate, Order, Align (SOA) principles optimize readability in text visualizations (e.g., grouping related terms, sorting by frequency).
  • Real-world applications include eSewa’s customer feedback analysis (word clouds for complaint trends) and NEPSE’s stock discussion sentiment tracking (parallel tag clouds).
  • Color and typography must align with cognitive load principles (e.g., high-contrast colors for critical terms, font size for hierarchy).
  • Exam focus: Compare techniques (e.g., word clouds vs. parallel tag clouds), explain Detroit dataset visualizations, and justify tool choices (e.g., Tableau for interactive text analytics).

1. Text Data: From Raw to Visualizable

Text data is unstructured (no predefined schema) but holds hidden patterns. To visualize it, we first represent it in structured forms:

Levels of Text Representation

Text data can be processed at different granularities:

Level Example Visualization Use Case
Character-level Emojis in tweets (😢, 😊) Bar charts of emoji frequency in customer reviews.
Word-level "Nepal", "earthquake", "2015" Word clouds, tag clouds.
Sentence/Paragraph "The earthquake destroyed 50% of Kathmandu." Parallel tag clouds for sentiment analysis.
Document-level Entire news articles or eSewa complaints Topic modeling (e.g., clustering similar complaints).
Hierarchical Email threads, code repositories Tree maps or dendrograms.

2. Core Techniques for Text Visualization

[object Object][object Object][object Object]Word CloudParallel Tag CloudsTopic ModelingHierarchical Text Visualization
Relationship between core text visualization techniques

A. Word Clouds: Frequency as Size

  • How it works: Words are sized proportionally to their frequency in a corpus.
  • Example: Visualizing eSewa customer complaints (2023 data).
    • Input: 10,000 complaints scraped from eSewa’s feedback portal.
    • Preprocessing: Remove stopwords (e.g., "the", "and"), stem words ("complaint" → "complain").
    • Output: Word cloud where "delay" (500 occurrences) is largest, followed by "payment" (300) and "customer" (200).
graph LR
    A["Raw Text\n(eSewa complaints)"] --> B["Preprocess\n(stopword removal, stemming)"]
    B --> C["Count Word Frequencies"]
    C --> D["Scale Font Sizes\n(by frequency)"]
    D --> E["Render Word Cloud\n(see IMAGE above)"]

Advantages:

  • Simple, intuitive for quick insights.
  • Works for single-document or corpus-level analysis.

Disadvantages:

  • No context: "Bank" could mean NMB or Global IME.
  • Overlap: Small words get lost.
  • No hierarchy: Cannot show relationships between terms.

B. Parallel Tag Clouds: Comparative Analysis

  • How it works: Multiple word clouds aligned for comparison (e.g., before/after an event).
  • Example: NEPSE stock discussions (2022 vs. 2023).
    • 2022: "growth", "dividend", "tech" (top terms).
    • 2023: "inflation", "interest rate", "recession" (new terms emerge).
    • Tool: Use Tableau or Python’s wordcloud library with matplotlib.

When to use:

  • Comparing two time periods (e.g., Pathao driver complaints pre- vs. post-lockdown).
  • Analyzing multiple categories (e.g., Daraz product reviews by star rating).

C. Topic Modeling: Thematic Clusters

  • How it works: Algorithms (e.g., Latent Dirichlet Allocation (LDA)) group words into topics.
  • Example: Detroit Dataset (car manufacturing complaints).
    • Input: 5,000 customer reviews from Ford/Mahindra.
    • Output: 5 topics:
      1. "Engine failure" (words: oil, leak, overheating)
      2. "Customer service" (words: rude, delay, manager)
      3. "Safety issues" (words: brake, seatbelt, crash)
    • Visualization: Radial dendrogram (shows topic similarity).
graph TD
    A["Detroit Dataset\n(Customer Reviews)"] --> B["Preprocess\n(TF-IDF weighting)"]
    B --> C["Apply LDA\n(5 topics)"]
    C --> D["Visualize as\nRadial Dendrogram"]
    D --> E["IMAGE: radial dendrogram of Detroit topics"]

Advantages:

  • Reveals latent themes (e.g., "safety" vs. "service").
  • Scalable for large corpora (e.g., Kathmandu Post archives).

Disadvantages:

  • Requires domain knowledge to interpret topics.
  • Black-box nature: Hard to debug topic assignments.

D. Hierarchical Text Visualization

For nested text data (e.g., email threads, code repositories), use:

  1. Tree Maps: Rectangles sized by frequency.
    • Example: GitHub repository structure (files as leaves, folders as branches).
  2. Dendrograms: Hierarchical clustering of documents.
    • Example: Ncell customer support tickets clustered by issue type.

3. Separate, Order, Align (SOA) Principles

To make text visualizations readable, apply:

  • Separate: Group related terms (e.g., enclose "bank" terms in a box).
  • Order: Sort by frequency, alphabetically, or semantically.
  • Align: Use grids or baselines for consistency.

Example: Word Cloud for Daraz Product Reviews

  • Bad: Random placement of "fast", "slow", "delivery".
  • Good: Group "delivery" terms together, sort by complaint severity.
graph LR
    A["Raw Reviews\n(Daraz)"] --> B["Separate\n(by category: delivery, quality)"]
    B --> C["Order\n(highest complaints first)"]
    C --> D["Align\n(grids for categories)"]
    D --> E["Final Visualization\n(see IMAGE below)"]

4. Real-World Applications in Nepal

Platform Text Data Visualization Technique Insight Gained
eSewa Customer complaints Word clouds + parallel tag clouds Identified "payment failure" as top issue.
NEPSE Stock discussion forums Topic modeling Tracked "inflation" as emerging concern.
Pathao Driver feedback Tree maps (by city) Kathmandu drivers complained more about traffic.
Ncell Customer support tickets Dendrograms Clustered "billing" and "network" issues.
Khalti Transaction logs Parallel tag clouds (success vs. failure) "Bank transfer" failures spiked in 2023.

5. The Detroit Dataset: A Case Study

Dataset: A collection of car manufacturing customer complaints (1990s–2000s) from Detroit, USA. Key Fields:

  • complaint_id, customer_name, vehicle_model, issue_description, resolution_status.

Visualization Approach:

  1. Word Cloud: For all complaints → "engine", "brake", "leak" dominate.
  2. Parallel Tag Clouds: Compare complaints by vehicle_model (e.g., Ford vs. Chevrolet).
  3. Topic Modeling: Identify 3 topics:
    • Mechanical failures (engine, transmission).
    • Customer service (rude, delay).
    • Safety (airbag, seatbelt).

Why This Matters for Exams:

  • Detroit dataset is a classic example of hierarchical text visualization.
  • Expected answer: Combine word clouds (frequency) + topic modeling (themes).

6. Tools for Text Visualization

Tool Best For Example Use Case
Tableau Interactive dashboards eSewa complaint trends over time.
Python (wordcloud) Static word clouds Quick analysis of NEPSE discussions.
Gephi Network graphs (e.g., co-occurrence) Mapping terms in Kathmandu Post articles.
VOSviewer Bibliometric maps Analyzing research papers on "AI in Nepal".
Raw Text0Preprocessed Text1Tokenized Words2TF-IDF Vectors3Visualized Output4
Text data pipeline stages (from raw to visualized)

7. Exam Tip: How to Score Full Marks

  1. For Detroit Dataset (2+3 marks):

    • Explain: It’s a car complaint corpus with fields like issue_description.
    • Visualization: Use word clouds for frequency + topic modeling for themes.
    • Example: "A word cloud would show 'engine' as largest, while topic modeling reveals 'mechanical' vs. 'service' clusters."
  2. For SOA Principles (6+4 marks):

    • Separate: "Group 'bank' terms in a box in a Khalti transaction word cloud."
    • Order: "Sort Daraz reviews by star rating (1-star first)."
    • Align: "Use a grid to align 'delivery' and 'quality' categories."
  3. For Hierarchical Data (1+4 marks):

    • Techniques:
      • Tree maps: For file structures (e.g., GitHub repos).
      • Dendrograms: For document clustering (e.g., Ncell tickets).
    • Example: "A tree map of a Python project shows src/ as largest, with main.py as a leaf."
  4. For Importance of Colors (3+2 marks):

    • Do: Use high contrast for critical terms (e.g., red for "failure" in Daraz reviews).
    • Avoid: Pastel colors (hard to read for colorblind users).
    • Example: "In a NEPSE discussion word cloud, 'inflation' in red stands out against 'growth' in green."

8. Worked Example: Analyzing Pathao Driver Complaints

Scenario: Pathao wants to visualize driver complaints in Kathmandu vs. Pokhara. Steps:

  1. Data: 2,000 complaints (2023), with fields: driver_id, city, complaint_text, rating.
  2. Preprocess: Remove stopwords, stem "traffic" → "trafic".
  3. Visualize:
    • Word Cloud: Combine both cities → "trafic", "passenger", "app".
    • Parallel Tag Clouds: Split by city.
      • Kathmandu: "trafic", "delay", "police".
      • Pokhara: "route", "navigat", "mountain".
  4. Insight: Traffic is a Kathmandu-specific issue; Pokhara drivers complain about navigation.

Based on the TU BCA syllabus for Data Analysis and Visualization (CACS455), unit 7.

Discussion

Loading…