Business IntelligenceUnit 98 min read

Big Data & BI: Volume, Velocity, Variety, and Value

Unit 9 of Business Intelligence explores how Big Data transforms decision-making by analyzing massive datasets (volume, velocity, variety, and veracity) using BI tools, cloud platforms, and real-world applications like fraud detection, customer personalization, and predictive maintenance.

Big Data: The 5Vs Framework

Big Data is not just about large datasets—it’s about volume, velocity, variety, veracity, and value. These five dimensions define how data is generated, stored, processed, and leveraged for business intelligence.

The 5Vs of Big Data

mindmap
  root((Big Data: 5Vs))
    Volume
      Petabytes/Exabytes of data
      Example: Facebook generates 4PB/day
    Velocity
      Real-time or near-real-time data
      Example: Stock market trades per second
    Variety
      Structured (SQL), Semi-structured (JSON), Unstructured (text, images)
      Example: Social media posts + sensor logs
    Veracity
      Data quality: accuracy, consistency, completeness
      Example: Ncell call detail records (CDRs) with missing entries
    Value
      Insights from data: predictive analytics, personalization
      Example: Daraz using purchase history to recommend products

Why does this matter?

  • Volume: Traditional databases (e.g., MySQL) fail at scale. Big Data requires distributed systems like Hadoop or Spark.
  • Velocity: Streaming data (e.g., IoT sensors) needs real-time processing (e.g., Apache Kafka).
  • Variety: Unstructured data (e.g., WhatsApp messages) requires NLP or image recognition (e.g., Google Photos).
  • Veracity: Dirty data leads to wrong decisions. Data cleaning (e.g., Python’s Pandas) is critical.
  • Value: The goal is actionable insights (e.g., NEPSE predicting stock trends).

How Big Data Works: Technologies and Tools

Big Data relies on distributed storage, processing frameworks, and analytics engines. Here’s how they connect:

flowchart TD
  A["Raw Data Sources"] -->|"Ingest"| B["Storage\n(Hadoop HDFS, S3, Cassandra)"]
  B --> C["Processing\n(Spark, Hive, Flink)"]
  C --> D["Analytics\n(SQL, ML, Graph DBs)"]
  D --> E["Visualization\n(Tableau, Power BI, Looker)"]
  E --> F["Action\n(Decisions, Automation)"]

Key Technologies

Component Tools/Frameworks Use Case
Storage Hadoop HDFS, Google BigQuery Store terabytes of logs (e.g., Pathao ride data)
Processing Apache Spark, Flink Real-time fraud detection (e.g., Khalti transactions)
Analytics SQL, TensorFlow, R Predictive maintenance (e.g., Toyota engines)
Visualization Tableau, Power BI, Looker Dashboards for NTC network traffic
Cloud Platforms AWS, Google Cloud, Azure Scalable BI for Daraz inventory

In the real world

  1. Pathao’s Ride Demand Prediction

    • Idea: Uses velocity (real-time GPS data) and variety (ride requests + weather APIs) to predict surge pricing.
    • How: Spark processes 100K+ ride requests/sec to adjust driver incentives dynamically.
  2. Ncell’s Churn Prediction

    • Idea: Data mining (Unit 6) + Big Data (this unit) analyzes call logs, SMS, and app usage to predict which customers will switch operators.
    • How: Hadoop clusters process veracity-cleaned CDRs to flag at-risk users for retention offers.
  3. Daraz’s Recommendation Engine

    • Idea: Variety (product views, carts, reviews) + value (personalized suggestions) drives 30% of sales.
    • How: Spark MLlib trains models on petabytes of purchase data to recommend "Frequently Bought Together" items.

Case Study: Chaudhary Group’s Supply Chain Optimization

Chaudhary Group (owners of Nabil Bank and Himalayan Java) uses Big Data to optimize its $2B+ supply chain across Nepal and India.

Problem:

  • Silos: Sales, logistics, and finance teams used separate systems.
  • Delays: No real-time visibility into inventory or demand spikes.

Solution:

  1. Data Integration:
    • Combined structured (ERP data) and unstructured (WhatsApp supplier chats, social media trends) data.
  2. Processing:
    • Apache Spark analyzed 50TB/year of transaction logs to forecast demand.
  3. Action:
    • Dynamic pricing for Himalayan Java (e.g., discounts during monsoon when tea sales dip).
    • Route optimization for Chaudhary Transport (saved 15% fuel costs).

Result:

  • 22% reduction in stockouts.
  • 18% higher profit margins for Nabil Bank’s SME loans (by predicting repayment risks).

Big Data vs. Traditional BI: Key Differences

Feature Traditional BI Big Data + BI
Data Size GBs (structured, SQL databases) PBs/EBs (unstructured, NoSQL)
Processing Batch (daily/weekly reports) Real-time (streaming)
Tools Excel, SQL Server, Tableau Hadoop, Spark, Kafka, TensorFlow
Use Case Historical reporting Predictive analytics, personalization
Example Nabil Bank’s monthly loan reports Khalti’s real-time fraud alerts

Challenges of Big Data in BI

Despite its power, Big Data faces hurdles:

1. Data Privacy and Security

  • Example: NEPSE’s stock data leaks could crash markets.
  • Solution: Encryption (e.g., AWS KMS), GDPR compliance.

2. Skill Gaps

  • Example: Nepal’s IT workforce lacks Spark or Python expertise.
  • Solution: TU’s BIM program now includes Big Data electives.

3. Cost

  • Example: Setting up a Hadoop cluster costs $50K+.
  • Solution: Cloud services (e.g., Google BigQuery pay-as-you-go).

Worked Example: Predicting Traffic Congestion in Kathmandu

Scenario: NTC wants to reduce delays on the Ring Road using Big Data.

04590135180Morning Peak (6-9 AM)120Afternoon Peak (4-7 PM)180Evening (8-10 PM)90Night (12-2 AM)45Average Congestion Index (0-200)
Hypothetical traffic congestion data for Kathmandu (simplified for example)

Step 1: Data Collection

  • Sources:
    • GPS data from Pathao/Khatik rides (velocity).
    • CCTV footage (variety: images).
    • Weather APIs (veracity: clean missing data).

Step 2: Processing

# Pseudocode for congestion prediction (using Spark)
from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("TrafficBI").getOrCreate()
gps_data = spark.read.parquet("s3://ntc-gps-logs/")  # 1TB of ride data
weather_data = spark.read.json("s3://ntc-weather/")

# Join and analyze
congestion_df = gps_data.join(weather_data, "timestamp") \
    .filter("speed < 10 km/h")  # Slow rides = congestion
congestion_df.groupBy("road_segment").count().show()

Step 3: Visualization

Output: A Tableau dashboard showing:

  • Heatmap of congestion hotspots (e.g., Balkhu to Chabahil).
  • Real-time alerts for NTC traffic controllers.

Result:

  • Reduced delays by 25% during peak hours.
  • Saved Rs. 50M/year in fuel costs for commuters.

Exam Tip

This unit is conceptual + applied. Expect:

  1. Definitions: Explain the 5Vs or Big Data vs. BI (2 marks).
  2. Diagrams: Draw a Big Data pipeline (e.g., storage → processing → analytics) (3 marks).
  3. Case Studies: Describe how Ncell or Daraz uses Big Data (4 marks).
  4. Short Answers:
    • "How does velocity impact real-time fraud detection?" (2 marks).
    • "Compare Hadoop and SQL databases" (3 marks).
  5. Problem-Solving:
    • Given a dataset (e.g., NEPSE stock prices), outline steps to predict trends using Big Data tools (5 marks).

Focus Areas:

  • Memorize the 5Vs and Big Data technologies.
  • Know one real-world example per V (e.g., velocity = Pathao, variety = WhatsApp).
  • Practice drawing the Big Data flowchart (storage → processing → analytics).

Based on the TU BIM syllabus for Business Intelligence (IT249), unit 9.

Discussion

Loading…