Business IntelligenceUnit 98 min read
Big Data & BI: Volume, Velocity, Variety, and Value
Unit 9 of Business Intelligence explores how Big Data transforms decision-making by analyzing massive datasets (volume, velocity, variety, and veracity) using BI tools, cloud platforms, and real-world applications like fraud detection, customer personalization, and predictive maintenance.
Big Data: The 5Vs Framework
Big Data is not just about large datasets—it’s about volume, velocity, variety, veracity, and value. These five dimensions define how data is generated, stored, processed, and leveraged for business intelligence.
The 5Vs of Big Data
mindmap
root((Big Data: 5Vs))
Volume
Petabytes/Exabytes of data
Example: Facebook generates 4PB/day
Velocity
Real-time or near-real-time data
Example: Stock market trades per second
Variety
Structured (SQL), Semi-structured (JSON), Unstructured (text, images)
Example: Social media posts + sensor logs
Veracity
Data quality: accuracy, consistency, completeness
Example: Ncell call detail records (CDRs) with missing entries
Value
Insights from data: predictive analytics, personalization
Example: Daraz using purchase history to recommend productsWhy does this matter?
- Volume: Traditional databases (e.g., MySQL) fail at scale. Big Data requires distributed systems like Hadoop or Spark.
- Velocity: Streaming data (e.g., IoT sensors) needs real-time processing (e.g., Apache Kafka).
- Variety: Unstructured data (e.g., WhatsApp messages) requires NLP or image recognition (e.g., Google Photos).
- Veracity: Dirty data leads to wrong decisions. Data cleaning (e.g., Python’s Pandas) is critical.
- Value: The goal is actionable insights (e.g., NEPSE predicting stock trends).
How Big Data Works: Technologies and Tools
Big Data relies on distributed storage, processing frameworks, and analytics engines. Here’s how they connect:
flowchart TD A["Raw Data Sources"] -->|"Ingest"| B["Storage\n(Hadoop HDFS, S3, Cassandra)"] B --> C["Processing\n(Spark, Hive, Flink)"] C --> D["Analytics\n(SQL, ML, Graph DBs)"] D --> E["Visualization\n(Tableau, Power BI, Looker)"] E --> F["Action\n(Decisions, Automation)"]
Key Technologies
| Component | Tools/Frameworks | Use Case |
|---|---|---|
| Storage | Hadoop HDFS, Google BigQuery | Store terabytes of logs (e.g., Pathao ride data) |
| Processing | Apache Spark, Flink | Real-time fraud detection (e.g., Khalti transactions) |
| Analytics | SQL, TensorFlow, R | Predictive maintenance (e.g., Toyota engines) |
| Visualization | Tableau, Power BI, Looker | Dashboards for NTC network traffic |
| Cloud Platforms | AWS, Google Cloud, Azure | Scalable BI for Daraz inventory |
In the real world
Pathao’s Ride Demand Prediction
- Idea: Uses velocity (real-time GPS data) and variety (ride requests + weather APIs) to predict surge pricing.
- How: Spark processes 100K+ ride requests/sec to adjust driver incentives dynamically.
Ncell’s Churn Prediction
- Idea: Data mining (Unit 6) + Big Data (this unit) analyzes call logs, SMS, and app usage to predict which customers will switch operators.
- How: Hadoop clusters process veracity-cleaned CDRs to flag at-risk users for retention offers.
Daraz’s Recommendation Engine
- Idea: Variety (product views, carts, reviews) + value (personalized suggestions) drives 30% of sales.
- How: Spark MLlib trains models on petabytes of purchase data to recommend "Frequently Bought Together" items.
Case Study: Chaudhary Group’s Supply Chain Optimization
Chaudhary Group (owners of Nabil Bank and Himalayan Java) uses Big Data to optimize its $2B+ supply chain across Nepal and India.
Problem:
- Silos: Sales, logistics, and finance teams used separate systems.
- Delays: No real-time visibility into inventory or demand spikes.
Solution:
- Data Integration:
- Combined structured (ERP data) and unstructured (WhatsApp supplier chats, social media trends) data.
- Processing:
- Apache Spark analyzed 50TB/year of transaction logs to forecast demand.
- Action:
- Dynamic pricing for Himalayan Java (e.g., discounts during monsoon when tea sales dip).
- Route optimization for Chaudhary Transport (saved 15% fuel costs).
Result:
- 22% reduction in stockouts.
- 18% higher profit margins for Nabil Bank’s SME loans (by predicting repayment risks).
Big Data vs. Traditional BI: Key Differences
| Feature | Traditional BI | Big Data + BI |
|---|---|---|
| Data Size | GBs (structured, SQL databases) | PBs/EBs (unstructured, NoSQL) |
| Processing | Batch (daily/weekly reports) | Real-time (streaming) |
| Tools | Excel, SQL Server, Tableau | Hadoop, Spark, Kafka, TensorFlow |
| Use Case | Historical reporting | Predictive analytics, personalization |
| Example | Nabil Bank’s monthly loan reports | Khalti’s real-time fraud alerts |
Challenges of Big Data in BI
Despite its power, Big Data faces hurdles:
1. Data Privacy and Security
- Example: NEPSE’s stock data leaks could crash markets.
- Solution: Encryption (e.g., AWS KMS), GDPR compliance.
2. Skill Gaps
- Example: Nepal’s IT workforce lacks Spark or Python expertise.
- Solution: TU’s BIM program now includes Big Data electives.
3. Cost
- Example: Setting up a Hadoop cluster costs $50K+.
- Solution: Cloud services (e.g., Google BigQuery pay-as-you-go).
Worked Example: Predicting Traffic Congestion in Kathmandu
Scenario: NTC wants to reduce delays on the Ring Road using Big Data.
Step 1: Data Collection
- Sources:
- GPS data from Pathao/Khatik rides (velocity).
- CCTV footage (variety: images).
- Weather APIs (veracity: clean missing data).
Step 2: Processing
# Pseudocode for congestion prediction (using Spark)
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("TrafficBI").getOrCreate()
gps_data = spark.read.parquet("s3://ntc-gps-logs/") # 1TB of ride data
weather_data = spark.read.json("s3://ntc-weather/")
# Join and analyze
congestion_df = gps_data.join(weather_data, "timestamp") \
.filter("speed < 10 km/h") # Slow rides = congestion
congestion_df.groupBy("road_segment").count().show()
Step 3: Visualization
Output: A Tableau dashboard showing:
- Heatmap of congestion hotspots (e.g., Balkhu to Chabahil).
- Real-time alerts for NTC traffic controllers.
Result:
- Reduced delays by 25% during peak hours.
- Saved Rs. 50M/year in fuel costs for commuters.
Exam Tip
This unit is conceptual + applied. Expect:
- Definitions: Explain the 5Vs or Big Data vs. BI (2 marks).
- Diagrams: Draw a Big Data pipeline (e.g., storage → processing → analytics) (3 marks).
- Case Studies: Describe how Ncell or Daraz uses Big Data (4 marks).
- Short Answers:
- "How does velocity impact real-time fraud detection?" (2 marks).
- "Compare Hadoop and SQL databases" (3 marks).
- Problem-Solving:
- Given a dataset (e.g., NEPSE stock prices), outline steps to predict trends using Big Data tools (5 marks).
Focus Areas:
- Memorize the 5Vs and Big Data technologies.
- Know one real-world example per V (e.g., velocity = Pathao, variety = WhatsApp).
- Practice drawing the Big Data flowchart (storage → processing → analytics).
Based on the TU BIM syllabus for Business Intelligence (IT249), unit 9.
Discussion
Loading…