Business IntelligenceUnit 910 min read
Big Data & BI: Technologies, Challenges & Applications
Unit 9 of Business Intelligence explores how big data technologies (Hadoop, Spark, NoSQL) integrate with BI tools to solve real-world problems, covering scalability challenges, real-time analytics, and case studies from Nepali and global companies.
TAKEAWAYS:
- Big data is not just large data—it’s the 3Vs (Volume, Velocity, Variety) that require distributed processing (Hadoop/Spark) and schema-flexible storage (NoSQL).
- BI + Big Data enables predictive analytics (e.g., Ncell’s churn prediction) and real-time dashboards (e.g., Daraz’s inventory optimization) by combining structured (SQL) and unstructured (text/social media) data.
- Challenges like data privacy (GDPR), bias in algorithms (WhatsApp’s recommendation systems), and infrastructure costs (Nepal’s NTC’s IoT sensor networks) must be addressed with governance frameworks.
- NoSQL databases (MongoDB, Cassandra) trade ACID for scalability—ideal for log data (Pathao’s ride history) or user profiles (eSewa’s transaction records), but require application-level transactions.
- Real-time analytics (e.g., YouTube’s trending videos, NEPSE’s stock price alerts) use stream processing (Flink, Kafka) to turn raw data into actionable insights within milliseconds.
- Case studies (e.g., Himalayan Java’s supply chain optimization, Google’s DeepMind for energy efficiency) show how big data + BI drive cost savings, personalization, and competitive advantage.
1. What Is Big Data? Beyond the Hype
Big data isn’t just about size—it’s about breaking traditional BI tools (like SQL databases) that struggle with:
- Volume: Petabytes of logs (e.g., NTC’s 5G network call records).
- Velocity: Millions of transactions/second (e.g., Khalti’s payment gateway).
- Variety: Unstructured data (emails, videos, sensor data from Daraz’s delivery drones).
- Veracity: Noise in social media (e.g., Nepal Police’s crime prediction from Twitter).
- Value: Turning raw data into profit, efficiency, or social impact.
How Big Data Differs from Traditional BI
| Feature | Traditional BI (SQL, OLAP) | Big Data (Hadoop, Spark) |
|---|---|---|
| Data Size | GBs (structured, tabular) | TBs/PBs (structured + unstructured) |
| Processing | Batch (daily/weekly) | Real-time or micro-batch |
| Schema | Fixed (relational) | Flexible (NoSQL: document, key-value) |
| Tools | Excel, Tableau, Power BI | Hadoop, Spark, Kafka, Flink |
| Use Case | Historical reporting | Predictive analytics, fraud detection |
2. Big Data Technologies: The Toolkit
A. Storage: HDFS and NoSQL
Hadoop Distributed File System (HDFS)
- Stores data across thousands of commodity servers (cheaper than enterprise storage).
- Example: Nepal Electricity Authority (NEA) uses HDFS to store smart meter readings from 5 million households.
- Weakness: Slow for real-time queries (latency ~seconds).
NoSQL Databases
- Types and Use Cases:
- Document (MongoDB): Stores JSON-like documents (e.g., Pathao’s ride data with driver, passenger, and location history).
- Key-Value (Redis): Caches session data (e.g., eSewa’s login tokens).
- Column-Family (Cassandra): Time-series data (e.g., NTC’s network traffic per tower).
- Graph (Neo4j): Relationships (e.g., Nepal Police’s crime hotspot mapping).
- Types and Use Cases:
B. Processing: Batch vs. Real-Time
Batch Processing (Hadoop MapReduce)
- How it works: Divides data into chunks, processes in parallel, and aggregates results.
- Example: Daraz’s nightly sales report (aggregates orders from 10M+ users).
- Limitation: Not suitable for real-time decisions (e.g., stock trading).
Real-Time Processing (Spark, Flink)
- Spark: In-memory processing (100x faster than Hadoop for iterative tasks).
- Example: NEPSE’s stock price alerts (analyzes trades in milliseconds).
- Flink: Event-time processing (critical for fraud detection in Khalti).
- Kafka: Distributed streaming platform (e.g., YouTube’s video upload pipeline).
- Spark: In-memory processing (100x faster than Hadoop for iterative tasks).
3. Big Data in Business Intelligence: The Bridge
BI tools (Tableau, Power BI) can’t handle raw big data directly. Instead:
ETL/ELT Pipelines
- Extract: Pull data from sources (e.g., WhatsApp’s message logs, Ncell’s call detail records).
- Load: Store in HDFS or NoSQL.
- Transform: Clean, aggregate, and model (using Spark SQL or Python).
- Example: Himalayan Java’s supply chain uses Spark to predict coffee bean demand from weather data.
Real-Time Dashboards
- Tools: Grafana (for metrics), Kibana (for logs), Tableau (for visualizations).
- Example: NTC’s network dashboard shows real-time congestion in Kathmandu (helps reroute traffic).
sequenceDiagram
participant User as Customer
participant App as Daraz App
participant Kafka as Kafka Stream
participant Spark as Spark Processing
participant DB as NoSQL DB
participant Dashboard as Real-Time Dashboard
User->>App: Places order
App->>Kafka: Sends event (order_id, user_id, items)
Kafka->>Spark: Streams data
Spark->>DB: Updates inventory
Spark->>Dashboard: Pushes metrics
Dashboard->>User: Shows "Order processing"4. Challenges of Big Data + BI
| Challenge | Cause | Solution | Nepali Example |
|---|---|---|---|
| Data Privacy | GDPR, PDPA (Nepal) | Anonymization, encryption | Nabil Bank’s loan data (masked before analytics) |
| Bias in Algorithms | Skewed training data | Fairness-aware ML models | Nepal Police’s crime prediction (adjusts for rural vs. urban bias) |
| Infrastructure Cost | Scaling storage/compute | Cloud (AWS, GCP) or hybrid models | NTC’s IoT sensors (uses edge computing to reduce cloud costs) |
| Data Quality | Noisy/unstructured data | Automated cleaning (Apache NiFi) | eSewa’s transaction logs (flags duplicates) |
5. Case Study: Himalayan Java’s Big Data-Driven Supply Chain
Problem: Coffee beans spoil if not stored at optimal humidity/temperature. Solution:
- IoT Sensors: Deployed in warehouses to track temperature, humidity, and CO₂ levels.
- Big Data Pipeline:
- Data → Kafka (streaming) → Spark (real-time alerts) → MongoDB (historical trends).
- BI Dashboard: Predicts spoilage risk and suggests reordering or temperature adjustments.
- Outcome:
- 20% reduction in waste.
- 15% faster response to supply chain disruptions.
mindmap
root((Himalayan Java: Big Data Supply Chain))
Data Sources
IoT Sensors
Weather APIs
Farmer Reports
Processing
Kafka (Streaming)
Spark (ML Models)
MongoDB (Storage)
BI Tools
Tableau (Dashboards)
Power BI (Alerts)
Outcomes
Waste Reduction
Demand Forecasting
Farmer Payout Optimization6. Big Data in Nepali Companies
| Company | Big Data Use Case | Technology Stack |
|---|---|---|
| Ncell | Churn prediction (customers switching to NTC) | Hadoop, Spark, Python (scikit-learn) |
| Daraz | Real-time inventory management | Kafka, Cassandra, Tableau |
| eSewa | Fraud detection in transactions | Flink, Redis, custom ML models |
| NTC | Network congestion analysis | Elasticsearch, Grafana, IoT sensors |
| Nepal Police | Crime hotspot prediction from social media | NLP (spaCy), MongoDB, Tableau |
7. Exam Tip: How to Score Full Marks
Define Clearly:
- Big data ≠ large data. Always mention 3Vs (or 5Vs).
- Example: "Big data in NEPSE refers to the velocity of stock trades (millions/day) and variety of data sources (news, social media, historical prices)."
Compare Technologies:
- Use tables to contrast Hadoop vs. Spark, SQL vs. NoSQL, or batch vs. real-time.
Link to Nepali Context:
- Ncell: Churn prediction using RFM analysis (Recency, Frequency, Monetary).
- Daraz: A/B testing on product recommendations (big data + BI).
- NTC: Predictive maintenance of cell towers using sensor data.
Diagrams Are Mandatory:
- Draw HDFS architecture, Spark workflow, or ETL pipeline in exams.
- Example question: "Explain how Khalti processes 10,000 transactions/sec using big data tools." → Answer with Kafka → Spark → NoSQL → Dashboard.
Challenges > Features:
- Examiners love real-world problems. Discuss:
- "How would you handle data privacy in Nabil Bank’s loan analytics?" → Anonymization + GDPR compliance.
- "Why can’t Daraz use traditional SQL for real-time inventory?" → Schema flexibility + scalability.
- Examiners love real-world problems. Discuss:
Case Study Format:
- Problem → Technology Used → Outcome.
- Example:
"Nepal Police uses big data from Twitter to predict crime. Problem: Noise in data. Solution: NLP (spaCy) + MongoDB. Outcome: 30% faster deployment of patrols."
Final Note: Big data + BI is about turning chaos into clarity. Whether it’s Ncell’s customer insights, Daraz’s logistics, or NEPSE’s trading alerts, the key is choosing the right tools for the job—and explaining why in your exam.
Based on the TU BITM syllabus for Business Intelligence (IT249), unit 9.
Discussion
Loading…