Big Data and AnalyticsUnit 27 min read
Big Data Architecture: Layers, Models, and Components
Unit 2 of Big Data and Analytics explores the foundational architecture of big data systems, covering the Lambda architecture, Kappa architecture, data lakes, data warehouses, and distributed storage/compute frameworks. It explains how these components interact to process, store, and analyze massive datasets efficientl
Key Concepts and Architectural Models
1. Lambda vs. Kappa Architecture: A Trade-off Between Batch and Stream Processing
Big data architectures are broadly classified into two models:
- Lambda Architecture: Combines batch processing (for historical data) and speed layers (for real-time data) to ensure accuracy and low latency.
- Kappa Architecture: Relies solely on stream processing (e.g., Apache Kafka + Spark Streaming) for real-time analytics, simplifying the system but requiring fault tolerance.
classDiagram
class LambdaArchitecture {
+Batch Layer (Hadoop MapReduce)
+Speed Layer (Storm/Spark Streaming)
+Serving Layer (Pre-computed views)
}
class KappaArchitecture {
+Stream Processing (Kafka + Spark)
+No Batch Layer
}
LambdaArchitecture --> KappaArchitecture : "Evolves into Kappa by removing batch layer"Why the choice matters:
- Lambda is used where historical accuracy is critical (e.g., fraud detection in banks like NMB).
- Kappa is preferred for real-time dashboards (e.g., Pathao’s dynamic pricing).
2. Data Lakes vs. Data Warehouses: Storage Strategies
| Feature | Data Lake | Data Warehouse |
|---|---|---|
| Structure | Schema-on-read (raw data stored as-is) | Schema-on-write (structured data) |
| Use Case | Exploratory analysis, AI/ML training | Reporting, BI dashboards |
| Tools | HDFS, S3, Delta Lake | Snowflake, Redshift, Google BigQuery |
| Example in Nepal | NTC’s raw call detail records (CDRs) | NEPSE’s structured stock market data |
Worked Example: NTC’s CDR Analysis NTC stores billions of call logs in a data lake (HDFS) for:
- Fraud detection (real-time stream processing with Spark).
- Network optimization (batch analysis of historical trends). Without a data lake, NTC would need to pre-process data into a warehouse, slowing down real-time alerts.
3. Distributed Storage and Compute: HDFS and YARN
Big data systems distribute workloads across clusters using:
- HDFS (Hadoop Distributed File System): Stores data across commodity servers with replication (default: 3 copies) for fault tolerance.
- YARN (Yet Another Resource Negotiator): Manages compute resources (CPU/memory) for applications like MapReduce or Spark.
graph TD
A["HDFS"] -->|"Stores Data"| B["DataNodes"]
A -->|"Manages Metadata"| C["NameNode"]
D["YARN"] -->|"Allocates Resources"| E["ResourceManager"]
E -->|"Schedules Tasks"| F["NodeManagers"]
F -->|"Executes Jobs"| G["MapReduce/Spark"]Real-World Tie-In: Daraz’s Inventory System Daraz uses HDFS + YARN to:
- Store product catalogs (100M+ items) across distributed nodes.
- Run real-time inventory updates via Spark jobs triggered by order queues (e.g., a user buying a phone in Kathmandu updates stock in Pokhara instantly).
4. Data Ingestion Layers: Batch vs. Stream
| Ingestion Type | Tools | Use Case | Example |
|---|---|---|---|
| Batch | Apache NiFi, Sqoop | Daily reports, ETL pipelines | NEPSE’s end-of-day stock data |
| Stream | Kafka, Flume | Real-time alerts, IoT sensor data | Pathao’s ride request processing |
Worked Example: Khalti’s Transaction Processing Khalti processes 50,000+ transactions/minute using:
- Kafka to ingest payment requests as streams.
- Spark Streaming to validate transactions in real-time.
- HDFS to store raw transaction logs for audits.
5. Serving Layer: Pre-Computed Views and APIs
The serving layer provides low-latency access to processed data via:
- Pre-computed views (e.g., "top 10 trending products" on Daraz).
- REST APIs (e.g., Ncell’s customer balance check).
How pre-aggregated data is served to users. (Image: Textractor, CC BY-SA 4.0, via Wikimedia Commons)
Real-World Example: eSewa’s Loan Approval eSewa uses a Lambda architecture to:
- Batch layer: Analyze historical loan data (stored in HDFS) to train a risk model (weekly).
- Speed layer: Use Spark to approve/reject loans in <2 seconds based on real-time credit scores.
In the Real World
Pathao’s Dynamic Pricing
- Idea Used: Kappa Architecture (stream processing with Kafka + Spark).
- How: Pathao ingests ride requests/second from drivers and passengers. Spark Streaming adjusts surge pricing in real-time based on demand (e.g., +50% during Kathmandu traffic jams). No batch layer is needed because pricing depends only on live data.
NTC’s Network Optimization
- Idea Used: Data Lake + Batch Processing.
- How: NTC stores raw CDR data (unstructured) in HDFS. Hadoop jobs run nightly to:
- Identify fraudulent call patterns (e.g., SIM boxes).
- Optimize cell tower placements in rural areas (e.g., using historical call density maps).
NEPSE’s Stock Market Analytics
- Idea Used: Data Warehouse + Batch ETL.
- How: NEPSE’s structured market data (prices, volumes) is loaded into a data warehouse (e.g., Snowflake) via Sqoop. Analysts run SQL queries to generate:
- Daily trading reports.
- Predictive models for stock trends (using historical batch data).
Exam Tip
Architecture Diagrams: Always draw Lambda vs. Kappa and HDFS/YARN in exams. Label:
- Batch/Speed layers in Lambda.
- Kafka + Spark in Kappa.
- NameNode/DataNode in HDFS.
Real-World Mapping: Link concepts to Nepalese examples:
- Data Lake → NTC/CDRs.
- Stream Processing → Pathao/Khalti.
- Batch Processing → NEPSE reports.
Short Answer Tricks:
- HDFS = "Distributed storage with replication."
- YARN = "Resource manager for Hadoop."
- Lambda = "Batch + Speed layers."
- Kappa = "Only stream processing."
Common Pitfalls:
- Don’t confuse data lakes (raw) with data warehouses (structured).
- Kappa is not a subset of Lambda—it’s an alternative.
- HDFS is for storage; YARN is for compute.
Visual Summary Table:
| Component | Purpose | Nepalese Example |
|---|---|---|
| HDFS | Distributed storage | NTC’s CDR data lake |
| YARN | Resource management | Daraz’s Spark jobs |
| Kafka | Stream ingestion | Khalti’s transactions |
| Lambda Architecture | Batch + Real-time | eSewa’s loan approval |
| Data Warehouse | Structured analytics | NEPSE’s stock data |
Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 2.
Discussion
Loading…