Big Data and AnalyticsUnit 29 min read

Big Data Architecture: Layers, Models, and Real-World Systems

Unit 2 of Big Data and Analytics explores the foundational architecture of big data systems, covering the Lambda architecture, Kappa architecture, data lakes vs. data warehouses, distributed storage, and processing layers (batch vs. real-time). It explains how companies like Google, Facebook, and Ncell design scalable

Core Concepts of Big Data Architecture

Big data architecture refers to the frameworks, tools, and layers designed to store, process, and analyze massive datasets efficiently. Unlike traditional databases, big data systems prioritize scalability, fault tolerance, and distributed processing. The architecture typically consists of:

  1. Data Ingestion Layer (collecting raw data from sources).
  2. Storage Layer (distributed storage like HDFS, S3, or data lakes).
  3. Processing Layer (batch processing with MapReduce/Spark or real-time with Kafka/Flume).
  4. Serving Layer (serving processed data to applications).
  5. Governance Layer (security, metadata management, and compliance).

1. Lambda vs. Kappa Architecture: Two Approaches to Big Data Processing

Big data systems use two dominant architectures: Lambda (hybrid batch + real-time) and Kappa (real-time only). The choice depends on latency requirements, cost, and use case.

Lambda Architecture (Batch + Real-Time)

Lambda architecture processes data in two parallel pipelines:

  • Batch Layer: Handles historical data using MapReduce/Spark (e.g., daily aggregations).
  • Speed Layer: Processes real-time data with streaming (e.g., Kafka + Storm).
  • Serving Layer: Merges results from both layers for queries.
graph LR
    A["Data Sources"] --> B["Batch Layer\n(MapReduce/Spark)"]
    A --> C["Speed Layer\n(Kafka + Storm)"]
    B --> D["Serving Layer\n(Merged Results)"]
    C --> D
    D --> E["Applications\n(eSewa Fraud Detection)"]

Example in Nepal:

  • eSewa uses Lambda architecture to detect fraudulent transactions in real-time (Speed Layer) while maintaining a historical record (Batch Layer) for audits.

Kappa Architecture (Real-Time Only)

Kappa simplifies Lambda by using only streaming (e.g., Kafka + Spark Streaming). All processing happens in real-time, reducing complexity.

Applications (Daraz Recommendations)Serving Layer (Real-Time Results)Stream Processing (Kafka + Spark Streaming)Data SourcesKappa Architecture (Real-Time Only)
Simplified real-time data flow in Kappa Architecture

Advantages of Kappa:

  • Simpler to maintain (no batch layer).
  • Lower latency (real-time only). Disadvantages:
  • Harder to reprocess historical data.
  • Requires high throughput for real-time.

Comparison Table: Lambda vs. Kappa

Feature Lambda Architecture Kappa Architecture
Processing Model Batch + Real-Time Real-Time Only
Complexity High (two pipelines) Low (single pipeline)
Latency Medium (batch delay) Low (real-time)
Use Case Fraud detection, analytics Real-time dashboards, IoT
Example in Nepal eSewa, Ncell billing Pathao ride analytics

2. Data Storage: Lakes vs. Warehouses

Big data storage is categorized into data lakes (raw, unstructured) and data warehouses (structured, optimized for queries).

Data Lake

  • Stores raw data in its native format (JSON, logs, images).
  • Uses HDFS, S3, or Delta Lake.
  • Example: Ncell stores raw call detail records (CDRs) in a data lake before processing.
flowchart TD
    A[Raw Data
    (JSON, Logs, Images)] --> B[Data Lake
    (HDFS/S3/Delta Lake)]
    B --> C[Processing
    (Spark/MapReduce)]
    C --> D[Analytics
    (Customer Churn Prediction)]
    D --> E["Ncell CDR Example"]
Data Lake workflow with Nepali example (Ncell CDRs)

Advantages:

  • Flexible schema (no upfront structure).
  • Cost-effective for unstructured data. Disadvantages:
  • Risk of "data swamp" (poor metadata).
  • Slower queries without optimization.

Data Warehouse

  • Stores structured, processed data (SQL tables).
  • Optimized for OLAP queries (e.g., sales reports).
  • Example: Daraz uses a data warehouse for product recommendations.
Example: NEPSE Stock AnalysisOLAP Cubes (Star Schema)Structured DataBatch Loading (Nightly)ETL ProcessesData Warehouse
Data Warehouse architecture with Nepali financial example

Comparison Table: Data Lake vs. Warehouse

Feature Data Lake Data Warehouse
Data Type Raw, unstructured Structured, processed
Storage Format HDFS, S3, Delta Lake SQL tables, Parquet
Query Speed Slower (unless optimized) Faster (indexed)
Use Case IoT, logs, multimedia Reporting, BI dashboards
Example in Nepal Ncell CDRs NEPSE stock analysis

3. Distributed Storage Systems

Big data relies on distributed storage to handle petabytes of data. Key systems:

  • HDFS (Hadoop Distributed File System): Stores data across clusters (used by Hadoop).
  • S3 (Amazon Simple Storage): Cloud-based, scalable storage.
  • Cassandra/ScyllaDB: NoSQL databases with distributed storage.

Example:

  • NTC (Nepal Telecom) uses HDFS to store call logs across multiple servers for network analytics.
00.751.52.253HDFS3S31Cassandra/ScyllaDB1
Nepal Telecom's distributed storage usage (HDFS for call logs, others for analytics)

Advantages of Distributed Storage:

  • Scalability: Add more nodes as data grows.
  • Fault Tolerance: Data replicated across nodes.
  • Cost-Effective: Cheaper than traditional databases.

4. Processing Layers: Batch vs. Real-Time

Big data processing happens in two modes:

  1. Batch Processing (e.g., MapReduce, Spark Batch):
    • Processes large datasets offline (e.g., nightly reports).
    • Example: NEPSE calculates daily stock market trends using batch processing.
  2. Real-Time Processing (e.g., Spark Streaming, Flink):
    • Processes data as it arrives (e.g., fraud detection).
    • Example: Khalti detects unusual transactions in milliseconds.

Comparison Table: Batch vs. Real-Time

Feature Batch Processing Real-Time Processing
Latency High (hours/days) Low (milliseconds)
Use Case Monthly reports, analytics Fraud detection, live dashboards
Tools MapReduce, Spark Batch Spark Streaming, Flink, Kafka
Example in Nepal NEPSE end-of-day analysis eSewa transaction monitoring

5. Governance and Security in Big Data

Big data architectures must ensure:

  • Data Privacy: Compliance with PDPA (Nepal’s data protection law).
  • Access Control: Role-based permissions (e.g., only analysts can query sales data).
  • Metadata Management: Tracking data lineage (who processed what and when).
2078 BSNepal Rastra Bankimplements GDPR-like d2080 BSeSewa introducesend-to-end encryption 2081 BSNTC deploysrole-based access cont
Key Nepali data governance milestones

Example:

  • Nepal Rastra Bank (NRB) enforces strict access controls on financial transaction data stored in big data systems.

In the Real World

  1. eSewa’s Fraud Detection

    • Uses Lambda architecture to combine real-time transaction monitoring (Speed Layer) with historical fraud patterns (Batch Layer).
    • How it works: If a user suddenly transfers ₹50,000 to an unknown account, the Speed Layer flags it instantly, while the Batch Layer checks if the user’s behavior matches past fraud cases.
  2. Daraz’s Recommendation Engine

    • Uses Kappa architecture with Spark Streaming to update product recommendations in real-time as users browse.
    • Example: If you view shoes repeatedly, Daraz’s system instantly pushes shoe deals to your feed—no batch delay.
  3. Ncell’s Network Analytics

    • Stores raw call logs in a data lake (HDFS) and processes them with Spark to predict network congestion.
    • Real-world impact: Helps Ncell optimize tower placements in Kathmandu’s busy areas.

Exam Tip

For TU/PU exams, focus on:

  1. Differentiating Lambda vs. Kappa: Know when to use each (e.g., Lambda for fraud detection, Kappa for real-time dashboards).
  2. Data Lake vs. Warehouse: Memorize the trade-offs (flexibility vs. query speed).
  3. Distributed Storage: Explain HDFS replication and why it’s fault-tolerant.
  4. Real-World Examples: Be ready to link concepts to Nepalese companies (e.g., eSewa + Lambda, NEPSE + batch processing).
  5. Diagrams: Draw the Lambda/Kappa architecture flow and HDFS block storage in exams—visuals score extra marks!

Based on the TU BIM syllabus for Big Data and Analytics (IT278), unit 2.

Discussion

Loading…