Big Data and AnalyticsUnit 27 min read

Big Data Architecture: Layers, Models, and Components

Unit 2 of Big Data and Analytics explores the foundational architecture of big data systems, covering the Lambda architecture, Kappa architecture, data lakes, data warehouses, and distributed storage/compute frameworks. It explains how these components interact to process, store, and analyze massive datasets efficientl

Key Concepts and Architectural Models

1. Lambda vs. Kappa Architecture: A Trade-off Between Batch and Stream Processing

Big data architectures are broadly classified into two models:

  • Lambda Architecture: Combines batch processing (for historical data) and speed layers (for real-time data) to ensure accuracy and low latency.
  • Kappa Architecture: Relies solely on stream processing (e.g., Apache Kafka + Spark Streaming) for real-time analytics, simplifying the system but requiring fault tolerance.
classDiagram
    class LambdaArchitecture {
        +Batch Layer (Hadoop MapReduce)
        +Speed Layer (Storm/Spark Streaming)
        +Serving Layer (Pre-computed views)
    }
    class KappaArchitecture {
        +Stream Processing (Kafka + Spark)
        +No Batch Layer
    }
    LambdaArchitecture --> KappaArchitecture : "Evolves into Kappa by removing batch layer"

Why the choice matters:

  • Lambda is used where historical accuracy is critical (e.g., fraud detection in banks like NMB).
  • Kappa is preferred for real-time dashboards (e.g., Pathao’s dynamic pricing).

2. Data Lakes vs. Data Warehouses: Storage Strategies

Feature Data Lake Data Warehouse
Structure Schema-on-read (raw data stored as-is) Schema-on-write (structured data)
Use Case Exploratory analysis, AI/ML training Reporting, BI dashboards
Tools HDFS, S3, Delta Lake Snowflake, Redshift, Google BigQuery
Example in Nepal NTC’s raw call detail records (CDRs) NEPSE’s structured stock market data

Worked Example: NTC’s CDR Analysis NTC stores billions of call logs in a data lake (HDFS) for:

  1. Fraud detection (real-time stream processing with Spark).
  2. Network optimization (batch analysis of historical trends). Without a data lake, NTC would need to pre-process data into a warehouse, slowing down real-time alerts.

3. Distributed Storage and Compute: HDFS and YARN

Big data systems distribute workloads across clusters using:

  • HDFS (Hadoop Distributed File System): Stores data across commodity servers with replication (default: 3 copies) for fault tolerance.
  • YARN (Yet Another Resource Negotiator): Manages compute resources (CPU/memory) for applications like MapReduce or Spark.
graph TD
    A["HDFS"] -->|"Stores Data"| B["DataNodes"]
    A -->|"Manages Metadata"| C["NameNode"]
    D["YARN"] -->|"Allocates Resources"| E["ResourceManager"]
    E -->|"Schedules Tasks"| F["NodeManagers"]
    F -->|"Executes Jobs"| G["MapReduce/Spark"]

Real-World Tie-In: Daraz’s Inventory System Daraz uses HDFS + YARN to:

  • Store product catalogs (100M+ items) across distributed nodes.
  • Run real-time inventory updates via Spark jobs triggered by order queues (e.g., a user buying a phone in Kathmandu updates stock in Pokhara instantly).

4. Data Ingestion Layers: Batch vs. Stream

Ingestion Type Tools Use Case Example
Batch Apache NiFi, Sqoop Daily reports, ETL pipelines NEPSE’s end-of-day stock data
Stream Kafka, Flume Real-time alerts, IoT sensor data Pathao’s ride request processing

Worked Example: Khalti’s Transaction Processing Khalti processes 50,000+ transactions/minute using:

  1. Kafka to ingest payment requests as streams.
  2. Spark Streaming to validate transactions in real-time.
  3. HDFS to store raw transaction logs for audits.

5. Serving Layer: Pre-Computed Views and APIs

The serving layer provides low-latency access to processed data via:

  • Pre-computed views (e.g., "top 10 trending products" on Daraz).
  • REST APIs (e.g., Ncell’s customer balance check).

lambda architecture serving layer diagram**How pre-aggregated data is served to users. (Image: Textractor, CC BY-SA 4.0, via Wikimedia Commons)

Real-World Example: eSewa’s Loan Approval eSewa uses a Lambda architecture to:

  1. Batch layer: Analyze historical loan data (stored in HDFS) to train a risk model (weekly).
  2. Speed layer: Use Spark to approve/reject loans in <2 seconds based on real-time credit scores.

In the Real World

  1. Pathao’s Dynamic Pricing

    • Idea Used: Kappa Architecture (stream processing with Kafka + Spark).
    • How: Pathao ingests ride requests/second from drivers and passengers. Spark Streaming adjusts surge pricing in real-time based on demand (e.g., +50% during Kathmandu traffic jams). No batch layer is needed because pricing depends only on live data.
  2. NTC’s Network Optimization

    • Idea Used: Data Lake + Batch Processing.
    • How: NTC stores raw CDR data (unstructured) in HDFS. Hadoop jobs run nightly to:
      • Identify fraudulent call patterns (e.g., SIM boxes).
      • Optimize cell tower placements in rural areas (e.g., using historical call density maps).
  3. NEPSE’s Stock Market Analytics

    • Idea Used: Data Warehouse + Batch ETL.
    • How: NEPSE’s structured market data (prices, volumes) is loaded into a data warehouse (e.g., Snowflake) via Sqoop. Analysts run SQL queries to generate:
      • Daily trading reports.
      • Predictive models for stock trends (using historical batch data).

Exam Tip

  1. Architecture Diagrams: Always draw Lambda vs. Kappa and HDFS/YARN in exams. Label:

    • Batch/Speed layers in Lambda.
    • Kafka + Spark in Kappa.
    • NameNode/DataNode in HDFS.
  2. Real-World Mapping: Link concepts to Nepalese examples:

    • Data Lake → NTC/CDRs.
    • Stream Processing → Pathao/Khalti.
    • Batch Processing → NEPSE reports.
  3. Short Answer Tricks:

    • HDFS = "Distributed storage with replication."
    • YARN = "Resource manager for Hadoop."
    • Lambda = "Batch + Speed layers."
    • Kappa = "Only stream processing."
  4. Common Pitfalls:

    • Don’t confuse data lakes (raw) with data warehouses (structured).
    • Kappa is not a subset of Lambda—it’s an alternative.
    • HDFS is for storage; YARN is for compute.

Visual Summary Table:

Component Purpose Nepalese Example
HDFS Distributed storage NTC’s CDR data lake
YARN Resource management Daraz’s Spark jobs
Kafka Stream ingestion Khalti’s transactions
Lambda Architecture Batch + Real-time eSewa’s loan approval
Data Warehouse Structured analytics NEPSE’s stock data

Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 2.

Discussion

Loading…