Big Data and AnalyticsUnit 512 min read

NoSQL Databases: Types, Models, Use Cases & Comparison

Unit 5 of Big Data and Analytics explores NoSQL databases—how they differ from SQL, their four core data models (document, key-value, column-family, graph), real-world applications in Nepalese and global tech, and when to use them over traditional relational databases.

What is NoSQL?

NoSQL (Not Only SQL) databases are non-relational databases designed to handle unstructured, semi-structured, or rapidly changing data at massive scale. Unlike SQL databases (e.g., MySQL, PostgreSQL), they do not enforce rigid schemas and prioritize scalability, flexibility, and performance over strict consistency.

Why NoSQL?

  • Schema-less design: Fields can vary per record (e.g., user profiles with optional fields like address or phone).
  • Horizontal scaling: Add more servers (nodes) to distribute load (vs. SQL’s vertical scaling).
  • High availability: Built for 24/7 uptime (e.g., social media, IoT sensors).
  • Flexible data models: Optimized for specific use cases (e.g., graphs for networks, documents for JSON).

In the real world

  1. eSewa (Nepal) uses a key-value store (like Redis) to cache frequently accessed user sessions and payment gateways. This reduces latency when millions of users log in during festival seasons (e.g., Dashain, Tihar).
  2. WhatsApp (Meta) relies on document databases (MongoDB) to store chat messages, media, and user metadata. Each message is a document with fields like sender_id, timestamp, media_url, and status (delivered/read), allowing flexible queries without a fixed schema.
  3. NTC’s smart meters (Nepal) use time-series databases (e.g., InfluxDB) to log electricity consumption every second. Traditional SQL would struggle with billions of rows per day; NoSQL handles this efficiently.

Types of NoSQL Databases

NoSQL databases are categorized into four main models, each optimized for specific data structures and access patterns. Below is a comparison table:

Model Data Structure Use Case Example Databases Query Language
Document JSON/XML (nested key-value pairs) User profiles, catalogs, content mgmt MongoDB, CouchDB MongoDB Query Language (MQL)
Key-Value key → value pairs Caching, sessions, real-time analytics Redis, DynamoDB Custom APIs (e.g., Redis CLI)
Column-Family Columns grouped by "families" Time-series, analytics, logs Cassandra, HBase CQL (Cassandra Query Language)
Graph Nodes + edges (relationships) Social networks, fraud detection Neo4j, Amazon Neptune Cypher (Neo4j)

1. Document Databases

Store data in JSON/XML-like documents, ideal for hierarchical data (e.g., nested user objects).

How it works:

  • Each document is a record with fields (like a row in SQL but flexible).
  • No fixed schema: Add/remove fields dynamically.
  • Example Document (MongoDB):
    {
      "_id": "user123",
      "name": "Rohan Shrestha",
      "email": "rohan@example.com",
      "orders": [
        { "order_id": "ord456", "amount": 1500, "status": "delivered" },
        { "order_id": "ord789", "amount": 800, "status": "pending" }
      ],
      "address": {
        "street": "Thapathali",
        "city": "Kathmandu"
      }
    }
    

Real-World Example: Daraz’s Order System

Daraz uses MongoDB to store order data. Each order is a document with fields like:

  • order_id, user_id, items (array of products), status, shipping_address.
  • Why NoSQL?
    • Orders have variable fields (e.g., some have gift_wrap, others don’t).
    • Scalability: Daraz handles millions of orders daily; MongoDB’s horizontal scaling helps.
    • Flexible queries: Find all orders from a city or with a specific product without altering the schema.

2. Key-Value Stores

The simplest NoSQL model: data is stored as key → value pairs, like a hash table.

How it works:

  • Keys are unique identifiers (e.g., user123).
  • Values can be strings, numbers, or binary data (e.g., session tokens, cached HTML).
  • No querying: Only retrieve by key (e.g., GET user123).

Real-World Example: Khalti’s Payment Caching

Khalti uses Redis (a key-value store) to cache:

  • User sessions: key = "session:user123", value = {token: "abc123", expiry: 3600}.
  • Frequent transactions: key = "txn:order456", value = {amount: 500, status: "completed"}.
  • Why?
    • Speed: Retrieving a cached session is microsecond-fast (vs. querying a SQL database).
    • Scalability: Redis clusters handle millions of concurrent users during festivals.

3. Column-Family Stores

Optimized for large-scale analytics and time-series data. Data is stored in columns (not rows), allowing efficient reads/writes for specific columns.

How it works:

  • Rows are grouped by columns (e.g., all timestamp data together).
  • Example: Storing sensor data from NTC’s smart meters.
    Column Family: "electricity_meters"
    Columns:
      - meter_id (partition key)
      - timestamp (clustering key)
      - consumption_kWh
      - is_peak_hour
    
  • Query: "Give me all consumption_kWh for meter_id=123 in the last hour."

Real-World Example: NTC’s Smart Grid Analytics

NTC uses Cassandra to store:

  • Billions of meter readings per day.
  • Why NoSQL?
    • Write-heavy: Meters send data every second; Cassandra handles high write throughput.
    • Time-based queries: "Show me consumption trends for the last 30 days" is fast because time-series columns are pre-grouped.

4. Graph Databases

Store data as nodes (entities) and edges (relationships), ideal for networked data.

How it works:

  • Nodes: Represent entities (e.g., users, products).
  • Edges: Represent relationships (e.g., "friends with," "ordered").
  • Properties: Key-value pairs on nodes/edges (e.g., user.age = 25).

Real-World Example: Pathao’s Ride-Matching

Pathao uses Neo4j to:

  • Match riders to drivers in real-time by analyzing:
    • Location nodes (with latitude, longitude).
    • Rider → Driver edges (with distance, estimated_time).
  • Fraud detection: Find suspicious patterns (e.g., a driver accepting too many rides in one area).
  • Why NoSQL?
    • Relationships matter: Pathao’s core logic is about who is near whom.
    • Fast traversals: Querying "Find all drivers within 500m of this rider" is millisecond-fast in Neo4j.

NoSQL vs. SQL: When to Use Which?

Use this decision table to choose between NoSQL and SQL:

Unstructured Data (e.g., JSON) (30%)Semi-Structured Data (e.g., XML) (25%)High Write Throughput (20%)Horizontal Scalability (25%)
Typical NoSQL use cases (approximate percentages)
Criteria NoSQL SQL (Relational)
Data Structure Unstructured/semi-structured (JSON, graphs) Structured (tables with fixed schemas)
Scalability Horizontal (add more servers) Vertical (upgrade single server)
Query Complexity Simple key lookups, flexible queries Complex joins, transactions (ACID)
Consistency Eventual consistency (BASE model) Strong consistency (ACID)
Use Cases Real-time analytics, IoT, social networks Banking, ERP, where transactions matter
Example Systems MongoDB, Cassandra, Neo4j MySQL, PostgreSQL, Oracle

Advantages and Disadvantages of NoSQL

✅ Advantages

  1. Flexibility: Add/remove fields without migrating data.
  2. Scalability: Handle petabytes of data across clusters.
  3. Performance: Optimized for specific access patterns (e.g., graphs for relationships).
  4. High Availability: Designed for 99.999% uptime (e.g., social media).
  5. Cost-Effective: Open-source options (MongoDB, Cassandra) reduce licensing costs.

❌ Disadvantages

  1. No Standard Query Language: Each database has its own syntax (e.g., MongoDB’s MQL vs. SQL).
  2. Limited Transactions: Most NoSQL databases lack full ACID support (except MongoDB 4.0+).
  3. Learning Curve: Requires understanding of distributed systems (e.g., sharding, replication).
  4. Data Integrity Risks: Schema-less design can lead to inconsistent data if not managed well.

How NoSQL Works Under the Hood

1. Data Distribution (Sharding)

NoSQL databases partition data across servers (shards) to avoid bottlenecks.

  • Example: MongoDB splits a users collection into shards by user_id range.
Server 1Shard 1: user_id 1-1,000,000Server 2Shard 2: user_id 1,000,001-2,000,000Server 3Shard 3: user_id 2,000,001-3,000,000MongoDB Cluster
Hierarchical sharding in MongoDB (range-based partitioning)

2. Replication for Fault Tolerance

Copies of data are stored on multiple servers to prevent loss.

  • Example: Cassandra replicates data across 3 nodes by default.
Data Center 1Replica 1 (Node 2)Data Center 2Replica 2 (Node 3)Primary Node (Data Center 1)
Cassandra's multi-data-center replication (3-node setup)

3. CAP Theorem Trade-offs

NoSQL databases must choose between:

  • Consistency (C): All nodes see the same data.
  • Availability (A): System remains operational.
  • Partition Tolerance (P): Works despite network failures.
  • Example: DynamoDB (Amazon’s NoSQL) prioritizes Availability and Partition Tolerance (AP), sacrificing some consistency.

Worked Example: Designing a NoSQL Database for a Nepalese E-Commerce Site

Scenario: Build a database for NepalBasket.com, an e-commerce site selling handmade products.

2023Product catalog(Document DB)2023User sessions(Key-Value Cache)2024Recommendationengine (Graph DB)
Phased NoSQL implementation timeline for Daraz-like system

Requirements:

  1. Store user profiles (name, email, address, order history).
  2. Handle product catalogs (name, price, images, categories).
  3. Track orders (items, status, shipping).
  4. Support real-time recommendations (e.g., "Users who bought X also bought Y").

Database Choice and Schema:

Data Type NoSQL Model Database Schema Example
User Profiles Document MongoDB { _id: "user1", name: "Sita", orders: [...] }
Product Catalog Document MongoDB { _id: "prod101", name: "Thapa Paper", price: 500, images: [...] }
Orders Document MongoDB { _id: "order456", user_id: "user1", items: [{prod_id: "prod101", qty: 2}], status: "shipped" }
Recommendations Graph Neo4j User --[BOUGHT]--> Product --[RECOMMENDED_FOR]--> User

Why This Design?

  • Flexibility: User profiles can have optional fields (e.g., phone only for some users).
  • Scalability: MongoDB can scale horizontally as traffic grows.
  • Recommendations: Neo4j’s graph model efficiently finds "frequently bought together" patterns.

Exam Tip

  1. Compare NoSQL Models: Be ready to explain when to use document vs. key-value vs. graph databases. Example:

    • Question: "Which NoSQL model would you use for a social media app’s friend network?"
    • Answer: Graph database (Neo4j) because relationships (friendships) are the core data.
  2. Real-World Applications: Link concepts to Nepalese companies:

    • eSewa: Key-value store for sessions.
    • Daraz: Document database for orders.
    • NTC: Column-family for meter data.
  3. CAP Theorem: Always mention trade-offs. Example:

    • Question: "Why does Cassandra prioritize Availability over Consistency?"
    • Answer: "Cassandra is designed for high availability (e.g., in data centers). It uses eventual consistency (BASE model) to ensure the system remains operational even if some nodes fail."
  4. Schema Design: Practice designing schemas for given scenarios (e.g., "Design a NoSQL database for a hospital patient records system"). Use document or graph models for hierarchical/social data.

  5. Common Pitfalls: Know the limitations:

    • NoSQL lacks joins (use denormalization or application-level joins).
    • No transactions in most NoSQL databases (except MongoDB with multi-document ACID).

Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 5.

Discussion

Loading…