Big Data and AnalyticsUnit 512 min read
NoSQL Databases: Types, Models, Use Cases & Comparison
Unit 5 of Big Data and Analytics explores NoSQL databases—how they differ from SQL, their four core data models (document, key-value, column-family, graph), real-world applications in Nepalese and global tech, and when to use them over traditional relational databases.
What is NoSQL?
NoSQL (Not Only SQL) databases are non-relational databases designed to handle unstructured, semi-structured, or rapidly changing data at massive scale. Unlike SQL databases (e.g., MySQL, PostgreSQL), they do not enforce rigid schemas and prioritize scalability, flexibility, and performance over strict consistency.
Why NoSQL?
- Schema-less design: Fields can vary per record (e.g., user profiles with optional fields like
addressorphone). - Horizontal scaling: Add more servers (nodes) to distribute load (vs. SQL’s vertical scaling).
- High availability: Built for 24/7 uptime (e.g., social media, IoT sensors).
- Flexible data models: Optimized for specific use cases (e.g., graphs for networks, documents for JSON).
In the real world
- eSewa (Nepal) uses a key-value store (like Redis) to cache frequently accessed user sessions and payment gateways. This reduces latency when millions of users log in during festival seasons (e.g., Dashain, Tihar).
- WhatsApp (Meta) relies on document databases (MongoDB) to store chat messages, media, and user metadata. Each message is a document with fields like
sender_id,timestamp,media_url, andstatus(delivered/read), allowing flexible queries without a fixed schema. - NTC’s smart meters (Nepal) use time-series databases (e.g., InfluxDB) to log electricity consumption every second. Traditional SQL would struggle with billions of rows per day; NoSQL handles this efficiently.
Types of NoSQL Databases
NoSQL databases are categorized into four main models, each optimized for specific data structures and access patterns. Below is a comparison table:
| Model | Data Structure | Use Case | Example Databases | Query Language |
|---|---|---|---|---|
| Document | JSON/XML (nested key-value pairs) | User profiles, catalogs, content mgmt | MongoDB, CouchDB | MongoDB Query Language (MQL) |
| Key-Value | key → value pairs |
Caching, sessions, real-time analytics | Redis, DynamoDB | Custom APIs (e.g., Redis CLI) |
| Column-Family | Columns grouped by "families" | Time-series, analytics, logs | Cassandra, HBase | CQL (Cassandra Query Language) |
| Graph | Nodes + edges (relationships) | Social networks, fraud detection | Neo4j, Amazon Neptune | Cypher (Neo4j) |
1. Document Databases
Store data in JSON/XML-like documents, ideal for hierarchical data (e.g., nested user objects).
How it works:
- Each document is a record with fields (like a row in SQL but flexible).
- No fixed schema: Add/remove fields dynamically.
- Example Document (MongoDB):
{ "_id": "user123", "name": "Rohan Shrestha", "email": "rohan@example.com", "orders": [ { "order_id": "ord456", "amount": 1500, "status": "delivered" }, { "order_id": "ord789", "amount": 800, "status": "pending" } ], "address": { "street": "Thapathali", "city": "Kathmandu" } }
Real-World Example: Daraz’s Order System
Daraz uses MongoDB to store order data. Each order is a document with fields like:
order_id,user_id,items(array of products),status,shipping_address.- Why NoSQL?
- Orders have variable fields (e.g., some have
gift_wrap, others don’t). - Scalability: Daraz handles millions of orders daily; MongoDB’s horizontal scaling helps.
- Flexible queries: Find all orders from a city or with a specific product without altering the schema.
- Orders have variable fields (e.g., some have
2. Key-Value Stores
The simplest NoSQL model: data is stored as key → value pairs, like a hash table.
How it works:
- Keys are unique identifiers (e.g.,
user123). - Values can be strings, numbers, or binary data (e.g., session tokens, cached HTML).
- No querying: Only retrieve by key (e.g.,
GET user123).
Real-World Example: Khalti’s Payment Caching
Khalti uses Redis (a key-value store) to cache:
- User sessions:
key = "session:user123",value = {token: "abc123", expiry: 3600}. - Frequent transactions:
key = "txn:order456",value = {amount: 500, status: "completed"}. - Why?
- Speed: Retrieving a cached session is microsecond-fast (vs. querying a SQL database).
- Scalability: Redis clusters handle millions of concurrent users during festivals.
3. Column-Family Stores
Optimized for large-scale analytics and time-series data. Data is stored in columns (not rows), allowing efficient reads/writes for specific columns.
How it works:
- Rows are grouped by columns (e.g., all
timestampdata together). - Example: Storing sensor data from NTC’s smart meters.
Column Family: "electricity_meters" Columns: - meter_id (partition key) - timestamp (clustering key) - consumption_kWh - is_peak_hour - Query: "Give me all
consumption_kWhformeter_id=123in the last hour."
Real-World Example: NTC’s Smart Grid Analytics
NTC uses Cassandra to store:
- Billions of meter readings per day.
- Why NoSQL?
- Write-heavy: Meters send data every second; Cassandra handles high write throughput.
- Time-based queries: "Show me consumption trends for the last 30 days" is fast because time-series columns are pre-grouped.
4. Graph Databases
Store data as nodes (entities) and edges (relationships), ideal for networked data.
How it works:
- Nodes: Represent entities (e.g., users, products).
- Edges: Represent relationships (e.g., "friends with," "ordered").
- Properties: Key-value pairs on nodes/edges (e.g.,
user.age = 25).
Real-World Example: Pathao’s Ride-Matching
Pathao uses Neo4j to:
- Match riders to drivers in real-time by analyzing:
Locationnodes (withlatitude,longitude).Rider → Driveredges (withdistance,estimated_time).
- Fraud detection: Find suspicious patterns (e.g., a driver accepting too many rides in one area).
- Why NoSQL?
- Relationships matter: Pathao’s core logic is about who is near whom.
- Fast traversals: Querying "Find all drivers within 500m of this rider" is millisecond-fast in Neo4j.
NoSQL vs. SQL: When to Use Which?
Use this decision table to choose between NoSQL and SQL:
| Criteria | NoSQL | SQL (Relational) |
|---|---|---|
| Data Structure | Unstructured/semi-structured (JSON, graphs) | Structured (tables with fixed schemas) |
| Scalability | Horizontal (add more servers) | Vertical (upgrade single server) |
| Query Complexity | Simple key lookups, flexible queries | Complex joins, transactions (ACID) |
| Consistency | Eventual consistency (BASE model) | Strong consistency (ACID) |
| Use Cases | Real-time analytics, IoT, social networks | Banking, ERP, where transactions matter |
| Example Systems | MongoDB, Cassandra, Neo4j | MySQL, PostgreSQL, Oracle |
Advantages and Disadvantages of NoSQL
✅ Advantages
- Flexibility: Add/remove fields without migrating data.
- Scalability: Handle petabytes of data across clusters.
- Performance: Optimized for specific access patterns (e.g., graphs for relationships).
- High Availability: Designed for 99.999% uptime (e.g., social media).
- Cost-Effective: Open-source options (MongoDB, Cassandra) reduce licensing costs.
❌ Disadvantages
- No Standard Query Language: Each database has its own syntax (e.g., MongoDB’s MQL vs. SQL).
- Limited Transactions: Most NoSQL databases lack full ACID support (except MongoDB 4.0+).
- Learning Curve: Requires understanding of distributed systems (e.g., sharding, replication).
- Data Integrity Risks: Schema-less design can lead to inconsistent data if not managed well.
How NoSQL Works Under the Hood
1. Data Distribution (Sharding)
NoSQL databases partition data across servers (shards) to avoid bottlenecks.
- Example: MongoDB splits a
userscollection into shards byuser_idrange.
2. Replication for Fault Tolerance
Copies of data are stored on multiple servers to prevent loss.
- Example: Cassandra replicates data across 3 nodes by default.
3. CAP Theorem Trade-offs
NoSQL databases must choose between:
- Consistency (C): All nodes see the same data.
- Availability (A): System remains operational.
- Partition Tolerance (P): Works despite network failures.
- Example: DynamoDB (Amazon’s NoSQL) prioritizes Availability and Partition Tolerance (AP), sacrificing some consistency.
Worked Example: Designing a NoSQL Database for a Nepalese E-Commerce Site
Scenario: Build a database for NepalBasket.com, an e-commerce site selling handmade products.
Requirements:
- Store user profiles (name, email, address, order history).
- Handle product catalogs (name, price, images, categories).
- Track orders (items, status, shipping).
- Support real-time recommendations (e.g., "Users who bought X also bought Y").
Database Choice and Schema:
| Data Type | NoSQL Model | Database | Schema Example |
|---|---|---|---|
| User Profiles | Document | MongoDB | { _id: "user1", name: "Sita", orders: [...] } |
| Product Catalog | Document | MongoDB | { _id: "prod101", name: "Thapa Paper", price: 500, images: [...] } |
| Orders | Document | MongoDB | { _id: "order456", user_id: "user1", items: [{prod_id: "prod101", qty: 2}], status: "shipped" } |
| Recommendations | Graph | Neo4j | User --[BOUGHT]--> Product --[RECOMMENDED_FOR]--> User |
Why This Design?
- Flexibility: User profiles can have optional fields (e.g.,
phoneonly for some users). - Scalability: MongoDB can scale horizontally as traffic grows.
- Recommendations: Neo4j’s graph model efficiently finds "frequently bought together" patterns.
Exam Tip
Compare NoSQL Models: Be ready to explain when to use document vs. key-value vs. graph databases. Example:
- Question: "Which NoSQL model would you use for a social media app’s friend network?"
- Answer: Graph database (Neo4j) because relationships (friendships) are the core data.
Real-World Applications: Link concepts to Nepalese companies:
- eSewa: Key-value store for sessions.
- Daraz: Document database for orders.
- NTC: Column-family for meter data.
CAP Theorem: Always mention trade-offs. Example:
- Question: "Why does Cassandra prioritize Availability over Consistency?"
- Answer: "Cassandra is designed for high availability (e.g., in data centers). It uses eventual consistency (BASE model) to ensure the system remains operational even if some nodes fail."
Schema Design: Practice designing schemas for given scenarios (e.g., "Design a NoSQL database for a hospital patient records system"). Use document or graph models for hierarchical/social data.
Common Pitfalls: Know the limitations:
- NoSQL lacks joins (use denormalization or application-level joins).
- No transactions in most NoSQL databases (except MongoDB with multi-document ACID).
Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 5.
Discussion
Loading…