Big Data and AnalyticsUnit 512 min read
NoSQL Databases: Types, Models, and Use Cases
Unit 5 of Big Data and Analytics explores NoSQL databases—how they differ from SQL, their four core data models (document, key-value, column-family, graph), real-world applications in Nepal (e.g., eSewa’s transaction logs, Daraz’s inventory scaling), and trade-offs in scalability, flexibility, and consistency. Includes
What is NoSQL?
NoSQL (Not Only SQL) databases are non-relational data stores designed to handle unstructured, semi-structured, or rapidly changing data at scale. Unlike traditional SQL databases (e.g., MySQL, PostgreSQL), NoSQL prioritizes horizontal scalability, flexible schemas, and high performance for distributed systems.
Key Characteristics
- Schema-less: Fields can vary per record (e.g., one user may have an
address, another may not). - Horizontal Scaling: Add more servers to handle growth (vs. vertical scaling in SQL).
- Distributed by Design: Built for cloud/deployments across multiple nodes.
- Diverse Data Models: No single "right" way to structure data.
Why NoSQL? Real-World Needs in Nepal
NoSQL solves problems SQL can’t:
- eSewa’s Transaction Logs: Millions of daily payments require high write throughput and low latency—NoSQL’s key-value stores (e.g., Redis) cache user sessions.
- Daraz’s Inventory: Product catalogs with variable attributes (e.g., some items have
size, others don’t) fit NoSQL’s document model (MongoDB). - Pathao’s Driver Locations: Real-time geospatial queries (e.g., "find nearest driver") use graph databases (Neo4j) to model driver-user routes.
NoSQL Data Models: Which One to Use?
NoSQL offers four primary models, each optimized for specific use cases. Below is a comparison table and visual breakdown.
Comparison Table
| Model | Structure | Best For | Example Databases | Nepali Use Case |
|---|---|---|---|---|
| Document | JSON/XML stored as records | Hierarchical data (e.g., user profiles) | MongoDB, CouchDB | eSewa user transaction histories |
| Key-Value | key → value pairs |
Caching, session storage | Redis, DynamoDB | Khalti’s fraud detection cache |
| Column-Family | Columns grouped by access | Large-scale analytics (e.g., logs) | Cassandra, HBase | NTC’s network traffic monitoring |
| Graph | Nodes + relationships | Connections (e.g., social networks) | Neo4j, Amazon Neptune | Ncell’s call-detail records |
1. Document Stores
How it works: Data stored as JSON/XML documents (e.g., a user record with nested fields like address.city). Ideal for hierarchical or semi-structured data.
graph LR
A["User Document"] --> B["_id: 123"]
A --> C["name: 'Ramesh'"]
A --> D["address: {city: 'Kathmandu', zip: '44600'}"]
A --> E["orders: [{product: 'laptop', price: 50000}]"]Example: Storing a Daraz order with variable fields (e.g., some orders have shipping_address, others don’t).
{
"order_id": "DARAZ-2024-001",
"user_id": "USER-456",
"items": [
{"product": "iPhone", "quantity": 1},
{"product": "charger", "quantity": 2}
],
"shipping": { "address": "Lalitpur", "tracked": true }
}
Advantages:
- Flexible schema (add fields without migration).
- Fast queries on nested data (e.g.,
find users in Kathmandu with orders > 50000).
Disadvantages:
- No native joins (must denormalize data).
- Can bloat storage if documents grow large.
2. Key-Value Stores
How it works: Simplest model—data stored as key → value pairs (e.g., user:123 → {name: "Sita", email: "sita@example.com"}). Optimized for speed and simplicity.
Example: Khalti’s fraud detection uses Redis to cache:
key:transaction:TXN-789value:{status: "pending", risk_score: 0.85, flags: ["high_value"]}
Advantages:
- Blazing fast reads/writes (microsecond latency).
- No schema management.
Disadvantages:
- Limited query flexibility (only key-based lookups).
- No support for complex relationships.
3. Column-Family Stores
How it works: Data organized by columns (not rows), with each column stored separately. Ideal for analytics and time-series data.
Example: NTC’s network traffic monitoring uses Cassandra to store:
- Rows:
sensor_id=NT-001, date=2024-05-01 - Columns:
time=09:00, value=1200Mbps,time=09:01, value=1500Mbps
Advantages:
- Efficient for read-heavy workloads (e.g., analytics).
- Scalable for large datasets (petabytes).
Disadvantages:
- Complex queries require careful schema design.
- Joins are expensive (avoid them).
4. Graph Databases
How it works: Data stored as nodes (entities) and edges (relationships). Perfect for connected data (e.g., social networks, fraud rings).
graph TD
A["Driver: Pathao-123"] -->|"rides"| B["User: USER-456"]
A -->|"rides"| C["User: USER-789"]
B -->|"orders"| D["Restaurant: FoodMandi"]
C -->|"orders"| DExample: Pathao’s driver-location system uses Neo4j to model:
- Nodes:
Driver,User,Restaurant - Edges:
rides(between Driver and User),delivers(to Restaurant)
Query Example:
MATCH (d:Driver)-[:rides]->(u:User)
WHERE u.location = "Kathmandu"
RETURN d.id, d.current_location
LIMIT 5
Advantages:
- Fast traversals (e.g., "find all drivers near a user").
- Natural for connected data.
Disadvantages:
- Overkill for simple CRUD apps.
- Harder to scale than document/key-value stores.
NoSQL vs. SQL: When to Choose Which?
Use this decision flowchart to pick the right database:
Worked Example: eSewa’s Payment System
Problem: eSewa processes 100,000+ transactions/day with variable data (e.g., some payments have recipient_address, others don’t).
Solution:
- Database: MongoDB (document store).
- Schema:
{ "transaction_id": "TXN-2024-001", "amount": 500, "sender": "USER-123", "recipient": "USER-456", "status": "completed", "metadata": { "note": "Rent", "address": "Lalitpur" } // Optional field }
Why Not SQL?
- SQL would require a rigid schema (e.g.,
metadatatable with nullable columns). - NoSQL’s flexibility handles ad-hoc fields without migrations.
In the Real World
eSewa (Nepal):
- Use Case: Transaction logs and user sessions.
- NoSQL Model: Key-Value (Redis) for caching session data (e.g.,
user:123 → {token: "abc123", expires: 2024-12-31}). - Why? Microsecond latency for session validation during payments.
Daraz (Nepal):
- Use Case: Product catalog with variable attributes.
- NoSQL Model: Document (MongoDB) stores products as JSON:
{ "_id": "PROD-789", "name": "Smartphone", "price": 45000, "specs": { "ram": "8GB", "storage": "128GB" }, // Optional "reviews": [{ "user": "USER-101", "rating": 5 }] } - Why? Avoids SQL’s rigid schema for products with optional fields (e.g., some have
specs, others don’t).
Pathao (Nepal):
- Use Case: Real-time driver-user matching.
- NoSQL Model: Graph (Neo4j) models:
- Nodes:
Driver,User,Restaurant - Edges:
rides(between Driver and User),delivers(to Restaurant)
- Nodes:
- Query: Find the nearest driver to a user’s location in milliseconds.
- Why? SQL would require expensive joins; graph databases excel at traversals.
NTC (Nepal):
- Use Case: Network traffic analytics.
- NoSQL Model: Column-Family (Cassandra) stores time-series data:
- Partition Key:
sensor_id + date - Columns:
time=09:00, value=1200Mbps
- Partition Key:
- Why? Handles petabytes of log data with low-latency reads.
Global Example: WhatsApp (Meta):
- Use Case: User messages and media.
- NoSQL Model: Document (MongoDB) stores conversations as JSON:
{ "conversation_id": "CONV-123", "participants": ["USER-1", "USER-2"], "messages": [ { "sender": "USER-1", "text": "Hi!", "timestamp": "2024-05-01T10:00" }, { "sender": "USER-2", "media": "photo.jpg", "timestamp": "2024-05-01T10:01" } ] } - Why? Flexible schema for messages with text, images, or voice notes.
Advantages and Disadvantages of NoSQL
| Advantage | Disadvantage | Mitigation |
|---|---|---|
| Scalability: Add nodes horizontally | No ACID: Eventual consistency | Use multi-document transactions (MongoDB) |
| Flexible Schema: Add fields without migration | Limited Querying: No SQL joins | Denormalize data or use graph traversals |
| High Performance: Optimized for reads/writes | Learning Curve: New tools/models | Start with document/key-value stores |
| Cost-Effective: Open-source options (e.g., Cassandra) | Data Duplication: Denormalization needed | Design schemas to minimize redundancy |
Common NoSQL Databases and Their Use Cases
| Database | Model | Open-Source? | Nepali Use Case | Global Use Case |
|---|---|---|---|---|
| MongoDB | Document | Yes | eSewa user profiles | Adobe Creative Cloud (user data) |
| Redis | Key-Value | Yes | Khalti fraud detection cache | Twitter (real-time analytics) |
| Cassandra | Column-Family | Yes | NTC network traffic logs | Netflix (recommendations) |
| Neo4j | Graph | No (Community Ed.) | Pathao driver-user routes | LinkedIn (social network) |
| DynamoDB | Key-Value | No (AWS) | Ncell customer billing | Airbnb (reservation system) |
Exam Tip
Define NoSQL Clearly:
- Start with: "NoSQL databases are non-relational stores designed for scalability, flexibility, and distributed data."
- Contrast with SQL: "Unlike SQL, NoSQL lacks fixed schemas and ACID transactions."
Model Comparisons:
- Memorize the 4 models (document, key-value, column-family, graph) and one example each.
- Draw a table in exams comparing them (like above) to score marks.
Real-World Applications:
- Nepal-focused: eSewa (document), Pathao (graph), Daraz (document), NTC (column-family).
- Global: WhatsApp (document), Twitter (key-value), Netflix (column-family).
Worked Examples:
- Always tie examples to Nepal (e.g., "How would you design Khalti’s transaction logs?" → key-value store).
- Show JSON/XML snippets for document stores.
Trade-offs:
- Weaknesses are exam bait: Know why NoSQL lacks joins or ACID, but how to work around it (e.g., denormalization).
Diagrams:
- Mermaid graphs for models (e.g., document structure, graph nodes).
- Real pictures for databases (e.g., MongoDB architecture, Cassandra nodes).
Pro Tip: If the exam asks "Which NoSQL model would you use for X?", follow this logic:
- Is data connected? → Graph.
- Need fast key lookups? → Key-Value.
- Hierarchical/semi-structured? → Document.
- Analytics on large logs? → Column-Family.
Based on the TU BIM syllabus for Big Data and Analytics (IT278), unit 5.
Discussion
Loading…