Internet of ThingsUnit 610 min read
IoT Data Storage & Processing: Techniques, Tools & Real-Time Analytics
Unit 6 of Internet of Things explores how IoT devices generate, store, process, and analyze massive data streams—from cloud databases to edge analytics—with real-world examples from Nepalese apps like eSewa and global platforms like Google Nest. Learn storage architectures (SQL vs. NoSQL), processing pipelines (batch v
Core Concepts: Why IoT Data Needs Special Handling
IoT devices generate exabytes of data daily (e.g., 100+ devices per second in smart cities). Unlike traditional databases, IoT data has unique traits:
mindmap
root((IoT Data Characteristics))
Voluminous
"Billions of devices → petabytes/day"
Velocity
"Real-time streams (e.g., traffic sensors every 100ms)"
Variety
"Structured (CSV), semi-structured (JSON), unstructured (video)"
Veracity
"Noisy data (sensor errors, missing values)"
Validity
"Expiry (e.g., weather data after 24h)"Key challenge: How to store, process, and analyze this data efficiently while ensuring low latency and scalability.
1. Data Storage Techniques in IoT
IoT storage must balance cost, speed, and reliability. Three main approaches:
A. Cloud Storage (Centralized)
Definition: Data is sent to remote servers (AWS, Google Cloud, Azure) for processing and storage. How it works:
- IoT devices (e.g., smart meters) send data via MQTT/HTTP to a cloud gateway.
- Gateway aggregates and forwards data to cloud databases (e.g., PostgreSQL, MongoDB).
- Analytics engines (e.g., Spark, Hadoop) process the data.
Example: eSewa’s Transaction Logs
- Problem: Millions of transactions daily need fraud detection in real-time.
- Solution: Uses AWS DynamoDB (NoSQL) to store transaction metadata + Amazon Kinesis for stream processing.
- Why NoSQL?
- Handles unstructured JSON logs (e.g.,
{user_id: "123", amount: 500, timestamp: "2024-05-20T14:30:00"}). - Auto-scaling for spikes in traffic (e.g., Dashain sales).
- Handles unstructured JSON logs (e.g.,
| Cloud Storage Type | Use Case | Pros | Cons |
|---|---|---|---|
| SQL (PostgreSQL) | Structured logs (e.g., NTC traffic cameras) | ACID compliance, complex queries | High latency for real-time analytics |
| NoSQL (MongoDB) | Unstructured sensor data (e.g., soil moisture) | Scalable, flexible schemas | No joins, eventual consistency |
| Object Storage (S3) | Raw media (e.g., CCTV footage) | Cheap, durable | Slow for frequent reads/writes |
B. Edge Storage (Decentralized)
Definition: Data is processed and stored locally (on IoT devices or edge gateways) before being sent to the cloud. Why?
- Reduces cloud bandwidth costs (critical for Nepal’s limited internet).
- Enables real-time decisions (e.g., traffic lights adjusting instantly).
Example: Pathao’s Ride-Hailing System
- Problem: Millions of GPS coordinates per second → cloud overload.
- Solution: Uses edge servers in Kathmandu/Pokhara to:
- Filter irrelevant data (e.g., discard GPS points if driver speed < 10 km/h).
- Store only aggregated stats (e.g., "1000 rides in Zone 3") in the cloud.
- Tech used: Apache Kafka (stream processing) + SQLite (local storage).
sequenceDiagram
participant Driver as Driver's Phone (Edge)
participant Gateway as Edge Gateway (Kathmandu)
participant Cloud as AWS Cloud
Driver->>Gateway: GPS (raw, 1Hz)
Gateway->>Gateway: Filter (remove duplicates)
Gateway->>Gateway: Aggregate (avg speed per block)
Gateway->>Cloud: Send summary (1/minute)
Cloud->>Driver: Optimized routeC. Hybrid Storage (Cloud + Edge)
Best of both worlds: Critical data stays local; analytics run in the cloud. Example: NTC’s Smart Traffic Management
- Edge: Traffic cameras store last 5 minutes of video locally (for immediate analysis).
- Cloud: Full dataset uploaded hourly for long-term trend analysis (e.g., "Thapathali junction congestion increases by 20% on Fridays").
Storage Comparison Table
| Feature | Cloud Storage | Edge Storage | Hybrid Storage |
|---|---|---|---|
| Latency | High (100ms–1s) | Ultra-low (<10ms) | Low (edge) + occasional cloud |
| Cost | High (bandwidth + storage) | Low (local SSD/HDD) | Moderate |
| Use Case | Historical analytics | Real-time alerts | Both (e.g., NTC traffic system) |
| Tech Examples | AWS S3, Google BigQuery | SQLite, Redis | Kafka + PostgreSQL |
2. Data Processing Techniques
IoT data must be cleaned, transformed, and analyzed before use.
A. Batch Processing
Definition: Data is processed in large chunks (e.g., daily reports). Tools: Hadoop, Spark. Example: Nepal Electricity Authority (NEA) processes monthly smart meter readings to detect fraud.
flowchart TD
A["Smart Meters"] -->|"Daily"| B["HDFS Storage"]
B --> C["Spark Job"]
C --> D["Fraud Detection Report"]Pros/Cons:
- ✅ Cheap for large datasets.
- ❌ Not real-time (e.g., NEA detects fraud after the bill is sent).
B. Stream Processing
Definition: Data is analyzed as it arrives (e.g., every second). Tools: Apache Kafka, Flink, AWS Kinesis. Example: Khalti’s Fraud Detection
- Problem: Detect fake transactions instantly (e.g., same card used in Kathmandu + Pokhara simultaneously).
- Solution:
- Transaction data streams into Kafka topics.
- Flink checks for anomalies (e.g., "same card, 50km apart in 1s").
- Block transaction before completion.
sequenceDiagram
participant Merchant as POS Terminal
participant Khalti as Khalti Server
participant Flink as Stream Processor
Merchant->>Khalti: Transaction (JSON)
Khalti->>Flink: {amount: 500, card: "1234", time: "14:30:00"}
Flink->>Flink: Check against blacklist
Flink->>Khalti: ALERT: Fraudulent!
Khalti->>Merchant: RejectPros/Cons:
- ✅ Real-time decisions (critical for security).
- ❌ Expensive to scale (requires powerful servers).
C. Edge Analytics
Definition: Processing happens on the device (e.g., a Raspberry Pi). Example: Google Nest Thermostat
- Problem: Sending temperature data to the cloud every second is wasteful.
- Solution:
- Pi reads local temperature.
- Runs ML model to predict heating needs.
- Only sends adjustment commands (e.g., "Turn on AC at 2 PM") to the cloud.
3. Big Data Tools for IoT
| Tool | Purpose | IoT Example |
|---|---|---|
| Hadoop | Distributed batch storage/processing | NEA’s monthly meter data analysis |
| Spark | Fast batch/stream processing | Khalti’s fraud detection |
| Kafka | Real-time data streaming | Pathao’s ride-matching system |
| Elasticsearch | Searching logs (e.g., "Find all failures in Zone 1") | NTC’s traffic incident logs |
| MongoDB | NoSQL database for unstructured data | Daraz’s customer behavior tracking |
4. Data Security Challenges
IoT data is a honey pot for hackers. Key risks:
- Unauthorized Access: Weak passwords on default IoT devices (e.g., cheap cameras).
- Data Tampering: Attackers alter sensor readings (e.g., hacking a glucose monitor to hide diabetes).
- Privacy Leaks: Location data from Pathao drivers sold to third parties.
Solutions:
- Encryption: TLS for data in transit, AES-256 for storage.
- Blockchain: Immutable logs (e.g., MedRec for hospital IoT devices).
- Zero-Trust Model: Verify every device before granting access.
Example: Ncell’s IoT Security
- Uses IoT-specific firewalls to block DDoS attacks on cell towers.
- Regular firmware updates to patch vulnerabilities (e.g., Mirai botnet exploits).
5. Worked Example: Smart Agriculture in Nepal
Scenario: A farmer in Chitwan uses IoT to monitor soil moisture and automate irrigation.
Data Flow
flowchart LR
A["Soil Sensor"] -->|"MQTT"| B["Raspberry Pi"]
B -->|"Filter"| C["Local SQLite DB"]
C -->|"Every 1h"| D["Cloud: AWS IoT Core"]
D -->|"ML Model"| E["Recommendation: 'Water at 6 PM'"]
E -->|"SMS"| F["Farmer's Phone"]Storage Choices
| Component | Storage Type | Why? |
|---|---|---|
| Raw sensor data | SQLite (edge) | Low power, no cloud dependency |
| Daily reports | PostgreSQL (cloud) | Structured queries for trends |
| Alerts | Redis (edge) | Fast retrieval for immediate actions |
Processing
- Edge: Pi runs a simple rule engine (e.g., "If moisture < 30%, water for 5 mins").
- Cloud: AWS Lambda analyzes weekly trends (e.g., "Soil dries faster in May").
In the Real World
eSewa’s Transaction Processing
- Idea Used: Stream processing + NoSQL storage.
- How: Uses Kafka to detect fraud in milliseconds by comparing transactions against blacklists. Data is stored in DynamoDB for auditing.
Pathao’s Ride Matching
- Idea Used: Edge filtering + hybrid storage.
- How: Driver locations are aggregated locally (e.g., "5 drivers in Zone 2") before sending to the cloud. Reduces cloud load by 90%.
Google Nest’s Energy Savings
- Idea Used: Edge analytics.
- How: The thermostat predicts when to turn on/off based on local weather (not cloud data). Saves 10% energy by avoiding unnecessary cloud trips.
Exam Tip
Compare Cloud vs. Edge Storage:
- Cloud: Best for historical analysis (e.g., "Which NTC route had the most accidents in 2023?").
- Edge: Best for real-time actions (e.g., "Turn off water pump if leak detected").
Tools Matter:
- Batch processing → Hadoop/Spark.
- Stream processing → Kafka/Flink.
- NoSQL → MongoDB/Cassandra (for unstructured data).
- SQL → PostgreSQL (for structured queries).
Real-World Tie-Ins:
- Always relate to Nepalese examples (eSewa, NTC, NEA) or global giants (Google Nest, Pathao).
- Example Answer Snippet:
"Like Pathao, Nepal’s Khalti uses Apache Kafka for stream processing to detect fraudulent transactions in real-time. Raw transaction data is stored in MongoDB for flexibility, while aggregated reports go to PostgreSQL for compliance audits."
Security is a Must:
- Expect 50% of questions to ask about encryption, blockchain, or zero-trust models.
- Example: "Explain how Ncell secures IoT data from DDoS attacks." → Mention firewalls, TLS, and regular updates.
Diagrams Save Marks:
- Draw sequence diagrams for data flows (e.g., sensor → edge → cloud).
- Use tables to compare storage/processing methods.
Based on the TU BCA syllabus for Internet of Things (CACS460), unit 6.
Discussion
Loading…