CACS460 Internet of Things

Internet of ThingsUnit 610 min read

IoT Data Storage & Processing: Techniques, Tools & Real-Time Analytics

Unit 6 of Internet of Things explores how IoT devices generate, store, process, and analyze massive data streams—from cloud databases to edge analytics—with real-world examples from Nepalese apps like eSewa and global platforms like Google Nest. Learn storage architectures (SQL vs. NoSQL), processing pipelines (batch v


Core Concepts: Why IoT Data Needs Special Handling

IoT devices generate exabytes of data daily (e.g., 100+ devices per second in smart cities). Unlike traditional databases, IoT data has unique traits:

mindmap
  root((IoT Data Characteristics))
    Voluminous
      "Billions of devices → petabytes/day"
    Velocity
      "Real-time streams (e.g., traffic sensors every 100ms)"
    Variety
      "Structured (CSV), semi-structured (JSON), unstructured (video)"
    Veracity
      "Noisy data (sensor errors, missing values)"
    Validity
      "Expiry (e.g., weather data after 24h)"

Key challenge: How to store, process, and analyze this data efficiently while ensuring low latency and scalability.


1. Data Storage Techniques in IoT

IoT storage must balance cost, speed, and reliability. Three main approaches:

A. Cloud Storage (Centralized)

Definition: Data is sent to remote servers (AWS, Google Cloud, Azure) for processing and storage. How it works:

  1. IoT devices (e.g., smart meters) send data via MQTT/HTTP to a cloud gateway.
  2. Gateway aggregates and forwards data to cloud databases (e.g., PostgreSQL, MongoDB).
  3. Analytics engines (e.g., Spark, Hadoop) process the data.

Example: eSewa’s Transaction Logs

  • Problem: Millions of transactions daily need fraud detection in real-time.
  • Solution: Uses AWS DynamoDB (NoSQL) to store transaction metadata + Amazon Kinesis for stream processing.
  • Why NoSQL?
    • Handles unstructured JSON logs (e.g., {user_id: "123", amount: 500, timestamp: "2024-05-20T14:30:00"}).
    • Auto-scaling for spikes in traffic (e.g., Dashain sales).
Cloud Storage Type Use Case Pros Cons
SQL (PostgreSQL) Structured logs (e.g., NTC traffic cameras) ACID compliance, complex queries High latency for real-time analytics
NoSQL (MongoDB) Unstructured sensor data (e.g., soil moisture) Scalable, flexible schemas No joins, eventual consistency
Object Storage (S3) Raw media (e.g., CCTV footage) Cheap, durable Slow for frequent reads/writes

B. Edge Storage (Decentralized)

Definition: Data is processed and stored locally (on IoT devices or edge gateways) before being sent to the cloud. Why?

  • Reduces cloud bandwidth costs (critical for Nepal’s limited internet).
  • Enables real-time decisions (e.g., traffic lights adjusting instantly).

Example: Pathao’s Ride-Hailing System

  • Problem: Millions of GPS coordinates per second → cloud overload.
  • Solution: Uses edge servers in Kathmandu/Pokhara to:
    1. Filter irrelevant data (e.g., discard GPS points if driver speed < 10 km/h).
    2. Store only aggregated stats (e.g., "1000 rides in Zone 3") in the cloud.
  • Tech used: Apache Kafka (stream processing) + SQLite (local storage).
sequenceDiagram
    participant Driver as Driver's Phone (Edge)
    participant Gateway as Edge Gateway (Kathmandu)
    participant Cloud as AWS Cloud
    Driver->>Gateway: GPS (raw, 1Hz)
    Gateway->>Gateway: Filter (remove duplicates)
    Gateway->>Gateway: Aggregate (avg speed per block)
    Gateway->>Cloud: Send summary (1/minute)
    Cloud->>Driver: Optimized route

C. Hybrid Storage (Cloud + Edge)

Best of both worlds: Critical data stays local; analytics run in the cloud. Example: NTC’s Smart Traffic Management

  • Edge: Traffic cameras store last 5 minutes of video locally (for immediate analysis).
  • Cloud: Full dataset uploaded hourly for long-term trend analysis (e.g., "Thapathali junction congestion increases by 20% on Fridays").

Storage Comparison Table

Feature Cloud Storage Edge Storage Hybrid Storage
Latency High (100ms–1s) Ultra-low (<10ms) Low (edge) + occasional cloud
Cost High (bandwidth + storage) Low (local SSD/HDD) Moderate
Use Case Historical analytics Real-time alerts Both (e.g., NTC traffic system)
Tech Examples AWS S3, Google BigQuery SQLite, Redis Kafka + PostgreSQL

2. Data Processing Techniques

IoT data must be cleaned, transformed, and analyzed before use.

A. Batch Processing

Definition: Data is processed in large chunks (e.g., daily reports). Tools: Hadoop, Spark. Example: Nepal Electricity Authority (NEA) processes monthly smart meter readings to detect fraud.

flowchart TD
    A["Smart Meters"] -->|"Daily"| B["HDFS Storage"]
    B --> C["Spark Job"]
    C --> D["Fraud Detection Report"]

Pros/Cons:

  • ✅ Cheap for large datasets.
  • ❌ Not real-time (e.g., NEA detects fraud after the bill is sent).

B. Stream Processing

Definition: Data is analyzed as it arrives (e.g., every second). Tools: Apache Kafka, Flink, AWS Kinesis. Example: Khalti’s Fraud Detection

  • Problem: Detect fake transactions instantly (e.g., same card used in Kathmandu + Pokhara simultaneously).
  • Solution:
    1. Transaction data streams into Kafka topics.
    2. Flink checks for anomalies (e.g., "same card, 50km apart in 1s").
    3. Block transaction before completion.
sequenceDiagram
    participant Merchant as POS Terminal
    participant Khalti as Khalti Server
    participant Flink as Stream Processor
    Merchant->>Khalti: Transaction (JSON)
    Khalti->>Flink: {amount: 500, card: "1234", time: "14:30:00"}
    Flink->>Flink: Check against blacklist
    Flink->>Khalti: ALERT: Fraudulent!
    Khalti->>Merchant: Reject

Pros/Cons:

  • ✅ Real-time decisions (critical for security).
  • ❌ Expensive to scale (requires powerful servers).

C. Edge Analytics

Definition: Processing happens on the device (e.g., a Raspberry Pi). Example: Google Nest Thermostat

  • Problem: Sending temperature data to the cloud every second is wasteful.
  • Solution:
    1. Pi reads local temperature.
    2. Runs ML model to predict heating needs.
    3. Only sends adjustment commands (e.g., "Turn on AC at 2 PM") to the cloud.

3. Big Data Tools for IoT

Tool Purpose IoT Example
Hadoop Distributed batch storage/processing NEA’s monthly meter data analysis
Spark Fast batch/stream processing Khalti’s fraud detection
Kafka Real-time data streaming Pathao’s ride-matching system
Elasticsearch Searching logs (e.g., "Find all failures in Zone 1") NTC’s traffic incident logs
MongoDB NoSQL database for unstructured data Daraz’s customer behavior tracking

4. Data Security Challenges

IoT data is a honey pot for hackers. Key risks:

  1. Unauthorized Access: Weak passwords on default IoT devices (e.g., cheap cameras).
  2. Data Tampering: Attackers alter sensor readings (e.g., hacking a glucose monitor to hide diabetes).
  3. Privacy Leaks: Location data from Pathao drivers sold to third parties.

Solutions:

  • Encryption: TLS for data in transit, AES-256 for storage.
  • Blockchain: Immutable logs (e.g., MedRec for hospital IoT devices).
  • Zero-Trust Model: Verify every device before granting access.

Example: Ncell’s IoT Security

  • Uses IoT-specific firewalls to block DDoS attacks on cell towers.
  • Regular firmware updates to patch vulnerabilities (e.g., Mirai botnet exploits).

5. Worked Example: Smart Agriculture in Nepal

Scenario: A farmer in Chitwan uses IoT to monitor soil moisture and automate irrigation.

Data Flow

flowchart LR
    A["Soil Sensor"] -->|"MQTT"| B["Raspberry Pi"]
    B -->|"Filter"| C["Local SQLite DB"]
    C -->|"Every 1h"| D["Cloud: AWS IoT Core"]
    D -->|"ML Model"| E["Recommendation: 'Water at 6 PM'"]
    E -->|"SMS"| F["Farmer's Phone"]

Storage Choices

Component Storage Type Why?
Raw sensor data SQLite (edge) Low power, no cloud dependency
Daily reports PostgreSQL (cloud) Structured queries for trends
Alerts Redis (edge) Fast retrieval for immediate actions

Processing

  • Edge: Pi runs a simple rule engine (e.g., "If moisture < 30%, water for 5 mins").
  • Cloud: AWS Lambda analyzes weekly trends (e.g., "Soil dries faster in May").

In the Real World

  1. eSewa’s Transaction Processing

    • Idea Used: Stream processing + NoSQL storage.
    • How: Uses Kafka to detect fraud in milliseconds by comparing transactions against blacklists. Data is stored in DynamoDB for auditing.
  2. Pathao’s Ride Matching

    • Idea Used: Edge filtering + hybrid storage.
    • How: Driver locations are aggregated locally (e.g., "5 drivers in Zone 2") before sending to the cloud. Reduces cloud load by 90%.
  3. Google Nest’s Energy Savings

    • Idea Used: Edge analytics.
    • How: The thermostat predicts when to turn on/off based on local weather (not cloud data). Saves 10% energy by avoiding unnecessary cloud trips.

Exam Tip

  1. Compare Cloud vs. Edge Storage:

    • Cloud: Best for historical analysis (e.g., "Which NTC route had the most accidents in 2023?").
    • Edge: Best for real-time actions (e.g., "Turn off water pump if leak detected").
  2. Tools Matter:

    • Batch processing → Hadoop/Spark.
    • Stream processing → Kafka/Flink.
    • NoSQL → MongoDB/Cassandra (for unstructured data).
    • SQL → PostgreSQL (for structured queries).
  3. Real-World Tie-Ins:

    • Always relate to Nepalese examples (eSewa, NTC, NEA) or global giants (Google Nest, Pathao).
    • Example Answer Snippet:

      "Like Pathao, Nepal’s Khalti uses Apache Kafka for stream processing to detect fraudulent transactions in real-time. Raw transaction data is stored in MongoDB for flexibility, while aggregated reports go to PostgreSQL for compliance audits."

  4. Security is a Must:

    • Expect 50% of questions to ask about encryption, blockchain, or zero-trust models.
    • Example: "Explain how Ncell secures IoT data from DDoS attacks." → Mention firewalls, TLS, and regular updates.
  5. Diagrams Save Marks:

    • Draw sequence diagrams for data flows (e.g., sensor → edge → cloud).
    • Use tables to compare storage/processing methods.

Based on the TU BCA syllabus for Internet of Things (CACS460), unit 6.

Discussion

Loading…