Internet of ThingsUnit 714 min read
Big Data Analytics & IoT: Tools, Techniques, and Real-Time Processing
Unit 7 of Internet of Things explores how IoT generates massive data streams, the tools (Hadoop, Spark, Kafka) and techniques (streaming, edge analytics) used to process them, and their applications in smart cities, healthcare, and industry—with real-world examples from Nepal and global tech giants.
TAKEAWAYS:
- IoT generates zettabytes of data daily; big data analytics tools (Hadoop, Spark, Flink) process it in real time or batch.
- Edge analytics reduces latency by processing data locally (e.g., smart traffic lights), while network analytics optimizes IoT traffic flows.
- Storage techniques (NoSQL, time-series databases) and processing models (lambda architecture) are tailored to IoT’s velocity, variety, and volume.
- Domain-specific analytics (e.g., predictive maintenance in factories, patient monitoring in hospitals) drive IoT’s value.
- Challenges include data privacy, scalability, and integrating analytics with legacy systems.
- Nepalese use cases: eSewa’s fraud detection (real-time transaction analytics), NTC’s smart grid monitoring (edge analytics), and Daraz’s supply chain optimization (big data).
1. Why IoT + Big Data? The Data Explosion
IoT devices (sensors, wearables, smart meters) generate data at unprecedented scales:
- Volume: A single smart city may have millions of sensors (traffic cameras, air quality monitors, utility meters) producing terabytes per day.
- Velocity: Data arrives in streams (e.g., a heartbeat monitor sends 100 readings/second).
- Variety: Structured (CSV logs), semi-structured (JSON from APIs), and unstructured (video/audio).
- Veracity: Noisy, incomplete, or inconsistent (e.g., a faulty temperature sensor in a greenhouse).
Example: In Nepal, NTC’s smart grid uses IoT sensors to monitor power usage every millisecond. Without big data tools, analyzing this would be impossible.
2. Big Data Analytics Tools for IoT
IoT analytics relies on distributed processing frameworks and specialized databases. Below are the key tools, categorized by function:
A. Data Processing Frameworks
| Tool | Type | Use Case in IoT | Example in Nepal |
|---|---|---|---|
| Apache Hadoop | Batch processing | Historical analysis (e.g., yearly energy consumption trends in NTC grids). | NTC’s post-facto load analysis. |
| Apache Spark | In-memory, batch/stream | Real-time fraud detection in eSewa/Khalti transactions. | Khalti’s instant transaction risk scoring. |
| Apache Flink | Stream processing | Real-time traffic routing in Pathao or Kathmandu Metro apps. | Pathao’s dynamic ride pricing. |
| Kafka | Event streaming | Aggregating sensor data from Nepal’s agricultural IoT (soil moisture, weather). | FarmLog’s real-time crop health alerts. |
B. Storage Systems
| Database | Type | IoT Use Case | Example |
|---|---|---|---|
| MongoDB | NoSQL (document) | Storing unstructured IoT data (e.g., JSON logs from Daraz’s delivery drones). | Daraz’s logistics sensor data. |
| InfluxDB | Time-series | Storing NTC’s smart meter readings (timestamped power usage). | NTC’s grid monitoring. |
| Cassandra | Wide-column | High-write scenarios (e.g., Ncell’s IoT SIM cards location tracking). | Ncell’s IoT device telemetry. |
MERMAID DIAGRAM: IoT Data Pipeline
flowchart TD
A["IoT Devices\n(Sensors, Actuators)"] -->|"Raw Data"| B["Edge Gateway\n(Filters/Preprocesses)"]
B --> C["Apache Kafka\n(Event Streaming)"]
C --> D1["Apache Spark\n(Real-Time Analytics)"]
C --> D2["Apache Hadoop\n(Batch Processing)"]
D1 --> E["NoSQL DB\n(MongoDB/InfluxDB)"]
D2 --> F["Data Warehouse\n(SQL for Reports)"]
E & F --> G["ML Models\n(Predictive Maintenance, Anomaly Detection)"]Caption: How data flows from sensors to analytics in a smart city (e.g., Kathmandu traffic management).
3. Key Analytics Techniques for IoT
A. Batch Processing (Offline Analytics)
- What it does: Processes historical data in large chunks (e.g., monthly reports).
- Tools: Hadoop MapReduce, Spark Batch.
- Example:
- Nepal Electricity Authority (NEA) uses batch analytics to identify peak demand patterns over a year, optimizing generator scheduling.
- Worked Example:
Suppose NEA collects 1TB of smart meter data/month. A Hadoop job could run weekly to:
- Aggregate data by hour/day.
- Apply linear regression to predict demand spikes.
- Generate a report for NEA engineers.
B. Stream Processing (Real-Time Analytics)
- What it does: Analyzes data as it arrives (e.g., fraud detection, live traffic rerouting).
- Tools: Spark Streaming, Flink, Kafka Streams.
- Example:
- eSewa/Khalti uses Spark Streaming to detect fraudulent transactions in <100ms:
- A user initiates a ₹50,000 transfer to an unknown account.
- Kafka ingests the transaction in real time.
- Spark checks:
- Is the IP address new? (Yes → flag).
- Is the device’s location consistent? (No → block).
- Alert sent to eSewa’s risk team.
- eSewa/Khalti uses Spark Streaming to detect fraudulent transactions in <100ms:
MERMAID DIAGRAM: Fraud Detection in eSewa
sequenceDiagram
participant User
participant eSewaApp
participant Kafka
participant Spark
participant RiskEngine
participant Database
User->>eSewaApp: Initiates ₹50K transfer
eSewaApp->>Kafka: Sends transaction event (JSON)
Kafka->>Spark: Streams event to fraud detector
Spark->>RiskEngine: Checks IP/device anomalies
RiskEngine-->>Spark: "High risk: new IP"
Spark->>Database: Logs fraud attempt
Spark->>eSewaApp: "Transaction blocked"
eSewaApp->>User: "Fraud detected. Contact support."Caption: Real-time fraud detection in Nepal’s digital wallets.
C. Edge Analytics (Processing at the Source)
- Why it matters: Reduces latency and bandwidth by processing data locally (e.g., on a Raspberry Pi).
- Use Cases:
- Smart traffic lights (e.g., Kathmandu Metro) adjust signals based on real-time camera feeds without sending data to a cloud server.
- Predictive maintenance in factories (e.g., Nepal’s textile mills use edge devices to detect machine failures before they happen).
Example:
- NTC’s smart grid uses edge analytics to:
- Monitor voltage fluctuations in a substation.
- If voltage drops >5%, the edge device automatically reroutes power to avoid blackouts.
- Only anomaly alerts (not raw data) are sent to the cloud.
4. Storage Techniques for IoT Data
IoT data has unique challenges:
- High write volume (e.g., a weather station logs data every second).
- Time-sensitive (e.g., stock market IoT needs sub-second updates).
- Schema flexibility (e.g., wearable health data may add new fields over time).
A. Time-Series Databases (TSDB)
- Best for: Metrics and events with timestamps (e.g., sensor readings).
- Examples: InfluxDB, TimescaleDB.
- Nepal Use Case:
- NTC stores smart meter data in InfluxDB to:
- Track hourly power consumption.
- Detect sudden spikes (e.g., a factory malfunction).
- Generate billing reports.
- NTC stores smart meter data in InfluxDB to:
B. NoSQL Databases
- Best for: Unstructured or semi-structured data (e.g., JSON from IoT APIs).
- Examples: MongoDB, Cassandra.
- Nepal Use Case:
- Daraz’s logistics uses MongoDB to store:
- Delivery drone telemetry (GPS, battery level, package status).
- Customer feedback (unstructured text + ratings).
- Daraz’s logistics uses MongoDB to store:
C. Data Lakes (Raw Storage)
- Best for: Long-term archival of raw IoT data (e.g., Nepal’s agricultural IoT storing 5 years of soil data).
- Tools: AWS S3, Azure Data Lake.
- Example:
- FarmLog (Nepal’s agri-IoT) stores raw sensor data in a lake to:
- Train ML models for crop disease prediction.
- Allow historical trend analysis (e.g., "How did monsoon rains affect yield in 2020?").
- FarmLog (Nepal’s agri-IoT) stores raw sensor data in a lake to:
MERMAID DIAGRAM: IoT Data Storage Options
mindmap
root((IoT Data Storage))
Time-Series DB
InfluxDB: Smart meters (NTC)
TimescaleDB: Industrial sensors
NoSQL
MongoDB: Daraz drone logs
Cassandra: Ncell IoT SIM tracking
Data Lake
AWS S3: FarmLog raw sensor data
Azure Data Lake: Historical weather + crop data
SQL (Legacy)
PostgreSQL: Structured reports (e.g., NEA billing)Caption: Storage choices depend on data type and access patterns.
5. Domain-Specific Analytics in IoT
Analytics are tailored to industries. Here’s how Nepal and global companies apply them:
| Domain | IoT Data Source | Analytics Technique | Nepal Example | Global Example |
|---|---|---|---|---|
| Smart Cities | Traffic cameras, air sensors | Computer vision + predictive modeling | Kathmandu Metro reroutes buses based on live congestion. | Singapore’s AI traffic lights. |
| Healthcare | Wearables, hospital sensors | Anomaly detection (ML) | KOC Hospital monitors ICU patients’ vitals in real time. | Apple Watch AFib detection. |
| Agriculture | Soil moisture, drones | Prescriptive analytics | FarmLog recommends irrigation schedules. | John Deere’s autonomous tractors. |
| Manufacturing | Factory sensors | Predictive maintenance (LSTM models) | Nepal’s textile mills predict loom failures. | Siemens’ smart factories. |
| Retail | POS, inventory sensors | Demand forecasting | Daraz optimizes warehouse stock levels. | Amazon’s warehouse robots. |
6. Challenges in IoT Big Data Analytics
| Challenge | Cause | Solution | Nepal-Specific Issue |
|---|---|---|---|
| Data Overload | Billions of IoT devices generating data. | Edge filtering + sampling. | NTC’s smart grid drowns in meter data. |
| Latency | Cloud processing delays. | Edge analytics. | Pathao’s ride app needs <1s response. |
| Privacy | Sensitive data (e.g., health wearables). | Federated learning, encryption. | eSewa/Khalti must protect transaction data. |
| Heterogeneous Data | Mix of structured/unstructured data. | Schema-less databases (MongoDB). | Daraz’s logistics combines GPS + text feedback. |
| Cost | Scaling storage/processing. | Hybrid cloud-edge models. | Nepal’s rural IoT (limited bandwidth). |
7. Worked Example: Predictive Maintenance in a Nepalese Textile Mill
Scenario: A textile mill in Dhaka (Nepal) uses IoT sensors on looms to monitor:
- Vibration levels (indicates wear).
- Temperature (overheating = failure).
- Oil pressure (lubrication issues).
Problem: Unplanned downtime costs ₹50,000/hour. The mill wants to predict failures before they happen.
Solution:
Data Collection:
- 100 sensors log data every 5 minutes → 240,000 readings/day.
- Stored in InfluxDB (time-series).
Edge Processing:
- A Raspberry Pi at each loom runs lightweight ML (e.g., Isolation Forest for anomaly detection).
- If vibration > threshold, it triggers an alert.
Cloud Analytics:
- Spark Streaming aggregates data from all looms.
- LSTM model (trained on historical data) predicts:
- "Loom #45 will fail in 12 hours (90% confidence)."
- Maintenance scheduled proactively.
Outcome:
- Downtime reduced by 60%.
- Cost savings: ₹20M/year.
MERMAID DIAGRAM: Predictive Maintenance Pipeline
stateDiagram-v2
[*] --> Idle
Idle --> Collecting: "Sensors log data"
Collecting --> EdgeFilter: "Raspberry Pi checks thresholds"
EdgeFilter --> Alert: "Vibration > threshold"
Alert --> Cloud: "Send to Spark"
Cloud --> TrainModel: "Update LSTM weekly"
TrainModel --> Predict: "Forecast failures"
Predict --> Maintenance: "Schedule repairs"
Maintenance --> [*]Caption: State transitions in the textile mill’s predictive maintenance system.
8. Exam Tip: How to Score Full Marks
Define Clearly:
- Always start with one-sentence definitions (e.g., "Big data analytics in IoT refers to the processing of high-velocity, heterogeneous data from sensors to extract actionable insights.").
Use Diagrams:
- Draw pipelines (e.g., IoT → Edge → Cloud → Analytics).
- Compare tools in tables (e.g., Hadoop vs. Spark for batch vs. stream).
Link to Nepal:
- Every example must tie to a local company (e.g., "NTC uses edge analytics to prevent blackouts").
- Mention data types: "eSewa’s Kafka streams handle JSON transaction logs."
Explain Trade-offs:
- "Edge analytics reduces latency but requires more local compute power."
- "NoSQL databases are flexible but lack ACID transactions."
Worked Examples:
- Show calculations (e.g., "If a smart meter sends 100 readings/second, daily volume = 8.64MB/day").
- Trace data flow (e.g., "Step 1: Sensor → Step 2: Edge filter → Step 3: Kafka → Step 4: Spark").
Common Pitfalls:
- ❌ "IoT and big data are the same." → ✅ "IoT generates data; big data analytics processes it."
- ❌ "All IoT data needs cloud processing." → ✅ "Edge analytics is critical for real-time systems."
- Data Sources (Sensors → Actuators)
- Processing (Edge → Cloud)
- Storage (TSDB → NoSQL → Data Lake)
- Tools (Spark → Kafka → Hadoop)
- Nepal Use Cases (NTC, eSewa, Daraz)."**
Final Checklist for Exams:
- Mention at least 2 Nepalese companies (e.g., NTC, eSewa, Daraz).
- Draw one pipeline diagram (IoT → Analytics → Action).
- Compare two tools (e.g., Spark vs. Flink for streaming).
- Explain one real-world trade-off (e.g., edge vs. cloud).
- Use one worked example (e.g., textile mill or smart grid).
Based on the TU BCA syllabus for Internet of Things (CACS460), unit 7.
Discussion
Loading…