Big Data and AnalyticsUnit 99 min read
Machine Learning on Big Data: Algorithms, Scalability & Applications
Unit 9 of Big Data and Analytics explores how traditional machine learning techniques are adapted for big data—scalable algorithms (e.g., stochastic gradient descent, ensemble methods), distributed training frameworks (Spark MLlib), and real-world use cases in fraud detection, recommendation systems, and predictive ana
Core Concepts & Key Definitions
Machine learning (ML) on big data refers to applying ML techniques to datasets that are volume-heavy (terabytes+), velocity-driven (streaming), or varied (unstructured/semi-structured). Unlike classical ML, big data ML requires:
- Scalability: Algorithms must distribute computations across clusters (e.g., Spark, Hadoop).
- Approximation: Exact solutions are often replaced by probabilistic or incremental methods.
- Feature Engineering: Extracting meaningful patterns from raw data (e.g., text, images, logs).
Why Traditional ML Fails on Big Data
| Challenge | Traditional ML Limitation | Big Data ML Solution |
|---|---|---|
| Data Volume | Fits in memory (RAM) | Out-of-core algorithms (e.g., SGD, mini-batch) |
| Velocity (Real-time) | Batch processing only | Online learning (e.g., Hoeffding Trees) |
| Variety (Unstructured) | Structured tabular data only | Feature hashing, embeddings (e.g., Word2Vec) |
| Veracity (Noise) | Assumes clean data | Robust models (e.g., Isolation Forest for outliers) |
1. Scalable Machine Learning Algorithms
Big data ML relies on algorithms designed for distributed computing or approximate solutions. Key categories:
A. Distributed Training Frameworks
- Stochastic Gradient Descent (SGD): Updates model weights using one sample at a time (or mini-batches). Scales horizontally via parameter servers (e.g., TensorFlow’s distributed SGD).
Example: Google’s RankBrain (search ranking) uses SGD to train on petabytes of query logs.
- MapReduce for ML: Frameworks like Spark MLlib implement algorithms (e.g., k-means, logistic regression) using MapReduce principles.
B. Approximate Algorithms
- Locality-Sensitive Hashing (LSH): For near-duplicate detection (e.g., plagiarism checkers like Copyscape).
- Count-Min Sketch: Approximates frequency counts in streaming data (used by Twitter’s trend detection).
- Hoeffding Trees: Decision trees for incremental learning (e.g., Ncell’s churn prediction).
C. Ensemble Methods for Big Data
- Distributed Random Forests: Trees trained on subsets of data (e.g., Amazon’s product recommendations).
- Gradient Boosting Machines (GBM): Implemented via XGBoost or LightGBM for scalability.
2. Feature Engineering for Big Data
Big data often includes unstructured data (text, images, logs). Techniques to extract features:
- Text: TF-IDF, Word2Vec, BERT embeddings (e.g., eSewa’s chatbot sentiment analysis).
- Images: CNN feature maps (e.g., Daraz’s product image search).
- Time-Series: Rolling statistics, Fourier transforms (e.g., NTC’s electricity demand forecasting).
Worked Example: Fraud Detection in Khalti
- Data: 10M transactions/day (structured + unstructured logs).
- Features:
- Time-based: Transaction frequency per user.
- Graph-based: Social network connections (e.g., sudden links to high-risk accounts).
- Text: SMS/email metadata (e.g., "URGENT: CLICK HERE" flags).
- Model: Isolation Forest (anomaly detection) trained on Spark.
- Output: Real-time fraud score (0–1) for each transaction.
3. Model Evaluation & Deployment
A. Challenges
- Bias-Variance Tradeoff: Big data can overfit if not regularized (e.g., dropout in neural nets).
- Latency: Models must predict in <100ms for apps like Pathao’s dynamic pricing.
- Concept Drift: Models degrade over time (e.g., NEPSE stock trends shift with policy changes).
B. Evaluation Metrics for Big Data
| Scenario | Metric | Example Use Case |
|---|---|---|
| Imbalanced Data | Precision-Recall Curve | Bank loan default prediction (99% non-defaults) |
| Streaming Data | AUC-ROC (real-time) | Ncell’s call-drop prediction |
| High-Dimensional Data | Silhouette Score | Khalti’s customer segmentation |
C. Deployment Architectures
- Batch: Pre-compute predictions (e.g., Daraz’s daily sales forecasts).
- Real-Time: Serve via Apache Flink or Kafka Streams (e.g., WhatsApp’s spam filtering).
sequenceDiagram
participant User as User
participant App as Pathao App
participant Model as Fraud Model (Flink)
participant DB as Redis Cache
```figure
{"type":"bar","labels":["Batch Processing","Stream Processing","Hybrid"],"values":[45,30,25],"caption":"Nepalese ML deployment preferences (2023 survey)"}User->>App: Request Ride
App->>Model: Predict Fraud Score
Model->>DB: Check Cache
DB-->>Model: Return Score (0.05)
Model-->>App: Approve (Score < 0.1)
App->>User: Confirm Ride
4. Real-World Applications in Nepal & Globally
In Nepal
eSewa’s Loan Approval
- Idea: Gradient Boosting (XGBoost) on transaction history, credit scores, and device fingerprinting.
- Scale: Processes 50K+ loan applications/month.
- Impact: Reduces manual review time by 70%.
NTC’s Smart Grid Optimization
- Idea: Reinforcement Learning (RL) to balance load across substations.
- Data: 100M+ smart meter readings/day.
- Tool: TensorFlow Extended (TFX) for distributed RL.
Pathao’s Dynamic Pricing
- Idea: Multi-Armed Bandit (MAB) algorithm to adjust fares in real-time.
- Features: Surge demand (from traffic data), driver availability, historical fares.
- Result: 25% higher driver earnings during peak hours.
Global Examples
Google’s DeepMind for Data Centers
- Idea: RL to optimize cooling systems (reduced energy use by 40%).
Netflix’s Recommendation Engine
- Idea: Collaborative Filtering + Deep Learning (Matrix Factorization + Neural Networks).
- Data: 1B+ user interactions/day.
Uber’s Surge Pricing
- Idea: Spatial-Temporal Models (LSTMs) to predict demand hotspots.
5. Tools & Frameworks
| Tool | Purpose | Nepal Use Case |
|---|---|---|
| Apache Spark MLlib | Distributed ML (Python/Scala) | Ncell’s customer churn analysis |
| TensorFlow Extended | Scalable ML pipelines | Nepal Rastra Bank’s fraud detection |
| H2O.ai | AutoML for big data | Daraz’s demand forecasting |
| Flink ML | Real-time ML | Khalti’s transaction monitoring |
Exam Tip
Algorithm Selection: Always justify why an algorithm (e.g., SGD vs. batch gradient descent) is chosen for big data. Example:
"For real-time fraud detection in Khalti, we use Hoeffding Trees because they support incremental learning and handle concept drift better than static decision trees."
Scalability Tradeoffs: Discuss accuracy vs. speed in exams. For instance:
- Approximate algorithms (LSH) sacrifice precision for speed.
- Distributed SGD may converge slower than batch methods but scales to petabytes.
Case Study Questions: Expect 20–30% of marks on applying concepts to Nepalese contexts (e.g., NTC’s grid optimization or eSewa’s loan risk). Use the CRISP-DM framework (Business Understanding → Data Prep → Modeling → Evaluation) to structure answers.
Visuals in Exams: Sketch data flow diagrams (e.g., how Spark MLlib distributes a k-means job) or feature importance plots (e.g., XGBoost’s SHAP values for loan approval).
Key Formulas to Memorize
Stochastic Gradient Descent Update Rule: Where = learning rate, = single training example.
Precision-Recall Tradeoff (for imbalanced data):
Isolation Forest Anomaly Score: Where = average path length to isolate , = normalization factor.
Common Pitfalls
- Ignoring Data Skew: Assuming uniform distribution in big data leads to biased models (e.g., NEPSE’s high-volume low-liquidity stocks).
- Over-Engineering Features: Extracting 1000+ features for a dataset with 1M samples causes overfitting.
- Neglecting Latency: A model with 99% accuracy is useless if it takes 5 seconds to predict (critical for Pathao’s ride matching).
Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 9.
Discussion
Loading…