Big Data and AnalyticsUnit 99 min read

Machine Learning on Big Data: Algorithms, Scalability & Applications

Unit 9 of Big Data and Analytics explores how traditional machine learning techniques are adapted for big data—scalable algorithms (e.g., stochastic gradient descent, ensemble methods), distributed training frameworks (Spark MLlib), and real-world use cases in fraud detection, recommendation systems, and predictive ana

Core Concepts & Key Definitions

Machine learning (ML) on big data refers to applying ML techniques to datasets that are volume-heavy (terabytes+), velocity-driven (streaming), or varied (unstructured/semi-structured). Unlike classical ML, big data ML requires:

  • Scalability: Algorithms must distribute computations across clusters (e.g., Spark, Hadoop).
  • Approximation: Exact solutions are often replaced by probabilistic or incremental methods.
  • Feature Engineering: Extracting meaningful patterns from raw data (e.g., text, images, logs).

Why Traditional ML Fails on Big Data

Challenge Traditional ML Limitation Big Data ML Solution
Data Volume Fits in memory (RAM) Out-of-core algorithms (e.g., SGD, mini-batch)
Velocity (Real-time) Batch processing only Online learning (e.g., Hoeffding Trees)
Variety (Unstructured) Structured tabular data only Feature hashing, embeddings (e.g., Word2Vec)
Veracity (Noise) Assumes clean data Robust models (e.g., Isolation Forest for outliers)

1. Scalable Machine Learning Algorithms

Big data ML relies on algorithms designed for distributed computing or approximate solutions. Key categories:

A. Distributed Training Frameworks

  • Stochastic Gradient Descent (SGD): Updates model weights using one sample at a time (or mini-batches). Scales horizontally via parameter servers (e.g., TensorFlow’s distributed SGD).
SGD Worker 1SGD Worker 2Mini-batch SamplerData Shards (HDFS)Converged ModelParameter ServerDistributed SGD Pipeline
Hierarchical flow of distributed SGD with parameter server aggregation

Example: Google’s RankBrain (search ranking) uses SGD to train on petabytes of query logs.

  • MapReduce for ML: Frameworks like Spark MLlib implement algorithms (e.g., k-means, logistic regression) using MapReduce principles.

B. Approximate Algorithms

  • Locality-Sensitive Hashing (LSH): For near-duplicate detection (e.g., plagiarism checkers like Copyscape).
  • Count-Min Sketch: Approximates frequency counts in streaming data (used by Twitter’s trend detection).
  • Hoeffding Trees: Decision trees for incremental learning (e.g., Ncell’s churn prediction).

C. Ensemble Methods for Big Data

  • Distributed Random Forests: Trees trained on subsets of data (e.g., Amazon’s product recommendations).
  • Gradient Boosting Machines (GBM): Implemented via XGBoost or LightGBM for scalability.

2. Feature Engineering for Big Data

Big data often includes unstructured data (text, images, logs). Techniques to extract features:

  • Text: TF-IDF, Word2Vec, BERT embeddings (e.g., eSewa’s chatbot sentiment analysis).
  • Images: CNN feature maps (e.g., Daraz’s product image search).
  • Time-Series: Rolling statistics, Fourier transforms (e.g., NTC’s electricity demand forecasting).

Worked Example: Fraud Detection in Khalti

  1. Data: 10M transactions/day (structured + unstructured logs).
  2. Features:
    • Time-based: Transaction frequency per user.
    • Graph-based: Social network connections (e.g., sudden links to high-risk accounts).
    • Text: SMS/email metadata (e.g., "URGENT: CLICK HERE" flags).
  3. Model: Isolation Forest (anomaly detection) trained on Spark.
  4. Output: Real-time fraud score (0–1) for each transaction.

3. Model Evaluation & Deployment

A. Challenges

  • Bias-Variance Tradeoff: Big data can overfit if not regularized (e.g., dropout in neural nets).
  • Latency: Models must predict in <100ms for apps like Pathao’s dynamic pricing.
  • Concept Drift: Models degrade over time (e.g., NEPSE stock trends shift with policy changes).

B. Evaluation Metrics for Big Data

Scenario Metric Example Use Case
Imbalanced Data Precision-Recall Curve Bank loan default prediction (99% non-defaults)
Streaming Data AUC-ROC (real-time) Ncell’s call-drop prediction
High-Dimensional Data Silhouette Score Khalti’s customer segmentation

C. Deployment Architectures

  • Batch: Pre-compute predictions (e.g., Daraz’s daily sales forecasts).
  • Real-Time: Serve via Apache Flink or Kafka Streams (e.g., WhatsApp’s spam filtering).
sequenceDiagram
    participant User as User
    participant App as Pathao App
    participant Model as Fraud Model (Flink)
    participant DB as Redis Cache

```figure
{"type":"bar","labels":["Batch Processing","Stream Processing","Hybrid"],"values":[45,30,25],"caption":"Nepalese ML deployment preferences (2023 survey)"}
User->>App: Request Ride
App->>Model: Predict Fraud Score
Model->>DB: Check Cache
DB-->>Model: Return Score (0.05)
Model-->>App: Approve (Score < 0.1)
App->>User: Confirm Ride

4. Real-World Applications in Nepal & Globally

In Nepal

  1. eSewa’s Loan Approval

    • Idea: Gradient Boosting (XGBoost) on transaction history, credit scores, and device fingerprinting.
    • Scale: Processes 50K+ loan applications/month.
    • Impact: Reduces manual review time by 70%.
  2. NTC’s Smart Grid Optimization

    • Idea: Reinforcement Learning (RL) to balance load across substations.
    • Data: 100M+ smart meter readings/day.
    • Tool: TensorFlow Extended (TFX) for distributed RL.
  3. Pathao’s Dynamic Pricing

    • Idea: Multi-Armed Bandit (MAB) algorithm to adjust fares in real-time.
    • Features: Surge demand (from traffic data), driver availability, historical fares.
    • Result: 25% higher driver earnings during peak hours.

Global Examples

  1. Google’s DeepMind for Data Centers

    • Idea: RL to optimize cooling systems (reduced energy use by 40%).
  2. Netflix’s Recommendation Engine

    • Idea: Collaborative Filtering + Deep Learning (Matrix Factorization + Neural Networks).
    • Data: 1B+ user interactions/day.
  3. Uber’s Surge Pricing

    • Idea: Spatial-Temporal Models (LSTMs) to predict demand hotspots.

5. Tools & Frameworks

Tool Purpose Nepal Use Case
Apache Spark MLlib Distributed ML (Python/Scala) Ncell’s customer churn analysis
TensorFlow Extended Scalable ML pipelines Nepal Rastra Bank’s fraud detection
H2O.ai AutoML for big data Daraz’s demand forecasting
Flink ML Real-time ML Khalti’s transaction monitoring

Exam Tip

  1. Algorithm Selection: Always justify why an algorithm (e.g., SGD vs. batch gradient descent) is chosen for big data. Example:

    "For real-time fraud detection in Khalti, we use Hoeffding Trees because they support incremental learning and handle concept drift better than static decision trees."

  2. Scalability Tradeoffs: Discuss accuracy vs. speed in exams. For instance:

    • Approximate algorithms (LSH) sacrifice precision for speed.
    • Distributed SGD may converge slower than batch methods but scales to petabytes.
  3. Case Study Questions: Expect 20–30% of marks on applying concepts to Nepalese contexts (e.g., NTC’s grid optimization or eSewa’s loan risk). Use the CRISP-DM framework (Business Understanding → Data Prep → Modeling → Evaluation) to structure answers.

  4. Visuals in Exams: Sketch data flow diagrams (e.g., how Spark MLlib distributes a k-means job) or feature importance plots (e.g., XGBoost’s SHAP values for loan approval).


Key Formulas to Memorize

  1. Stochastic Gradient Descent Update Rule: Where = learning rate, = single training example.

  2. Precision-Recall Tradeoff (for imbalanced data):

  3. Isolation Forest Anomaly Score: Where = average path length to isolate , = normalization factor.


Common Pitfalls

  • Ignoring Data Skew: Assuming uniform distribution in big data leads to biased models (e.g., NEPSE’s high-volume low-liquidity stocks).
  • Over-Engineering Features: Extracting 1000+ features for a dataset with 1M samples causes overfitting.
  • Neglecting Latency: A model with 99% accuracy is useless if it takes 5 seconds to predict (critical for Pathao’s ride matching).

Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 9.

Discussion

Loading…