Big Data and AnalyticsUnit 811 min read

Big Data Analytics: Techniques, Tools & Applications

Unit 8 of Big Data and Analytics explores core techniques for extracting insights from massive datasets, including descriptive, predictive, and prescriptive analytics, along with tools like regression, clustering, classification, and association rules. It covers real-world applications in Nepal (e.g., NTC’s network opt

Core Techniques in Big Data Analytics

Big Data Analytics transforms raw data into actionable insights using statistical, machine learning, and optimization methods. These techniques are categorized into three broad types:

1. Descriptive Analytics: Summarizing the Past

Descriptive analytics answers "What happened?" by aggregating historical data to identify trends, patterns, or anomalies. Common methods include:

  • Data aggregation (sum, average, count).
  • Data mining (frequent pattern mining, association rules).
  • OLAP (Online Analytical Processing) for multidimensional analysis.

How It Works

  1. Data Collection: Gather structured/unstructured data (e.g., customer transactions, social media posts).
  2. Cleaning & Preprocessing: Handle missing values, outliers, and noise.
  3. Summarization: Use statistical measures (mean, median, standard deviation) or visualizations (charts, dashboards).
  4. Pattern Discovery: Apply algorithms like Apriori (for market basket analysis) or FP-Growth (faster alternative).

Worked Example: Daraz’s "Frequently Bought Together"

Daraz uses association rule mining to suggest products like: "Customers who bought a smartphone also bought a screen protector (support=5000, confidence=60%)." Steps:

  1. Transaction Data: Scan 1 million orders to find co-occurring items.
  2. Apriori Algorithm:
    • Generate frequent itemsets (e.g., {smartphone, screen protector} appears in 5000 transactions).
    • Calculate support (frequency) and confidence (probability).
  3. Result: Display rules like "If X, then Y" on product pages to boost sales.
graph LR
    A["Raw Transaction Data"] --> B["Preprocessing: Remove Noise"]
    B --> C["Apriori Algorithm: Find Frequent Itemsets"]
    C --> D["Calculate Support & Confidence"]
    D --> E["Generate Rules: {X} → {Y}"]
    E --> F["Display on Daraz Website"]

Real-World Application: NTC’s Network Traffic Analysis

Nepal Telecom (NTC) uses descriptive analytics to:

  • Track peak usage hours (e.g., 7–9 PM) to optimize server load.
  • Identify regions with high call drops via geospatial heatmaps.
  • Tool: Apache Spark + Tableau for real-time dashboards.

2. Predictive Analytics: Forecasting the Future

Predictive analytics answers "What will happen?" using historical data and statistical/machine learning models. Key techniques:

  • Regression (linear, logistic, polynomial).
  • Classification (decision trees, SVM, Naive Bayes).
  • Time Series Forecasting (ARIMA, exponential smoothing).

How It Works

  1. Data Preparation: Split data into training/test sets.
  2. Model Training: Fit algorithms to historical data (e.g., train a decision tree on loan approvals).
  3. Prediction: Apply the model to new data (e.g., predict if a customer will default).

Worked Example: Ncell’s Churn Prediction

Ncell loses 15% of customers yearly. To reduce churn, they use logistic regression:

  • Features: Call duration, data usage, customer service complaints.
  • Target: Binary (1 = churns, 0 = stays).
  • Model: P(churn) = 1 / (1 + e^(-(β0 + β1*calls + β2*complaints + ...)))
  • Output: Customers with P(churn) > 0.7 get discounts.
graph TD
    A["Customer Data"] --> B["Feature Engineering: Calls, Complaints"]
    B --> C["Train Logistic Regression Model"]
    C --> D["Predict Churn Probability"]
    D --> E["Send Retention Offers to High-Risk Users"]

Real-World Application: Pathao’s Demand Forecasting

Pathao uses time series forecasting to:

  • Predict rider demand in Kathmandu’s busy routes (e.g., Thapathali to Patan).
  • Tool: Prophet (by Meta) to handle seasonality (e.g., higher demand on weekends).
  • Impact: Optimizes driver allocation, reducing wait times by 20%.

3. Prescriptive Analytics: Recommending Actions

Prescriptive analytics answers "What should we do?" by optimizing decisions. Techniques:

  • Optimization Algorithms (linear programming, genetic algorithms).
  • Simulation (Monte Carlo for risk analysis).
  • Reinforcement Learning (dynamic decision-making).

How It Works

  1. Define Objectives: Maximize profit, minimize cost.
  2. Constraints: Budget, time, regulations.
  3. Solve: Use solvers (e.g., Python’s scipy.optimize).

Worked Example: Khalti’s Fraud Detection

Khalti processes 500,000 transactions/day. To detect fraud:

  1. Anomaly Detection: Use Isolation Forest to flag unusual transactions (e.g., $500 transfer at 3 AM).
  2. Prescriptive Action: Block transaction if P(fraud) > 0.95; notify user for verification.
  3. Feedback Loop: Update model with new fraud patterns.
classDiagram
    class Transaction {
        +amount: float
        +time: datetime
        +location: str
        +user_id: str
    }
    class FraudModel {
        +train(data: list[Transaction]) void
        +predict(tx: Transaction) float
    }
    class ActionEngine {
        +block(tx: Transaction) void
        +notify(user: str) void
    }
    Transaction --> FraudModel : "P(fraud)"
    FraudModel --> ActionEngine : "if P > 0.95"

Real-World Application: NEPSE’s Portfolio Optimization

Nepal Stock Exchange (NEPSE) uses portfolio optimization to:

  • Balance risk/reward for investors using Markowitz Model.
  • Constraints: Budget ≤ Rs. 1,000,000; max 20% in one stock.
  • Tool: Python’s cvxpy for linear programming.
  • Result: Suggests a diversified portfolio (e.g., 30% NMB, 25% Global IME, 15% NTC).

4. Advanced Techniques

A. Clustering: Segmenting Data

Unsupervised learning to group similar data points. Methods:

  • K-Means: Partition data into K clusters (e.g., customer segmentation).
  • DBSCAN: Handles irregular shapes (e.g., anomaly detection).

Example: E-Sewa clusters users by spending habits to target ads.

graph LR
    A["User Data: Age, Location, Spend"] --> B["K-Means: K=4"]
    B --> C["Cluster 1: High Spenders"]
    B --> D["Cluster 2: Budget Users"]
    D --> E["Send Discounts to Cluster 2"]

B. Natural Language Processing (NLP)

Extract insights from text (e.g., WhatsApp’s spam detection).

  • Tokenization: Split text into words.
  • Sentiment Analysis: Classify reviews as positive/negative (e.g., Daraz product reviews).

Example: NTC analyzes customer complaints on Twitter to improve service.

C. Graph Analytics

Analyze relationships (e.g., social networks, fraud rings).

  • PageRank: Used by Google to rank websites.
  • Community Detection: Identify groups in Pathao’s rider network.

Tools for Big Data Analytics

Tool Type Use Case Example in Nepal
Apache Spark Batch/Stream Processing Large-scale ML (e.g., NTC traffic) NTC’s network analytics
Hadoop HDFS Storage Store raw data (e.g., Daraz orders) Daraz’s petabyte-scale data lake
Tableau Visualization Dashboards (e.g., Khalti transactions) Khalti’s fraud monitoring dashboard
Python (Scikit-learn) ML Predictive models (e.g., Ncell churn) Ncell’s customer analytics
TensorFlow Deep Learning Image recognition (e.g., OCR for NRS) Nepal Rastra Bank’s document processing

In the Real World

  1. Google Search Ranking

    • Technique: Predictive + Prescriptive Analytics
    • How: Uses collaborative filtering (like Netflix) and PageRank to rank results. Predicts user intent (e.g., "best laptop in Nepal") and prescribes the most relevant ads.
  2. WhatsApp’s Spam Detection

    • Technique: NLP + Classification
    • How: Trains a Random Forest model on labeled messages (spam/ham). Flags messages with P(spam) > 0.9 and moves them to "Spam" folder.
  3. NTC’s 5G Network Optimization

    • Technique: Time Series + Simulation
    • How: Uses ARIMA models to forecast data demand spikes during festivals (e.g., Dashain). Simulates server loads to pre-allocate bandwidth in high-traffic areas (e.g., Thamel).
  4. Daraz’s "Because You Viewed" Recommendations

    • Technique: Collaborative Filtering (CF)
    • How: If User A and B buy similar products, recommend B’s purchases to A. Uses matrix factorization to handle sparse data (millions of users, few interactions).
  5. Khalti’s Anti-Money Laundering (AML)

    • Technique: Graph Analytics + Anomaly Detection
    • How: Builds a transaction graph where nodes = users, edges = transfers. Flags suspicious patterns (e.g., rapid transfers between unrelated accounts) using community detection.

Exam Tip

What Examiners Look For

  1. Definitions: Know the difference between descriptive, predictive, and prescriptive analytics. For example:

    • Descriptive: "What happened?" (e.g., "Sales dropped 10% in Q2.")
    • Predictive: "What will happen?" (e.g., "Sales will drop another 5% if ads stop.")
    • Prescriptive: "What should we do?" (e.g., "Run a 20% discount campaign.")
  2. Algorithms: Be ready to explain how Apriori works (step-by-step) or how logistic regression calculates probabilities. Use pseudocode or Mermaid diagrams in exams if allowed.

  3. Real-World Mapping: Always tie techniques to Nepalese examples. For instance:

    • Question: "How would you reduce Ncell’s churn?"
    • Answer: Use logistic regression on call logs/complaints → predict churn → offer discounts to high-risk users.
  4. Tools vs. Techniques:

    • Tool: Hadoop (storage), Spark (processing), Tableau (visualization).
    • Technique: K-Means (clustering), Random Forest (classification).
    • Exam Pitfall: Don’t confuse tools (e.g., Python) with techniques (e.g., decision trees).
  5. Diagrams: Draw flowcharts for processes (e.g., Apriori algorithm) or class diagrams for systems (e.g., Khalti’s fraud detection). Use timelines for predictive models (e.g., ARIMA steps).

  6. Math Shortcuts: Memorize key formulas:

    • Support (A → B): P(A ∩ B) / P(total transactions)
    • Confidence (A → B): P(B|A) = P(A ∩ B) / P(A)
    • Logistic Regression: P(Y=1) = σ(β0 + β1X)
  7. Case Study Questions: Expect questions like:

    • "How would you analyze Daraz’s customer data to increase sales?" Answer:
      1. Descriptive: Use Apriori to find frequent itemsets.
      2. Predictive: Train a decision tree to predict high-value customers.
      3. Prescriptive: Recommend "Buy Together" bundles and offer discounts.

Common Mistakes to Avoid

  • Overcomplicating: Examiners prefer clear steps over jargon. For example, instead of "We’ll use a neural network," say: "We’ll train a simple decision tree with features like ‘call duration’ and ‘complaints’ to predict churn."
  • Ignoring Assumptions: Always state assumptions. For example: "Apriori assumes transactions are independent (no sequential patterns)."
  • Skipping Validation: Never forget to split data into training/test sets in predictive analytics.
  • Visuals Without Labels: If you draw a flowchart, label every box (e.g., "Step 1: Data Collection").

Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 8.

Discussion

Loading…