Big Data and AnalyticsUnit 89 min read
Big Data Analytics: Techniques, Tools & Real-World Impact
Unit 8 of Big Data and Analytics explores core techniques for extracting insights from massive datasets—descriptive, predictive, prescriptive analytics, and their tools (SQL, R, Python, Tableau). Covers real-world applications in Nepal (e.g., Ncell’s churn prediction, Daraz’s demand forecasting) and global platforms (G
Core Techniques in Big Data Analytics
Big Data Analytics transforms raw data into actionable insights using three primary approaches:
1. Descriptive Analytics: "What Happened?"
Definition: Summarizes historical data to understand past trends, patterns, or performance. Uses aggregation, visualization, and basic statistics. Key Tools: SQL, Excel, Tableau, Power BI, Hadoop (for large-scale data). Example: Ncell’s monthly call-duration reports to identify peak usage hours.
graph LR
A["Raw Data\n(Customer Calls, SMS, Data Usage)"] --> B["Data Cleaning\n(Filter noise, handle missing values)"]
B --> C["Aggregation\n(Group by time, user, region)"]
C --> D["Visualization\n(Charts, dashboards)"]
D --> E["Insight\n'Peak hours: 7–9 PM'"]Real-World Tie-In:
- eSewa: Uses descriptive analytics to track transaction volumes by district, helping identify high-demand areas for server scaling.
- Daraz: Analyzes order data to show best-selling products by season (e.g., umbrellas before monsoon).
Worked Example: Kathmandu Traffic Congestion Problem: NTC wants to reduce gridlock in Thapathali. They collect GPS data from 50,000 taxis via Pathao’s API. Steps:
- Clean data: Remove outliers (e.g., a taxi stuck in a parking lot for 2 hours).
- Aggregate: Group by hour/day/week to find congestion hotspots.
- Visualize:
pie title Traffic Congestion by Time (Thapathali) "7–9 AM" : 35% "5–7 PM" : 40% "Other" : 25% - Insight: 75% of delays occur during rush hours. Action: NTC prioritizes signal timing adjustments for these slots.
2. Predictive Analytics: "What Will Happen?"
Definition: Uses statistical models, machine learning (ML), and data mining to forecast future trends based on historical data. Key Tools: R, Python (Scikit-learn, TensorFlow), Spark MLlib, SAS. Example: NEPSE predicts stock price movements using past trading data + economic indicators.
graph TD
A["Historical Data\n(Stock prices, volume, news sentiment)"] --> B["Feature Engineering\n(Technical indicators: moving avg, RSI)"]
B --> C["Model Training\n(Linear Regression, Random Forest)"]
C --> D["Prediction\n'NEPSE index: +1.2% tomorrow'"]
D --> E["Alert\n(Trader gets SMS: 'Buy 100 shares of NABIL')"]Real-World Tie-In:
- Khalti: Uses predictive analytics to flag fraudulent transactions before they complete (e.g., sudden large transfers from a new device).
- Pathao: Predicts driver demand in a zone to optimize fleet deployment (reduces empty rides by 20%).
Worked Example: Bank Loan Default Prediction (Global IME Bank) Problem: Global IME wants to reduce bad loans. They have data on 10,000 past applicants (income, credit score, loan amount, repayment history). Steps:
- Data Prep: Handle missing values (e.g., impute average income for 5% missing records).
- Feature Selection: Use correlation analysis to pick top predictors (e.g.,
credit_score,loan_to_income_ratio). - Model: Train a Random Forest Classifier (handles non-linear relationships).
- Output:
Applicant Predicted Default Risk Action Ram 85% Reject loan Sita 12% Approve (low APR) Hari 45% Approve (high APR)
Advantages/Disadvantages:
| Pros | Cons |
|---|---|
| Reduces risk (e.g., loan defaults) | Requires clean, high-quality data |
| Automates decision-making | Models degrade over time (retrain needed) |
| Scalable for large datasets | Black-box models lack interpretability |
3. Prescriptive Analytics: "What Should We Do?"
Definition: Goes beyond prediction to recommend optimal actions using optimization algorithms, simulation, and business rules. Key Tools: Python (PuLP, SciPy), Excel Solver, IBM ILOG CPLEX. Example: Daraz’s dynamic pricing engine adjusts product prices in real-time based on demand and competitor prices.
graph LR
A["Current State\n(Inventory: 500 units, Demand forecast: 800)"]
B["Constraints\n(Cost: $10/unit, Max budget: $6,000)"]
C["Objective\n(Maximize profit)"]
A & B & C --> D["Optimization Algorithm\n(Solve for: Order 300 more units)"]
D --> E["Action\n'Place bulk order from supplier X'"]Real-World Tie-In:
- NTC: Uses prescriptive analytics to dynamically adjust electricity tariffs in real-time based on grid load and weather forecasts (reduces blackouts).
- Google Ads: Recommends ad bids and placements to maximize clicks while staying within budget.
Worked Example: Ncell’s Network Tower Placement Problem: Ncell wants to expand 4G coverage in rural Nepal. They have:
- 50 existing towers.
- 200 potential sites (cost: $50K–$200K each).
- Coverage radius: 10 km per tower.
- Budget: $8M.
Steps:
- Define Objective: Maximize coverage (km²) within budget.
- Constraints:
- No tower closer than 8 km to another (interference).
- Terrain data (hills reduce coverage).
- Model: Use Integer Linear Programming (ILP) in Python:
from pulp import LpProblem, LpMaximize, LpVariable, LpBinary model = LpProblem("Ncell_Tower_Placement", LpMaximize) x = LpVariable.dicts("Tower", range(200), cat='Binary') # 1 if placed, 0 else model += lpSum([coverage[i] * x[i] for i in range(200)]) # Maximize coverage for i in range(200): for j in range(i+1, 200): model += x[i] + x[j] <= 1 # No two towers too close model += lpSum([cost[i] * x[i] for i in range(200)]) <= 8_000_000 - Solution: Optimal placement of 32 towers covering 95% of target areas.
Key Analytics Techniques & Tools
| Technique | Tools | Use Case | Example in Nepal |
|---|---|---|---|
| Statistical Analysis | R, Python (StatsModels), Excel | Trend analysis, hypothesis testing | NEPSE analyzing stock volatility |
| Data Mining | Weka, Orange, RapidMiner | Pattern discovery (e.g., customer segmentation) | Daraz identifying high-value shoppers |
| Machine Learning | Scikit-learn, TensorFlow, Spark ML | Classification, regression, clustering | Khalti’s fraud detection (XGBoost) |
| Text Mining | NLTK, spaCy, Apache OpenNLP | Sentiment analysis, topic modeling | NTC analyzing customer complaints on Twitter |
| Network Analysis | Gephi, NetworkX, GraphFrames | Social network mapping, fraud rings | Pathao detecting driver collusion |
| Optimization | PuLP, Gurobi, Excel Solver | Resource allocation, routing | NTC optimizing electricity distribution |
In the Real World
- WhatsApp (Global): Uses predictive analytics to detect spam messages by analyzing sender behavior (e.g., sudden high-volume messages from a new number). Tool: Custom ML models trained on labeled spam/ham datasets.
- YouTube (Global): Employs prescriptive analytics to recommend videos. The system:
- Predicts user interest (collaborative filtering).
- Optimizes watch time (A/B tests video thumbnails).
- Adjusts ad placements in real-time.
- Ncell (Nepal): Deploys descriptive + predictive analytics for:
- Churn prediction: Identifies customers likely to switch to NTC (using RFM analysis: Recency, Frequency, Monetary value).
- Network optimization: Predicts traffic spikes during cricket matches to pre-allocate bandwidth.
Worked Example Tie-In:
- Ncell’s Churn Prediction:
- Data: 1M customers with 12 months of usage data.
- Model: Logistic Regression (predicts probability of churn).
- Output: Top 10% at-risk customers get a discount offer.
- Result: Reduced churn by 15%.
Visualizing Big Data Analytics
1. The Analytics Maturity Model
flowchart TD
A["Descriptive\n'What happened?'"] --> B["Diagnostic\n'Why did it happen?'"]
B --> C["Predictive\n'What will happen?'"]
C --> D["Prescriptive\n'What should we do?'"]
D --> E["Cognitive\n'AI-driven autonomy'"]Caption: Most organizations start at Descriptive (e.g., eSewa reports) and progress to Predictive (e.g., Khalti fraud alerts).
2. Real Analytics Tools in Action
Exam Tip
Focus on the 3 "A"s:
- Descriptive → Aggregation, visualization (SQL, Tableau).
- Predictive → Models (Regression, Decision Trees, Neural Nets).
- Prescriptive → Optimization (ILP, Simulation).
Common Exam Questions:
- Scenario-based: "How would NTC use analytics to reduce power outages?" → Link to predictive (load forecasting) + prescriptive (dynamic pricing).
- Tool comparison: "When would you use R vs. Python for analytics?" → R for stats, Python for scalability (Spark).
- Math: Expect 2–3 marks on linear regression or optimization constraints.
Avoid:
- Vague answers like "Big Data is used everywhere." Specify the technique (e.g., "Khalti uses anomaly detection in prescriptive analytics to block fraud").
- Ignoring constraints in optimization problems (e.g., budget, capacity).
Diagram Shortcut:
- For any analytics workflow, draw a 4-step flow:
- Data Collection (e.g., Ncell’s call logs).
- Processing (cleaning, feature engineering).
- Modeling (tool + technique).
- Action (decision or alert).
- For any analytics workflow, draw a 4-step flow:
Key Formula to Remember:
- Linear Regression:
- R² (Goodness of Fit):
- Optimization Objective: Maximize/Minimize subject to .
Based on the TU BIM syllabus for Big Data and Analytics (IT278), unit 8.
Discussion
Loading…