Business IntelligenceUnit 78 min read
Business Analytics & Predictive Modelling: Techniques, Tools & Applications
Unit 7 of Business Intelligence explores how organizations use historical data to predict trends, optimize decisions, and automate actions—covering descriptive, predictive, and prescriptive analytics, machine learning models (regression, classification, clustering), and real-world case studies from Nepalese and global
Core Concepts: What is Business Analytics?
Business Analytics (BA) transforms raw data into actionable insights through statistical analysis, modeling, and optimization. It bridges the gap between data and decision-making by answering three key questions:
- Descriptive Analytics: "What happened?" (e.g., sales trends, customer behavior)
- Predictive Analytics: "What will happen?" (e.g., demand forecasting, churn prediction)
- Prescriptive Analytics: "What should we do?" (e.g., dynamic pricing, route optimization)
mindmap
root((Business Analytics))
Descriptive
"Summarizes past data (e.g., reports, dashboards)"
Tools: SQL, Excel, Tableau
Predictive
"Uses ML to forecast outcomes (e.g., 'Will this customer buy again?')"
Models: Regression, Decision Trees, Neural Networks
Prescriptive
"Recommends optimal actions (e.g., 'Set price at $X to maximize profit')"
Techniques: Optimization, Simulation, AIHow It Works: The Analytics Pipeline
- Data Collection: Gather structured/unstructured data (e.g., transaction logs, social media, sensors).
- Data Cleaning: Handle missing values, outliers, and noise (e.g., using Python’s
pandasor SQL’sCASE WHEN). - Model Training: Apply algorithms (e.g., linear regression for sales forecasting).
- Validation: Test accuracy (e.g., confusion matrix for classification).
- Deployment: Integrate into business processes (e.g., automated alerts for fraud).
Predictive Modelling: Algorithms and Applications
Predictive models use historical data to forecast future events. Key techniques include:
1. Regression Analysis
Definition: Predicts continuous outcomes (e.g., sales, temperature) using linear/non-linear relationships. Example: Nepal Electricity Authority (NEA) uses regression to forecast electricity demand based on weather and past usage.
Worked Example: Daraz’s Inventory Forecasting
- Problem: Overstocking leads to waste; understocking loses sales.
- Solution: Daraz trains a multiple linear regression model on:
- Historical sales data (dependent variable:
units_sold) - Independent variables:
season,promotions,competitor_prices,holidays.
- Historical sales data (dependent variable:
- Outcome: Reduces excess inventory by 15% by adjusting orders dynamically.
| Metric | Before Analytics | After Analytics |
|---|---|---|
| Inventory Turnover | 4 times/year | 6 times/year |
| Waste Reduction | 8% | 22% |
| Customer Satisfaction | 78% | 89% |
2. Classification Models
Definition: Predicts categorical outcomes (e.g., "Will a loan default?" or "Is this email spam?"). Algorithms:
- Decision Trees: Simple, interpretable (e.g., "If income > $50k AND credit_score > 700 → Approve Loan").
- Random Forest: Ensemble of trees for higher accuracy.
- Logistic Regression: Probabilistic classification (e.g., "85% chance of churn").
Real-World Use: Nabil Bank’s Loan Approval System
- Problem: Manual loan approvals were slow and biased.
- Solution: Nabil Bank built a Random Forest classifier trained on:
- Applicant data:
income,employment_history,credit_score,loan_amount. - Outcome:
default(1) orno_default(0).
- Applicant data:
- Result: Approval time reduced from 3 days → 2 hours; default rate dropped by 12%.
3. Clustering (Unsupervised Learning)
Definition: Groups similar data points without predefined labels (e.g., customer segmentation). Algorithms:
- K-Means: Partitions data into K clusters (e.g., "Find 4 customer segments").
- Hierarchical Clustering: Builds a tree of clusters (dendrogram).
Example: Pathao’s Driver Optimization
- Problem: Drivers in Kathmandu had uneven earnings due to route imbalances.
- Solution: Pathao used K-Means clustering to group:
- Cluster 1: High-demand zones (e.g., Thapathali, Lakshmi Marg).
- Cluster 2: Low-demand zones (e.g., rural areas).
- Action: Dynamically adjusted surge pricing and driver incentives.
- Impact: Driver earnings increased by 20% in high-demand clusters.
4. Time-Series Forecasting
Definition: Predicts future values based on past trends (e.g., stock prices, weather). Models:
- ARIMA: AutoRegressive Integrated Moving Average (e.g., "Forecast NEPSE index").
- Prophet (Facebook): Handles seasonality (e.g., "Predict eSewa transactions during Dashain").
Case Study: eSewa’s Holiday Traffic Prediction
- Problem: Server crashes during Dashain due to unexpected user spikes.
- Solution: eSewa used ARIMA to forecast:
- Dependent Variable:
daily_transactions. - Independent Variables:
day_of_week,holiday_flag,promotions.
- Dependent Variable:
- Outcome: Scaled servers dynamically, reducing downtime by 90%.
Business Analytics in Nepal: Case Studies
| Company | Analytics Technique | Business Impact | Tools Used |
|---|---|---|---|
| Ncell | Churn Prediction (Logistic Regression) | Reduced customer loss by 18% by targeting at-risk users with discounts. | Python (scikit-learn), SQL |
| Daraz | Demand Forecasting (ARIMA) | Cut warehouse costs by 25% via optimized inventory. | R, Tableau |
| Nabil Bank | Fraud Detection (Random Forest) | Blocked $5M/year in fraudulent transactions. | SAS, Python |
| NEA | Energy Demand Prediction (Regression) | Balanced grid load, reducing blackouts by 40%. | MATLAB, Excel |
| Pathao | Driver Clustering (K-Means) | Increased driver earnings by 20% in high-demand zones. | Google BigQuery, Python |
Exam Tip: How to Score Full Marks
- Define Clearly: Always start with definitions (e.g., "Predictive modelling uses historical data to forecast future probabilities using algorithms like regression or decision trees.").
- Use Real Examples: Examiners love Nepalese cases. Link models to companies (e.g., "Like Nabil Bank’s loan approval system, a retail firm could use Random Forest to...").
- Show Maths Where Needed:
- For regression, write the equation and interpret coefficients.
- For classification, draw a confusion matrix and calculate accuracy.
- Compare Models: Use tables to contrast techniques (e.g., "Decision Trees vs. Random Forest").
- Visualize: Sketch a decision tree, time-series graph, or cluster plot in your answer. Even a rough diagram earns partial credit.
- Link to Business Value: End each answer with an impact statement (e.g., "This reduces costs by X% and improves Y by Z%").
Common Pitfalls to Avoid:
- Confusing descriptive (what happened) with predictive (what will happen) analytics.
- Forgetting to validate models (always mention "train-test split" or "cross-validation").
- Overcomplicating answers—stick to 1–2 models per question unless asked for a comparison.
Based on the TU BITM syllabus for Business Intelligence (IT249), unit 7.
Discussion
Loading…