Business IntelligenceUnit 714 min read
Business Analytics & Predictive Modelling: Techniques, Tools & Applications
Unit 7 of Business Intelligence explores how organizations use historical data to predict future trends, optimize decisions, and automate processes through statistical modeling, machine learning, and prescriptive analytics—with real-world case studies from Nepali and global businesses.
TAKEAWAYS:
- Business analytics (BA) turns raw data into actionable insights using descriptive, predictive, and prescriptive techniques.
- Predictive modeling uses statistical algorithms (regression, classification, clustering) to forecast outcomes like customer churn or sales trends.
- Machine learning (supervised/unsupervised) enables systems to learn patterns from data without explicit programming.
- Tools like Excel, Python (scikit-learn), R, Tableau, and Power BI are essential for implementing BA solutions.
- Nepali companies (e.g., Nabil Bank, Daraz, NTC) use BA for risk assessment, demand forecasting, and fraud detection.
- Ethical challenges (bias, privacy) and model validation are critical for real-world deployment.
1. What Is Business Analytics?
Business Analytics (BA) is the process of transforming data into insights to drive business decisions. It bridges the gap between raw data and strategic actions. BA is categorized into three types:
Key Components:
- Data Collection: Gather structured/unstructured data (e.g., transaction logs, social media, sensors).
- Data Processing: Clean, integrate, and transform data (ETL processes from Unit 4).
- Modeling: Apply statistical/machine learning techniques.
- Visualization: Present insights via dashboards (Unit 8).
- Action: Implement decisions (e.g., pricing adjustments, resource allocation).
2. Predictive Modeling: How It Works
Predictive modeling uses historical data to forecast future events. It relies on:
- Supervised Learning: Models trained on labeled data (e.g., "customer X churned" → predict churn for customer Y).
- Unsupervised Learning: Identifies hidden patterns (e.g., customer segmentation).
- Reinforcement Learning: Optimizes actions via trial-and-error (e.g., dynamic pricing).
Common Techniques:
| Technique | Use Case | Example in Nepal |
|---|---|---|
| Linear Regression | Sales forecasting | Daraz predicting demand for Diwali sales |
| Decision Trees | Credit scoring | Nabil Bank approving loans |
| Clustering (K-Means) | Customer segmentation | NTC grouping subscribers by usage patterns |
| Neural Networks | Fraud detection | Khalti identifying suspicious transactions |
| Time Series Analysis | Stock price prediction | NEPSE forecasting share trends |
WORKED EXAMPLE: Nabil Bank’s Loan Default Prediction Scenario: Nabil Bank wants to predict which loan applicants are likely to default. Data Used:
- Historical loan data (2018–2023): 50,000 records with features like:
- Credit score (numeric)
- Income level (categorical: Low/Medium/High)
- Loan amount (numeric)
- Employment status (categorical)
- Target variable: Default (Yes/No).
Steps:
Data Preprocessing:
- Handle missing values (e.g., impute income for 5% missing records).
- Encode categorical variables (e.g., "Low" → 0, "Medium" → 1, "High" → 2).
- Split data into training (70%) and testing (30%) sets.
Model Selection:
- Train a Logistic Regression model (for binary classification).
- Compare with Random Forest (handles non-linear relationships).
Evaluation:
- Metrics: Accuracy (85%), Precision (88%), Recall (82%).
- Confusion Matrix:
| | Predicted No Default | Predicted Default | |----------------|----------------------|-------------------| | Actual No | 35,000 (True Neg.) | 2,000 (False Pos.)| | Actual Default | 1,500 (False Neg.) | 1,500 (True Pos.) |
Deployment:
- Integrate model into Nabil Bank’s loan approval system.
- Flag high-risk applicants for manual review.
Why This Matters:
- Reduces default rates by 20% (saving ₹50M/year).
- Automates 60% of loan decisions, cutting processing time by 40%.
3. Machine Learning in Business Analytics
Machine learning (ML) enables systems to learn from data without explicit programming. Key algorithms:
A. Supervised Learning (Labeled Data)
Example: Pathao’s Ride Demand Prediction
- Problem: Predict peak hours for driver shortages in Kathmandu.
- Data: 1M ride requests with features:
- Time of day, day of week, weather, location (latitude/longitude).
- Model: Gradient Boosting (XGBoost) trained on 6 months of data.
- Output: Predicts demand spikes with 92% accuracy, helping Pathao deploy drivers efficiently.
B. Unsupervised Learning (Unlabeled Data)
Example: Daraz’s Customer Segmentation
- Problem: Group customers by purchasing behavior to personalize marketing.
- Data: 100K users with features:
- Purchase frequency, average order value, product categories bought.
- Model: K-Means Clustering (K=5) identifies:
- High-value buyers (20% of users, 60% of revenue).
- Bargain hunters (30% of users, low AOV).
- Occasional shoppers (50% of users, sporadic purchases).
- Action: Daraz sends discount coupons to bargain hunters and VIP offers to high-value buyers, increasing repeat purchases by 15%.
4. Prescriptive Analytics: From "What If" to "What Should We Do"
While predictive analytics answers "What will happen?", prescriptive analytics answers "What should we do?" using optimization techniques.
Example: NTC’s Network Optimization
- Problem: Reduce network congestion during festivals (e.g., Dashain).
- Data: Historical call volumes, tower locations, user mobility patterns.
- Tools:
- Linear Programming to allocate bandwidth.
- Simulation to test "what-if" scenarios (e.g., adding 10 towers in Thapathali).
- Outcome:
- Reduced call drops by 35% during peak hours.
- Saved ₹200M in infrastructure costs by optimizing tower placements.
5. Tools for Business Analytics
| Tool Category | Tools | Use Case |
|---|---|---|
| Statistical Software | R, Python (Pandas, SciKit-Learn) | Model building, hypothesis testing |
| BI Dashboards | Tableau, Power BI, Qlik | Visualizing KPIs (e.g., sales trends) |
| ETL/Big Data | SQL, Apache Spark, Talend | Data cleaning and integration |
| Cloud Platforms | AWS SageMaker, Google Vertex AI | Scalable ML model deployment |
| Open-Source | TensorFlow, Keras | Deep learning for complex patterns |
Example: eSewa’s Fraud Detection
- Tools: Python (scikit-learn), AWS Lambda.
- Process:
- Collect transaction data (amount, time, location).
- Train an Isolation Forest model to detect anomalies.
- Flag suspicious transactions in real-time (e.g., ₹50K transfer at 3 AM).
- Impact: Reduced fraud losses by 40% in 2023.
6. Challenges and Ethical Considerations
| Challenge | Example in Nepal | Mitigation Strategy |
|---|---|---|
| Data Privacy | Ncell collecting location data without consent | Comply with PDPA (2018); anonymize data |
| Bias in Models | Bank loan models favoring urban over rural applicants | Audit models for fairness; diversify training data |
| Overfitting | A model predicting NEPSE shares with 100% training accuracy but fails in testing | Use cross-validation; simplify models |
| Interpretability | Black-box ML models (e.g., deep learning) | Use SHAP values to explain predictions |
Case Study: Himalayan Java’s Supply Chain Analytics
- Problem: Predict coffee bean shortages due to weather.
- Challenge: Limited historical weather data for remote farms.
- Solution:
- Partnered with IMD (India Meteorological Department) for satellite data.
- Used ensemble models (combining regression + time series) to forecast yields.
- Outcome: Reduced stockouts by 25% and optimized procurement costs.
7. Real-World Applications in Nepal
| Company | BA Application | Impact |
|---|---|---|
| Nabil Bank | Loan default prediction (Logistic Regression) | 20% reduction in defaults |
| Daraz | Demand forecasting (XGBoost) | 15% increase in sales during Diwali |
| NTC | Network congestion optimization | 35% fewer call drops during festivals |
| Pathao | Ride demand prediction (Gradient Boosting) | 20% reduction in driver shortages |
| eSewa | Fraud detection (Isolation Forest) | 40% lower fraud losses |
| NEPSE | Stock price prediction (ARIMA) | Helps traders make informed decisions |
8. Step-by-Step: Building a Predictive Model (Case Study)
Scenario: Predicting Khalti’s transaction fraud using Python.
Step 1: Data Collection
- Features:
- Transaction amount, time, location, device type, user history.
- Target:
is_fraud(1 = fraud, 0 = legitimate).
Step 2: Exploratory Data Analysis (EDA)
import pandas as pd
import matplotlib.pyplot as plt
# Load data
data = pd.read_csv("khalti_transactions.csv")
# Check fraud distribution
plt.figure(figsize=(8, 5))
data["is_fraud"].value_counts().plot(kind="bar")
plt.title("Fraud vs Legitimate Transactions")
Output:
Fraud: 2% of transactions
Legitimate: 98% of transactions
→ Problem: Imbalanced dataset (fraud is rare). Solution: Use SMOTE (Synthetic Minority Oversampling).
Step 3: Model Training
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
# Split data
X = data.drop("is_fraud", axis=1)
y = data["is_fraud"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3)
# Train model
model = RandomForestClassifier(n_estimators=100)
model.fit(X_train, y_train)
Step 4: Evaluation
from sklearn.metrics import classification_report
y_pred = model.predict(X_test)
print(classification_report(y_test, y_pred))
Output:
precision recall f1-score support
0 0.99 1.00 0.99 2970
1 0.85 0.75 0.80 58
→ Recall for fraud (0.75) means the model catches 75% of fraudulent transactions.
Step 5: Deployment
- Deploy model as an API using Flask/FastAPI.
- Integrate with Khalti’s backend to flag transactions in real-time.
9. Exam Tips for TU/PU/NEB
Understand the Difference:
- Descriptive: "What happened?" (e.g., sales reports).
- Predictive: "What will happen?" (e.g., regression models).
- Prescriptive: "What should we do?" (e.g., optimization).
Common Exam Questions:
- Short Answer:
- Define "business analytics." (1 mark)
- Difference between supervised and unsupervised learning. (2 marks)
- Long Answer (5–10 marks):
- Explain how Nabil Bank can use predictive modeling to reduce loan defaults. (Include steps: data collection → model training → evaluation → deployment.)
- Compare decision trees and neural networks for classification. (Use a table with pros/cons.)
- Case Study (10–15 marks):
- Given a dataset (e.g., customer churn), build a predictive model in steps. Assume you have Excel/Python.
- Critique a real-world BI failure (e.g., Facebook’s Cambridge Analytica scandal).
- Short Answer:
Visuals Are Key:
- Always draw flowcharts for processes (e.g., BA workflow).
- Use tables to compare techniques (e.g., regression vs. classification).
- Label diagrams clearly (e.g., "K-Means Clustering: 3 Customer Segments").
Practical Tips:
- Memorize tools: Python (scikit-learn), R, Tableau, SQL.
- Know metrics: Accuracy, precision, recall, F1-score.
- Relate to Nepal: Use examples from banks, e-commerce, telecom, or NEPSE.
Avoid Common Mistakes:
- ❌ Confusing correlation (e.g., ice cream sales vs. drowning) with causation.
- ❌ Ignoring data preprocessing (e.g., not handling missing values).
- ❌ Overcomplicating models (e.g., using deep learning when linear regression suffices).
10. Quick Revision Checklist
Before the exam, ensure you can: ✅ Define business analytics, predictive modeling, and prescriptive analytics. ✅ List 3 supervised and 2 unsupervised learning algorithms. ✅ Explain how Nabil Bank or Daraz uses BA in their operations. ✅ Draw a flowchart of the BA process. ✅ Compare 2 techniques (e.g., regression vs. decision trees) in a table. ✅ Describe one ethical challenge in BA (e.g., bias, privacy). ✅ Write Python/R pseudocode for a simple predictive model.
Based on the TU BIM syllabus for Business Intelligence (IT249), unit 7.
Discussion
Loading…