Data Warehousing and Data MiningUnit 87 min read
Prediction & Regression: Models, Algorithms & Applications
Unit 8 of Data Warehousing and Data Mining explores prediction vs. regression, linear/logistic regression mechanics, evaluation metrics (RMSE, R²), and real-world applications in finance, healthcare, and e-commerce—with worked examples using NEPSE stock prices, Daraz delivery delays, and bank loan approvals.
Core Concepts
1. Prediction vs. Regression: Key Differences
Prediction and regression are both supervised learning techniques, but they differ in their goals and outputs:
| Aspect | Prediction | Regression |
|---|---|---|
| Output Type | Discrete (categorical) values | Continuous (numerical) values |
| Example Tasks | Spam detection, loan approval | House price estimation, stock trends |
| Algorithms | Decision Trees, Naïve Bayes, SVM | Linear Regression, Polynomial Regression, Ridge/Lasso |
Why it matters:
- Prediction answers "What category does this belong to?"
- Regression answers "What numerical value does this correspond to?"
2. Linear Regression: The Workhorse of Prediction
Linear regression models the relationship between a dependent variable (Y) and one or more independent variables (X) using a straight-line equation: where:
- = intercept
- = coefficients
- = error term
How It Works: Step-by-Step
- Data Collection: Gather historical data (e.g., NEPSE stock prices vs. inflation rates).
- Model Training: Use Ordinary Least Squares (OLS) to minimize the sum of squared errors (SSE).
- Equation Derivation: Solve for coefficients using calculus or matrix algebra.
- Prediction: Plug new values into the equation to estimate .
Worked Example: Predicting NEPSE Index
Suppose we model the NEPSE index () based on monthly inflation rate () and global oil prices (). Given data:
| Month | Inflation (%) | Oil Price ($/barrel) | NEPSE Index |
|---|---|---|---|
| Jan | 4.2 | 65 | 1850 |
| Feb | 4.5 | 68 | 1880 |
| Mar | 4.8 | 70 | 1920 |
Step 1: Fit the model (using software or manual calculation): Step 2: Predict April’s NEPSE index if inflation = 5.0% and oil = $72: Visualization of Fit:
graph LR
A["Actual NEPSE (1850, 1880, 1920)"] --> B["Linear Regression Line"]
B --> C["Predicted: 1950 for April"]3. Logistic Regression: For Binary Outcomes
While linear regression predicts continuous values, logistic regression predicts probabilities (0 to 1) for binary outcomes (e.g., loan approval: Yes/No).
Sigmoid Function
The output is squashed using the logistic function: where = probability of the positive class (e.g., "approved").
Worked Example: Bank Loan Approval
A bank uses credit score () and income () to predict loan approval. Given:
- If and , the model outputs: Decision Rule: Approve if .
Visualization of Sigmoid Curve:
graph TD
A["Linear Combination: Z = β₀ + β₁X"] --> B["Sigmoid: P(Y=1) = 1/(1+e⁻ᶻ)"]
B --> C["Output: 0 (No) to 1 (Yes)"]4. Evaluation Metrics
| Metric | Formula | When to Use |
|---|---|---|
| RMSE | Regression (lower = better) | |
| R² (R-squared) | Explains variance (0 to 1, higher = better) | |
| Accuracy | Classification (binary) | |
| AUC-ROC | Area under ROC curve | Class imbalance (e.g., fraud detection) |
Example: If a model predicts house prices with RMSE = $15,000, it’s off by $15K on average.
In the Real World
NEPSE Stock Predictions
- Company: NEPSE (Nepal Stock Exchange)
- Idea Used: Linear Regression
- How: Analysts use historical data (inflation, GDP growth) to predict future index movements. For example, a model trained on 2018–2022 data might forecast a 5% rise in 2024 if inflation stays below 6%.
Daraz Delivery Time Estimation
- Company: Daraz (Alibaba Group)
- Idea Used: Polynomial Regression
- How: Daraz estimates delivery delays based on:
- Distance from warehouse ()
- Number of orders in queue ()
- Weather conditions ()
- Example: If , , and rain = "Yes," the model predicts a 48-hour delay.
Khalti Fraud Detection
- Company: Khalti (Nepal’s fintech)
- Idea Used: Logistic Regression
- How: Khalti flags suspicious transactions by calculating:
- Example: A $500 transfer from Kathmandu to Pokhara at 3 AM might trigger .
Advanced Regression Techniques
1. Polynomial Regression
Extends linear regression by adding non-linear terms (e.g., , ): When to Use: When the relationship between and is curved (e.g., Daraz delivery times vs. distance).
Example:
graph LR
A["Distance (km)"] --> B["Delivery Time (hours)"]
B --> C["Quadratic Fit: Y = 2 + 0.5X + 0.01X²"]2. Regularization: Ridge vs. Lasso
| Method | Penalty Term | Use Case |
|---|---|---|
| Ridge | Multicollinearity (many correlated features) | |
| Lasso | Feature selection (sparse models) |
Example: Predicting house prices with 50 features (e.g., room count, school proximity). Lasso might zero out irrelevant features like "number of mailboxes."
Exam Tip
- Always compare prediction vs. regression in questions asking for "supervised learning techniques."
- For linear regression:
- Know the equation and how to interpret coefficients.
- Practice manual calculations (even with small datasets).
- Remember: R² = 1 means perfect fit; RMSE = 0 means no error.
- For logistic regression:
- Draw the sigmoid curve and explain how it maps linear output to probabilities.
- Know the decision threshold (usually 0.5).
- Real-world applications:
- Link NEPSE to regression, Khalti to classification, and Daraz to polynomial regression.
- Expect worked examples with small datasets (3–5 rows).
- Avoid common mistakes:
- Don’t confuse regression (continuous) with classification (discrete).
- Don’t assume linear regression works for non-linear data (use polynomial/non-linear models instead).
Visual Summary of Key Models
mindmap
root((Prediction & Regression))
Linear Regression
Equation: Y = β₀ + β₁X
Goal: Minimize SSE
Logistic Regression
Equation: P(Y=1) = 1/(1+e⁻ᶻ)
Goal: Binary classification
Polynomial Regression
Extends to X², X³
For curved relationships
Regularization
Ridge: L2 penalty
Lasso: L1 penaltyBased on the TU BITM syllabus for Data Warehousing and Data Mining (IT274), unit 8.
Discussion
Loading…