Machine LearningUnit 710 min read
Ensemble Methods & Advanced ML Topics: Boosting, Bagging, Stacking, and Model Optimization
Unit 7 of Machine Learning: Explores advanced techniques like ensemble methods (bagging, boosting, stacking), model optimization, hyperparameter tuning, and real-world applications of ML systems in Nepal and globally.
TAKEAWAYS:
- Ensemble methods combine multiple models to improve accuracy, with bagging (e.g., Random Forest) reducing variance and boosting (e.g., AdaBoost, XGBoost) reducing bias.
- Stacking hierarchically combines predictions from diverse models, while boosting iteratively corrects errors by weighting misclassified samples.
- Hyperparameter tuning (e.g., grid search, random search) optimizes model performance, with cross-validation ensuring robust evaluation.
- Advanced topics like ensemble pruning, neural architecture search, and explainable AI address scalability and interpretability in real-world systems.
- Nepalese apps like Daraz (recommendation systems) and Ncell (fraud detection) use ensemble methods for reliability.
- Exam focus: Compare bagging vs. boosting, explain confusion matrix (from Unit 2), and apply hyperparameter tuning to a given dataset.
1. Ensemble Methods: Combining Models for Better Performance
Ensemble methods leverage multiple models to improve generalization, reduce overfitting, and enhance predictive accuracy. They are widely used in Nepal’s Ncell fraud detection (combining isolation forests and logistic regression) and Daraz’s recommendation engines (ensemble of collaborative filtering and deep learning).
1.1 Bagging (Bootstrap Aggregating)
Bagging reduces variance by training multiple models on random subsets of data and averaging their predictions. The most popular bagging algorithm is Random Forest.
How it works:
- Bootstrap sampling: Create multiple datasets by sampling with replacement from the original dataset.
- Train models: Fit a base model (e.g., decision tree) on each subset.
- Aggregate predictions: Combine predictions via voting (classification) or averaging (regression).
Example: Random Forest for Loan Approval Suppose a bank (e.g., NMB) uses a Random Forest to approve loans. The dataset includes features like income, credit score, and loan amount. Each tree in the forest makes a prediction, and the final decision is based on majority voting.
graph TD
A["Original Dataset"] --> B["Bootstrap Sample 1"]
A --> C["Bootstrap Sample 2"]
B --> D["Decision Tree 1"]
C --> E["Decision Tree 2"]
D --> F["Predict: Approve"]
E --> G["Predict: Reject"]
F & G --> H["Majority Vote: Approve"]Advantages:
- Reduces overfitting.
- Handles high-dimensional data well.
- Works with noisy data.
Disadvantages:
- Computationally expensive.
- Less interpretable than single models.
1.2 Boosting: Sequential Error Correction
Boosting reduces bias by iteratively training models on weighted samples, focusing on misclassified instances. Popular algorithms include AdaBoost, XGBoost, and LightGBM.
How it works:
- Weight initialization: Assign equal weights to all samples.
- Train model: Fit a weak learner (e.g., decision stump) on weighted data.
- Update weights: Increase weights of misclassified samples.
- Repeat: Train subsequent models on updated weights and combine predictions.
Example: XGBoost for Traffic Prediction (Kathmandu) Suppose NTC uses XGBoost to predict traffic congestion. The model starts with uniform weights, then focuses more on areas with historical delays, improving accuracy over time.
graph TD
A["Initial Dataset"] --> B["Train Weak Learner 1: Misclassified Samples"]
B --> C["Update Weights: Focus on Errors"]
C --> D["Train Weak Learner 2: Corrects Errors"]
D --> E["Combine Predictions: Weighted Average"]
E --> F["Final Prediction: Improved Accuracy"]
C -->|"Weight Update"| DAdvantages:
- High predictive accuracy.
- Handles imbalanced datasets well.
- Works well with structured data.
Disadvantages:
- Prone to overfitting if not tuned properly.
- Computationally intensive for large datasets.
Comparison Table: Bagging vs. Boosting
| Feature | Bagging (Random Forest) | Boosting (AdaBoost/XGBoost) |
|---|---|---|
| Goal | Reduce variance | Reduce bias |
| Training | Parallel | Sequential |
| Error Handling | Averages predictions | Focuses on misclassified samples |
| Use Case | High variance, noisy data | Structured data, imbalanced classes |
| Example | Ncell fraud detection | Daraz recommendation system |
1.3 Stacking: Hierarchical Model Combination
Stacking combines predictions from multiple models using a meta-learner (e.g., logistic regression or neural network). It is more flexible than bagging or boosting but computationally expensive.
How it works:
- Base models: Train diverse models (e.g., SVM, decision tree, neural network).
- Meta-learner: Use their predictions as input to a second-level model.
Example: Stacking for NEPSE Stock Prediction Suppose an investor uses stacking to predict stock prices. The base models could be:
- Linear regression (trend analysis).
- SVM (classification of price movements).
- Neural network (pattern recognition).
The meta-learner (e.g., logistic regression) combines these predictions for the final forecast.
graph TD
A["Base Model 1: Linear Regression"] --> B["Meta-Learner: Logistic Regression"]
C["Base Model 2: SVM"] --> B
D["Base Model 3: Neural Network"] --> B
B --> E["Final Prediction: Optimized Decision"]Advantages:
- High flexibility and accuracy.
- Can combine diverse model types.
Disadvantages:
- Complex to implement.
- Risk of overfitting if not regularized.
2. Hyperparameter Tuning and Model Optimization
Hyperparameters (e.g., learning rate, tree depth) are not learned from data but set before training. Tuning them improves model performance.
2.1 Grid Search vs. Random Search
- Grid search: Exhaustively tests all combinations of hyperparameters.
- Random search: Samples hyperparameters randomly, often more efficient.
Example: Tuning XGBoost for Loan Default Prediction Suppose a bank uses XGBoost to predict loan defaults. The hyperparameters to tune might include:
max_depth: Tree depth (e.g., 3, 5, 7).learning_rate: Step size (e.g., 0.01, 0.1, 0.2).n_estimators: Number of trees (e.g., 50, 100, 200).
Using random search, the algorithm might test:
max_depth=5,learning_rate=0.1,n_estimators=100.max_depth=7,learning_rate=0.01,n_estimators=200.
The best combination is selected based on cross-validation performance.
2.2 Cross-Validation
Cross-validation splits data into folds to evaluate model robustness. Common methods:
- k-fold CV: Splits data into
kfolds, trains onk-1folds, validates on the remaining fold. - Stratified CV: Ensures class balance in each fold (useful for imbalanced datasets).
Example: 5-Fold CV for Daraz Product Recommendations Suppose Daraz uses a Random Forest to recommend products. The dataset is split into 5 folds:
- Train on folds 1-4, validate on fold 5.
- Train on folds 2-5, validate on fold 1.
- Repeat for all folds. The average performance across folds gives a robust estimate.
3. Advanced Topics in Machine Learning
3.1 Ensemble Pruning
Pruning removes weak or redundant models from an ensemble to improve efficiency without sacrificing accuracy. For example, in a Random Forest, pruning can remove trees that contribute little to the final prediction.
3.2 Neural Architecture Search (NAS)
NAS automates the design of neural networks by searching for optimal architectures (e.g., layer sizes, connections). Companies like Google use NAS to optimize models for tasks like image recognition.
3.3 Explainable AI (XAI)
XAI techniques (e.g., SHAP values, LIME) make ML models interpretable. For example, Pathao might use XAI to explain why a ride price is recommended, ensuring transparency for users.
4. Real-World Applications in Nepal
4.1 Daraz: Recommendation Systems
Daraz uses ensemble methods to combine collaborative filtering (user-item interactions) and content-based filtering (product features) to recommend products. Boosting helps correct biases in user preferences.
4.2 Ncell: Fraud Detection
Ncell employs ensemble methods (e.g., isolation forests + logistic regression) to detect fraudulent transactions. Bagging reduces false positives, while boosting focuses on high-risk patterns.
4.3 NEPSE: Stock Market Prediction
Investors use stacking to combine technical analysis (e.g., moving averages) with fundamental data (e.g., P/E ratios) to predict stock trends. Hyperparameter tuning ensures the model adapts to market volatility.
5. Worked Example: Confusion Matrix and Hyperparameter Tuning
Problem: Given a dataset of loan approvals (approved/rejected), fit a model and tune hyperparameters to maximize accuracy.
Step 1: Confusion Matrix Assume a classifier predicts loan approvals with the following results:
- True Positives (TP): 80 (correctly approved).
- False Positives (FP): 10 (incorrectly approved).
- False Negatives (FN): 5 (incorrectly rejected).
- True Negatives (TN): 75 (correctly rejected).
The confusion matrix:
| Predicted: Approve | Predicted: Reject
----------|-------------------|-------------------
Actual: Approve | 80 (TP) | 5 (FN)
Actual: Reject | 10 (FP) | 75 (TN)
Step 2: Hyperparameter Tuning with Random Forest
Tune max_depth and n_estimators using 5-fold cross-validation:
max_depth=5,n_estimators=100→ Accuracy: 92%max_depth=7,n_estimators=200→ Accuracy: 94%max_depth=3,n_estimators=50→ Accuracy: 88%
The best combination is max_depth=7, n_estimators=200.
6. Exam Tips
- Compare bagging and boosting: Know their goals (variance vs. bias reduction), training methods (parallel vs. sequential), and real-world examples.
- Confusion matrix: Understand TP, FP, FN, TN, and metrics like precision, recall, and F1-score.
- Hyperparameter tuning: Explain grid search vs. random search and the role of cross-validation.
- Ensemble pruning: Mention how it improves efficiency.
- Real-world tie-ins: Relate to Nepalese apps (e.g., Daraz, Ncell) or global examples (e.g., Google’s NAS).
- Worked examples: Practice tuning hyperparameters on small datasets and interpreting confusion matrices.
Based on the TU BCA syllabus for Machine Learning (CACS486), unit 7.
Discussion
Loading…