CACS486 Machine Learning

Machine LearningUnit 710 min read

Ensemble Methods & Advanced ML Topics: Boosting, Bagging, Stacking, and Model Optimization

Unit 7 of Machine Learning: Explores advanced techniques like ensemble methods (bagging, boosting, stacking), model optimization, hyperparameter tuning, and real-world applications of ML systems in Nepal and globally.

TAKEAWAYS:

  • Ensemble methods combine multiple models to improve accuracy, with bagging (e.g., Random Forest) reducing variance and boosting (e.g., AdaBoost, XGBoost) reducing bias.
  • Stacking hierarchically combines predictions from diverse models, while boosting iteratively corrects errors by weighting misclassified samples.
  • Hyperparameter tuning (e.g., grid search, random search) optimizes model performance, with cross-validation ensuring robust evaluation.
  • Advanced topics like ensemble pruning, neural architecture search, and explainable AI address scalability and interpretability in real-world systems.
  • Nepalese apps like Daraz (recommendation systems) and Ncell (fraud detection) use ensemble methods for reliability.
  • Exam focus: Compare bagging vs. boosting, explain confusion matrix (from Unit 2), and apply hyperparameter tuning to a given dataset.

1. Ensemble Methods: Combining Models for Better Performance

Ensemble methods leverage multiple models to improve generalization, reduce overfitting, and enhance predictive accuracy. They are widely used in Nepal’s Ncell fraud detection (combining isolation forests and logistic regression) and Daraz’s recommendation engines (ensemble of collaborative filtering and deep learning).

1.1 Bagging (Bootstrap Aggregating)

Bagging reduces variance by training multiple models on random subsets of data and averaging their predictions. The most popular bagging algorithm is Random Forest.

How it works:

  1. Bootstrap sampling: Create multiple datasets by sampling with replacement from the original dataset.
  2. Train models: Fit a base model (e.g., decision tree) on each subset.
  3. Aggregate predictions: Combine predictions via voting (classification) or averaging (regression).

Example: Random Forest for Loan Approval Suppose a bank (e.g., NMB) uses a Random Forest to approve loans. The dataset includes features like income, credit score, and loan amount. Each tree in the forest makes a prediction, and the final decision is based on majority voting.

graph TD
    A["Original Dataset"] --> B["Bootstrap Sample 1"]
    A --> C["Bootstrap Sample 2"]
    B --> D["Decision Tree 1"]
    C --> E["Decision Tree 2"]
    D --> F["Predict: Approve"]
    E --> G["Predict: Reject"]
    F & G --> H["Majority Vote: Approve"]

Advantages:

  • Reduces overfitting.
  • Handles high-dimensional data well.
  • Works with noisy data.

Disadvantages:

  • Computationally expensive.
  • Less interpretable than single models.

1.2 Boosting: Sequential Error Correction

Boosting reduces bias by iteratively training models on weighted samples, focusing on misclassified instances. Popular algorithms include AdaBoost, XGBoost, and LightGBM.

How it works:

  1. Weight initialization: Assign equal weights to all samples.
  2. Train model: Fit a weak learner (e.g., decision stump) on weighted data.
  3. Update weights: Increase weights of misclassified samples.
  4. Repeat: Train subsequent models on updated weights and combine predictions.

Example: XGBoost for Traffic Prediction (Kathmandu) Suppose NTC uses XGBoost to predict traffic congestion. The model starts with uniform weights, then focuses more on areas with historical delays, improving accuracy over time.

graph TD
    A["Initial Dataset"] --> B["Train Weak Learner 1: Misclassified Samples"]
    B --> C["Update Weights: Focus on Errors"]
    C --> D["Train Weak Learner 2: Corrects Errors"]
    D --> E["Combine Predictions: Weighted Average"]
    E --> F["Final Prediction: Improved Accuracy"]
    C -->|"Weight Update"| D

Advantages:

  • High predictive accuracy.
  • Handles imbalanced datasets well.
  • Works well with structured data.

Disadvantages:

  • Prone to overfitting if not tuned properly.
  • Computationally intensive for large datasets.

Comparison Table: Bagging vs. Boosting

Feature Bagging (Random Forest) Boosting (AdaBoost/XGBoost)
Goal Reduce variance Reduce bias
Training Parallel Sequential
Error Handling Averages predictions Focuses on misclassified samples
Use Case High variance, noisy data Structured data, imbalanced classes
Example Ncell fraud detection Daraz recommendation system

1.3 Stacking: Hierarchical Model Combination

Stacking combines predictions from multiple models using a meta-learner (e.g., logistic regression or neural network). It is more flexible than bagging or boosting but computationally expensive.

How it works:

  1. Base models: Train diverse models (e.g., SVM, decision tree, neural network).
  2. Meta-learner: Use their predictions as input to a second-level model.

Example: Stacking for NEPSE Stock Prediction Suppose an investor uses stacking to predict stock prices. The base models could be:

  • Linear regression (trend analysis).
  • SVM (classification of price movements).
  • Neural network (pattern recognition).

The meta-learner (e.g., logistic regression) combines these predictions for the final forecast.

graph TD
    A["Base Model 1: Linear Regression"] --> B["Meta-Learner: Logistic Regression"]
    C["Base Model 2: SVM"] --> B
    D["Base Model 3: Neural Network"] --> B
    B --> E["Final Prediction: Optimized Decision"]

Advantages:

  • High flexibility and accuracy.
  • Can combine diverse model types.

Disadvantages:

  • Complex to implement.
  • Risk of overfitting if not regularized.

2. Hyperparameter Tuning and Model Optimization

Hyperparameters (e.g., learning rate, tree depth) are not learned from data but set before training. Tuning them improves model performance.

  • Grid search: Exhaustively tests all combinations of hyperparameters.
  • Random search: Samples hyperparameters randomly, often more efficient.
0255075100Grid Search100Random Search50Computational Cost (Relative Units)
Comparison of computational efficiency between exhaustive grid search and random search for hyperparameter tuning.

Example: Tuning XGBoost for Loan Default Prediction Suppose a bank uses XGBoost to predict loan defaults. The hyperparameters to tune might include:

  • max_depth: Tree depth (e.g., 3, 5, 7).
  • learning_rate: Step size (e.g., 0.01, 0.1, 0.2).
  • n_estimators: Number of trees (e.g., 50, 100, 200).

Using random search, the algorithm might test:

  • max_depth=5, learning_rate=0.1, n_estimators=100.
  • max_depth=7, learning_rate=0.01, n_estimators=200.

The best combination is selected based on cross-validation performance.

2.2 Cross-Validation

Cross-validation splits data into folds to evaluate model robustness. Common methods:

  • k-fold CV: Splits data into k folds, trains on k-1 folds, validates on the remaining fold.
  • Stratified CV: Ensures class balance in each fold (useful for imbalanced datasets).

Example: 5-Fold CV for Daraz Product Recommendations Suppose Daraz uses a Random Forest to recommend products. The dataset is split into 5 folds:

  1. Train on folds 1-4, validate on fold 5.
  2. Train on folds 2-5, validate on fold 1.
  3. Repeat for all folds. The average performance across folds gives a robust estimate.

3. Advanced Topics in Machine Learning

3.1 Ensemble Pruning

Pruning removes weak or redundant models from an ensemble to improve efficiency without sacrificing accuracy. For example, in a Random Forest, pruning can remove trees that contribute little to the final prediction.

1005015030701202060110
Random Forest tree pruning example: Highlighted nodes represent pruned weak predictors retained for efficiency.

3.2 Neural Architecture Search (NAS)

NAS automates the design of neural networks by searching for optimal architectures (e.g., layer sizes, connections). Companies like Google use NAS to optimize models for tasks like image recognition.

3.3 Explainable AI (XAI)

XAI techniques (e.g., SHAP values, LIME) make ML models interpretable. For example, Pathao might use XAI to explain why a ride price is recommended, ensuring transparency for users.


4. Real-World Applications in Nepal

4.1 Daraz: Recommendation Systems

Daraz uses ensemble methods to combine collaborative filtering (user-item interactions) and content-based filtering (product features) to recommend products. Boosting helps correct biases in user preferences.

4.2 Ncell: Fraud Detection

Ncell employs ensemble methods (e.g., isolation forests + logistic regression) to detect fraudulent transactions. Bagging reduces false positives, while boosting focuses on high-risk patterns.

4.3 NEPSE: Stock Market Prediction

Investors use stacking to combine technical analysis (e.g., moving averages) with fundamental data (e.g., P/E ratios) to predict stock trends. Hyperparameter tuning ensures the model adapts to market volatility.


5. Worked Example: Confusion Matrix and Hyperparameter Tuning

Problem: Given a dataset of loan approvals (approved/rejected), fit a model and tune hyperparameters to maximize accuracy.

Step 1: Confusion Matrix Assume a classifier predicts loan approvals with the following results:

  • True Positives (TP): 80 (correctly approved).
  • False Positives (FP): 10 (incorrectly approved).
  • False Negatives (FN): 5 (incorrectly rejected).
  • True Negatives (TN): 75 (correctly rejected).

The confusion matrix:

          | Predicted: Approve | Predicted: Reject
----------|-------------------|-------------------
Actual: Approve | 80 (TP)          | 5 (FN)
Actual: Reject   | 10 (FP)          | 75 (TN)

Step 2: Hyperparameter Tuning with Random Forest Tune max_depth and n_estimators using 5-fold cross-validation:

  • max_depth=5, n_estimators=100 → Accuracy: 92%
  • max_depth=7, n_estimators=200 → Accuracy: 94%
  • max_depth=3, n_estimators=50 → Accuracy: 88%

The best combination is max_depth=7, n_estimators=200.


6. Exam Tips

  1. Compare bagging and boosting: Know their goals (variance vs. bias reduction), training methods (parallel vs. sequential), and real-world examples.
  2. Confusion matrix: Understand TP, FP, FN, TN, and metrics like precision, recall, and F1-score.
  3. Hyperparameter tuning: Explain grid search vs. random search and the role of cross-validation.
  4. Ensemble pruning: Mention how it improves efficiency.
  5. Real-world tie-ins: Relate to Nepalese apps (e.g., Daraz, Ncell) or global examples (e.g., Google’s NAS).
  6. Worked examples: Practice tuning hyperparameters on small datasets and interpreting confusion matrices.

Based on the TU BCA syllabus for Machine Learning (CACS486), unit 7.

Discussion

Loading…