Machine LearningUnit 114 min read
ML Basics: Definitions, Types, and Real-World Applications
Unit 1 of Machine Learning introduces core concepts like supervised vs. unsupervised learning, key ML tasks (classification, regression, clustering), and real-world applications in apps like eSewa, Pathao, and NEPSE. It covers ML vs. AI, problem formulation, and ethical considerations.
TAKEAWAYS:
- Machine Learning (ML) is a subset of AI where systems learn from data to make predictions or decisions without explicit programming.
- ML tasks include supervised learning (labeled data), unsupervised learning (unlabeled data), and reinforcement learning (learning from rewards).
- Real-world examples: eSewa uses supervised learning for fraud detection, Pathao applies clustering for dynamic pricing, and NEPSE relies on regression for stock trend analysis.
- Key challenges in ML include data quality, overfitting, bias, and scalability.
- Ethical considerations (e.g., fairness, privacy) are critical in deploying ML systems.
- The ML pipeline follows: data → preprocessing → model selection → training → evaluation → deployment.
1. What is Machine Learning?
Machine Learning (ML) is a branch of Artificial Intelligence (AI) where algorithms learn patterns from data to make decisions or predictions. Unlike traditional programming, ML models improve with experience (more data).
ML vs. AI vs. Deep Learning
| Term | Definition | Example |
|---|---|---|
| AI | Simulates human intelligence (reasoning, problem-solving). | Chatbots, self-driving cars. |
| ML | Subset of AI where systems learn from data. | Spam detection, recommendation systems. |
| Deep Learning | Subset of ML using neural networks with many layers (deep architectures). | Image recognition (Google Photos), voice assistants (Siri). |
2. Types of Machine Learning
ML is classified into three main types based on data labeling and learning approach:
A. Supervised Learning
- Uses labeled data (input-output pairs).
- Tasks: Classification (discrete outputs) and Regression (continuous outputs).
Example:
- eSewa uses supervised learning to classify transactions as fraudulent (1) or legitimate (0).
- NEPSE predicts stock prices (regression) using historical data.
Worked Example: Spam Classification (Binary Classification) Suppose we train a model to classify emails as spam (1) or not spam (0).
| Email Feature | Spam (1) | Not Spam (0) |
|---|---|---|
| Contains "FREE" | 1 | 0 |
| Sender unknown | 1 | 0 |
| Has attachments | 1 | 0 |
Model Prediction:
- If an email has "FREE" + unknown sender, the model predicts spam (1) with high confidence.
B. Unsupervised Learning
- Uses unlabeled data to find hidden patterns.
- Tasks: Clustering, Dimensionality Reduction, Association Rule Mining.
Example:
- Pathao uses clustering to group similar rider locations for dynamic pricing.
- NTC analyzes customer behavior to segment users (e.g., heavy vs. light data users).
Worked Example: Customer Segmentation (K-Means Clustering) Suppose a telecom company (Ncell) wants to segment customers based on monthly usage (minutes) and data consumption (GB).
| Customer | Minutes Used | Data Used (GB) |
|---|---|---|
| A | 300 | 5 |
| B | 1000 | 20 |
| C | 500 | 10 |
K-Means Clustering (k=2):
- Randomly assign 2 centroids (e.g., (300,5) and (1000,20)).
- Assign each customer to the nearest centroid.
- Recalculate centroids and repeat until convergence.
Result:
- Cluster 1 (Light Users): Low minutes, low data (e.g., Customer A).
- Cluster 2 (Heavy Users): High minutes, high data (e.g., Customer B).
C. Reinforcement Learning (RL)
- Learns by trial and error (agent-reward-environment interaction).
- Used in robotics, game AI, and autonomous systems.
Example:
- Pathao’s delivery optimization uses RL to find the fastest route while avoiding traffic.
- NEPSE’s algorithmic trading learns to buy/sell stocks based on market rewards.
Worked Example: Grid World Navigation (Simple RL) An agent (robot) navigates a grid to reach a goal (G) while avoiding obstacles (X).
S → A → B → G
| | |
X X .
- States (S, A, B, G): Positions on the grid.
- Actions: Move up, down, left, right.
- Rewards:
- +10 for reaching G.
- -1 for hitting X.
- 0 for other moves.
Q-Learning Algorithm:
- Start at S, choose an action (e.g., right → A).
- Receive reward (0), update Q-table.
- Repeat until G is reached.
3. Key ML Tasks
| Task | Definition | Example | Algorithm |
|---|---|---|---|
| Classification | Predicts discrete labels (categories). | Spam detection, disease diagnosis. | Logistic Regression, SVM, Decision Trees. |
| Regression | Predicts continuous values. | Stock price prediction, house price estimation. | Linear Regression, Polynomial Regression. |
| Clustering | Groups similar data points (unsupervised). | Customer segmentation, image compression. | K-Means, Hierarchical Clustering. |
| Dimensionality Reduction | Reduces feature space while retaining info. | Face recognition, PCA for image compression. | PCA, t-SNE. |
| Association Rule Mining | Finds relationships between variables. | Market basket analysis (e.g., "People who buy X also buy Y"). | Apriori, FP-Growth. |
4. The Machine Learning Pipeline
Every ML project follows these steps:
flowchart LR
A[Data Collection
(NEPSE Stock Data)] --> B[Preprocessing
(Normalization)] --> C[Model Selection
(Linear Regression)] --> D[Training
(1000 Epochs)] --> E[Evaluation
(RMSE: 0.5)] --> F[Deployment
(API for Traders)]NEPSE stock prediction pipeline (simplified)flowchart LR
A["Data Collection"] --> B["Data Preprocessing"]
B --> C["Model Selection"]
C --> D["Training"]
D --> E["Evaluation"]
E -->|"If Poor Performance"| F["Hyperparameter Tuning"]
E -->|"If Good Performance"| G["Deployment"]
G --> H["Monitoring & Maintenance"]Real-World Example: eSewa Fraud Detection
- Data Collection: Transaction logs (amount, time, user ID).
- Preprocessing: Clean missing values, encode user IDs.
- Model Selection: Logistic Regression (for binary classification).
- Training: Fit model on historical fraud/non-fraud data.
- Evaluation: Test on unseen data (accuracy = 95%).
- Deployment: Integrate model into eSewa’s backend.
- Monitoring: Retrain weekly with new fraud patterns.
5. Challenges in Machine Learning
| Challenge | Description | Solution |
|---|---|---|
| Overfitting | Model performs well on training data but poorly on test data. | Use cross-validation, regularization, or more data. |
| Underfitting | Model is too simple to capture patterns. | Use more complex models or feature engineering. |
| Bias & Fairness | Model performs poorly on certain groups (e.g., gender, race). | Use fairness-aware algorithms, diverse datasets. |
| Data Quality | Missing, noisy, or irrelevant data. | Cleaning, imputation, feature selection. |
| Scalability | Model struggles with large datasets. | Use distributed computing (Spark, Hadoop), deep learning. |
Worked Example: Overfitting in a Polynomial Regression Suppose we fit a 10th-degree polynomial to predict house prices based on size.
# Overfitted model (high variance)
from sklearn.preprocessing import PolynomialFeatures
poly = PolynomialFeatures(degree=10)
X_poly = poly.fit_transform(X)
model = LinearRegression().fit(X_poly, y)
Problem: The model fits training data perfectly but fails on test data.
Solution: Use Ridge Regression (L2 regularization) to penalize large coefficients.
6. Ethical Considerations in ML
ML systems can perpetuate bias, privacy violations, and misuse. Key ethical concerns:
- Bias: If training data is skewed (e.g., mostly male faces in a facial recognition dataset), the model may perform poorly for women.
- Privacy: Models trained on personal data (e.g., health records) must comply with GDPR/Nepal’s Data Privacy Act.
- Accountability: Who is responsible if an ML model makes a harmful decision (e.g., loan denial)?
Example: NEPSE’s Algorithmic Trading Bias
- If historical stock data favors certain companies, the model may exclude small-cap stocks, leading to unfair market predictions.
- Solution: Use diverse datasets and fairness-aware algorithms.
7. Real-World Applications in Nepal
| Company/App | ML Technique Used | Application |
|---|---|---|
| eSewa | Supervised Learning (Classification) | Fraud detection in transactions. |
| Pathao | Clustering, Reinforcement Learning | Dynamic pricing, route optimization. |
| Ncell | Clustering, Time-Series Forecasting | Customer segmentation, network traffic prediction. |
| NTC | Anomaly Detection | Identifying unusual data usage (potential hacking). |
| NEPSE | Regression, Time-Series Analysis | Stock price prediction, trend analysis. |
| Daraz | Recommendation Systems (Collaborative Filtering) | Product recommendations based on user behavior. |
Worked Example: Pathao’s Dynamic Pricing Pathao uses clustering + reinforcement learning to adjust fares in real-time.
- Clustering: Groups riders into high-demand (peak hours) and low-demand (off-peak) clusters.
- RL Agent: Learns to increase prices during high demand to balance supply-demand.
- Result: Faster matches for riders and higher earnings for drivers.
8. Exam Tip: How to Score Full Marks
- Define Clearly:
- Always start with definitions (e.g., "Supervised learning is a type of ML where the model learns from labeled data...").
- Use Diagrams:
- Draw ML pipeline flowcharts, decision trees, or clustering visualizations in exams.
- Relate to Real-World Examples:
- Mention eSewa, Pathao, or NEPSE wherever possible (examiners love local examples).
- Show Mathematical Steps (Where Applicable):
- For regression/classification, write equations (e.g., cost function for linear regression).
- Discuss Limitations:
- No ML technique is perfect—mention bias, overfitting, or data dependency in your answer.
- Practice Short-Answer Questions:
- Common exam questions:
- "Differentiate between supervised and unsupervised learning."
- "How does Pathao use ML for dynamic pricing?"
- "What is overfitting? How can it be avoided?"
- Common exam questions:
Sample Exam Question & Answer: Q: Explain how NEPSE can use time-series forecasting to predict stock prices. Include a step-by-step approach.
A:
- Data Collection: Gather historical stock prices (opening, closing, volume) for the last 5 years.
- Preprocessing:
- Handle missing values (impute with mean).
- Normalize data (Min-Max scaling).
- Model Selection:
- Use ARIMA (AutoRegressive Integrated Moving Average) for time-series forecasting.
- Training:
- Split data into training (80%) and testing (20%).
- Fit ARIMA model on training data.
- Evaluation:
- Metrics: RMSE (Root Mean Squared Error), MAE (Mean Absolute Error).
- Prediction:
- Forecast next 30 days’ stock prices.
- Deployment:
- Integrate predictions into NEPSE’s trading dashboard.
Visual:
graph LR
A["Historical Data"] --> B["Preprocessing"]
B --> C["ARIMA Model"]
C --> D["Train-Test Split"]
D --> E["Forecasting"]
E --> F["NEPSE Dashboard"]Final Note: This unit is foundational—master the definitions, types, and real-world links to excel in exams. Always connect theory to Nepalese examples (e.g., eSewa, Pathao) to stand out!
Based on the PU BE Computer (PU) syllabus for Machine Learning (CMP364), unit 1.
Discussion
Loading…