Machine LearningUnit 114 min read

ML Basics: Definitions, Types, and Real-World Applications

Unit 1 of Machine Learning introduces core concepts like supervised vs. unsupervised learning, key ML tasks (classification, regression, clustering), and real-world applications in apps like eSewa, Pathao, and NEPSE. It covers ML vs. AI, problem formulation, and ethical considerations.

TAKEAWAYS:

  • Machine Learning (ML) is a subset of AI where systems learn from data to make predictions or decisions without explicit programming.
  • ML tasks include supervised learning (labeled data), unsupervised learning (unlabeled data), and reinforcement learning (learning from rewards).
  • Real-world examples: eSewa uses supervised learning for fraud detection, Pathao applies clustering for dynamic pricing, and NEPSE relies on regression for stock trend analysis.
  • Key challenges in ML include data quality, overfitting, bias, and scalability.
  • Ethical considerations (e.g., fairness, privacy) are critical in deploying ML systems.
  • The ML pipeline follows: data → preprocessing → model selection → training → evaluation → deployment.

1. What is Machine Learning?

Machine Learning (ML) is a branch of Artificial Intelligence (AI) where algorithms learn patterns from data to make decisions or predictions. Unlike traditional programming, ML models improve with experience (more data).

ML vs. AI vs. Deep Learning

Term Definition Example
AI Simulates human intelligence (reasoning, problem-solving). Chatbots, self-driving cars.
ML Subset of AI where systems learn from data. Spam detection, recommendation systems.
Deep Learning Subset of ML using neural networks with many layers (deep architectures). Image recognition (Google Photos), voice assistants (Siri).

2. Types of Machine Learning

ML is classified into three main types based on data labeling and learning approach:

-5-4-3-2-112345-55101520xyLoss Function (MSE)Gradient Descent PathMinimum
Loss function optimization (NEPSE stock prediction)
11111SABGX1X2
Pathao’s RL grid world (explored path marked)
eSewa Fraud DetectionClassificationNEPSE Stock PredictionRegressionSupervised Learning
Supervised learning tasks with Nepali examples (highlighted)

A. Supervised Learning

  • Uses labeled data (input-output pairs).
  • Tasks: Classification (discrete outputs) and Regression (continuous outputs).

Example:

  • eSewa uses supervised learning to classify transactions as fraudulent (1) or legitimate (0).
  • NEPSE predicts stock prices (regression) using historical data.

Worked Example: Spam Classification (Binary Classification) Suppose we train a model to classify emails as spam (1) or not spam (0).

Email Feature Spam (1) Not Spam (0)
Contains "FREE" 1 0
Sender unknown 1 0
Has attachments 1 0

Model Prediction:

  • If an email has "FREE" + unknown sender, the model predicts spam (1) with high confidence.

B. Unsupervised Learning

  • Uses unlabeled data to find hidden patterns.
  • Tasks: Clustering, Dimensionality Reduction, Association Rule Mining.

Example:

  • Pathao uses clustering to group similar rider locations for dynamic pricing.
  • NTC analyzes customer behavior to segment users (e.g., heavy vs. light data users).

Worked Example: Customer Segmentation (K-Means Clustering) Suppose a telecom company (Ncell) wants to segment customers based on monthly usage (minutes) and data consumption (GB).

Customer Minutes Used Data Used (GB)
A 300 5
B 1000 20
C 500 10

K-Means Clustering (k=2):

  1. Randomly assign 2 centroids (e.g., (300,5) and (1000,20)).
  2. Assign each customer to the nearest centroid.
  3. Recalculate centroids and repeat until convergence.

Result:

  • Cluster 1 (Light Users): Low minutes, low data (e.g., Customer A).
  • Cluster 2 (Heavy Users): High minutes, high data (e.g., Customer B).

C. Reinforcement Learning (RL)

  • Learns by trial and error (agent-reward-environment interaction).
  • Used in robotics, game AI, and autonomous systems.

Example:

  • Pathao’s delivery optimization uses RL to find the fastest route while avoiding traffic.
  • NEPSE’s algorithmic trading learns to buy/sell stocks based on market rewards.

Worked Example: Grid World Navigation (Simple RL) An agent (robot) navigates a grid to reach a goal (G) while avoiding obstacles (X).

S → A → B → G
|    |    |
X    X    .
  • States (S, A, B, G): Positions on the grid.
  • Actions: Move up, down, left, right.
  • Rewards:
    • +10 for reaching G.
    • -1 for hitting X.
    • 0 for other moves.

Q-Learning Algorithm:

  1. Start at S, choose an action (e.g., right → A).
  2. Receive reward (0), update Q-table.
  3. Repeat until G is reached.

3. Key ML Tasks

Task Definition Example Algorithm
Classification Predicts discrete labels (categories). Spam detection, disease diagnosis. Logistic Regression, SVM, Decision Trees.
Regression Predicts continuous values. Stock price prediction, house price estimation. Linear Regression, Polynomial Regression.
Clustering Groups similar data points (unsupervised). Customer segmentation, image compression. K-Means, Hierarchical Clustering.
Dimensionality Reduction Reduces feature space while retaining info. Face recognition, PCA for image compression. PCA, t-SNE.
Association Rule Mining Finds relationships between variables. Market basket analysis (e.g., "People who buy X also buy Y"). Apriori, FP-Growth.

4. The Machine Learning Pipeline

Every ML project follows these steps:

flowchart LR
    A[Data Collection
    (NEPSE Stock Data)] --> B[Preprocessing
    (Normalization)] --> C[Model Selection
    (Linear Regression)] --> D[Training
    (1000 Epochs)] --> E[Evaluation
    (RMSE: 0.5)] --> F[Deployment
    (API for Traders)]
NEPSE stock prediction pipeline (simplified)
flowchart LR
    A["Data Collection"] --> B["Data Preprocessing"]
    B --> C["Model Selection"]
    C --> D["Training"]
    D --> E["Evaluation"]
    E -->|"If Poor Performance"| F["Hyperparameter Tuning"]
    E -->|"If Good Performance"| G["Deployment"]
    G --> H["Monitoring & Maintenance"]

Real-World Example: eSewa Fraud Detection

  1. Data Collection: Transaction logs (amount, time, user ID).
  2. Preprocessing: Clean missing values, encode user IDs.
  3. Model Selection: Logistic Regression (for binary classification).
  4. Training: Fit model on historical fraud/non-fraud data.
  5. Evaluation: Test on unseen data (accuracy = 95%).
  6. Deployment: Integrate model into eSewa’s backend.
  7. Monitoring: Retrain weekly with new fraud patterns.

5. Challenges in Machine Learning

Challenge Description Solution
Overfitting Model performs well on training data but poorly on test data. Use cross-validation, regularization, or more data.
Underfitting Model is too simple to capture patterns. Use more complex models or feature engineering.
Bias & Fairness Model performs poorly on certain groups (e.g., gender, race). Use fairness-aware algorithms, diverse datasets.
Data Quality Missing, noisy, or irrelevant data. Cleaning, imputation, feature selection.
Scalability Model struggles with large datasets. Use distributed computing (Spark, Hadoop), deep learning.

Worked Example: Overfitting in a Polynomial Regression Suppose we fit a 10th-degree polynomial to predict house prices based on size.

# Overfitted model (high variance)
from sklearn.preprocessing import PolynomialFeatures
poly = PolynomialFeatures(degree=10)
X_poly = poly.fit_transform(X)
model = LinearRegression().fit(X_poly, y)

Problem: The model fits training data perfectly but fails on test data.

Solution: Use Ridge Regression (L2 regularization) to penalize large coefficients.


6. Ethical Considerations in ML

ML systems can perpetuate bias, privacy violations, and misuse. Key ethical concerns:

  • Bias: If training data is skewed (e.g., mostly male faces in a facial recognition dataset), the model may perform poorly for women.
  • Privacy: Models trained on personal data (e.g., health records) must comply with GDPR/Nepal’s Data Privacy Act.
  • Accountability: Who is responsible if an ML model makes a harmful decision (e.g., loan denial)?

Example: NEPSE’s Algorithmic Trading Bias

  • If historical stock data favors certain companies, the model may exclude small-cap stocks, leading to unfair market predictions.
  • Solution: Use diverse datasets and fairness-aware algorithms.

7. Real-World Applications in Nepal

Company/App ML Technique Used Application
eSewa Supervised Learning (Classification) Fraud detection in transactions.
Pathao Clustering, Reinforcement Learning Dynamic pricing, route optimization.
Ncell Clustering, Time-Series Forecasting Customer segmentation, network traffic prediction.
NTC Anomaly Detection Identifying unusual data usage (potential hacking).
NEPSE Regression, Time-Series Analysis Stock price prediction, trend analysis.
Daraz Recommendation Systems (Collaborative Filtering) Product recommendations based on user behavior.

Worked Example: Pathao’s Dynamic Pricing Pathao uses clustering + reinforcement learning to adjust fares in real-time.

  1. Clustering: Groups riders into high-demand (peak hours) and low-demand (off-peak) clusters.
  2. RL Agent: Learns to increase prices during high demand to balance supply-demand.
  3. Result: Faster matches for riders and higher earnings for drivers.

8. Exam Tip: How to Score Full Marks

  1. Define Clearly:
    • Always start with definitions (e.g., "Supervised learning is a type of ML where the model learns from labeled data...").
  2. Use Diagrams:
    • Draw ML pipeline flowcharts, decision trees, or clustering visualizations in exams.
  3. Relate to Real-World Examples:
    • Mention eSewa, Pathao, or NEPSE wherever possible (examiners love local examples).
  4. Show Mathematical Steps (Where Applicable):
    • For regression/classification, write equations (e.g., cost function for linear regression).
  5. Discuss Limitations:
    • No ML technique is perfect—mention bias, overfitting, or data dependency in your answer.
  6. Practice Short-Answer Questions:
    • Common exam questions:
      • "Differentiate between supervised and unsupervised learning."
      • "How does Pathao use ML for dynamic pricing?"
      • "What is overfitting? How can it be avoided?"

Sample Exam Question & Answer: Q: Explain how NEPSE can use time-series forecasting to predict stock prices. Include a step-by-step approach.

A:

  1. Data Collection: Gather historical stock prices (opening, closing, volume) for the last 5 years.
  2. Preprocessing:
    • Handle missing values (impute with mean).
    • Normalize data (Min-Max scaling).
  3. Model Selection:
    • Use ARIMA (AutoRegressive Integrated Moving Average) for time-series forecasting.
  4. Training:
    • Split data into training (80%) and testing (20%).
    • Fit ARIMA model on training data.
  5. Evaluation:
    • Metrics: RMSE (Root Mean Squared Error), MAE (Mean Absolute Error).
  6. Prediction:
    • Forecast next 30 days’ stock prices.
  7. Deployment:
    • Integrate predictions into NEPSE’s trading dashboard.

Visual:

graph LR
    A["Historical Data"] --> B["Preprocessing"]
    B --> C["ARIMA Model"]
    C --> D["Train-Test Split"]
    D --> E["Forecasting"]
    E --> F["NEPSE Dashboard"]

Final Note: This unit is foundational—master the definitions, types, and real-world links to excel in exams. Always connect theory to Nepalese examples (e.g., eSewa, Pathao) to stand out!

Based on the PU BE Computer (PU) syllabus for Machine Learning (CMP364), unit 1.

Discussion

Loading…