Artificial IntelligenceUnit 612 min read
Bayes’ Theorem, Probabilistic Reasoning & Decision Trees
Unit 6 of Artificial Intelligence covers probabilistic reasoning under uncertainty, Bayes’ theorem, conditional probability, decision trees, and their applications in real-world AI systems like medical diagnosis, spam filtering, and financial risk assessment.
TAKEAWAYS:
- Bayes’ theorem updates probabilities when new evidence arrives, enabling AI to make smarter decisions under uncertainty.
- Conditional probability () measures how likely event is given event , forming the backbone of probabilistic reasoning.
- Decision trees model choices and outcomes, balancing risk and reward in AI-driven decisions (e.g., loan approvals, medical tests).
- Naïve Bayes simplifies calculations by assuming feature independence, widely used in spam detection and sentiment analysis.
- Expected utility theory helps AI agents choose actions that maximize long-term benefits, even with incomplete information.
- Real-world applications include fraud detection (Khalti), disease prediction (healthcare apps), and route optimization (Pathao).
1. Why Probability in AI?
AI systems often operate in uncertain environments—where data is noisy, incomplete, or probabilistic. For example:
- A spam filter cannot be 100% sure an email is spam.
- A doctor cannot diagnose a disease with absolute certainty.
- A self-driving car must make decisions despite sensor errors.
Probabilistic reasoning allows AI to handle uncertainty by assigning degrees of belief (probabilities) to possible outcomes.
Key Concepts
- Probability (): Likelihood of event occurring.
- Conditional Probability (): Probability of given has occurred.
- Joint Probability (): Probability of both and happening together.
- Independent Events: (e.g., rolling a die and flipping a coin).
2. Bayes’ Theorem: Updating Beliefs with Evidence
Bayes’ theorem formalizes how to update probabilities when new evidence arrives. It is the foundation of probabilistic reasoning in AI.
Formula
- : Posterior probability (belief after seeing evidence ).
- : Likelihood (how likely is if is true).
- : Prior probability (initial belief before seeing ).
- : Marginal probability (total probability of , often computed as ).
Worked Example: Medical Diagnosis
Suppose a disease affects 1% of the population (). A test is 95% accurate:
- If a person has the disease, the test is positive 95% of the time ().
- If a person does not have the disease, the test is negative 90% of the time ().
Question: If a randomly selected person tests positive, what is the probability they actually have the disease?
Solution
Define events:
- : Person has the disease.
- : Test is positive.
Given:
- (prior).
- (likelihood if diseased).
- (false positive rate).
Compute (marginal probability):
Apply Bayes’ theorem:
Surprising Result: Even with a positive test, the probability of having the disease is only ~8.76%! This is because the disease is rare (), and false positives dominate.
3. Conditional Probability and Independence
Conditional Probability Table Example
Consider a spam detection system that checks two features:
- Feature 1 (F1): Email contains the word "FREE".
- Feature 2 (F2): Email has urgent language.
| Spam () | Not Spam () | |
|---|---|---|
| 0.8 | 0.1 | |
| 0.7 | 0.2 | |
| 0.6 | 0.02 |
Question: What is (probability email is spam given both features)?
Solution:
Independence Check
Two events and are independent if: Example: Rolling a die and flipping a coin are independent. Non-example: Rain and carrying an umbrella are not independent (if it rains, you’re more likely to carry an umbrella).
4. Naïve Bayes Classifier: Simplifying Probabilistic Models
Naïve Bayes assumes features are conditionally independent given the class. This simplifies calculations while remaining effective.
Formula
For a classification problem with features and classes :
Worked Example: Sentiment Analysis
A sentiment analysis model classifies tweets as Positive () or Negative () based on words:
- Words: "love", "hate".
- Data:
- , .
- , .
- , .
Tweet: "I love this product but hate the delivery." Question: What is the probability this tweet is Positive?
Solution
- Extract features: "love" (present), "hate" (present).
- Compute likelihoods:
- .
- .
- Apply Naïve Bayes:
- Normalize:
Conclusion: The tweet is more likely Negative (43.3%) than Positive (56.7%), despite containing "love". This shows how context matters even in Naïve Bayes!
5. Decision Trees for Risk-Based Decisions
Decision trees model sequential decisions under uncertainty, balancing risks and rewards.
Key Components
- Root Node: Starting point (e.g., "Should I invest?").
- Branches: Possible actions or outcomes.
- Leaf Nodes: Final outcomes with associated probabilities.
- Utility Values: Numerical values representing desirability (e.g., profit, cost).
Example: Loan Approval System (Like NMB Bank)
A bank uses a decision tree to approve loans based on:
- Credit Score (High/Low).
- Income Level (High/Low).
Data:
- , .
- , .
- Default rates:
- High Credit + High Income: 5% default.
- High Credit + Low Income: 15% default.
- Low Credit + High Income: 30% default.
- Low Credit + Low Income: 60% default.
Decision Tree:
Question: Should the bank approve a loan for a customer with High Credit + Low Income?
Solution:
- Expected Loss if approved:
- Expected Profit if rejected: 0 (no business).
- Decision: Approve if expected profit > expected loss (e.g., if loan amount is small).
6. Expected Utility Theory
AI agents should maximize expected utility, not just probability. Utility combines:
- Probability of outcomes.
- Value (utility) of outcomes.
Formula
Example: Investment Decision
An investor has two options:
- Safe Investment: 100% chance of $10,000 profit.
- Risky Investment: 50% chance of $50,000 profit, 50% chance of $0.
Question: Which should they choose if they are risk-neutral?
Solution:
- Safe: .
- Risky: . Decision: Choose the risky investment (higher expected utility).
In the Real World
Khalti Fraud Detection
- Uses Bayes’ theorem to calculate .
- Example: If a transaction has high amount + unusual location, the system flags it as likely fraud ().
Pathao Driver Routing
- Uses decision trees to balance:
- Distance to pickup.
- Traffic conditions.
- Driver’s current location.
- Example: If , the system reroutes to avoid delays.
- Uses decision trees to balance:
Nepal Electricity Load Prediction (NTC)
- Uses Naïve Bayes to predict peak demand based on:
- Temperature.
- Time of day.
- Historical usage.
- Example: If , NTC prepares extra power.
- Uses Naïve Bayes to predict peak demand based on:
7. Comparing Probabilistic Models
| Model | Assumptions | Use Case | Pros | Cons |
|---|---|---|---|---|
| Bayes’ Theorem | No independence required | Medical diagnosis, spam filtering | Accurate with evidence | Computationally expensive |
| Naïve Bayes | Features independent given class | Sentiment analysis, text classification | Fast, works well with high dimensions | Overly simplistic |
| Decision Trees | No probability assumptions | Loan approval, risk management | Interpretable, no training needed | Sensitive to data variations |
| Expected Utility | Utility values defined | Investment, resource allocation | Considers risk-reward tradeoff | Requires utility quantification |
Exam Tip
Bayes’ Theorem Questions:
- Always define events clearly (, ).
- Compute using the law of total probability.
- Common pitfall: Ignoring and directly computing .
Naïve Bayes:
- Assume independence to simplify calculations.
- Use log probabilities to avoid underflow with small numbers.
Decision Trees:
- Draw the tree explicitly for step-by-step marks.
- Calculate expected values at each node.
Real-World Applications:
- Link concepts to Nepali examples (e.g., Khalti fraud, Pathao routing).
- Explain why probability is needed (uncertainty in real data).
8. Common Mistakes to Avoid
- Ignoring prior probabilities: Always use , not just .
- Assuming independence when none exists: Naïve Bayes fails if features are correlated.
- Misapplying Bayes’ theorem: .
- Overcomplicating decision trees: Start with a simple binary tree before adding complexity.
9. Practice Problems
Medical Testing:
- A test for a rare disease (prevalence = 2%) has 99% true positive rate and 95% true negative rate.
- If a person tests positive, what is ?
Spam Filter:
- Words "FREE" and "WIN" appear in 80% of spam and 5% of non-spam.
- If an email contains both, what is ?
Investment Decision:
- Option A: 70% chance of $10,000, 30% chance of $0.
- Option B: 100% chance of $7,000.
- Which maximizes expected utility?
10. Summary Visual
Final Note: Probabilistic reasoning is everywhere in AI—from fraud detection to healthcare. Master Bayes’ theorem, Naïve Bayes, and decision trees, and you’ll excel in both exams and real-world applications!
Based on the TU BITM syllabus for Artificial Intelligence (IT228), unit 6.
Discussion
Loading…