Artificial IntelligenceUnit 912 min read
Reinforcement Learning & Robotics: Agents, RL Algorithms, Robotics Vision
Unit 9 of Artificial Intelligence explores how agents learn from rewards in unknown environments (reinforcement learning) and how robots use sensors, vision, and RL to navigate and act autonomously—with real-world examples from Nepalese apps and industrial robots.
TAKEAWAYS:
- Reinforcement learning (RL) trains agents via trial-and-error using reward signals, not labeled data (unlike supervised learning).
- Markov Decision Processes (MDPs) model RL environments with states, actions, rewards, and transition probabilities.
- Temporal Difference (TD) learning (e.g., Q-learning) updates value estimates by comparing predicted vs. actual rewards.
- Robotics integrates machine vision (e.g., edge detection, depth sensing) with RL for tasks like autonomous delivery (Pathao) or warehouse sorting.
- Exploration vs. exploitation trade-offs (e.g., ε-greedy policies) balance learning new actions vs. using known ones.
- Real-world RL applications include Nepalese traffic optimization (NTC’s adaptive signal systems) and Khalti’s fraud detection via reward-based anomaly scoring.
1. Reinforcement Learning: The Core Idea
Reinforcement learning (RL) is a learning paradigm where an agent interacts with an environment to maximize cumulative reward. Unlike supervised learning (where data is labeled), RL agents learn by exploring actions, receiving rewards/penalties, and updating their policy (strategy) over time.
Key Components of RL
- Agent: The learner (e.g., a robot, AI model, or app algorithm).
- Environment: The world the agent acts in (e.g., a game, a warehouse, or a traffic network).
- State (S): Current situation (e.g., robot’s position, user’s location in Pathao).
- Action (A): Possible moves (e.g., "turn left," "deliver package").
- Reward (R): Immediate feedback (e.g., +10 for successful delivery, -5 for collision).
- Policy (π): Strategy mapping states to actions (e.g., "if at junction, turn right 60% of the time").
Example: Teaching a Drone to Deliver Packages (Pathao-Style)
Suppose a drone must navigate from Point A (warehouse) to Point B (customer) while avoiding obstacles. The RL agent:
- Starts at A, takes an action (e.g., "fly north").
- Receives reward:
- +5 if closer to B,
- -10 if hits a building,
- 0 if neutral.
- Updates its policy to prefer actions that maximize long-term reward.
2. Markov Decision Processes (MDPs): The Mathematical Model
RL problems are often modeled as MDPs, where:
- Markov Property: Future states depend only on the current state and action (no memory needed).
- Transition Probabilities: : Probability of moving to state from after action .
- Reward Function: : Immediate reward for transition .
Formal Definitions
| Term | Definition | Example (Drone Delivery) |
|---|---|---|
| State (S) | Discrete set of possible situations. | Coordinates (x,y), battery level, weather. |
| Action (A) | Possible moves the agent can take. | {North, South, East, West, Hover}. |
| Transition | : Probability of given and . | 80% chance of moving north if no obstacle. |
| Reward (R) | Immediate scalar feedback. | +10 for reaching destination, -5 for crash. |
| Discount Factor | : How much future rewards matter (0 ≤ γ ≤ 1). | γ=0.9: Future rewards count 90% as much. |
Worked Example: Grid World Navigation
Consider a 3×3 grid where the agent starts at (0,0) and must reach (2,2). Rewards:
- +10 for reaching (2,2),
- -1 for each step,
- -100 for hitting a wall.
Policy Trace:
- State: (0,0), Action: Right → (0,1), Reward: -1.
- State: (0,1), Action: Down → (1,1), Reward: -1.
- State: (1,1), Action: Down → (2,1), Reward: -1.
- State: (2,1), Action: Right → (2,2), Reward: +10. Total Reward: -3 + 10 = +7.
3. RL Algorithms: How Agents Learn
A. Value-Based Methods (Q-Learning)
- Goal: Learn the Q-value : Expected total reward of taking action in state .
- Update Rule (Bellman Equation):
- : Learning rate (e.g., 0.1).
- : Discount factor (e.g., 0.9).
Example: Traffic Light Control (NTC) NTC’s adaptive traffic signals use Q-learning to:
- Observe current traffic density (state).
- Try green/red timings (actions).
- Receive reward (e.g., +5 for smooth flow, -3 for congestion).
- Update Q-table to prefer timings that maximize throughput.
B. Policy Gradient Methods
- Directly optimize the policy (probability of taking action in state ).
- Used when actions are continuous (e.g., robot arm movements).
- Example: Teaching a robot arm to pick objects (see Section 4).
C. Exploration vs. Exploitation
Agents must balance:
- Exploration: Trying new actions to learn.
- Exploitation: Using known good actions. Strategies:
- ε-greedy: With probability , explore (random action); else, exploit (best known action).
- Upper Confidence Bound (UCB): Explore actions with high uncertainty.
Example: Pathao’s Delivery Routes
- Exploitation: Use the fastest known route to a customer.
- Exploration: Occasionally try a new route to discover a shortcut (even if slower initially).
4. Robotics: Sensors, Vision, and RL
Robots use RL to autonomously perform tasks like navigation, manipulation, and decision-making. Key components:
A. Machine Vision in Robotics
Robots perceive their environment using:
- Cameras: Capture images, detect edges/objects.
- LIDAR: Measures distances (used in self-driving cars).
- Depth Sensors: Create 3D maps (e.g., Kinect).
Example: Edge Detection for Obstacle Avoidance
# Simplified edge detection (Sobel operator)
import cv2
sobel_x = cv2.Sobel(image, cv2.CV_64F, 1, 0, ksize=5)
edges = cv2.Canny(sobel_x, 50, 150)
Output:
The robot’s RL agent uses these edges to:
- Avoid collisions (reward: +1 for safe path, -10 for crash).
- Navigate to targets (reward: +5 for reaching goal).
B. Robot Arm Control with RL
Task: Pick a cup from a table.
- State: Camera image + arm position.
- Action: Move arm (x,y,z angles).
- Reward:
- +10 if cup is grasped,
- -1 for each collision,
- -0.1 per second (encourage speed).
Policy Learning:
- Start with random movements (exploration).
- Gradually learn to move toward the cup (exploitation).
5. Real-World Applications in Nepal
A. NTC’s Adaptive Traffic Signals
- RL Idea: Q-learning to adjust signal timings based on real-time traffic.
- How:
- Sensors detect vehicle count (state).
- Signals try different timings (actions).
- Reward = reduction in wait time.
- Result: 15–20% faster traffic flow in Kathmandu.
B. Khalti’s Fraud Detection
- RL Idea: Anomaly detection as a reward-based learning problem.
- How:
- State: Transaction features (amount, time, location).
- Action: Flag as "fraud" or "legitimate."
- Reward:
- +1 for catching fraud,
- -10 for false positives (angry users).
- Outcome: Reduces fraudulent transactions by 30%.
C. Pathao’s Delivery Optimization
- RL Idea: Dynamic routing with exploration.
- How:
- State: Rider location, traffic, weather.
- Action: Choose next route segment.
- Reward: Delivery speed + rider safety.
- Example Trace:
- State: Near Thapathali, heavy rain.
- Action 1: Take Ring Road (exploration).
- Reward: -2 (slow due to rain).
- Action 2: Next time, takes shorter route (exploitation).
D. NEPSE Stock Trading Bots
- RL Idea: Portfolio management as an MDP.
- How:
- State: Stock prices, market trends.
- Action: Buy/sell/hold.
- Reward: Profit (or loss).
- Challenge: High exploration cost (real money!).
6. Comparison: RL vs. Other Learning Paradigms
| Feature | Reinforcement Learning | Supervised Learning | Unsupervised Learning |
|---|---|---|---|
| Data Needed | Environment interactions | Labeled data (input-output) | Unlabeled data |
| Feedback | Delayed, cumulative reward | Immediate labels | No labels |
| Example Tasks | Game playing, robotics | Image classification | Clustering, dimensionality reduction |
| Exploration | Critical (must try actions) | Not needed | Not needed |
| Example in Nepal | Pathao routing, NTC signals | Ncell’s spam detection | Daraz’s product recommendations |
7. Challenges in RL and Robotics
- Sample Inefficiency: RL requires many trials (e.g., a robot may crash 100 times before learning).
- Credit Assignment: Determining which actions led to a reward (e.g., in a long sequence).
- Scalability: High-dimensional states (e.g., raw images) need function approximation (e.g., deep RL).
- Safety: Ensuring robots don’t harm humans during learning (e.g., a self-driving car testing).
Solution Approaches:
- Deep RL: Use neural networks to approximate Q-values or policies (e.g., DQN for Atari games).
- Simulators: Train in virtual environments first (e.g., NVIDIA’s Isaac Sim for robots).
- Human Feedback: Guide exploration (e.g., teleoperated robots).
Exam Tip
What Examiners Want to See:
- Definitions: Clearly define MDP, policy, Q-learning, and exploration-exploitation.
- Math: Show the Bellman equation for Q-learning and explain each term.
- Examples: Relate RL to Nepalese apps (e.g., NTC, Pathao) or everyday scenarios (e.g., teaching a drone).
- Diagrams: Draw:
- An MDP state-action-reward loop.
- A Q-table for a simple problem (e.g., grid world).
- A robot’s sensor setup (camera + LIDAR).
- Algorithms: Write pseudocode for ε-greedy or Q-learning updates.
- Real-World Tie: Always link theory to one Nepalese application (e.g., "Like Pathao’s routing, RL agents balance speed and safety").
Common Mistakes to Avoid:
- Confusing supervised learning (labeled data) with RL (rewards).
- Forgetting the discount factor in reward calculations.
- Describing RL as "trial and error" without mentioning reward maximization.
- Ignoring exploration strategies (e.g., ε-greedy).
Sample Exam Question & Answer: Q: "Explain how Q-learning can be used to optimize traffic signal timings at a busy intersection in Kathmandu. Include a Q-table example." A:
- Model as MDP:
- States: Traffic density (low/medium/high).
- Actions: Green duration (30s/45s/60s).
- Rewards: Negative wait time for vehicles.
- Q-Table Example:
State\Action 30s Green 45s Green 60s Green Low Traffic 0.8 0.7 0.5 High Traffic 0.3 0.9 0.6 - Learning:
- Start with random timings (exploration).
- Update Q-values using .
- Result: Signals adapt to real-time traffic, reducing congestion.
Final Visual Summary:
Based on the TU BSc CSIT syllabus for Artificial Intelligence (CSC266), unit 9.
Discussion
Loading…