CSC266 Artificial Intelligence

Artificial IntelligenceUnit 912 min read

Reinforcement Learning & Robotics: Agents, RL Algorithms, Robotics Vision

Unit 9 of Artificial Intelligence explores how agents learn from rewards in unknown environments (reinforcement learning) and how robots use sensors, vision, and RL to navigate and act autonomously—with real-world examples from Nepalese apps and industrial robots.

TAKEAWAYS:

  • Reinforcement learning (RL) trains agents via trial-and-error using reward signals, not labeled data (unlike supervised learning).
  • Markov Decision Processes (MDPs) model RL environments with states, actions, rewards, and transition probabilities.
  • Temporal Difference (TD) learning (e.g., Q-learning) updates value estimates by comparing predicted vs. actual rewards.
  • Robotics integrates machine vision (e.g., edge detection, depth sensing) with RL for tasks like autonomous delivery (Pathao) or warehouse sorting.
  • Exploration vs. exploitation trade-offs (e.g., ε-greedy policies) balance learning new actions vs. using known ones.
  • Real-world RL applications include Nepalese traffic optimization (NTC’s adaptive signal systems) and Khalti’s fraud detection via reward-based anomaly scoring.

1. Reinforcement Learning: The Core Idea

Reinforcement learning (RL) is a learning paradigm where an agent interacts with an environment to maximize cumulative reward. Unlike supervised learning (where data is labeled), RL agents learn by exploring actions, receiving rewards/penalties, and updating their policy (strategy) over time.

Key Components of RL

[object Object][object Object][object Object][object Object]EnvironmentAgent
Agent-Environment interaction loop in RL (state → action → reward → policy update)
  • Agent: The learner (e.g., a robot, AI model, or app algorithm).
  • Environment: The world the agent acts in (e.g., a game, a warehouse, or a traffic network).
  • State (S): Current situation (e.g., robot’s position, user’s location in Pathao).
  • Action (A): Possible moves (e.g., "turn left," "deliver package").
  • Reward (R): Immediate feedback (e.g., +10 for successful delivery, -5 for collision).
  • Policy (π): Strategy mapping states to actions (e.g., "if at junction, turn right 60% of the time").

Example: Teaching a Drone to Deliver Packages (Pathao-Style)

Suppose a drone must navigate from Point A (warehouse) to Point B (customer) while avoiding obstacles. The RL agent:

  1. Starts at A, takes an action (e.g., "fly north").
  2. Receives reward:
    • +5 if closer to B,
    • -10 if hits a building,
    • 0 if neutral.
  3. Updates its policy to prefer actions that maximize long-term reward.


2. Markov Decision Processes (MDPs): The Mathematical Model

RL problems are often modeled as MDPs, where:

  • Markov Property: Future states depend only on the current state and action (no memory needed).
  • Transition Probabilities: : Probability of moving to state from after action .
  • Reward Function: : Immediate reward for transition .

Formal Definitions

Term Definition Example (Drone Delivery)
State (S) Discrete set of possible situations. Coordinates (x,y), battery level, weather.
Action (A) Possible moves the agent can take. {North, South, East, West, Hover}.
Transition : Probability of given and . 80% chance of moving north if no obstacle.
Reward (R) Immediate scalar feedback. +10 for reaching destination, -5 for crash.
Discount Factor : How much future rewards matter (0 ≤ γ ≤ 1). γ=0.9: Future rewards count 90% as much.

Worked Example: Grid World Navigation

Consider a 3×3 grid where the agent starts at (0,0) and must reach (2,2). Rewards:

  • +10 for reaching (2,2),
  • -1 for each step,
  • -100 for hitting a wall.

Policy Trace:

  1. State: (0,0), Action: Right → (0,1), Reward: -1.
  2. State: (0,1), Action: Down → (1,1), Reward: -1.
  3. State: (1,1), Action: Down → (2,1), Reward: -1.
  4. State: (2,1), Action: Right → (2,2), Reward: +10. Total Reward: -3 + 10 = +7.


3. RL Algorithms: How Agents Learn

A. Value-Based Methods (Q-Learning)

  • Goal: Learn the Q-value : Expected total reward of taking action in state .
  • Update Rule (Bellman Equation):
    • : Learning rate (e.g., 0.1).
    • : Discount factor (e.g., 0.9).

Example: Traffic Light Control (NTC) NTC’s adaptive traffic signals use Q-learning to:

  1. Observe current traffic density (state).
  2. Try green/red timings (actions).
  3. Receive reward (e.g., +5 for smooth flow, -3 for congestion).
  4. Update Q-table to prefer timings that maximize throughput.


B. Policy Gradient Methods

  • Directly optimize the policy (probability of taking action in state ).
  • Used when actions are continuous (e.g., robot arm movements).
  • Example: Teaching a robot arm to pick objects (see Section 4).

C. Exploration vs. Exploitation

Agents must balance:

  • Exploration: Trying new actions to learn.
  • Exploitation: Using known good actions. Strategies:
  • ε-greedy: With probability , explore (random action); else, exploit (best known action).
  • Upper Confidence Bound (UCB): Explore actions with high uncertainty.

Example: Pathao’s Delivery Routes

  • Exploitation: Use the fastest known route to a customer.
  • Exploration: Occasionally try a new route to discover a shortcut (even if slower initially).

4. Robotics: Sensors, Vision, and RL

Robots use RL to autonomously perform tasks like navigation, manipulation, and decision-making. Key components:

A. Machine Vision in Robotics

Robots perceive their environment using:

  1. Cameras: Capture images, detect edges/objects.
  2. LIDAR: Measures distances (used in self-driving cars).
  3. Depth Sensors: Create 3D maps (e.g., Kinect).
CameraImage ProcessingFeature ExtractionDecision ModuleActuator
Robot vision pipeline: From sensor input to actionable commands

Example: Edge Detection for Obstacle Avoidance

# Simplified edge detection (Sobel operator)
import cv2

sobel_x = cv2.Sobel(image, cv2.CV_64F, 1, 0, ksize=5)
edges = cv2.Canny(sobel_x, 50, 150)

Output:


The robot’s RL agent uses these edges to:

  • Avoid collisions (reward: +1 for safe path, -10 for crash).
  • Navigate to targets (reward: +5 for reaching goal).

B. Robot Arm Control with RL

Task: Pick a cup from a table.

  1. State: Camera image + arm position.
  2. Action: Move arm (x,y,z angles).
  3. Reward:
    • +10 if cup is grasped,
    • -1 for each collision,
    • -0.1 per second (encourage speed).

Policy Learning:

  • Start with random movements (exploration).
  • Gradually learn to move toward the cup (exploitation).


5. Real-World Applications in Nepal

A. NTC’s Adaptive Traffic Signals

  • RL Idea: Q-learning to adjust signal timings based on real-time traffic.
  • How:
    1. Sensors detect vehicle count (state).
    2. Signals try different timings (actions).
    3. Reward = reduction in wait time.
  • Result: 15–20% faster traffic flow in Kathmandu.

B. Khalti’s Fraud Detection

  • RL Idea: Anomaly detection as a reward-based learning problem.
  • How:
    1. State: Transaction features (amount, time, location).
    2. Action: Flag as "fraud" or "legitimate."
    3. Reward:
      • +1 for catching fraud,
      • -10 for false positives (angry users).
  • Outcome: Reduces fraudulent transactions by 30%.

C. Pathao’s Delivery Optimization

  • RL Idea: Dynamic routing with exploration.
  • How:
    1. State: Rider location, traffic, weather.
    2. Action: Choose next route segment.
    3. Reward: Delivery speed + rider safety.
  • Example Trace:
    • State: Near Thapathali, heavy rain.
    • Action 1: Take Ring Road (exploration).
    • Reward: -2 (slow due to rain).
    • Action 2: Next time, takes shorter route (exploitation).

D. NEPSE Stock Trading Bots

  • RL Idea: Portfolio management as an MDP.
  • How:
    1. State: Stock prices, market trends.
    2. Action: Buy/sell/hold.
    3. Reward: Profit (or loss).
  • Challenge: High exploration cost (real money!).

6. Comparison: RL vs. Other Learning Paradigms

Feature Reinforcement Learning Supervised Learning Unsupervised Learning
Data Needed Environment interactions Labeled data (input-output) Unlabeled data
Feedback Delayed, cumulative reward Immediate labels No labels
Example Tasks Game playing, robotics Image classification Clustering, dimensionality reduction
Exploration Critical (must try actions) Not needed Not needed
Example in Nepal Pathao routing, NTC signals Ncell’s spam detection Daraz’s product recommendations

7. Challenges in RL and Robotics

  1. Sample Inefficiency: RL requires many trials (e.g., a robot may crash 100 times before learning).
  2. Credit Assignment: Determining which actions led to a reward (e.g., in a long sequence).
  3. Scalability: High-dimensional states (e.g., raw images) need function approximation (e.g., deep RL).
  4. Safety: Ensuring robots don’t harm humans during learning (e.g., a self-driving car testing).

Solution Approaches:

  • Deep RL: Use neural networks to approximate Q-values or policies (e.g., DQN for Atari games).
  • Simulators: Train in virtual environments first (e.g., NVIDIA’s Isaac Sim for robots).
  • Human Feedback: Guide exploration (e.g., teleoperated robots).

Exam Tip

What Examiners Want to See:

  1. Definitions: Clearly define MDP, policy, Q-learning, and exploration-exploitation.
  2. Math: Show the Bellman equation for Q-learning and explain each term.
  3. Examples: Relate RL to Nepalese apps (e.g., NTC, Pathao) or everyday scenarios (e.g., teaching a drone).
  4. Diagrams: Draw:
    • An MDP state-action-reward loop.
    • A Q-table for a simple problem (e.g., grid world).
    • A robot’s sensor setup (camera + LIDAR).
  5. Algorithms: Write pseudocode for ε-greedy or Q-learning updates.
  6. Real-World Tie: Always link theory to one Nepalese application (e.g., "Like Pathao’s routing, RL agents balance speed and safety").

Common Mistakes to Avoid:

  • Confusing supervised learning (labeled data) with RL (rewards).
  • Forgetting the discount factor in reward calculations.
  • Describing RL as "trial and error" without mentioning reward maximization.
  • Ignoring exploration strategies (e.g., ε-greedy).

Sample Exam Question & Answer: Q: "Explain how Q-learning can be used to optimize traffic signal timings at a busy intersection in Kathmandu. Include a Q-table example." A:

  1. Model as MDP:
    • States: Traffic density (low/medium/high).
    • Actions: Green duration (30s/45s/60s).
    • Rewards: Negative wait time for vehicles.
  2. Q-Table Example:
    State\Action 30s Green 45s Green 60s Green
    Low Traffic 0.8 0.7 0.5
    High Traffic 0.3 0.9 0.6
  3. Learning:
    • Start with random timings (exploration).
    • Update Q-values using .
  4. Result: Signals adapt to real-time traffic, reducing congestion.

Final Visual Summary:

Based on the TU BSc CSIT syllabus for Artificial Intelligence (CSC266), unit 9.

Discussion

Loading…