Video 2: Q Learning Overview
Updated: Apr 10, 2025
Reinforcement learning (RL) enables AI agents to learn through trial and error. One of its most fundamental algorithms is Q-learning, which helps agents make optimal decisions by estimating future rewards. In this post, we'll break down how Q-learning works using a simple grid-world example.
Key Concepts in Q-Learning
1. States, Actions, and Q-Values
State (S): What the agent observes (e.g., position, velocity).
Action (A): What the agent can do (e.g., move left, jump).
Q-Value: The expected reward for taking an action in a given state.
A Q-table stores these values, where:
Rows = States
Columns = Actions
(In the video this is flipped to make the Q table fit on slides better)
Initially, all Q-values are 0 because the agent hasn’t learned yet.
How Q-Learning Works
1. The Agent Explores the Environment
The agent starts in a state (e.g., S1) and takes an action (e.g., Move Down).
It receives a reward (e.g., -1 for moving to a non-reward space, +10 for reaching the goal).
2. Updating the Q-Table
The agent updates its Q-table using the Bellman Equation:
Copy
New Q(S,A) = Old Q(S,A) + α [Reward + γ * Max Q(S',A') - Old Q(S,A)]Where:
α (Learning Rate): Controls how much new info overrides old (e.g., 0.01).
γ (Discount Factor): How much future rewards matter (e.g., 0.5).
Example Update:
From S7 → S8 (Move Right):
Reward = -1
Best future Q-value (from S8) = 0.1
Update:
Q(S7, Right) = 0 + 0.01 [-1 + (0.5)(0.1) - 0] = -0.0095
This slight improvement helps the agent learn over time.
Balancing Exploration & Exploitation: Epsilon-Greedy Strategy
A key challenge is ensuring the agent explores new actions while also exploiting known rewards.
How Epsilon-Greedy Works:
Epsilon (ε): Probability of taking a random action (e.g., start at 1.0 = 100% random).
Decay Over Time: After each episode, ε reduces (e.g., ε = max(ε - decay, final_ε)).
Action Selection:
Random Move (Exploration):
If rand() < ε, pick a random action.
Best Known Move (Exploitation):
Otherwise, pick the action with the highest Q-value.
This ensures the agent:
✅ Learns from past successes (exploitation).
✅ Discovers better strategies (exploration).
Why Use Q-Learning?
Simple & Effective: Works well in small, discrete environments (e.g., grid worlds).
No Prior Knowledge Needed: Learns purely from rewards.
Flexible: Can adapt to changing environments.
Limitations:
Curse of Dimensionality: Q-tables grow exponentially with more states/actions.
Struggles with Continuous Spaces: Better suited for discrete problems.
Key Takeaways
🔹 Q-learning helps agents estimate the best actions using a Q-table.
🔹 The Bellman Equation updates Q-values based on rewards and future expectations.
🔹 Epsilon-greedy balances exploration and exploitation.
🔹 Works best in small, discrete environments (for large problems, consider Deep Q-Networks).
Slides: https://shorturl.at/C2KLV



Comments