top of page
Search

Video 2: Q Learning Overview

Mar 2, 2025
2 min read

Updated: Apr 10, 2025


Reinforcement learning (RL) enables AI agents to learn through trial and error. One of its most fundamental algorithms is Q-learning, which helps agents make optimal decisions by estimating future rewards. In this post, we'll break down how Q-learning works using a simple grid-world example.


Key Concepts in Q-Learning

1. States, Actions, and Q-Values

  • State (S): What the agent observes (e.g., position, velocity).

  • Action (A): What the agent can do (e.g., move left, jump).

  • Q-Value: The expected reward for taking an action in a given state.

A Q-table stores these values, where:

  • Rows = States

  • Columns = Actions

  • (In the video this is flipped to make the Q table fit on slides better)

Initially, all Q-values are 0 because the agent hasn’t learned yet.


How Q-Learning Works

1. The Agent Explores the Environment

  • The agent starts in a state (e.g., S1) and takes an action (e.g., Move Down).

  • It receives a reward (e.g., -1 for moving to a non-reward space, +10 for reaching the goal).

2. Updating the Q-Table

The agent updates its Q-table using the Bellman Equation:

Copy

New Q(S,A) = Old Q(S,A) + α [Reward + γ * Max Q(S',A') - Old Q(S,A)]

Where:

  • α (Learning Rate): Controls how much new info overrides old (e.g., 0.01).

  • γ (Discount Factor): How much future rewards matter (e.g., 0.5).

Example Update:

  • From S7 → S8 (Move Right):

    • Reward = -1

    • Best future Q-value (from S8) = 0.1

    • Update:

      Q(S7, Right) = 0 + 0.01 [-1 + (0.5)(0.1) - 0] = -0.0095

    • This slight improvement helps the agent learn over time.


Balancing Exploration & Exploitation: Epsilon-Greedy Strategy

A key challenge is ensuring the agent explores new actions while also exploiting known rewards.

How Epsilon-Greedy Works:

  • Epsilon (ε): Probability of taking a random action (e.g., start at 1.0 = 100% random).

  • Decay Over Time: After each episode, ε reduces (e.g., ε = max(ε - decay, final_ε)).

Action Selection:

  1. Random Move (Exploration):

    • If rand() < ε, pick a random action.

  2. Best Known Move (Exploitation):

    • Otherwise, pick the action with the highest Q-value.

This ensures the agent:

✅ Learns from past successes (exploitation).

✅ Discovers better strategies (exploration).


Why Use Q-Learning?

  • Simple & Effective: Works well in small, discrete environments (e.g., grid worlds).

  • No Prior Knowledge Needed: Learns purely from rewards.

  • Flexible: Can adapt to changing environments.

Limitations:

  • Curse of Dimensionality: Q-tables grow exponentially with more states/actions.

  • Struggles with Continuous Spaces: Better suited for discrete problems.


Key Takeaways

🔹 Q-learning helps agents estimate the best actions using a Q-table.

🔹 The Bellman Equation updates Q-values based on rewards and future expectations.

🔹 Epsilon-greedy balances exploration and exploitation.

🔹 Works best in small, discrete environments (for large problems, consider Deep Q-Networks).






 
 
 

Comments


bottom of page