top of page
Search

Video 8: Training the Agent & Troubleshooting

Apr 10, 2025
3 min read

Updated: Apr 16, 2025


In this blog post, we’ll explore how a reinforcement learning (RL) agent is trained using episodes, how the Q-network updates its policy, and common troubleshooting issues like overfitting and underfitting. This builds on our previous discussion of action execution, reward calculation, and experience replay.


The Training Loop: How Episodes Shape Learning

The agent learns through episodes—self-contained training sessions where the environment resets periodically. In this game:

  • An episode ends when:

    • The agent collects all coins (success).

    • The agent takes too long to collect a coin (timeout).

Why Use Episodes?

  1. Encourages Repetition & Learning

    • The agent gets multiple attempts at the same task, reinforcing good strategies (e.g., collecting closer coins first).

    • Without episodes, the agent might only learn short-term behaviors.

  2. Prevents Long-Term Penalties

    • Mistakes (like falling off the map) don’t carry over—each episode is a fresh start.

  3. Tracks Progress Over Time

    • Metrics like time per episode, coins collected, and total rewards help measure improvement.

Resetting the Environment

At the start of each episode:

  • The agent’s position, velocity, and jump count (current_jumps) reset.

  • All coins reappear (reset_collectables()).

  • The exploration rate (epsilon) decays slightly, reducing randomness over time.


How the Agent Trains: The train_dqn() Function

The Deep Q-Network (DQN) learns by sampling past experiences from the replay buffer and adjusting its predictions. Here’s how it works:

Step 1: Sampling Experiences

  • A batch of past experiences (state, action, reward, next state) is randomly selected.

Step 2: Calculating Target Q-Values

For each experience:

  1. The Q-network predicts the best action for the next state.

  2. The target network (a more stable version of the Q-network) estimates the future reward (double_q_value).

  3. The target Q-value is computed as:

    target_q_value = reward + (gamma * double_q_value)

    • gamma (discount factor) determines how much future rewards matter (typically 0.9-0.99).

Step 3: Updating the Q-Network

  • The Q-network adjusts its predictions to minimize the difference between its estimates and the target Q-values.

  • The target network syncs with the Q-network every 1000 steps to stabilize learning.


Key Hyperparameters

Parameter

Role

Recommended Adjustments

epsilon

Exploration rate (start high, decay over time)

Start at 0.9, decay slowly

gamma

Future reward discount (0.9-0.99)

Higher = agent values long-term rewards more

max_buffer_size

Replay buffer capacity

Larger = more diverse experiences

batch_size

Training sample size

32-128 (balance speed & stability)

train_frequency

How often the agent trains (e.g., every 5 steps)

More frequent = faster learning (but unstable)


Troubleshooting: Overfitting vs. Underfitting

1. Overfitting (Agent Gets Stuck in Suboptimal Policy)

  • Symptoms:

    • The agent repeats the same ineffective actions (e.g., jumping in place instead of collecting coins).

  • Solutions:

    • Reduce neural network complexity (fewer hidden nodes).

    • Lower the learning rate (smaller policy updates).

    • Add randomness (e.g., random obstacles or coin placements).

2. Underfitting (Agent Acts Randomly Even After Training)

  • Symptoms:

    • The agent fails to develop a clear strategy, even with low epsilon.

  • Solutions:

    • Increase neural network size (more hidden nodes).

    • Slow epsilon decay (more exploration time).

    • Increase learning rate (faster policy adjustments).


Final Thoughts & Next Steps

This pipeline—episodic training, experience replay, and Q-network updates—helps the agent gradually improve its strategy. Key takeaways:

✅ Episodes allow the agent to learn from repeated attempts.

✅ Experience replay prevents catastrophic forgetting by reusing past experiences.

✅ Target networks stabilize training by decoupling prediction and learning.

✅ Hyperparameter tuning is crucial to avoid overfitting/underfitting.

The agent may need hundreds (or thousands) of episodes to master the game. Monitoring metrics like average reward per episode helps track progress.


 
 
 

Comments


bottom of page