top of page
Search

Video 7: Action Execution, Reward Calculation & Experience Replay

Apr 10, 2025
3 min read

In this blog post, we’ll break down the key concepts covered in Video 7: Action Execution, Reward Calculation & Experience Replay, where we explore how an AI agent learns to navigate a game environment, collects rewards, and improves its decision-making over time.


The Role of _physics_process()

The _physics_process() function is the backbone of the agent’s learning loop. It runs every frame and handles:

  • Applying gravity to the agent if it’s not grounded.

  • Updating timers and tracking the agent’s position relative to the nearest coin.

  • Executing actions (moving left, right, or jumping) based on the agent’s policy.


How Actions Are Executed

The execute_action() function determines how the agent moves:

  • If the agent is on the ground, its jump counter (current_jumps) resets to 0.

  • A small deceleration force is applied if no movement action is taken.

  • Action 0 moves right, Action 1 moves left, and Action 2 triggers a jump (if allowed).

The agent can only perform a double jump before landing again.


Reward Calculation: Encouraging Smart Behavior

Rewards shape the agent’s behavior by reinforcing good decisions. The get_reward() function calculates rewards based on:

  1. Coin Collection:

    • If the agent collects a coin, it receives a high reward (100 - 0.5 * time taken).

    • This incentivizes quick coin collection.

  2. Distance to Coin:

    • If no coin is collected, the reward depends on whether the agent moved closer to or farther from the nearest coin.

    • The reward is scaled down (multiplied by 0.01) and clamped between -0.2 and 1 to prevent drastic policy shifts.

    • A small penalty is applied to discourage wasting time.

This reward structure ensures the agent learns efficient navigation strategies without overreacting to single experiences.


Experience Replay: Learning from Past Actions

The agent stores its experiences in a replay buffer via store_experience(). This buffer holds:

  • Previous states

  • Actions taken

  • Rewards received

  • Next states

Why Use a Replay Buffer?

  1. Efficient Learning:

    • The agent reuses past experiences by sampling randomly from the buffer, making training more data-efficient.

  2. Stable Training:

    • Mixing older and newer experiences prevents the agent from overfitting to recent actions, leading to more robust policies.

When training, the agent randomly samples a batch of experiences, allowing it to learn from a diverse set of past decisions.


Episode Management & Policy Updates

At the end of each episode (when all coins are collected or time runs out), several steps occur:

  1. Epsilon Decay:

    • The exploration rate (epsilon) decreases, reducing random actions over time.

  2. Debugging & Logging:

    • Key metrics (rewards, steps taken) are logged for analysis.

  3. Agent Reset:

    • The environment resets, and the agent starts a new episode.

Training the Agent

If enough experiences are stored, the agent undergoes training:

  • The Q-network (policy) is updated using sampled experiences.

  • Periodically, the target network (used for stable Q-value estimation) syncs with the Q-network.


Action Selection: The Epsilon-Greedy Strategy

The choose_action() function implements the epsilon-greedy strategy:

  1. Exploration (Random Action):

    • If a random number is less than epsilon, the agent takes a random action.

    • This ensures the agent explores new strategies.

  2. Exploitation (Best Action):

    • Otherwise, the agent picks the action with the highest Q-value (predicted reward).

    • If multiple actions have the same Q-value, one is chosen randomly.

This balance between exploration and exploitation helps the agent refine its policy over time.


Conclusion

By breaking down the _physics_process() loop, reward calculation, and experience replay, we see how reinforcement learning agents:

✅ Execute actions in a controlled environment.

✅ Learn from rewards to optimize behavior.

✅ Improve gradually through experience replay.

This structured approach ensures the agent becomes more efficient at navigating its environment while avoiding suboptimal policies.


 
 
 

Comments


bottom of page