PPO LunarLander-v3

This model was trained as part of the Hugging Face Deep Reinforcement Learning Course — Unit 1.

Environment

LunarLander-v3

Algorithm

Proximal Policy Optimization (PPO)

Training

  • Framework: Stable-Baselines3
  • Timesteps: 1,000,000
  • Parallel environments: 8
  • Seed: 42

Evaluation

  • Mean reward: 275.11
  • Standard deviation: 19.77
  • Mean minus standard deviation: 255.34

Important

This model uses the current Gymnasium environment:

LunarLander-v3

The older LunarLander-v2 environment is deprecated in current Gymnasium versions.

Downloads last month
9
Video Preview
loading