hamim-87's picture
Submit PPO LunarLander-v2 agent for Deep RL Course Unit 8
c39b9dd verified
|
Raw History Blame Contribute Delete
1.24 kB
metadata
tags:
  - LunarLander-v2
  - ppo
  - deep-reinforcement-learning
  - reinforcement-learning
  - custom-implementation
  - deep-rl-course
model-index:
  - name: PPO
    results:
      - task:
          type: reinforcement-learning
          name: reinforcement-learning
        dataset:
          name: LunarLander-v2
          type: LunarLander-v2
        metrics:
          - type: mean_reward
            value: '-76.70 +/- 25.91'
            name: mean_reward
            verified: false

PPO Agent Playing LunarLander-v2

This is my Hugging Face Deep Reinforcement Learning Course Unit 8 Part 1 submission. The PPO agent was implemented from scratch with PyTorch and trained on the exact course environment, LunarLander-v2.

Evaluation

  • Mean reward: -76.70
  • Standard deviation: 25.91
  • Episodes: 20

Hyperparameters

{
  "env_id": "LunarLander-v2",
  "seed": 1,
  "total_timesteps": 500000,
  "learning_rate": 0.0003,
  "num_envs": 8,
  "num_steps": 256,
  "num_minibatches": 8,
  "update_epochs": 4,
  "gamma": 0.99,
  "gae_lambda": 0.95,
  "clip_coef": 0.2,
  "ent_coef": 0.01,
  "vf_coef": 0.5,
  "max_grad_norm": 0.5,
  "anneal_lr": true,
  "norm_adv": true,
  "clip_vloss": true,
  "target_kl": 0.03,
  "save_every_updates": 20,
  "resume": true
}