|
Download README.md from hamim-87/ppo-LunarLander-v2-unit8: direct link, hf CLI and curl.
- Browser
- Download file 1.24 kB
-
https://huggingface.co/hamim-87/ppo-LunarLander-v2-unit8/resolve/main/README.md
- Command line
-
hf download hf://hamim-87/ppo-LunarLander-v2-unit8/README.md
-
curl -L -o README.md https://huggingface.co/hamim-87/ppo-LunarLander-v2-unit8/resolve/main/README.md
1.24 kB
metadata
tags:
- LunarLander-v2
- ppo
- deep-reinforcement-learning
- reinforcement-learning
- custom-implementation
- deep-rl-course
model-index:
- name: PPO
results:
- task:
type: reinforcement-learning
name: reinforcement-learning
dataset:
name: LunarLander-v2
type: LunarLander-v2
metrics:
- type: mean_reward
value: '-76.70 +/- 25.91'
name: mean_reward
verified: false
PPO Agent Playing LunarLander-v2
This is my Hugging Face Deep Reinforcement Learning Course Unit 8 Part 1 submission.
The PPO agent was implemented from scratch with PyTorch and trained on the exact
course environment, LunarLander-v2.
Evaluation
- Mean reward: -76.70
- Standard deviation: 25.91
- Episodes: 20
Hyperparameters
{
"env_id": "LunarLander-v2",
"seed": 1,
"total_timesteps": 500000,
"learning_rate": 0.0003,
"num_envs": 8,
"num_steps": 256,
"num_minibatches": 8,
"update_epochs": 4,
"gamma": 0.99,
"gae_lambda": 0.95,
"clip_coef": 0.2,
"ent_coef": 0.01,
"vf_coef": 0.5,
"max_grad_norm": 0.5,
"anneal_lr": true,
"norm_adv": true,
"clip_vloss": true,
"target_kl": 0.03,
"save_every_updates": 20,
"resume": true
}