Aadit-032 commited on
Commit
226af2e
·
verified ·
1 Parent(s): a525a42

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +16 -27
README.md CHANGED
@@ -1,37 +1,26 @@
1
  ---
 
2
  library_name: stable-baselines3
3
  tags:
4
- - LunarLander-v3
5
  - deep-reinforcement-learning
6
- - reinforcement-learning
7
- - stable-baselines3
8
- model-index:
9
- - name: PPO
10
- results:
11
- - task:
12
- type: reinforcement-learning
13
- name: reinforcement-learning
14
- dataset:
15
- name: LunarLander-v3
16
- type: LunarLander-v3
17
- metrics:
18
- - type: mean_reward
19
- value: 275.37 +/- 23.18
20
- name: mean_reward
21
- verified: false
22
  ---
23
 
24
- # **PPO** Agent playing **LunarLander-v3**
25
- This is a trained model of a **PPO** agent playing **LunarLander-v3**
26
- using the [stable-baselines3 library](https://github.com/DLR-RM/stable-baselines3).
27
 
28
- ## Usage (with Stable-baselines3)
29
- TODO: Add your code
30
 
 
 
 
 
 
31
 
32
- ```python
33
- from stable_baselines3 import ...
34
- from huggingface_sb3 import load_from_hub
35
 
36
- ...
37
- ```
 
1
  ---
2
+ license: mit
3
  library_name: stable-baselines3
4
  tags:
 
5
  - deep-reinforcement-learning
6
+ - gymnasium
7
+ - lunar-lander
8
+ - ppo
9
+ - sb3
 
 
 
 
 
 
 
 
 
 
 
 
10
  ---
11
 
12
+ # PPO LunarLander-v3
 
 
13
 
14
+ This repository contains a Stable-Baselines3 PPO agent trained to solve the Gymnasium LunarLander environment.
 
15
 
16
+ ## Training setup
17
+ - Algorithm: PPO
18
+ - Environment: LunarLander-v3
19
+ - Policy: MlpPolicy
20
+ - Training timesteps: 1,000,000
21
 
22
+ ## Evaluation
23
+ The agent was evaluated on the LunarLander environment with deterministic rollout settings.
 
24
 
25
+ ## Notes
26
+ This model is intended for experimentation and educational purposes.