Reinforcement Learning
stable-baselines3
4d-snake
4d-snake-4x4
snake
deep-reinforcement-learning
sb3-contrib
maskable-ppo
Eval Results (legacy)
Instructions to use BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- stable-baselines3
How to use BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc with stable-baselines3:
from huggingface_sb3 import load_from_hub checkpoint = load_from_hub( repo_id="BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc", filename="{MODEL FILENAME}.zip", ) - Notebooks
- Google Colab
- Kaggle
exp05b_ppo_4x4_from_bc: checkpoint, config, evaluation files, model card
Browse files
README.md
CHANGED
|
@@ -42,7 +42,7 @@ model-index:
|
|
| 42 |
|
| 43 |
# 4d-snake-exp05b-ppo-4x4-from-bc
|
| 44 |
|
| 45 |
-
A `512x512` MLP policy for **4-dimensional snake on the 4^4 board** (256 cells, 8 moves): MaskablePPO fine-tuned from the behaviour-cloned network, no curriculum, 20,000,000 environment steps. From a length-1 start it completes the board in every deterministic evaluation episode, evaluated with the protocol of [docs/evaluation.md](https://github.com/BurnyCoder/4d-snake-
|
| 46 |
|
| 47 |
## Results (`eval/summary.json`)
|
| 48 |
|
|
@@ -53,10 +53,10 @@ A `512x512` MLP policy for **4-dimensional snake on the 4^4 board** (256 cells,
|
|
| 53 |
|
| 54 |
## How to use
|
| 55 |
|
| 56 |
-
The observation is this repository's `4*C + 2` float vector and the action space its `2*ndim` masked moves ([docs/game_rules.md](https://github.com/BurnyCoder/4d-snake-
|
| 57 |
|
| 58 |
```bash
|
| 59 |
-
git clone https://github.com/BurnyCoder/4d-snake-
|
| 60 |
hf download BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc best_model.zip --local-dir weights
|
| 61 |
uv run snake4d evaluate --set model_path=weights/best_model.zip --set size=4 --set ndim=4
|
| 62 |
```
|
|
@@ -80,7 +80,7 @@ Use deterministic mode: the cloned policy follows a fixed Hamiltonian cycle and
|
|
| 80 |
|
| 81 |
## Training
|
| 82 |
|
| 83 |
-
- Phase `train`; experiment file `experiments/exp05b_ppo_4x4_from_bc.env`; write-up: https://github.com/BurnyCoder/4d-snake-
|
| 84 |
- Resolved configuration (`config.json`):
|
| 85 |
|
| 86 |
```json
|
|
@@ -138,9 +138,9 @@ Use deterministic mode: the cloned policy follows a fixed Hamiltonian cycle and
|
|
| 138 |
|
| 139 |
## Provenance
|
| 140 |
|
| 141 |
-
- Code: https://github.com/BurnyCoder/4d-snake-
|
| 142 |
- Library versions (`versions.json`): torch 2.14.0+cu130, gymnasium 1.3.0, stable-baselines3 2.9.0, sb3-contrib 2.9.0, numpy 2.5.2, pygame-ce 2.5.8, cuda_device NVIDIA GeForce RTX 5070 Laptop GPU.
|
| 143 |
-
- `eval/summary.json` and `eval/eval_episodes.csv` are the files the repository's reports quote; every evaluated network is compared in [reports/networks.md](https://github.com/BurnyCoder/4d-snake-
|
| 144 |
- Collection: https://huggingface.co/collections/BurnyCoder/4d-snake-rl-all-evaluated-networks-6a9d0a0a66c7efcd101b7741
|
| 145 |
|
| 146 |
## Files
|
|
|
|
| 42 |
|
| 43 |
# 4d-snake-exp05b-ppo-4x4-from-bc
|
| 44 |
|
| 45 |
+
A `512x512` MLP policy for **4-dimensional snake on the 4^4 board** (256 cells, 8 moves): MaskablePPO fine-tuned from the behaviour-cloned network, no curriculum, 20,000,000 environment steps. From a length-1 start it completes the board in every deterministic evaluation episode, evaluated with the protocol of [docs/evaluation.md](https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/docs/evaluation.md) (100 episodes x 3 seeds, masked `evaluate_policy`).
|
| 46 |
|
| 47 |
## Results (`eval/summary.json`)
|
| 48 |
|
|
|
|
| 53 |
|
| 54 |
## How to use
|
| 55 |
|
| 56 |
+
The observation is this repository's `4*C + 2` float vector and the action space its `2*ndim` masked moves ([docs/game_rules.md](https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/docs/game_rules.md)), so the checkpoint runs inside `snake4d`'s environment:
|
| 57 |
|
| 58 |
```bash
|
| 59 |
+
git clone https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent.git && cd 4d-snake-reinforcement-learning-agent && uv sync
|
| 60 |
hf download BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc best_model.zip --local-dir weights
|
| 61 |
uv run snake4d evaluate --set model_path=weights/best_model.zip --set size=4 --set ndim=4
|
| 62 |
```
|
|
|
|
| 80 |
|
| 81 |
## Training
|
| 82 |
|
| 83 |
+
- Phase `train`; experiment file `experiments/exp05b_ppo_4x4_from_bc.env`; write-up: https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/reports/experiments/exp05_bc_4x4.md.
|
| 84 |
- Resolved configuration (`config.json`):
|
| 85 |
|
| 86 |
```json
|
|
|
|
| 138 |
|
| 139 |
## Provenance
|
| 140 |
|
| 141 |
+
- Code: https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent at commit `016b5dc97583b7c086ad172226997e6aeea92dca`.
|
| 142 |
- Library versions (`versions.json`): torch 2.14.0+cu130, gymnasium 1.3.0, stable-baselines3 2.9.0, sb3-contrib 2.9.0, numpy 2.5.2, pygame-ce 2.5.8, cuda_device NVIDIA GeForce RTX 5070 Laptop GPU.
|
| 143 |
+
- `eval/summary.json` and `eval/eval_episodes.csv` are the files the repository's reports quote; every evaluated network is compared in [reports/networks.md](https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/reports/networks.md).
|
| 144 |
- Collection: https://huggingface.co/collections/BurnyCoder/4d-snake-rl-all-evaluated-networks-6a9d0a0a66c7efcd101b7741
|
| 145 |
|
| 146 |
## Files
|