BurnyCoder commited on
Commit
0b20f90
·
verified ·
1 Parent(s): 8e391d2

exp05b_ppo_4x4_from_bc: checkpoint, config, evaluation files, model card

Browse files
Files changed (1) hide show
  1. README.md +6 -6
README.md CHANGED
@@ -42,7 +42,7 @@ model-index:
42
 
43
  # 4d-snake-exp05b-ppo-4x4-from-bc
44
 
45
- A `512x512` MLP policy for **4-dimensional snake on the 4^4 board** (256 cells, 8 moves): MaskablePPO fine-tuned from the behaviour-cloned network, no curriculum, 20,000,000 environment steps. From a length-1 start it completes the board in every deterministic evaluation episode, evaluated with the protocol of [docs/evaluation.md](https://github.com/BurnyCoder/4d-snake-rl/blob/main/docs/evaluation.md) (100 episodes x 3 seeds, masked `evaluate_policy`).
46
 
47
  ## Results (`eval/summary.json`)
48
 
@@ -53,10 +53,10 @@ A `512x512` MLP policy for **4-dimensional snake on the 4^4 board** (256 cells,
53
 
54
  ## How to use
55
 
56
- The observation is this repository's `4*C + 2` float vector and the action space its `2*ndim` masked moves ([docs/game_rules.md](https://github.com/BurnyCoder/4d-snake-rl/blob/main/docs/game_rules.md)), so the checkpoint runs inside `snake4d`'s environment:
57
 
58
  ```bash
59
- git clone https://github.com/BurnyCoder/4d-snake-rl.git && cd 4d-snake-rl && uv sync
60
  hf download BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc best_model.zip --local-dir weights
61
  uv run snake4d evaluate --set model_path=weights/best_model.zip --set size=4 --set ndim=4
62
  ```
@@ -80,7 +80,7 @@ Use deterministic mode: the cloned policy follows a fixed Hamiltonian cycle and
80
 
81
  ## Training
82
 
83
- - Phase `train`; experiment file `experiments/exp05b_ppo_4x4_from_bc.env`; write-up: https://github.com/BurnyCoder/4d-snake-rl/blob/main/reports/experiments/exp05_bc_4x4.md.
84
  - Resolved configuration (`config.json`):
85
 
86
  ```json
@@ -138,9 +138,9 @@ Use deterministic mode: the cloned policy follows a fixed Hamiltonian cycle and
138
 
139
  ## Provenance
140
 
141
- - Code: https://github.com/BurnyCoder/4d-snake-rl at commit `016b5dc97583b7c086ad172226997e6aeea92dca`.
142
  - Library versions (`versions.json`): torch 2.14.0+cu130, gymnasium 1.3.0, stable-baselines3 2.9.0, sb3-contrib 2.9.0, numpy 2.5.2, pygame-ce 2.5.8, cuda_device NVIDIA GeForce RTX 5070 Laptop GPU.
143
- - `eval/summary.json` and `eval/eval_episodes.csv` are the files the repository's reports quote; every evaluated network is compared in [reports/networks.md](https://github.com/BurnyCoder/4d-snake-rl/blob/main/reports/networks.md).
144
  - Collection: https://huggingface.co/collections/BurnyCoder/4d-snake-rl-all-evaluated-networks-6a9d0a0a66c7efcd101b7741
145
 
146
  ## Files
 
42
 
43
  # 4d-snake-exp05b-ppo-4x4-from-bc
44
 
45
+ A `512x512` MLP policy for **4-dimensional snake on the 4^4 board** (256 cells, 8 moves): MaskablePPO fine-tuned from the behaviour-cloned network, no curriculum, 20,000,000 environment steps. From a length-1 start it completes the board in every deterministic evaluation episode, evaluated with the protocol of [docs/evaluation.md](https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/docs/evaluation.md) (100 episodes x 3 seeds, masked `evaluate_policy`).
46
 
47
  ## Results (`eval/summary.json`)
48
 
 
53
 
54
  ## How to use
55
 
56
+ The observation is this repository's `4*C + 2` float vector and the action space its `2*ndim` masked moves ([docs/game_rules.md](https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/docs/game_rules.md)), so the checkpoint runs inside `snake4d`'s environment:
57
 
58
  ```bash
59
+ git clone https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent.git && cd 4d-snake-reinforcement-learning-agent && uv sync
60
  hf download BurnyCoder/4d-snake-exp05b-ppo-4x4-from-bc best_model.zip --local-dir weights
61
  uv run snake4d evaluate --set model_path=weights/best_model.zip --set size=4 --set ndim=4
62
  ```
 
80
 
81
  ## Training
82
 
83
+ - Phase `train`; experiment file `experiments/exp05b_ppo_4x4_from_bc.env`; write-up: https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/reports/experiments/exp05_bc_4x4.md.
84
  - Resolved configuration (`config.json`):
85
 
86
  ```json
 
138
 
139
  ## Provenance
140
 
141
+ - Code: https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent at commit `016b5dc97583b7c086ad172226997e6aeea92dca`.
142
  - Library versions (`versions.json`): torch 2.14.0+cu130, gymnasium 1.3.0, stable-baselines3 2.9.0, sb3-contrib 2.9.0, numpy 2.5.2, pygame-ce 2.5.8, cuda_device NVIDIA GeForce RTX 5070 Laptop GPU.
143
+ - `eval/summary.json` and `eval/eval_episodes.csv` are the files the repository's reports quote; every evaluated network is compared in [reports/networks.md](https://github.com/BurnyCoder/4d-snake-reinforcement-learning-agent/blob/main/reports/networks.md).
144
  - Collection: https://huggingface.co/collections/BurnyCoder/4d-snake-rl-all-evaluated-networks-6a9d0a0a66c7efcd101b7741
145
 
146
  ## Files