beshy3752's picture
Add seed-2 teacher checkpoint + benchmark (public)
ea13efe verified
|
Raw History Blame Contribute Delete
2.7 kB
---
license: mit
tags:
- reinforcement-learning
- quadruped
- locomotion
- parkour
- unitree-go2
- isaac-lab
---
# Eurekaverse Go2 Parkour Policy — Seed 2
A Unitree **Go2** quadruped parkour locomotion policy trained with an **LLM-generated environment curriculum** (Eurekaverse-style), reproduced on **Isaac Lab 3.0**.
This is the final policy of **seed 2** of a 3-seed reproduction run (independent RL seed + independent LLM-generated curriculum from seed 1).
> ⚠️ **This is the privileged _teacher_ policy** (it consumes terrain **scandots** — a height scan around the robot). It is **NOT deployable standalone**: without scandots it just collapses. To run on a real robot / another sim, **distill a depth-based student** first (`train.py --use_camera ...`). See the code repo for setup.
## What this is
- **Task:** parkour locomotion — traverse obstacle courses toward 8 sequential goals (ramps, boxes, stepping stones, stairs, gaps, beams, …).
- **Training:** 5 curriculum iterations × 8 parallel policy runs × 2000 PPO steps/env, resumed from a 1000-step flat-ground walk pretrain. Terrains generated each iteration by `gpt-4o-2024-08-06`. This checkpoint is the iteration-4 winner (parallel run 2, lineage [2,4,4,6]).
- **Checkpoint:** `model_11000.pt` (final). `model_0.pt` is the iteration-4 starting point.
## Benchmark performance
Held-out benchmark of **20 parkour tasks × 10 difficulty levels**, metric = **number of goals reached (out of 8)**:
- **This policy (seed 2): 4.52 / 8**
- For reference, seed 1's final policy scored **4.36 / 8** — the two independent seeds land within ~0.2, i.e. the reproduction is seed-stable.
Per-task numbers are in `benchmark_results.txt`.
## Files
| file | description |
|------|-------------|
| `model_11000.pt` | final policy checkpoint |
| `model_0.pt` | iteration-4 starting checkpoint |
| `legged_robot_config.pkl` | pickled `(env_cfg, train_cfg)` — **required** to load the policy |
| `final_iteration_terrain.py` | the LLM-generated terrain this policy trained on (iter 4) |
| `benchmark_results.txt` | per-task benchmark evaluation |
## How to load / use
Not a standalone / `transformers` model. Requires **Isaac Lab 3.0** + the `extreme-parkour` / `legged_gym` env used for training. Place the checkpoint + `legged_robot_config.pkl` under `logs/<proj>/<exptid>/` and load via the repo's `task_registry` (see `scripts/evaluate.py`). This is a **reproduction/fork ported to Isaac Lab 3.0**, so absolute numbers may differ from the original Eurekaverse paper.
## Attribution
Built on **Eurekaverse** (Liang et al., CoRL 2024) and **extreme-parkour**. Preserve upstream licenses when redistributing.