maze2d_easy_native256_noncot_chunk_k10_world_model_20260707_perseg (EMA, bf16)

Fine-tuned BAGEL-7B-MoT action-conditioned visual world model — Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split). Step-5000 EMA weights, cast fp32 → bfloat16 (inference runs in bf16 autocast — lossless for inference). No optimizer state; not a training-resume checkpoint.

  • Load into the BAGEL-7B-MoT architecture (base: ByteDance-Seed/BAGEL-7B-MoT).
  • Training data: ultrastar111/maze2d_easy_native256_noncot_chunk_k10_20260707_perseg
  • Recipe: cold start from vanilla BAGEL, lr 2e-5, cosine, warmup 300, 5000 steps, 40k-token packing (maze2d cot-K1: 24k), EMA 0.993.
  • Eval contract: CoT ckpts need the per-segment decoder (BAGEL_DECODER=perseg) + CFG off.
  • License: CC-BY-NC-4.0 (research use).
training_data_manifest.txt
STUDY:    Maze2D K-action-chunk feedback-interval (noncot, K=10, difficulty=easy; GT env-feedback re-grounding except K=inf)
data:     /data/home/raychai/hf_datasets/maze2d_easy_native256_noncot_chunk_k10_20260707_perseg/training (q95 baked in; maze2d instruction incl. total-budget + cadence baked in by remap)
config:   ./data/configs/vlm_gym_maze2d_imagined_cot_envfb_train.yaml (envfb loader 'interleaved_cot_envfb', data-dir overridden)
init:     VANILLA /home/jiaxin/unified_world_model/pretrained/BAGEL-7B-MoT (EMA); hparams lr2e-5/warmup100/cosine/total5000/tokens50000/ema0.993/save1000; probe OFF
git_rev:  0c4fc0786a9c30bfcb8d442261717c0e8d46760c
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support