--- license: apache-2.0 base_model: allenai/MolmoAct2-DROID tags: [robotics, lerobot, molmoact2, crimson, genome] --- # Crimson right-arm MolmoAct2 checkpoints Six MolmoAct2 checkpoints fine-tuned on `CrimsonRobot/crimson-first-dateset` (right arm, "pick and place the bottle"). Prepared per the Crimson/Genome model-integration contract. Full contract: `model-integration/model-contract.yaml`. | path in repo | seed | step | CI-MSE (std, lower=better) | rank | |---|---|---|---|---| | `s2000/001000` | 2000 | 1000 | 0.283 | 1 | | `s1000/001500` | 1000 | 1500 | 0.295 | 2 | | `s1000/000500` | 1000 | 500 | 0.322 | 3 | | `s2000/000500` | 2000 | 500 | 0.347 | 4 | | `s2000/001500` | 2000 | 1500 | 0.352 | 5 | | `s1000/001000` | 1000 | 1000 | 0.373 | 6 | Each folder is a LeRobot `pretrained_model` (config.json, model.safetensors 12.7 GB, pre/post-processor with normalization stats, train_config.json). CI-MSE = offline metric on 10 held-out episodes (40–49); it is a *ranking hypothesis*. The point of deployment is to measure real success rate per checkpoint and check whether this ranking holds. ## 1. What the model expects / returns (summary of the contract) | | | |---|---| | Arm | right only | | State in (8) | `rarm_j1..rarm_j7` measured rad, `right_gripper` measured servo rad (dataset `observation.state`, unchanged) | | Cameras in | `top` = cam_high, `right_wrist` = cam_wrist; RGB 640×480; no crop/rotation at inference | | Task | `"pick and place the bottle"` (only trained string) | | Action out (10×8) | absolute commanded rad `rarm_j1..j7` + gripper score (≥0.5 → close, else open); **already denormalized** | | Timing | 10 rows at 0.1 s; row 0 = command for the observation time; execute all 10 then re-query (1 s replanning) | | History | none (1 image per cam + 1 state) | | Conventions | identical to the dataset export: no sign flips, offsets or wrapping | Do not re-apply normalization. Do not substitute a missing camera. Gripper conversion score→servo command is the adapter's job. ## 2. Install (inference machine, ≥16 GB NVIDIA GPU, tested on RTX 5090 / target RTX 5080 16 GB) ```bash uv venv .venv --python 3.12 && source .venv/bin/activate uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128 git clone https://github.com/huggingface/lerobot && (cd lerobot && git checkout 3f2c29ef7e44b1ddccbcda3b6a63939e53639e9e) uv pip install -e "lerobot[molmoact2,dataset]" fastapi uvicorn pillow scipy "huggingface_hub[cli,hf_transfer]" hf auth login # token with read access to this repo HF_HUB_ENABLE_HF_TRANSFER=1 hf download Kavin60606/crimson-ma2-ckpts --local-dir ckpts # all 6 (76 GB) — or --include "s2000/001000/*" hf download CrimsonRobot/crimson-first-dateset --repo-type dataset --include "meta/*" --local-dir dataset # feature definitions (tiny) cp -r ckpts/model-integration/* . ``` ## 3. Run the inference API (no motor access) ```bash HF_HUB_OFFLINE=1 python serve_policy.py --ckpt ckpts/s2000/001000 --dataset-meta-root dataset --port 8090 --cuda-graph python inference_example.py --url http://localhost:8090 --seed 0 # runs sample/observation.json, checks against sample/prediction.json ``` Endpoints: `GET /health`, `POST /reset` (call at every episode start), `POST /predict` Request: `{obs_id, t_obs, state:[8], images:{top:, right_wrist:}, task, seed?}` Response: `{obs_id, session, seq, actions:[10][8], inference_ms, denormalized:true, sample_interval_s:0.1, ...}` Errors: HTTP 400 on wrong state length / non-finite / missing camera / wrong image size / blank image; HTTP 500 on non-finite output. One request at a time; a stale result is never returned. Measured latency (RTX 5090, bf16, CUDA graph, bs=1): inference p50 227 ms, p95 232 ms; warm-up first call ~1.3 s (send 3 dummy requests). Keep one HTTP connection open; JPEG quality 75 is enough (120 KB/request). On a local RTX 5080 expect ~250–300 ms end-to-end (estimate, not measured). ## 4. Connect Genome to the model (adapter side, owned by Genome) At 10 Hz: 1. grab latest `cam_high`, `cam_wrist` (RGB 640×480 JPEG) + measured joint state in dataset order; 2. `POST /predict` with a fresh `obs_id`; 3. apply Genome limits / rate limits / collision policy; gripper: `actions[k][7] >= 0.5 → close`; 4. execute rows 0..9 at 0.1 s spacing (or, with ~0.3 s latency, skip the first `ceil(latency/0.1)` rows and re-request when 5 rows remain); 5. `POST /reset` at every new episode, takeover or tool change. Adapter must reject a chunk older than 1 s and never re-use one. ## 5. Measure success rate (this is the deliverable) **Protocol** (same for every checkpoint): - Scene: 1 bottle, 5 marked start positions on the table (P1–P5) spanning the training distribution; fixed place target; same lighting/camera poses as the pilot recording. - Start pose: dataset start pose, gripper open. `POST /reset` before each rollout. - Timeout 30 s per rollout. Operator stops on unsafe motion (counts as failure, note the reason). - **10 rollouts per checkpoint** = P1–P5 × 2 reps → 60 rollouts total. Interleave checkpoints (do P1 for all 6 ckpts, then P2 …) so drift/fatigue/lighting average out. `model-integration/rollout_order.csv` gives the exact order. - Score each rollout with 4 binary stages (standard for pick-and-place reporting): | stage | criterion | |---|---| | reach | gripper fingers reach the bottle (within ~2 cm) | | grasp | gripper closes on the bottle and lifts it ≥ 3 cm | | carry | bottle transported above the place target | | place = **success** | bottle released at the target, upright, within the target marking, arm retreats | - Record: checkpoint, position, rep, the 4 stage flags, failure reason (free text), and if possible the episode as a LeRobot recording. **Helper**: `python dashboard.py --ckpt-root ckpts --results-root results --dataset-meta-root dataset --port 8091 --serve-port 8090` → web page on `localhost:8091` to switch the served checkpoint (restarts the API on 8090) and log rollouts to `rollouts.csv`; shows per-checkpoint SR next to the CI-MSE rank and a live Spearman. Optional; a spreadsheet with the same columns is fine. **Report** (per checkpoint): SR = successes/10 with Wilson 95% interval; stage-wise rates (reach/grasp/carry/place); dominant failure mode. Overall: Spearman between CI-MSE rank and SR rank across the 6 checkpoints, plus whether the CI-MSE top-1/top-2 contains the SR top-1. With 10 rollouts a difference of < 2 successes between checkpoints is not significant; treat it as a tie. Send back `rollouts.csv` (or the sheet). `code/correlate_crimson.py` in `Kavin60606/crimson-cimse-experiment` computes the correlation report from it. ## 6. Limitations Single task, single object, single scene, 40 training episodes; right arm only; no left arm, no language generalization. Per-joint dataset ranges (rad): j1 [−0.36,0.49] j2 [−0.35,0.21] j3 [−0.38,0.30] j4 [−1.54,−0.87] j5 [−0.27,0.21] j6 [−0.20,0.24] j7 [−0.69,−0.23] — context only, not a safe workspace. Limits, clipping and collision policy stay in Genome.