Crimson right-arm MolmoAct2 checkpoints

Six MolmoAct2 checkpoints fine-tuned on CrimsonRobot/crimson-first-dateset (right arm, "pick and place the bottle"). Prepared per the Crimson/Genome model-integration contract. Full contract: model-integration/model-contract.yaml.

path in repo seed step CI-MSE (std, lower=better) rank
s2000/001000 2000 1000 0.283 1
s1000/001500 1000 1500 0.295 2
s1000/000500 1000 500 0.322 3
s2000/000500 2000 500 0.347 4
s2000/001500 2000 1500 0.352 5
s1000/001000 1000 1000 0.373 6

Each folder is a LeRobot pretrained_model (config.json, model.safetensors 12.7 GB, pre/post-processor with normalization stats, train_config.json). CI-MSE = offline metric on 10 held-out episodes (40–49); it is a ranking hypothesis. The point of deployment is to measure real success rate per checkpoint and check whether this ranking holds.

1. What the model expects / returns (summary of the contract)

Arm right only
State in (8) rarm_j1..rarm_j7 measured rad, right_gripper measured servo rad (dataset observation.state, unchanged)
Cameras in top = cam_high, right_wrist = cam_wrist; RGB 640×480; no crop/rotation at inference
Task "pick and place the bottle" (only trained string)
Action out (10×8) absolute commanded rad rarm_j1..j7 + gripper score (≥0.5 → close, else open); already denormalized
Timing 10 rows at 0.1 s; row 0 = command for the observation time; execute all 10 then re-query (1 s replanning)
History none (1 image per cam + 1 state)
Conventions identical to the dataset export: no sign flips, offsets or wrapping

Do not re-apply normalization. Do not substitute a missing camera. Gripper conversion score→servo command is the adapter's job.

2. Install (inference machine, ≥16 GB NVIDIA GPU, tested on RTX 5090 / target RTX 5080 16 GB)

uv venv .venv --python 3.12 && source .venv/bin/activate
uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
git clone https://github.com/huggingface/lerobot && (cd lerobot && git checkout 3f2c29ef7e44b1ddccbcda3b6a63939e53639e9e)
uv pip install -e "lerobot[molmoact2,dataset]" fastapi uvicorn pillow scipy "huggingface_hub[cli,hf_transfer]"
hf auth login   # token with read access to this repo
HF_HUB_ENABLE_HF_TRANSFER=1 hf download Kavin60606/crimson-ma2-ckpts --local-dir ckpts            # all 6 (76 GB) — or --include "s2000/001000/*"
hf download CrimsonRobot/crimson-first-dateset --repo-type dataset --include "meta/*" --local-dir dataset   # feature definitions (tiny)
cp -r ckpts/model-integration/* .

3. Run the inference API (no motor access)

HF_HUB_OFFLINE=1 python serve_policy.py --ckpt ckpts/s2000/001000 --dataset-meta-root dataset --port 8090 --cuda-graph
python inference_example.py --url http://localhost:8090 --seed 0      # runs sample/observation.json, checks against sample/prediction.json

Endpoints: GET /health, POST /reset (call at every episode start), POST /predict Request: {obs_id, t_obs, state:[8], images:{top:<b64 jpeg>, right_wrist:<b64 jpeg>}, task, seed?} Response: {obs_id, session, seq, actions:[10][8], inference_ms, denormalized:true, sample_interval_s:0.1, ...} Errors: HTTP 400 on wrong state length / non-finite / missing camera / wrong image size / blank image; HTTP 500 on non-finite output. One request at a time; a stale result is never returned.

Measured latency (RTX 5090, bf16, CUDA graph, bs=1): inference p50 227 ms, p95 232 ms; warm-up first call ~1.3 s (send 3 dummy requests). Keep one HTTP connection open; JPEG quality 75 is enough (120 KB/request). On a local RTX 5080 expect ~250–300 ms end-to-end (estimate, not measured).

4. Connect Genome to the model (adapter side, owned by Genome)

At 10 Hz:

  1. grab latest cam_high, cam_wrist (RGB 640×480 JPEG) + measured joint state in dataset order;
  2. POST /predict with a fresh obs_id;
  3. apply Genome limits / rate limits / collision policy; gripper: actions[k][7] >= 0.5 → close;
  4. execute rows 0..9 at 0.1 s spacing (or, with ~0.3 s latency, skip the first ceil(latency/0.1) rows and re-request when 5 rows remain);
  5. POST /reset at every new episode, takeover or tool change.

Adapter must reject a chunk older than 1 s and never re-use one.

5. Measure success rate (this is the deliverable)

Protocol (same for every checkpoint):

  • Scene: 1 bottle, 5 marked start positions on the table (P1–P5) spanning the training distribution; fixed place target; same lighting/camera poses as the pilot recording.
  • Start pose: dataset start pose, gripper open. POST /reset before each rollout.
  • Timeout 30 s per rollout. Operator stops on unsafe motion (counts as failure, note the reason).
  • 10 rollouts per checkpoint = P1–P5 × 2 reps → 60 rollouts total. Interleave checkpoints (do P1 for all 6 ckpts, then P2 …) so drift/fatigue/lighting average out. model-integration/rollout_order.csv gives the exact order.
  • Score each rollout with 4 binary stages (standard for pick-and-place reporting):
stage criterion
reach gripper fingers reach the bottle (within ~2 cm)
grasp gripper closes on the bottle and lifts it ≥ 3 cm
carry bottle transported above the place target
place = success bottle released at the target, upright, within the target marking, arm retreats
  • Record: checkpoint, position, rep, the 4 stage flags, failure reason (free text), and if possible the episode as a LeRobot recording.

Helper: python dashboard.py --ckpt-root ckpts --results-root results --dataset-meta-root dataset --port 8091 --serve-port 8090 → web page on localhost:8091 to switch the served checkpoint (restarts the API on 8090) and log rollouts to rollouts.csv; shows per-checkpoint SR next to the CI-MSE rank and a live Spearman. Optional; a spreadsheet with the same columns is fine.

Report (per checkpoint): SR = successes/10 with Wilson 95% interval; stage-wise rates (reach/grasp/carry/place); dominant failure mode. Overall: Spearman between CI-MSE rank and SR rank across the 6 checkpoints, plus whether the CI-MSE top-1/top-2 contains the SR top-1. With 10 rollouts a difference of < 2 successes between checkpoints is not significant; treat it as a tie.

Send back rollouts.csv (or the sheet). code/correlate_crimson.py in Kavin60606/crimson-cimse-experiment computes the correlation report from it.

6. Limitations

Single task, single object, single scene, 40 training episodes; right arm only; no left arm, no language generalization. Per-joint dataset ranges (rad): j1 [−0.36,0.49] j2 [−0.35,0.21] j3 [−0.38,0.30] j4 [−1.54,−0.87] j5 [−0.27,0.21] j6 [−0.20,0.24] j7 [−0.69,−0.23] — context only, not a safe workspace. Limits, clipping and collision policy stay in Genome.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for Kavin60606/crimson-ma2-ckpts

Finetuned
(1)
this model