Instructions to use Kavin60606/crimson-ma2-ckpts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Kavin60606/crimson-ma2-ckpts with LeRobot:
- Notebooks
- Google Colab
- Kaggle
- Crimson right-arm MolmoAct2 checkpoints
- 1. What the model expects / returns (summary of the contract)
- 2. Install (inference machine, ≥16 GB NVIDIA GPU, tested on RTX 5090 / target RTX 5080 16 GB)
- 3. Run the inference API (no motor access)
- 4. Connect Genome to the model (adapter side, owned by Genome)
- 5. Measure success rate (this is the deliverable)
- 6. Limitations
- 1. What the model expects / returns (summary of the contract)
Crimson right-arm MolmoAct2 checkpoints
Six MolmoAct2 checkpoints fine-tuned on CrimsonRobot/crimson-first-dateset (right arm, "pick and place the bottle").
Prepared per the Crimson/Genome model-integration contract. Full contract: model-integration/model-contract.yaml.
| path in repo | seed | step | CI-MSE (std, lower=better) | rank |
|---|---|---|---|---|
s2000/001000 |
2000 | 1000 | 0.283 | 1 |
s1000/001500 |
1000 | 1500 | 0.295 | 2 |
s1000/000500 |
1000 | 500 | 0.322 | 3 |
s2000/000500 |
2000 | 500 | 0.347 | 4 |
s2000/001500 |
2000 | 1500 | 0.352 | 5 |
s1000/001000 |
1000 | 1000 | 0.373 | 6 |
Each folder is a LeRobot pretrained_model (config.json, model.safetensors 12.7 GB, pre/post-processor with normalization stats, train_config.json).
CI-MSE = offline metric on 10 held-out episodes (40–49); it is a ranking hypothesis. The point of deployment is to measure real success rate per checkpoint and check whether this ranking holds.
1. What the model expects / returns (summary of the contract)
| Arm | right only |
| State in (8) | rarm_j1..rarm_j7 measured rad, right_gripper measured servo rad (dataset observation.state, unchanged) |
| Cameras in | top = cam_high, right_wrist = cam_wrist; RGB 640×480; no crop/rotation at inference |
| Task | "pick and place the bottle" (only trained string) |
| Action out (10×8) | absolute commanded rad rarm_j1..j7 + gripper score (≥0.5 → close, else open); already denormalized |
| Timing | 10 rows at 0.1 s; row 0 = command for the observation time; execute all 10 then re-query (1 s replanning) |
| History | none (1 image per cam + 1 state) |
| Conventions | identical to the dataset export: no sign flips, offsets or wrapping |
Do not re-apply normalization. Do not substitute a missing camera. Gripper conversion score→servo command is the adapter's job.
2. Install (inference machine, ≥16 GB NVIDIA GPU, tested on RTX 5090 / target RTX 5080 16 GB)
uv venv .venv --python 3.12 && source .venv/bin/activate
uv pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
git clone https://github.com/huggingface/lerobot && (cd lerobot && git checkout 3f2c29ef7e44b1ddccbcda3b6a63939e53639e9e)
uv pip install -e "lerobot[molmoact2,dataset]" fastapi uvicorn pillow scipy "huggingface_hub[cli,hf_transfer]"
hf auth login # token with read access to this repo
HF_HUB_ENABLE_HF_TRANSFER=1 hf download Kavin60606/crimson-ma2-ckpts --local-dir ckpts # all 6 (76 GB) — or --include "s2000/001000/*"
hf download CrimsonRobot/crimson-first-dateset --repo-type dataset --include "meta/*" --local-dir dataset # feature definitions (tiny)
cp -r ckpts/model-integration/* .
3. Run the inference API (no motor access)
HF_HUB_OFFLINE=1 python serve_policy.py --ckpt ckpts/s2000/001000 --dataset-meta-root dataset --port 8090 --cuda-graph
python inference_example.py --url http://localhost:8090 --seed 0 # runs sample/observation.json, checks against sample/prediction.json
Endpoints: GET /health, POST /reset (call at every episode start), POST /predict
Request: {obs_id, t_obs, state:[8], images:{top:<b64 jpeg>, right_wrist:<b64 jpeg>}, task, seed?}
Response: {obs_id, session, seq, actions:[10][8], inference_ms, denormalized:true, sample_interval_s:0.1, ...}
Errors: HTTP 400 on wrong state length / non-finite / missing camera / wrong image size / blank image; HTTP 500 on non-finite output. One request at a time; a stale result is never returned.
Measured latency (RTX 5090, bf16, CUDA graph, bs=1): inference p50 227 ms, p95 232 ms; warm-up first call ~1.3 s (send 3 dummy requests). Keep one HTTP connection open; JPEG quality 75 is enough (120 KB/request). On a local RTX 5080 expect ~250–300 ms end-to-end (estimate, not measured).
4. Connect Genome to the model (adapter side, owned by Genome)
At 10 Hz:
- grab latest
cam_high,cam_wrist(RGB 640×480 JPEG) + measured joint state in dataset order; POST /predictwith a freshobs_id;- apply Genome limits / rate limits / collision policy; gripper:
actions[k][7] >= 0.5 → close; - execute rows 0..9 at 0.1 s spacing (or, with ~0.3 s latency, skip the first
ceil(latency/0.1)rows and re-request when 5 rows remain); POST /resetat every new episode, takeover or tool change.
Adapter must reject a chunk older than 1 s and never re-use one.
5. Measure success rate (this is the deliverable)
Protocol (same for every checkpoint):
- Scene: 1 bottle, 5 marked start positions on the table (P1–P5) spanning the training distribution; fixed place target; same lighting/camera poses as the pilot recording.
- Start pose: dataset start pose, gripper open.
POST /resetbefore each rollout. - Timeout 30 s per rollout. Operator stops on unsafe motion (counts as failure, note the reason).
- 10 rollouts per checkpoint = P1–P5 × 2 reps → 60 rollouts total. Interleave checkpoints (do P1 for all 6 ckpts, then P2 …) so drift/fatigue/lighting average out.
model-integration/rollout_order.csvgives the exact order. - Score each rollout with 4 binary stages (standard for pick-and-place reporting):
| stage | criterion |
|---|---|
| reach | gripper fingers reach the bottle (within ~2 cm) |
| grasp | gripper closes on the bottle and lifts it ≥ 3 cm |
| carry | bottle transported above the place target |
| place = success | bottle released at the target, upright, within the target marking, arm retreats |
- Record: checkpoint, position, rep, the 4 stage flags, failure reason (free text), and if possible the episode as a LeRobot recording.
Helper: python dashboard.py --ckpt-root ckpts --results-root results --dataset-meta-root dataset --port 8091 --serve-port 8090
→ web page on localhost:8091 to switch the served checkpoint (restarts the API on 8090) and log rollouts to rollouts.csv; shows per-checkpoint SR next to the CI-MSE rank and a live Spearman. Optional; a spreadsheet with the same columns is fine.
Report (per checkpoint): SR = successes/10 with Wilson 95% interval; stage-wise rates (reach/grasp/carry/place); dominant failure mode. Overall: Spearman between CI-MSE rank and SR rank across the 6 checkpoints, plus whether the CI-MSE top-1/top-2 contains the SR top-1. With 10 rollouts a difference of < 2 successes between checkpoints is not significant; treat it as a tie.
Send back rollouts.csv (or the sheet). code/correlate_crimson.py in Kavin60606/crimson-cimse-experiment computes the correlation report from it.
6. Limitations
Single task, single object, single scene, 40 training episodes; right arm only; no left arm, no language generalization. Per-joint dataset ranges (rad): j1 [−0.36,0.49] j2 [−0.35,0.21] j3 [−0.38,0.30] j4 [−1.54,−0.87] j5 [−0.27,0.21] j6 [−0.20,0.24] j7 [−0.69,−0.23] — context only, not a safe workspace. Limits, clipping and collision policy stay in Genome.
Model tree for Kavin60606/crimson-ma2-ckpts
Base model
allenai/MolmoAct2-DROID