--- license: apache-2.0 language: - en - zh tags: - robotics - grasping - vision-language-model - locate-anything - joint - multi-task pipeline_tag: image-text-to-text --- # Grasp Anything 2D — Checkpoint 2500 (joint 联合) 基于 NVIDIA [LocateAnything-3B](https://huggingface.co/nvidia/LocateAnything-3B) 的语言引导二维抓取模型,joint 联合训练产物。单 checkpoint 同时支持 `grasp_contact` 和 `grasp_rect` 两种任务。 ## 任务 `grasp_contact` + `grasp_rect` 联合 — 单 checkpoint 双任务,共享 LLM backbone。 - contact: 输出平行夹爪两个接触点 `(x1, y1, x2, y2)` - rect: 输出矩形抓取 `(cx, cy, θ, width)` ## 评测结果(RealVLG 官方 mini633) | Split | n | gAcc_corrected_strict | mIoU_strict | center_err_median (px) | |---|---|---|---|---| | seen | 243 | **61.73%** | 55.07% | 11.11 | | similar | 226 | **53.10%** | 46.85% | 15.00 | | novel | 164 | 26.83% | 25.93% | 25.36 | - 评测协议: `evaluate_realvlg_contact.py`,fast 模式 - **similar split gAcc 53.1% — 所有 checkpoint 中最高**(双任务正则带来更强泛化) - 训练量仅 contact 专用的 1/8,性能却与 5x 训练量的 SFT checkpoint 持平 ## 使用方法 ```python from locate_anything_service.model import LocateAnythingRuntime from locate_anything_service.config import Settings # joint checkpoint 两种模式都能跑 settings = Settings(model_id="charlesH777/grasp-anything-joint-2500") runtime = LocateAnythingRuntime(settings) # contact 任务 result = runtime.predict("image.jpg", "抓取红色杯子", mode="grasp_contact") # rect 任务 result = runtime.predict("image.jpg", "抓取红色杯子", mode="grasp_rect") ``` ## 训练配置 - 阶段: joint geometry (contact + grasp_rect 双任务联合) - LLM LoRA: rank 32, target Qwen2.5-3B-Instruct - 训练量: seen_contact_blocks = 19,941, seen_grasp_rect_blocks = 20,005 - contact_pair_weight: 1.0, grasp_rect_pose_weight: 1.0