Image-Text-to-Text
Transformers
Safetensors
qwen3_vl
vision-language
multimodal
grpo
reinforcement-learning
medical
cardiac
mri
vqa
easyr1
conversational
Instructions to use ai-mind-lab/CineMR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ai-mind-lab/CineMR with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ai-mind-lab/CineMR") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ai-mind-lab/CineMR") model = AutoModelForMultimodalLM.from_pretrained("ai-mind-lab/CineMR", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ai-mind-lab/CineMR with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ai-mind-lab/CineMR" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-mind-lab/CineMR", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ai-mind-lab/CineMR
- SGLang
How to use ai-mind-lab/CineMR with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ai-mind-lab/CineMR" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-mind-lab/CineMR", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ai-mind-lab/CineMR" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-mind-lab/CineMR", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ai-mind-lab/CineMR with Docker Model Runner:
docker model run hf.co/ai-mind-lab/CineMR
Add files using upload-large-folder tool
Browse files- README.md +15 -12
- config.json +1 -1
- generation_config.json +1 -1
- model.safetensors +1 -1
README.md
CHANGED
|
@@ -115,10 +115,10 @@ with torch.no_grad():
|
|
| 115 |
print(processor.decode(out[0], skip_special_tokens=True))
|
| 116 |
```
|
| 117 |
|
| 118 |
-
Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
|
| 119 |
-
|
| 120 |
> **Decoding note.** Pure greedy decoding (`do_sample=False`, no repetition penalty) can drive this checkpoint into **repetition collapse** (a single reasoning sentence repeated until the token cap, with no `\boxed{}` answer or `<tool_call>` emitted). The evaluation numbers below were produced with `do_sample=True`, `temperature=0.7`, `repetition_penalty=1.15`, `no_repeat_ngram_size=0`, `max_new_tokens=2048`, and 4 sampled rollouts per prompt. Use a repetition penalty (≈1.1–1.2) for stable outputs.
|
| 121 |
|
|
|
|
|
|
|
| 122 |
## Training procedure (summary)
|
| 123 |
|
| 124 |
1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
|
|
@@ -132,26 +132,29 @@ LoRA weights are merged into the base checkpoint for Hub deployment.
|
|
| 132 |
|
| 133 |
## Evaluation
|
| 134 |
|
| 135 |
-
Evaluated on the **CineMR test split** (1,191
|
| 136 |
|
| 137 |
| Metric | Value |
|
| 138 |
|---|---|
|
| 139 |
-
| Mean rollout accuracy | 0.
|
| 140 |
-
| pass@4 (any correct) |
|
| 141 |
-
|
|
| 142 |
-
|
|
| 143 |
-
|
|
|
|
|
|
|
|
| 144 |
|
| 145 |
-
**Accuracy by reasoning layer (
|
| 146 |
|
| 147 |
-
**Tool use:** tool-decision accuracy
|
| 148 |
|
| 149 |
-
Numbers are from this GRPO checkpoint evaluated with the project `eval_sft` pipeline.
|
| 150 |
|
| 151 |
## Limitations
|
| 152 |
|
| 153 |
- Trained on public cardiac MRI challenge-style corpora (ACDC, M&Ms, M&Ms-2); generalization to other scanners, sequences, or pathologies is not guaranteed.
|
| 154 |
-
-
|
|
|
|
| 155 |
- Tool-use formatting in outputs may be inconsistent unless prompts and decoding match training.
|
| 156 |
|
| 157 |
## License
|
|
|
|
| 115 |
print(processor.decode(out[0], skip_special_tokens=True))
|
| 116 |
```
|
| 117 |
|
|
|
|
|
|
|
| 118 |
> **Decoding note.** Pure greedy decoding (`do_sample=False`, no repetition penalty) can drive this checkpoint into **repetition collapse** (a single reasoning sentence repeated until the token cap, with no `\boxed{}` answer or `<tool_call>` emitted). The evaluation numbers below were produced with `do_sample=True`, `temperature=0.7`, `repetition_penalty=1.15`, `no_repeat_ngram_size=0`, `max_new_tokens=2048`, and 4 sampled rollouts per prompt. Use a repetition penalty (≈1.1–1.2) for stable outputs.
|
| 119 |
|
| 120 |
+
Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
|
| 121 |
+
|
| 122 |
## Training procedure (summary)
|
| 123 |
|
| 124 |
1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
|
|
|
|
| 132 |
|
| 133 |
## Evaluation
|
| 134 |
|
| 135 |
+
Evaluated on the **CineMR held-out test split** (n=3,320; distinct from the 1,191-sample validation split) with the training prompt template, 4 sampled rollouts per prompt (`temperature=0.7`, `repetition_penalty=1.15`, `max_new_tokens=2048`). `pass@4` is the fraction of items with ≥1 correct rollout; `mean rollout acc` averages correctness over all 4 rollouts (this is the metric reported as "Acc." in the paper). All figures below are the mean ± SD over 3 independent evaluation runs of this exact checkpoint (same weights, stochastic decoding only — this captures rollout-sampling variance, not training-seed variance).
|
| 136 |
|
| 137 |
| Metric | Value |
|
| 138 |
|---|---|
|
| 139 |
+
| Mean rollout accuracy | 39.24% ± 0.09pp |
|
| 140 |
+
| pass@4 (any correct) | 55.98% |
|
| 141 |
+
| Ground-truth satisfied (pass@1) | 38.83% |
|
| 142 |
+
| ROUGE-L$^\dagger$ | 0.620 |
|
| 143 |
+
| BERTScore F1$^\dagger$ | 0.974 |
|
| 144 |
+
|
| 145 |
+
$^\dagger$ROUGE-L and BERTScore F1 have not been re-measured on this checkpoint; these two figures carry over from the original submission's checkpoint.
|
| 146 |
|
| 147 |
+
**Accuracy by reasoning layer (mean rollout acc):** L1 10.87%, L2 73.22%, L3 66.90%, L4 55.84%, L5 11.46%, L6 37.25%. By clinical stage: L3–4 (clinical-criteria) 65.30% ± 0.66pp, L5–6 (full-diagnosis) 17.02% ± 1.41pp.
|
| 148 |
|
| 149 |
+
**Tool use:** tool-decision accuracy 89.61% (precision 100%), tool recall on required items 99.99%, redundant tool calls on optional items 0.00%, trace/JSON format validity 99.98%, tool-name set-match 88.28%, argument accuracy 88.28%, predicted names ⊆ expected 100%.
|
| 150 |
|
| 151 |
+
Numbers are from this exact GRPO checkpoint (`global_step_153`, from `sft_v6_live_tool`) evaluated with the project `eval_sft` pipeline. See the project eval scripts under `<cardiac_cine_repo>/cardiac_cine/cine-cogito/` for reference pipelines.
|
| 152 |
|
| 153 |
## Limitations
|
| 154 |
|
| 155 |
- Trained on public cardiac MRI challenge-style corpora (ACDC, M&Ms, M&Ms-2); generalization to other scanners, sequences, or pathologies is not guaranteed.
|
| 156 |
+
- The evaluation above is held-out test-set accuracy (mean ± SD over 3 stochastic-decoding runs of this one checkpoint), which captures rollout-sampling variance but not training-seed variance — true multi-seed (independently trained) variance has not been measured.
|
| 157 |
+
- A non-VLM control that runs all six computational tools unconditionally and routes the outputs through the same clinical decision tree, with no VLM at all, currently exceeds this checkpoint's own diagnostic accuracy on the same test set. This checkpoint is a proof-of-concept that tool-integrated reasoning is necessary for this task, not evidence that it is the best way to obtain it — see the paper for the full discussion.
|
| 158 |
- Tool-use formatting in outputs may be inconsistent unless prompts and decoding match training.
|
| 159 |
|
| 160 |
## License
|
config.json
CHANGED
|
@@ -48,7 +48,7 @@
|
|
| 48 |
"vocab_size": 151936
|
| 49 |
},
|
| 50 |
"tie_word_embeddings": false,
|
| 51 |
-
"transformers_version": "5.
|
| 52 |
"video_token_id": 151656,
|
| 53 |
"vision_config": {
|
| 54 |
"deepstack_visual_indexes": [
|
|
|
|
| 48 |
"vocab_size": 151936
|
| 49 |
},
|
| 50 |
"tie_word_embeddings": false,
|
| 51 |
+
"transformers_version": "5.13.0",
|
| 52 |
"video_token_id": 151656,
|
| 53 |
"vision_config": {
|
| 54 |
"deepstack_visual_indexes": [
|
generation_config.json
CHANGED
|
@@ -9,5 +9,5 @@
|
|
| 9 |
"temperature": 0.7,
|
| 10 |
"top_k": 20,
|
| 11 |
"top_p": 0.8,
|
| 12 |
-
"transformers_version": "5.
|
| 13 |
}
|
|
|
|
| 9 |
"temperature": 0.7,
|
| 10 |
"top_k": 20,
|
| 11 |
"top_p": 0.8,
|
| 12 |
+
"transformers_version": "5.13.0"
|
| 13 |
}
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 17534340584
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1cc472f8e5af19b91d59c21ed6e1c344e8251e5ac667d8f53263cd8b8dc03238
|
| 3 |
size 17534340584
|