Image-Text-to-Text
Transformers
Safetensors
qwen3_vl
vision-language
multimodal
grpo
reinforcement-learning
medical
cardiac
mri
vqa
easyr1
conversational
Instructions to use ai-mind-lab/CineMR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ai-mind-lab/CineMR with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ai-mind-lab/CineMR") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ai-mind-lab/CineMR") model = AutoModelForMultimodalLM.from_pretrained("ai-mind-lab/CineMR", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ai-mind-lab/CineMR with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ai-mind-lab/CineMR" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-mind-lab/CineMR", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ai-mind-lab/CineMR
- SGLang
How to use ai-mind-lab/CineMR with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ai-mind-lab/CineMR" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-mind-lab/CineMR", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ai-mind-lab/CineMR" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-mind-lab/CineMR", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ai-mind-lab/CineMR with Docker Model Runner:
docker model run hf.co/ai-mind-lab/CineMR
Add files using upload-large-folder tool
Browse files- README.md +37 -4
- model.safetensors +1 -1
README.md
CHANGED
|
@@ -26,6 +26,17 @@ This Hub release is the **merged full-weight** GRPO export — a single `model.s
|
|
| 26 |
|
| 27 |
**Training data:** [`ai-mind-lab/CineMR`](https://huggingface.co/datasets/ai-mind-lab/CineMR) (ACDC, M&Ms, M&Ms-2 cardiac MRI VQA with optional tool-use supervision).
|
| 28 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
## Model summary
|
| 30 |
|
| 31 |
| | |
|
|
@@ -93,13 +104,21 @@ text = processor.apply_chat_template(messages, tokenize=False, add_generation_pr
|
|
| 93 |
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
|
| 94 |
|
| 95 |
with torch.no_grad():
|
| 96 |
-
out = model.generate(
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 97 |
|
| 98 |
print(processor.decode(out[0], skip_special_tokens=True))
|
| 99 |
```
|
| 100 |
|
| 101 |
Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
|
| 102 |
|
|
|
|
|
|
|
| 103 |
## Training procedure (summary)
|
| 104 |
|
| 105 |
1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
|
|
@@ -113,7 +132,21 @@ LoRA weights are merged into the base checkpoint for Hub deployment.
|
|
| 113 |
|
| 114 |
## Evaluation
|
| 115 |
|
| 116 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 117 |
|
| 118 |
## Limitations
|
| 119 |
|
|
@@ -131,8 +164,8 @@ If you use CineMR, please cite the base Qwen3-VL model and acknowledge the CineM
|
|
| 131 |
|
| 132 |
```bibtex
|
| 133 |
@misc{cinemr_qwen3vl8b_grpo,
|
| 134 |
-
title = {CineMR:
|
| 135 |
-
author = {
|
| 136 |
year = {2026},
|
| 137 |
howpublished = {\url{https://huggingface.co/ai-mind-lab/CineMR}},
|
| 138 |
note = {GRPO checkpoint; dataset at huggingface.co/datasets/ai-mind-lab/CineMR},
|
|
|
|
| 26 |
|
| 27 |
**Training data:** [`ai-mind-lab/CineMR`](https://huggingface.co/datasets/ai-mind-lab/CineMR) (ACDC, M&Ms, M&Ms-2 cardiac MRI VQA with optional tool-use supervision).
|
| 28 |
|
| 29 |
+
## Authors
|
| 30 |
+
|
| 31 |
+
Kunyang Li<sup>1,†</sup>, Hai Nguyen<sup>1,2,†</sup>, Joshua Lowe<sup>1,†</sup>, Chenguang Zhao<sup>3</sup>, Peace C. Madueme<sup>3</sup>, Mehdi Hedjazi Moghari<sup>4</sup>, Mubarak Shah<sup>1,§</sup>, Pegah Khosravi<sup>1,2,§</sup>, Yuzhang Shang<sup>1,§</sup>
|
| 32 |
+
|
| 33 |
+
<sup>1</sup> Institute for Artificial Intelligence, University of Central Florida
|
| 34 |
+
<sup>2</sup> Department of Clinical Sciences, College of Medicine, University of Central Florida
|
| 35 |
+
<sup>3</sup> Nemours Children's Health, Orlando, Florida
|
| 36 |
+
<sup>4</sup> West Virginia University Medicine Children's Hospital, Morgantown, West Virginia
|
| 37 |
+
|
| 38 |
+
<sup>†</sup> Co-first author · <sup>§</sup> Corresponding author
|
| 39 |
+
|
| 40 |
## Model summary
|
| 41 |
|
| 42 |
| | |
|
|
|
|
| 104 |
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
|
| 105 |
|
| 106 |
with torch.no_grad():
|
| 107 |
+
out = model.generate(
|
| 108 |
+
**inputs,
|
| 109 |
+
max_new_tokens=2048,
|
| 110 |
+
do_sample=True,
|
| 111 |
+
temperature=0.7,
|
| 112 |
+
repetition_penalty=1.15,
|
| 113 |
+
)
|
| 114 |
|
| 115 |
print(processor.decode(out[0], skip_special_tokens=True))
|
| 116 |
```
|
| 117 |
|
| 118 |
Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
|
| 119 |
|
| 120 |
+
> **Decoding note.** Pure greedy decoding (`do_sample=False`, no repetition penalty) can drive this checkpoint into **repetition collapse** (a single reasoning sentence repeated until the token cap, with no `\boxed{}` answer or `<tool_call>` emitted). The evaluation numbers below were produced with `do_sample=True`, `temperature=0.7`, `repetition_penalty=1.15`, `no_repeat_ngram_size=0`, `max_new_tokens=2048`, and 4 sampled rollouts per prompt. Use a repetition penalty (≈1.1–1.2) for stable outputs.
|
| 121 |
+
|
| 122 |
## Training procedure (summary)
|
| 123 |
|
| 124 |
1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
|
|
|
|
| 132 |
|
| 133 |
## Evaluation
|
| 134 |
|
| 135 |
+
Evaluated on the **CineMR test split** (1,191 samples) with the training prompt template, 4 sampled rollouts per prompt (`temperature=0.7`, `repetition_penalty=1.15`, `max_new_tokens=2048`). `pass@k` is the fraction of items with ≥1 correct rollout; `mean rollout acc` averages correctness over all rollouts.
|
| 136 |
+
|
| 137 |
+
| Metric | Value |
|
| 138 |
+
|---|---|
|
| 139 |
+
| Mean rollout accuracy | 0.378 |
|
| 140 |
+
| pass@4 (any correct) | 0.553 |
|
| 141 |
+
| ROUGE-L | 0.620 |
|
| 142 |
+
| BERTScore F1 | 0.974 |
|
| 143 |
+
| Ground-truth satisfied | 0.370 |
|
| 144 |
+
|
| 145 |
+
**Accuracy by reasoning layer (pass@4):** L1 0.306, L2 0.771, L3 0.906, L4 0.614, L5 0.314, L6 0.563. By clinical stage: phenotype (L1–L4) pass@4 0.572, etiology (L5–L6) pass@4 0.353.
|
| 146 |
+
|
| 147 |
+
**Tool use:** tool-decision accuracy 0.891 (precision 0.999), tool recall on required items 99.8% (predicted names ⊆ expected 100%), trace/JSON format validity 99.4%, tool-name set-match 0.862, argument accuracy 0.879.
|
| 148 |
+
|
| 149 |
+
Numbers are from this GRPO checkpoint evaluated with the project `eval_sft` pipeline. Re-run on a held-out test split before drawing conclusions; the small GRPO validation set is used only for checkpoint tracking.
|
| 150 |
|
| 151 |
## Limitations
|
| 152 |
|
|
|
|
| 164 |
|
| 165 |
```bibtex
|
| 166 |
@misc{cinemr_qwen3vl8b_grpo,
|
| 167 |
+
title = {CineMR: Augmenting Vision-Language Models with Tool-Integrated Reasoning for Quantitative Cardiac MRI Diagnosis},
|
| 168 |
+
author = {Li, Kunyang and Nguyen, Hai and Lowe, Joshua and Zhao, Chenguang and Madueme, Peace C. and Moghari, Mehdi Hedjazi and Shah, Mubarak and Khosravi, Pegah and Shang, Yuzhang},
|
| 169 |
year = {2026},
|
| 170 |
howpublished = {\url{https://huggingface.co/ai-mind-lab/CineMR}},
|
| 171 |
note = {GRPO checkpoint; dataset at huggingface.co/datasets/ai-mind-lab/CineMR},
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 17534340584
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e069a4035458cc235f9b782900a5e17a4ad6d8f379c66cdf0759b0ed7fe5564b
|
| 3 |
size 17534340584
|