viennh2012 commited on
Commit
ca2a637
·
verified ·
1 Parent(s): 353e5e8

Add files using upload-large-folder tool

Browse files
Files changed (2) hide show
  1. README.md +37 -4
  2. model.safetensors +1 -1
README.md CHANGED
@@ -26,6 +26,17 @@ This Hub release is the **merged full-weight** GRPO export — a single `model.s
26
 
27
  **Training data:** [`ai-mind-lab/CineMR`](https://huggingface.co/datasets/ai-mind-lab/CineMR) (ACDC, M&Ms, M&Ms-2 cardiac MRI VQA with optional tool-use supervision).
28
 
 
 
 
 
 
 
 
 
 
 
 
29
  ## Model summary
30
 
31
  | | |
@@ -93,13 +104,21 @@ text = processor.apply_chat_template(messages, tokenize=False, add_generation_pr
93
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
94
 
95
  with torch.no_grad():
96
- out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
 
 
 
 
 
 
97
 
98
  print(processor.decode(out[0], skip_special_tokens=True))
99
  ```
100
 
101
  Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
102
 
 
 
103
  ## Training procedure (summary)
104
 
105
  1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
@@ -113,7 +132,21 @@ LoRA weights are merged into the base checkpoint for Hub deployment.
113
 
114
  ## Evaluation
115
 
116
- Evaluate on the CineMR test split with the same frame paths and prompt template as training. Report metrics on extracted `\boxed{}` answers and, if applicable, tool-call correctness. See the project eval scripts under `CardiacCine/cardiac_cine/cine-cogito/` for reference pipelines.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
117
 
118
  ## Limitations
119
 
@@ -131,8 +164,8 @@ If you use CineMR, please cite the base Qwen3-VL model and acknowledge the CineM
131
 
132
  ```bibtex
133
  @misc{cinemr_qwen3vl8b_grpo,
134
- title = {CineMR: Cardiac MRI Vision-Language Model (Qwen3-VL-8B, GRPO)},
135
- author = {AI Mind Lab},
136
  year = {2026},
137
  howpublished = {\url{https://huggingface.co/ai-mind-lab/CineMR}},
138
  note = {GRPO checkpoint; dataset at huggingface.co/datasets/ai-mind-lab/CineMR},
 
26
 
27
  **Training data:** [`ai-mind-lab/CineMR`](https://huggingface.co/datasets/ai-mind-lab/CineMR) (ACDC, M&Ms, M&Ms-2 cardiac MRI VQA with optional tool-use supervision).
28
 
29
+ ## Authors
30
+
31
+ Kunyang Li<sup>1,†</sup>, Hai Nguyen<sup>1,2,†</sup>, Joshua Lowe<sup>1,†</sup>, Chenguang Zhao<sup>3</sup>, Peace C. Madueme<sup>3</sup>, Mehdi Hedjazi Moghari<sup>4</sup>, Mubarak Shah<sup>1,§</sup>, Pegah Khosravi<sup>1,2,§</sup>, Yuzhang Shang<sup>1,§</sup>
32
+
33
+ <sup>1</sup> Institute for Artificial Intelligence, University of Central Florida
34
+ <sup>2</sup> Department of Clinical Sciences, College of Medicine, University of Central Florida
35
+ <sup>3</sup> Nemours Children's Health, Orlando, Florida
36
+ <sup>4</sup> West Virginia University Medicine Children's Hospital, Morgantown, West Virginia
37
+
38
+ <sup>†</sup> Co-first author &nbsp;·&nbsp; <sup>§</sup> Corresponding author
39
+
40
  ## Model summary
41
 
42
  | | |
 
104
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
105
 
106
  with torch.no_grad():
107
+ out = model.generate(
108
+ **inputs,
109
+ max_new_tokens=2048,
110
+ do_sample=True,
111
+ temperature=0.7,
112
+ repetition_penalty=1.15,
113
+ )
114
 
115
  print(processor.decode(out[0], skip_special_tokens=True))
116
  ```
117
 
118
  Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
119
 
120
+ > **Decoding note.** Pure greedy decoding (`do_sample=False`, no repetition penalty) can drive this checkpoint into **repetition collapse** (a single reasoning sentence repeated until the token cap, with no `\boxed{}` answer or `<tool_call>` emitted). The evaluation numbers below were produced with `do_sample=True`, `temperature=0.7`, `repetition_penalty=1.15`, `no_repeat_ngram_size=0`, `max_new_tokens=2048`, and 4 sampled rollouts per prompt. Use a repetition penalty (≈1.1–1.2) for stable outputs.
121
+
122
  ## Training procedure (summary)
123
 
124
  1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
 
132
 
133
  ## Evaluation
134
 
135
+ Evaluated on the **CineMR test split** (1,191 samples) with the training prompt template, 4 sampled rollouts per prompt (`temperature=0.7`, `repetition_penalty=1.15`, `max_new_tokens=2048`). `pass@k` is the fraction of items with ≥1 correct rollout; `mean rollout acc` averages correctness over all rollouts.
136
+
137
+ | Metric | Value |
138
+ |---|---|
139
+ | Mean rollout accuracy | 0.378 |
140
+ | pass@4 (any correct) | 0.553 |
141
+ | ROUGE-L | 0.620 |
142
+ | BERTScore F1 | 0.974 |
143
+ | Ground-truth satisfied | 0.370 |
144
+
145
+ **Accuracy by reasoning layer (pass@4):** L1 0.306, L2 0.771, L3 0.906, L4 0.614, L5 0.314, L6 0.563. By clinical stage: phenotype (L1–L4) pass@4 0.572, etiology (L5–L6) pass@4 0.353.
146
+
147
+ **Tool use:** tool-decision accuracy 0.891 (precision 0.999), tool recall on required items 99.8% (predicted names ⊆ expected 100%), trace/JSON format validity 99.4%, tool-name set-match 0.862, argument accuracy 0.879.
148
+
149
+ Numbers are from this GRPO checkpoint evaluated with the project `eval_sft` pipeline. Re-run on a held-out test split before drawing conclusions; the small GRPO validation set is used only for checkpoint tracking.
150
 
151
  ## Limitations
152
 
 
164
 
165
  ```bibtex
166
  @misc{cinemr_qwen3vl8b_grpo,
167
+ title = {CineMR: Augmenting Vision-Language Models with Tool-Integrated Reasoning for Quantitative Cardiac MRI Diagnosis},
168
+ author = {Li, Kunyang and Nguyen, Hai and Lowe, Joshua and Zhao, Chenguang and Madueme, Peace C. and Moghari, Mehdi Hedjazi and Shah, Mubarak and Khosravi, Pegah and Shang, Yuzhang},
169
  year = {2026},
170
  howpublished = {\url{https://huggingface.co/ai-mind-lab/CineMR}},
171
  note = {GRPO checkpoint; dataset at huggingface.co/datasets/ai-mind-lab/CineMR},
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:455f92d2f358a15d62a4f934970a876d66a3edd7618d677d7275b5b0e131ba53
3
  size 17534340584
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e069a4035458cc235f9b782900a5e17a4ad6d8f379c66cdf0759b0ed7fe5564b
3
  size 17534340584