viennh2012 commited on
Commit
09d3184
·
verified ·
1 Parent(s): ca2a637

Add files using upload-large-folder tool

Browse files
Files changed (4) hide show
  1. README.md +15 -12
  2. config.json +1 -1
  3. generation_config.json +1 -1
  4. model.safetensors +1 -1
README.md CHANGED
@@ -115,10 +115,10 @@ with torch.no_grad():
115
  print(processor.decode(out[0], skip_special_tokens=True))
116
  ```
117
 
118
- Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
119
-
120
  > **Decoding note.** Pure greedy decoding (`do_sample=False`, no repetition penalty) can drive this checkpoint into **repetition collapse** (a single reasoning sentence repeated until the token cap, with no `\boxed{}` answer or `<tool_call>` emitted). The evaluation numbers below were produced with `do_sample=True`, `temperature=0.7`, `repetition_penalty=1.15`, `no_repeat_ngram_size=0`, `max_new_tokens=2048`, and 4 sampled rollouts per prompt. Use a repetition penalty (≈1.1–1.2) for stable outputs.
121
 
 
 
122
  ## Training procedure (summary)
123
 
124
  1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
@@ -132,26 +132,29 @@ LoRA weights are merged into the base checkpoint for Hub deployment.
132
 
133
  ## Evaluation
134
 
135
- Evaluated on the **CineMR test split** (1,191 samples) with the training prompt template, 4 sampled rollouts per prompt (`temperature=0.7`, `repetition_penalty=1.15`, `max_new_tokens=2048`). `pass@k` is the fraction of items with ≥1 correct rollout; `mean rollout acc` averages correctness over all rollouts.
136
 
137
  | Metric | Value |
138
  |---|---|
139
- | Mean rollout accuracy | 0.378 |
140
- | pass@4 (any correct) | 0.553 |
141
- | ROUGE-L | 0.620 |
142
- | BERTScore F1 | 0.974 |
143
- | Ground-truth satisfied | 0.370 |
 
 
144
 
145
- **Accuracy by reasoning layer (pass@4):** L1 0.306, L2 0.771, L3 0.906, L4 0.614, L5 0.314, L6 0.563. By clinical stage: phenotype (L1–L4) pass@4 0.572, etiology (L5–L6) pass@4 0.353.
146
 
147
- **Tool use:** tool-decision accuracy 0.891 (precision 0.999), tool recall on required items 99.8% (predicted names ⊆ expected 100%), trace/JSON format validity 99.4%, tool-name set-match 0.862, argument accuracy 0.879.
148
 
149
- Numbers are from this GRPO checkpoint evaluated with the project `eval_sft` pipeline. Re-run on a held-out test split before drawing conclusions; the small GRPO validation set is used only for checkpoint tracking.
150
 
151
  ## Limitations
152
 
153
  - Trained on public cardiac MRI challenge-style corpora (ACDC, M&Ms, M&Ms-2); generalization to other scanners, sequences, or pathologies is not guaranteed.
154
- - GRPO training used a small validation set for checkpoint tracking; prefer held-out test evaluation before drawing conclusions.
 
155
  - Tool-use formatting in outputs may be inconsistent unless prompts and decoding match training.
156
 
157
  ## License
 
115
  print(processor.decode(out[0], skip_special_tokens=True))
116
  ```
117
 
 
 
118
  > **Decoding note.** Pure greedy decoding (`do_sample=False`, no repetition penalty) can drive this checkpoint into **repetition collapse** (a single reasoning sentence repeated until the token cap, with no `\boxed{}` answer or `<tool_call>` emitted). The evaluation numbers below were produced with `do_sample=True`, `temperature=0.7`, `repetition_penalty=1.15`, `no_repeat_ngram_size=0`, `max_new_tokens=2048`, and 4 sampled rollouts per prompt. Use a repetition penalty (≈1.1–1.2) for stable outputs.
119
 
120
+ Use the same `trust_remote_code=True` and bfloat16 settings as in training. For evaluation, match the CineMR prompt template and decoding settings used in your eval script.
121
+
122
  ## Training procedure (summary)
123
 
124
  1. **SFT** on CineMR JSONL (train split) starting from Qwen3-VL-8B-Instruct; weights merged to a full `transformers` checkpoint.
 
132
 
133
  ## Evaluation
134
 
135
+ Evaluated on the **CineMR held-out test split** (n=3,320; distinct from the 1,191-sample validation split) with the training prompt template, 4 sampled rollouts per prompt (`temperature=0.7`, `repetition_penalty=1.15`, `max_new_tokens=2048`). `pass@4` is the fraction of items with ≥1 correct rollout; `mean rollout acc` averages correctness over all 4 rollouts (this is the metric reported as "Acc." in the paper). All figures below are the mean ± SD over 3 independent evaluation runs of this exact checkpoint (same weights, stochastic decoding only — this captures rollout-sampling variance, not training-seed variance).
136
 
137
  | Metric | Value |
138
  |---|---|
139
+ | Mean rollout accuracy | 39.24% ± 0.09pp |
140
+ | pass@4 (any correct) | 55.98% |
141
+ | Ground-truth satisfied (pass@1) | 38.83% |
142
+ | ROUGE-L$^\dagger$ | 0.620 |
143
+ | BERTScore F1$^\dagger$ | 0.974 |
144
+
145
+ $^\dagger$ROUGE-L and BERTScore F1 have not been re-measured on this checkpoint; these two figures carry over from the original submission's checkpoint.
146
 
147
+ **Accuracy by reasoning layer (mean rollout acc):** L1 10.87%, L2 73.22%, L3 66.90%, L4 55.84%, L5 11.46%, L6 37.25%. By clinical stage: L3–4 (clinical-criteria) 65.30% ± 0.66pp, L5–6 (full-diagnosis) 17.02% ± 1.41pp.
148
 
149
+ **Tool use:** tool-decision accuracy 89.61% (precision 100%), tool recall on required items 99.99%, redundant tool calls on optional items 0.00%, trace/JSON format validity 99.98%, tool-name set-match 88.28%, argument accuracy 88.28%, predicted names ⊆ expected 100%.
150
 
151
+ Numbers are from this exact GRPO checkpoint (`global_step_153`, from `sft_v6_live_tool`) evaluated with the project `eval_sft` pipeline. See the project eval scripts under `<cardiac_cine_repo>/cardiac_cine/cine-cogito/` for reference pipelines.
152
 
153
  ## Limitations
154
 
155
  - Trained on public cardiac MRI challenge-style corpora (ACDC, M&Ms, M&Ms-2); generalization to other scanners, sequences, or pathologies is not guaranteed.
156
+ - The evaluation above is held-out test-set accuracy (mean ± SD over 3 stochastic-decoding runs of this one checkpoint), which captures rollout-sampling variance but not training-seed variance — true multi-seed (independently trained) variance has not been measured.
157
+ - A non-VLM control that runs all six computational tools unconditionally and routes the outputs through the same clinical decision tree, with no VLM at all, currently exceeds this checkpoint's own diagnostic accuracy on the same test set. This checkpoint is a proof-of-concept that tool-integrated reasoning is necessary for this task, not evidence that it is the best way to obtain it — see the paper for the full discussion.
158
  - Tool-use formatting in outputs may be inconsistent unless prompts and decoding match training.
159
 
160
  ## License
config.json CHANGED
@@ -48,7 +48,7 @@
48
  "vocab_size": 151936
49
  },
50
  "tie_word_embeddings": false,
51
- "transformers_version": "5.8.0",
52
  "video_token_id": 151656,
53
  "vision_config": {
54
  "deepstack_visual_indexes": [
 
48
  "vocab_size": 151936
49
  },
50
  "tie_word_embeddings": false,
51
+ "transformers_version": "5.13.0",
52
  "video_token_id": 151656,
53
  "vision_config": {
54
  "deepstack_visual_indexes": [
generation_config.json CHANGED
@@ -9,5 +9,5 @@
9
  "temperature": 0.7,
10
  "top_k": 20,
11
  "top_p": 0.8,
12
- "transformers_version": "5.8.0"
13
  }
 
9
  "temperature": 0.7,
10
  "top_k": 20,
11
  "top_p": 0.8,
12
+ "transformers_version": "5.13.0"
13
  }
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e069a4035458cc235f9b782900a5e17a4ad6d8f379c66cdf0759b0ed7fe5564b
3
  size 17534340584
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1cc472f8e5af19b91d59c21ed6e1c344e8251e5ac667d8f53263cd8b8dc03238
3
  size 17534340584