Image-Text-to-Text
PEFT
Safetensors
English
decision-model
system-one
calibration
typesafe
decision-circuits
lora
vision
Instructions to use jbarney/circuit-vl-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use jbarney/circuit-vl-4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
v1.2: options encoded side by side; answers do not depend on option order
Browse files- README.md +3 -0
- adapter/adapter_model.safetensors +1 -1
- config.json +21 -7
- head.pt +1 -1
README.md
CHANGED
|
@@ -18,6 +18,8 @@ tags:
|
|
| 18 |
|
| 19 |
# circuit-vl-4b
|
| 20 |
|
|
|
|
|
|
|
| 21 |
**v1.1.** Same base, head and recipe as v1.0, with real photographs (Open Images V7, human-verified labels) added to the training data. "Vision grid v2" below names the dataset, not the model. See [Versions](#versions).
|
| 22 |
|
| 23 |
A **System One** decision model for images: a state that carries one or
|
|
@@ -84,6 +86,7 @@ threshold.
|
|
| 84 |
|
| 85 |
| tag | date | what changed |
|
| 86 |
|---|---|---|
|
|
|
|
| 87 |
| `v1.1` | 2026-09 | Trains on vision grid v2: the rendered cells plus three cells of real photographs. |
|
| 88 |
| `v1.0` | 2026-09 | First release, rendered or scripted data only. |
|
| 89 |
|
|
|
|
| 18 |
|
| 19 |
# circuit-vl-4b
|
| 20 |
|
| 21 |
+
**v1.2** (2026-09-22). The options of a choice question are encoded side by side, so the answer cannot depend on the order they are listed in: 0.0% of answers change under reordering on its grid (v1.1: 3.6%); grid .967 / ECE .021, POPE .907 / .069. `config.json` carries `"parallel_options": true`; serve it with the circuit repo's scorer, which applies the mask. See [Versions](#versions).
|
| 22 |
+
|
| 23 |
**v1.1.** Same base, head and recipe as v1.0, with real photographs (Open Images V7, human-verified labels) added to the training data. "Vision grid v2" below names the dataset, not the model. See [Versions](#versions).
|
| 24 |
|
| 25 |
A **System One** decision model for images: a state that carries one or
|
|
|
|
| 86 |
|
| 87 |
| tag | date | what changed |
|
| 88 |
|---|---|---|
|
| 89 |
+
| `v1.2` | 2026-09-22 | Options encoded side by side: order-stable answers. Same data and recipe as v1.1 otherwise. |
|
| 90 |
| `v1.1` | 2026-09 | Trains on vision grid v2: the rendered cells plus three cells of real photographs. |
|
| 91 |
| `v1.0` | 2026-09 | First release, rendered or scripted data only. |
|
| 92 |
|
adapter/adapter_model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 132195448
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9f221f558ffb8b2116bc038632b5775bd625c88de50476152d07957b32ea7a89
|
| 3 |
size 132195448
|
config.json
CHANGED
|
@@ -12,30 +12,44 @@
|
|
| 12 |
"<|box_end|>",
|
| 13 |
"<|fim_middle|>"
|
| 14 |
],
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
"best": {
|
| 16 |
-
"ece": 0.
|
| 17 |
"accuracy": 0.9671052631578947,
|
| 18 |
-
"kl": 0.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
"step": 600
|
| 20 |
},
|
| 21 |
"args": {
|
| 22 |
"model": "Qwen/Qwen3-VL-4B-Instruct",
|
| 23 |
"data": "data/vision/grid/train.jsonl",
|
| 24 |
-
"out": "runs/circuit-vl-4b-
|
| 25 |
"epochs": 2,
|
| 26 |
-
"batch":
|
| 27 |
"lr": 0.0001,
|
| 28 |
"head_lr": 0.001,
|
| 29 |
"rank": 16,
|
| 30 |
"max_length": 1536,
|
| 31 |
"val_frac": 0.1,
|
| 32 |
-
"eval_every":
|
| 33 |
"limit": null,
|
| 34 |
"seed": 0,
|
| 35 |
"dtype": "bfloat16",
|
| 36 |
-
"
|
|
|
|
|
|
|
|
|
|
| 37 |
"wandb": "s1proto",
|
| 38 |
-
"run_name": "circuit-vl-4b-
|
| 39 |
"resume": false,
|
| 40 |
"head": "pointer",
|
| 41 |
"load_4bit": false,
|
|
|
|
| 12 |
"<|box_end|>",
|
| 13 |
"<|fim_middle|>"
|
| 14 |
],
|
| 15 |
+
"parallel_options": true,
|
| 16 |
+
"temperatures": {
|
| 17 |
+
"noul": 1.75,
|
| 18 |
+
"choice": 1.75,
|
| 19 |
+
"score": 1.0
|
| 20 |
+
},
|
| 21 |
"best": {
|
| 22 |
+
"ece": 0.03482257006199742,
|
| 23 |
"accuracy": 0.9671052631578947,
|
| 24 |
+
"kl": 0.09835513901895927,
|
| 25 |
+
"ece_by_type": {
|
| 26 |
+
"noul": 0.04722936921998076,
|
| 27 |
+
"choice": 0.031651588013538984
|
| 28 |
+
},
|
| 29 |
+
"ece_worst": 0.04722936921998076,
|
| 30 |
"step": 600
|
| 31 |
},
|
| 32 |
"args": {
|
| 33 |
"model": "Qwen/Qwen3-VL-4B-Instruct",
|
| 34 |
"data": "data/vision/grid/train.jsonl",
|
| 35 |
+
"out": "runs/circuit-vl-4b-par",
|
| 36 |
"epochs": 2,
|
| 37 |
+
"batch": 4,
|
| 38 |
"lr": 0.0001,
|
| 39 |
"head_lr": 0.001,
|
| 40 |
"rank": 16,
|
| 41 |
"max_length": 1536,
|
| 42 |
"val_frac": 0.1,
|
| 43 |
+
"eval_every": 100,
|
| 44 |
"limit": null,
|
| 45 |
"seed": 0,
|
| 46 |
"dtype": "bfloat16",
|
| 47 |
+
"acc_floor": 0.01,
|
| 48 |
+
"parallel_options": true,
|
| 49 |
+
"micro": 0,
|
| 50 |
+
"grad_checkpoint": true,
|
| 51 |
"wandb": "s1proto",
|
| 52 |
+
"run_name": "circuit-vl-4b-par",
|
| 53 |
"resume": false,
|
| 54 |
"head": "pointer",
|
| 55 |
"load_4bit": false,
|
head.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 5244685
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:fb3e5b652c3a044b9156e995f75fcf7d09a6d87e26f07ac193b162784a8ecf87
|
| 3 |
size 5244685
|