pi-dal commited on
Commit
a180642
·
verified ·
1 Parent(s): 8fa4530

Upload folder using huggingface_hub

Browse files
Files changed (3) hide show
  1. README.md +110 -0
  2. adapter.safetensors +3 -0
  3. linnaeus.json +40 -0
README.md ADDED
@@ -0,0 +1,110 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-2B
4
+ library_name: transformers
5
+ tags:
6
+ - lora
7
+ - vision-language
8
+ - decision-model
9
+ - rlcd
10
+ ---
11
+
12
+ # Linnaeus-0.1.0-2B
13
+
14
+ Built on Qwen/Qwen3.5-2B, this checkpoint returns decision distributions from text and images. It supports candidate selection (`choice`), truth estimates (`noul`), and ordered scores (`score`) without generating reasoning or free-form responses.
15
+
16
+ ## Model details
17
+
18
+ | Property | Value |
19
+ | --- | --- |
20
+ | Base model | Qwen/Qwen3.5-2B |
21
+ | Selection | Seed 42, update 2,800 |
22
+ | Inference | merged LoRA, BF16, fused operations, shared-prefix parallel candidate scoring |
23
+ | Inputs | Text and one PIL image; 2–128 candidates; 4,096 tokens per question |
24
+
25
+ Questions reuse a shared input prefix and compute their suffixes in parallel. Additional questions still require computation.
26
+
27
+ ## Training
28
+
29
+ The training recipe combines RLCD and auxiliary cross-entropy, following the pinned Laya and Laya Vision references. Temperature calibration uses an independent partition after LoRA merging. Calibration quality is measured below.
30
+
31
+ [Training recipe](../recipe.json) · [RLCD implementation and upstream attribution](../../../docs/rlcd.md)
32
+
33
+ ## Evaluation
34
+
35
+ Held-out macro accuracy: 80.78%.
36
+
37
+ | Dataset | N | Accuracy | NLL before / after calibration | ECE before / after |
38
+ | --- | ---: | ---: | ---: | ---: |
39
+ | ag_news | 7600 | 90.04% | 0.8728 / 0.5378 | 0.0832 / 0.0710 |
40
+ | aokvqa | 1138 | 83.66% | 0.6347 / 0.4746 | 0.0930 / 0.0513 |
41
+ | banking77 | 3080 | 73.70% | 1.0536 / 1.1711 | 0.0275 / 0.2001 |
42
+ | boolq | 3270 | 87.71% | 0.3854 / 0.3288 | 0.0610 / 0.0648 |
43
+ | clevr_attribute | 53734 | 98.92% | 0.1687 / 0.1230 | 0.0100 / 0.0095 |
44
+ | clevr_count | 35422 | 89.72% | 0.7159 / 0.2652 | 0.0793 / 0.0077 |
45
+ | clevr_exist | 20196 | 98.58% | 0.2526 / 0.0955 | 0.0137 / 0.0127 |
46
+ | contract_nli | 1173 | 85.59% | 0.5014 / 0.4020 | 0.0888 / 0.0338 |
47
+ | emotion | 2000 | 77.00% | 0.7237 / 0.6812 | 0.0777 / 0.0488 |
48
+ | esci_es | 1482 | 60.26% | 0.9241 / 0.9775 | 0.0342 / 0.1266 |
49
+ | esci_jp | 1633 | 63.81% | 0.8878 / 0.9531 | 0.0349 / 0.1410 |
50
+ | esci_us | 1145 | 57.64% | 0.9185 / 0.9712 | 0.0489 / 0.0868 |
51
+ | mail_phishing | 1050 | 98.95% | 0.1327 / 0.0467 | 0.0097 / 0.0094 |
52
+ | mail_spam | 854 | 98.95% | 0.1658 / 0.0580 | 0.0105 / 0.0098 |
53
+ | massive_en-US | 2970 | 78.89% | 0.7713 / 0.7944 | 0.0532 / 0.1141 |
54
+ | massive_zh-CN | 2921 | 76.82% | 0.8801 / 0.8552 | 0.0664 / 0.0904 |
55
+ | scienceqa | 2017 | 92.66% | 0.2942 / 0.2049 | 0.0489 / 0.0342 |
56
+ | screenqa_choice | 848 | 22.05% | 2.3429 / 2.4277 | 0.0464 / 0.0603 |
57
+ | screenqa_noul | 2148 | 72.35% | 0.5747 / 0.6116 | 0.0341 / 0.1288 |
58
+ | sharc | 8276 | 73.19% | 0.6671 / 0.6700 | 0.0705 / 0.0547 |
59
+ | sms_spam | 794 | 99.37% | 0.0910 / 0.0319 | 0.0062 / 0.0061 |
60
+ | typed_decisions | 2000 | 73.05% | 0.9068 / 0.9944 | 0.1073 / 0.2750 |
61
+ | vqav2_yesno | 8102 | 85.88% | 0.3962 / 0.4835 | 0.0248 / 0.1920 |
62
+ | wikiqa | 6160 | 96.17% | 0.4844 / 0.1710 | 0.0374 / 0.0322 |
63
+ | xnli_en | 5009 | 87.20% | 0.3602 / 0.3727 | 0.0307 / 0.0601 |
64
+ | xnli_zh | 5009 | 78.20% | 0.5986 / 0.5582 | 0.0824 / 0.0297 |
65
+
66
+ ### Inference speed
67
+
68
+ Warm RX 7900 XTX end-to-end predict latency, including preprocessing and transfers.
69
+ Three warmups and 20 synchronized repetitions; network and queueing excluded.
70
+
71
+ | Engine | Workload | p50 ms | p95 ms | Decisions/s |
72
+ | --- | --- | ---: | ---: | ---: |
73
+ | Linnaeus | vision_protocol_text_1q | 44.53 | 45.29 | 22.4 |
74
+ | Linnaeus | vision_protocol_text_3q | 47.81 | 48.29 | 62.9 |
75
+ | Linnaeus | vision_protocol_image_1q | 49.88 | 51.27 | 19.9 |
76
+ | Linnaeus | vision_protocol_image_3q | 100.83 | 102.87 | 29.6 |
77
+ | Linnaeus | distinct_text_1q | 46.83 | 48.00 | 21.3 |
78
+ | Linnaeus | distinct_text_5q | 94.32 | 95.89 | 52.8 |
79
+ | Linnaeus | distinct_text_10q | 94.83 | 96.16 | 105.1 |
80
+ | Linnaeus | distinct_text_50q | 122.42 | 124.17 | 407.7 |
81
+
82
+ ### Benchmark coverage
83
+
84
+
85
+ ## Use
86
+
87
+ ```python
88
+ from linnaeus.predictor import Predictor
89
+ model = Predictor.from_checkpoint("runs/2b/checkpoint")
90
+ ```
91
+
92
+ [Installation, question definitions and response fields](../../../docs/inference.md).
93
+
94
+ ## Limitations
95
+
96
+ For `choice` and `score`, `confidence` is `1 - H(p) / log(K)`. For `noul`, it is `max(p, 1 - p)`. These distribution summaries are not empirical correctness guarantees; calibration metrics use maximum probability and observed correctness.
97
+ Raw predictions, resource samples, calibration bins and quality diagnostics accompany this report.
98
+
99
+ - One seed; variation across seeds is unmeasured. Development selects weights; calibration fits temperatures; test never selects either.
100
+ - Laya references use their own templates, FP32 CPU weights and published temperatures; Linnaeus uses merged BF16 weights.
101
+ - Laya Vision has VQAv2 source-pool and A-OKVQA selection exposure. Backbone pretraining exposure is unverified.
102
+ - Bub acceptance verifies local decision-tool calls, not autonomous planning quality.
103
+ - Source model, dataset and image terms apply; this report does not assign a new weight license.
104
+
105
+ ## Reproducibility
106
+
107
+ Base revision: `15852e8c16360a2fea060d615a32b45270f8a8fc`.
108
+ Weight SHA-256: `0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602`.
109
+ Recipe SHA-256: `d0017a70f1b6c789318700339cd454a74a05a3cac9f0687c1c322bd6eb031c62`.
110
+ [Recorded source-file hashes](../source-sha256.json) identify the code snapshot used for the run. A training Git revision is not recorded.
adapter.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602
3
+ size 33706120
linnaeus.json ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format_version": 1,
3
+ "project": "linnaeus",
4
+ "distribution": "linnaeus",
5
+ "version": "0.1.0",
6
+ "model_id": "Linnaeus-0.1.0-2B",
7
+ "base_model": "Qwen/Qwen3.5-2B",
8
+ "base_path": ".cache/models/Qwen3.5-2B",
9
+ "base_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
10
+ "adapter": "qwen3.5",
11
+ "lora_rank": 8,
12
+ "image_pixels": 262144,
13
+ "max_length": 2048,
14
+ "backend": {
15
+ "linear_patch": true,
16
+ "triton_convolution": true,
17
+ "experimental_rocm_sdpa": true,
18
+ "fused_norm_and_swiglu": true,
19
+ "shared_prefix": true,
20
+ "frozen_vision_cache_MiB": 128
21
+ },
22
+ "inference": "merged LoRA, BF16, fused operations, shared-prefix parallel candidate scoring",
23
+ "source_hashes": {
24
+ "train": "4e8b258c9e5d5a98810868df5d572a824a9d1d6c5aa557ab5123cb7aab5de373",
25
+ "dev": "b736750efff3e30ae8a7772c482fa248899b8a1cf69ace104310e2540671d552",
26
+ "calibration": "815cfe0855dd5494f9def19ddb2a6b12e4116c3ae24d2775db620021dd07fb28",
27
+ "test": "fdac64d536d67c011d55d00352f0be71fa3bdfb7811c0be0e05a5d66cec98095"
28
+ },
29
+ "selected_run": "runs/2b/seed-42",
30
+ "selected_step": 2800,
31
+ "selection": "maximum development macro accuracy within the run; no test selection",
32
+ "seed": 42,
33
+ "dev_macro_accuracy": 0.802734375,
34
+ "temperatures": {
35
+ "choice": 1.6947466135025024,
36
+ "noul": 2.9416117668151855,
37
+ "score": 4.854694843292236
38
+ },
39
+ "weights_sha256": "0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602"
40
+ }