Instructions to use pi-dal/Linnaeus-0.1.0-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pi-dal/Linnaeus-0.1.0-2B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("pi-dal/Linnaeus-0.1.0-2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Upload folder using huggingface_hub
Browse files- README.md +110 -0
- adapter.safetensors +3 -0
- linnaeus.json +40 -0
README.md
ADDED
|
@@ -0,0 +1,110 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3.5-2B
|
| 4 |
+
library_name: transformers
|
| 5 |
+
tags:
|
| 6 |
+
- lora
|
| 7 |
+
- vision-language
|
| 8 |
+
- decision-model
|
| 9 |
+
- rlcd
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# Linnaeus-0.1.0-2B
|
| 13 |
+
|
| 14 |
+
Built on Qwen/Qwen3.5-2B, this checkpoint returns decision distributions from text and images. It supports candidate selection (`choice`), truth estimates (`noul`), and ordered scores (`score`) without generating reasoning or free-form responses.
|
| 15 |
+
|
| 16 |
+
## Model details
|
| 17 |
+
|
| 18 |
+
| Property | Value |
|
| 19 |
+
| --- | --- |
|
| 20 |
+
| Base model | Qwen/Qwen3.5-2B |
|
| 21 |
+
| Selection | Seed 42, update 2,800 |
|
| 22 |
+
| Inference | merged LoRA, BF16, fused operations, shared-prefix parallel candidate scoring |
|
| 23 |
+
| Inputs | Text and one PIL image; 2–128 candidates; 4,096 tokens per question |
|
| 24 |
+
|
| 25 |
+
Questions reuse a shared input prefix and compute their suffixes in parallel. Additional questions still require computation.
|
| 26 |
+
|
| 27 |
+
## Training
|
| 28 |
+
|
| 29 |
+
The training recipe combines RLCD and auxiliary cross-entropy, following the pinned Laya and Laya Vision references. Temperature calibration uses an independent partition after LoRA merging. Calibration quality is measured below.
|
| 30 |
+
|
| 31 |
+
[Training recipe](../recipe.json) · [RLCD implementation and upstream attribution](../../../docs/rlcd.md)
|
| 32 |
+
|
| 33 |
+
## Evaluation
|
| 34 |
+
|
| 35 |
+
Held-out macro accuracy: 80.78%.
|
| 36 |
+
|
| 37 |
+
| Dataset | N | Accuracy | NLL before / after calibration | ECE before / after |
|
| 38 |
+
| --- | ---: | ---: | ---: | ---: |
|
| 39 |
+
| ag_news | 7600 | 90.04% | 0.8728 / 0.5378 | 0.0832 / 0.0710 |
|
| 40 |
+
| aokvqa | 1138 | 83.66% | 0.6347 / 0.4746 | 0.0930 / 0.0513 |
|
| 41 |
+
| banking77 | 3080 | 73.70% | 1.0536 / 1.1711 | 0.0275 / 0.2001 |
|
| 42 |
+
| boolq | 3270 | 87.71% | 0.3854 / 0.3288 | 0.0610 / 0.0648 |
|
| 43 |
+
| clevr_attribute | 53734 | 98.92% | 0.1687 / 0.1230 | 0.0100 / 0.0095 |
|
| 44 |
+
| clevr_count | 35422 | 89.72% | 0.7159 / 0.2652 | 0.0793 / 0.0077 |
|
| 45 |
+
| clevr_exist | 20196 | 98.58% | 0.2526 / 0.0955 | 0.0137 / 0.0127 |
|
| 46 |
+
| contract_nli | 1173 | 85.59% | 0.5014 / 0.4020 | 0.0888 / 0.0338 |
|
| 47 |
+
| emotion | 2000 | 77.00% | 0.7237 / 0.6812 | 0.0777 / 0.0488 |
|
| 48 |
+
| esci_es | 1482 | 60.26% | 0.9241 / 0.9775 | 0.0342 / 0.1266 |
|
| 49 |
+
| esci_jp | 1633 | 63.81% | 0.8878 / 0.9531 | 0.0349 / 0.1410 |
|
| 50 |
+
| esci_us | 1145 | 57.64% | 0.9185 / 0.9712 | 0.0489 / 0.0868 |
|
| 51 |
+
| mail_phishing | 1050 | 98.95% | 0.1327 / 0.0467 | 0.0097 / 0.0094 |
|
| 52 |
+
| mail_spam | 854 | 98.95% | 0.1658 / 0.0580 | 0.0105 / 0.0098 |
|
| 53 |
+
| massive_en-US | 2970 | 78.89% | 0.7713 / 0.7944 | 0.0532 / 0.1141 |
|
| 54 |
+
| massive_zh-CN | 2921 | 76.82% | 0.8801 / 0.8552 | 0.0664 / 0.0904 |
|
| 55 |
+
| scienceqa | 2017 | 92.66% | 0.2942 / 0.2049 | 0.0489 / 0.0342 |
|
| 56 |
+
| screenqa_choice | 848 | 22.05% | 2.3429 / 2.4277 | 0.0464 / 0.0603 |
|
| 57 |
+
| screenqa_noul | 2148 | 72.35% | 0.5747 / 0.6116 | 0.0341 / 0.1288 |
|
| 58 |
+
| sharc | 8276 | 73.19% | 0.6671 / 0.6700 | 0.0705 / 0.0547 |
|
| 59 |
+
| sms_spam | 794 | 99.37% | 0.0910 / 0.0319 | 0.0062 / 0.0061 |
|
| 60 |
+
| typed_decisions | 2000 | 73.05% | 0.9068 / 0.9944 | 0.1073 / 0.2750 |
|
| 61 |
+
| vqav2_yesno | 8102 | 85.88% | 0.3962 / 0.4835 | 0.0248 / 0.1920 |
|
| 62 |
+
| wikiqa | 6160 | 96.17% | 0.4844 / 0.1710 | 0.0374 / 0.0322 |
|
| 63 |
+
| xnli_en | 5009 | 87.20% | 0.3602 / 0.3727 | 0.0307 / 0.0601 |
|
| 64 |
+
| xnli_zh | 5009 | 78.20% | 0.5986 / 0.5582 | 0.0824 / 0.0297 |
|
| 65 |
+
|
| 66 |
+
### Inference speed
|
| 67 |
+
|
| 68 |
+
Warm RX 7900 XTX end-to-end predict latency, including preprocessing and transfers.
|
| 69 |
+
Three warmups and 20 synchronized repetitions; network and queueing excluded.
|
| 70 |
+
|
| 71 |
+
| Engine | Workload | p50 ms | p95 ms | Decisions/s |
|
| 72 |
+
| --- | --- | ---: | ---: | ---: |
|
| 73 |
+
| Linnaeus | vision_protocol_text_1q | 44.53 | 45.29 | 22.4 |
|
| 74 |
+
| Linnaeus | vision_protocol_text_3q | 47.81 | 48.29 | 62.9 |
|
| 75 |
+
| Linnaeus | vision_protocol_image_1q | 49.88 | 51.27 | 19.9 |
|
| 76 |
+
| Linnaeus | vision_protocol_image_3q | 100.83 | 102.87 | 29.6 |
|
| 77 |
+
| Linnaeus | distinct_text_1q | 46.83 | 48.00 | 21.3 |
|
| 78 |
+
| Linnaeus | distinct_text_5q | 94.32 | 95.89 | 52.8 |
|
| 79 |
+
| Linnaeus | distinct_text_10q | 94.83 | 96.16 | 105.1 |
|
| 80 |
+
| Linnaeus | distinct_text_50q | 122.42 | 124.17 | 407.7 |
|
| 81 |
+
|
| 82 |
+
### Benchmark coverage
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
## Use
|
| 86 |
+
|
| 87 |
+
```python
|
| 88 |
+
from linnaeus.predictor import Predictor
|
| 89 |
+
model = Predictor.from_checkpoint("runs/2b/checkpoint")
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
[Installation, question definitions and response fields](../../../docs/inference.md).
|
| 93 |
+
|
| 94 |
+
## Limitations
|
| 95 |
+
|
| 96 |
+
For `choice` and `score`, `confidence` is `1 - H(p) / log(K)`. For `noul`, it is `max(p, 1 - p)`. These distribution summaries are not empirical correctness guarantees; calibration metrics use maximum probability and observed correctness.
|
| 97 |
+
Raw predictions, resource samples, calibration bins and quality diagnostics accompany this report.
|
| 98 |
+
|
| 99 |
+
- One seed; variation across seeds is unmeasured. Development selects weights; calibration fits temperatures; test never selects either.
|
| 100 |
+
- Laya references use their own templates, FP32 CPU weights and published temperatures; Linnaeus uses merged BF16 weights.
|
| 101 |
+
- Laya Vision has VQAv2 source-pool and A-OKVQA selection exposure. Backbone pretraining exposure is unverified.
|
| 102 |
+
- Bub acceptance verifies local decision-tool calls, not autonomous planning quality.
|
| 103 |
+
- Source model, dataset and image terms apply; this report does not assign a new weight license.
|
| 104 |
+
|
| 105 |
+
## Reproducibility
|
| 106 |
+
|
| 107 |
+
Base revision: `15852e8c16360a2fea060d615a32b45270f8a8fc`.
|
| 108 |
+
Weight SHA-256: `0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602`.
|
| 109 |
+
Recipe SHA-256: `d0017a70f1b6c789318700339cd454a74a05a3cac9f0687c1c322bd6eb031c62`.
|
| 110 |
+
[Recorded source-file hashes](../source-sha256.json) identify the code snapshot used for the run. A training Git revision is not recorded.
|
adapter.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602
|
| 3 |
+
size 33706120
|
linnaeus.json
ADDED
|
@@ -0,0 +1,40 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format_version": 1,
|
| 3 |
+
"project": "linnaeus",
|
| 4 |
+
"distribution": "linnaeus",
|
| 5 |
+
"version": "0.1.0",
|
| 6 |
+
"model_id": "Linnaeus-0.1.0-2B",
|
| 7 |
+
"base_model": "Qwen/Qwen3.5-2B",
|
| 8 |
+
"base_path": ".cache/models/Qwen3.5-2B",
|
| 9 |
+
"base_revision": "15852e8c16360a2fea060d615a32b45270f8a8fc",
|
| 10 |
+
"adapter": "qwen3.5",
|
| 11 |
+
"lora_rank": 8,
|
| 12 |
+
"image_pixels": 262144,
|
| 13 |
+
"max_length": 2048,
|
| 14 |
+
"backend": {
|
| 15 |
+
"linear_patch": true,
|
| 16 |
+
"triton_convolution": true,
|
| 17 |
+
"experimental_rocm_sdpa": true,
|
| 18 |
+
"fused_norm_and_swiglu": true,
|
| 19 |
+
"shared_prefix": true,
|
| 20 |
+
"frozen_vision_cache_MiB": 128
|
| 21 |
+
},
|
| 22 |
+
"inference": "merged LoRA, BF16, fused operations, shared-prefix parallel candidate scoring",
|
| 23 |
+
"source_hashes": {
|
| 24 |
+
"train": "4e8b258c9e5d5a98810868df5d572a824a9d1d6c5aa557ab5123cb7aab5de373",
|
| 25 |
+
"dev": "b736750efff3e30ae8a7772c482fa248899b8a1cf69ace104310e2540671d552",
|
| 26 |
+
"calibration": "815cfe0855dd5494f9def19ddb2a6b12e4116c3ae24d2775db620021dd07fb28",
|
| 27 |
+
"test": "fdac64d536d67c011d55d00352f0be71fa3bdfb7811c0be0e05a5d66cec98095"
|
| 28 |
+
},
|
| 29 |
+
"selected_run": "runs/2b/seed-42",
|
| 30 |
+
"selected_step": 2800,
|
| 31 |
+
"selection": "maximum development macro accuracy within the run; no test selection",
|
| 32 |
+
"seed": 42,
|
| 33 |
+
"dev_macro_accuracy": 0.802734375,
|
| 34 |
+
"temperatures": {
|
| 35 |
+
"choice": 1.6947466135025024,
|
| 36 |
+
"noul": 2.9416117668151855,
|
| 37 |
+
"score": 4.854694843292236
|
| 38 |
+
},
|
| 39 |
+
"weights_sha256": "0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602"
|
| 40 |
+
}
|