How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
# Run inference directly in the terminal:
llama cli -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
# Run inference directly in the terminal:
llama cli -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
# Run inference directly in the terminal:
./llama-cli -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
Use Docker
docker model run hf.co/DreamBlooms/Linnaeus-0.1.0-2B-GGUF:
Quick Links

Linnaeus-0.1.0-2B

Built on Qwen/Qwen3.5-2B, this checkpoint returns decision distributions from text and images. It supports candidate selection (choice), truth estimates (noul), and ordered scores (score) without generating reasoning or free-form responses.

Benchmarks (measured)

  • Held-out macro accuracy 80.78% (26 groups, 180k examples; upstream Dohnuts-0.1.0-0.8B: 78.21%)
  • JevBench v1.2.2 public 231 tasks: 73.16% (easy 100 / standard 91.7 / hard 49.6) — top of the ~2B local-model class, ahead of all sub-1B classifiers
  • Laya app suites: beats upstream on 7/11 groups (spam, routing, typed-decisions strongest)
  • XNLI 15-language: 76.0% — recovers the suite upstream lost to Laya multilingual (73.8%)

On-device exports (Apple Silicon)

Merged + quantized builds for macOS/iOS (text MLX builds are text-only; the VLM builds keep the vision tower):

Repo Size JevBench Target
Linnaeus-0.1.0-2B-merged 4.3 GB 71.0% (MPS) Mac dev / conversion source
Linnaeus-0.1.0-2B-MLX-8bit 1.9 GB 70.56% Mac + iPhone
Linnaeus-0.1.0-2B-MLX-4bit 1.0 GB 67.53% iPhone size-optimized
Linnaeus-0.1.0-2B-MLX-VLM-8bit 2.5 GB text 70.56% + images Mac + iPhone multimodal
Linnaeus-0.1.0-2B-MLX-VLM-4bit 1.6 GB text ~67% + images iPhone multimodal, size pick

The merged checkpoint embeds the decision head as an extra lm_head row (score_row_id), so any stock LM runtime produces decision scores at <|fim_suffix|> marker positions. See linnaeus-runtime.json and github.com/pi-dal/Linnaeus src/linnaeus/mlx_predictor.py.

Model details

Property Value
Base model Qwen/Qwen3.5-2B
Selection Seed 42, update 2,800
Inference merged LoRA, BF16, fused operations, shared-prefix parallel candidate scoring
Inputs Text and one PIL image; 2–128 candidates; 4,096 tokens per question

Questions reuse a shared input prefix and compute their suffixes in parallel. Additional questions still require computation.

Training

The training recipe combines RLCD and auxiliary cross-entropy, following the pinned Laya and Laya Vision references. Temperature calibration uses an independent partition after LoRA merging. Calibration quality is measured below.

Training recipe · RLCD implementation and upstream attribution

Evaluation

Held-out macro accuracy: 80.78%.

Dataset N Accuracy NLL before / after calibration ECE before / after
ag_news 7600 90.04% 0.8728 / 0.5378 0.0832 / 0.0710
aokvqa 1138 83.66% 0.6347 / 0.4746 0.0930 / 0.0513
banking77 3080 73.70% 1.0536 / 1.1711 0.0275 / 0.2001
boolq 3270 87.71% 0.3854 / 0.3288 0.0610 / 0.0648
clevr_attribute 53734 98.92% 0.1687 / 0.1230 0.0100 / 0.0095
clevr_count 35422 89.72% 0.7159 / 0.2652 0.0793 / 0.0077
clevr_exist 20196 98.58% 0.2526 / 0.0955 0.0137 / 0.0127
contract_nli 1173 85.59% 0.5014 / 0.4020 0.0888 / 0.0338
emotion 2000 77.00% 0.7237 / 0.6812 0.0777 / 0.0488
esci_es 1482 60.26% 0.9241 / 0.9775 0.0342 / 0.1266
esci_jp 1633 63.81% 0.8878 / 0.9531 0.0349 / 0.1410
esci_us 1145 57.64% 0.9185 / 0.9712 0.0489 / 0.0868
mail_phishing 1050 98.95% 0.1327 / 0.0467 0.0097 / 0.0094
mail_spam 854 98.95% 0.1658 / 0.0580 0.0105 / 0.0098
massive_en-US 2970 78.89% 0.7713 / 0.7944 0.0532 / 0.1141
massive_zh-CN 2921 76.82% 0.8801 / 0.8552 0.0664 / 0.0904
scienceqa 2017 92.66% 0.2942 / 0.2049 0.0489 / 0.0342
screenqa_choice 848 22.05% 2.3429 / 2.4277 0.0464 / 0.0603
screenqa_noul 2148 72.35% 0.5747 / 0.6116 0.0341 / 0.1288
sharc 8276 73.19% 0.6671 / 0.6700 0.0705 / 0.0547
sms_spam 794 99.37% 0.0910 / 0.0319 0.0062 / 0.0061
typed_decisions 2000 73.05% 0.9068 / 0.9944 0.1073 / 0.2750
vqav2_yesno 8102 85.88% 0.3962 / 0.4835 0.0248 / 0.1920
wikiqa 6160 96.17% 0.4844 / 0.1710 0.0374 / 0.0322
xnli_en 5009 87.20% 0.3602 / 0.3727 0.0307 / 0.0601
xnli_zh 5009 78.20% 0.5986 / 0.5582 0.0824 / 0.0297

Inference speed

Warm RTX 4090 end-to-end predict latency, including preprocessing and transfers. Three warmups and 20 synchronized repetitions; network and queueing excluded.

Engine Workload p50 ms p95 ms Decisions/s
Linnaeus vision_protocol_text_1q 44.53 45.29 22.4
Linnaeus vision_protocol_text_3q 47.81 48.29 62.9
Linnaeus vision_protocol_image_1q 49.88 51.27 19.9
Linnaeus vision_protocol_image_3q 100.83 102.87 29.6
Linnaeus distinct_text_1q 46.83 48.00 21.3
Linnaeus distinct_text_5q 94.32 95.89 52.8
Linnaeus distinct_text_10q 94.83 96.16 105.1
Linnaeus distinct_text_50q 122.42 124.17 407.7

Benchmark coverage

Use

from linnaeus.predictor import Predictor
model = Predictor.from_checkpoint("runs/2b/checkpoint")

Installation, question definitions and response fields.

Limitations

For choice and score, confidence is 1 - H(p) / log(K). For noul, it is max(p, 1 - p). These distribution summaries are not empirical correctness guarantees; calibration metrics use maximum probability and observed correctness. Raw predictions, resource samples, calibration bins and quality diagnostics accompany this report.

  • One seed; variation across seeds is unmeasured. Development selects weights; calibration fits temperatures; test never selects either.
  • Laya references use their own templates, FP32 CPU weights and published temperatures; Linnaeus uses merged BF16 weights.
  • Laya Vision has VQAv2 source-pool and A-OKVQA selection exposure. Backbone pretraining exposure is unverified.
  • Bub acceptance verifies local decision-tool calls, not autonomous planning quality.
  • Source model, dataset and image terms apply; this report does not assign a new weight license.

Reproducibility

Base revision: 15852e8c16360a2fea060d615a32b45270f8a8fc. Weight SHA-256: 0424555feea3126ee02d8d1b12636fad8cf491a5964a00a04b2fec67bcc0f602. Recipe SHA-256: d0017a70f1b6c789318700339cd454a74a05a3cac9f0687c1c322bd6eb031c62. Recorded source-file hashes identify the code snapshot used for the run. A training Git revision is not recorded.

GGUF

CPU inference is available through dohnuts.cpp, a native C++ port on llama.cpp. It runs the model on CPU, reads the scalar decision head at each answer slot, and serves the same POST /v1/systemone wire format. No GPU or Python runtime is needed. The checkpoint is quantized from pi-dal/Linnaeus-0.1.0-2B (Qwen/Qwen3.5-2B plus a rank-8 LoRA) and uses the dohnuts profile.

File Contents Size
Linnaeus-0.1.0-2B-F16.gguf f16 weights, reference 3.8 GB
Linnaeus-0.1.0-2B-Q8_0.gguf Q8_0, conservative default 2.0 GB
Linnaeus-0.1.0-2B-Q4_K_M.gguf Q4_K_M, smallest 1.3 GB
mmproj-Linnaeus-0.1.0-2B-bf16.gguf vision tower (bf16), for image input 671 MB
head.f32 scalar decision head, required at runtime 8.2 kB
linnaeus.json profile, calibrated temperatures, flags 1.3 kB

The three text GGUFs are text-only; pair mmproj-Linnaeus-0.1.0-2B-bf16.gguf with any of them to enable image inputs. head.f32 and linnaeus.json are required at runtime; linnaeus.json carries the calibrated choice / noul / score temperatures from the upstream checkpoint.

This conversion is structural and has been smoke-tested with dohnuts-cli: it loads and answers choice / noul / score, and the F16, Q8_0 and Q4_K_M builds agree on every argmax on the sample set. It has not been benchmarked against the PyTorch reference. Q8_0 is the conservative default.

git clone https://github.com/DreamBlooms/dohnuts.cpp
cd dohnuts.cpp
git submodule update --init --depth 1
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON
cmake --build build -j --target dohnuts-cli

build/dohnuts-cli --server --port 8080 \
  --model Linnaeus-0.1.0-2B-Q8_0.gguf \
  --head head.f32 --metadata linnaeus.json

Add --mmproj mmproj-Linnaeus-0.1.0-2B-bf16.gguf for image inputs.

Ask one state several questions (same request shape as the PyTorch server):

curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
  -d '{"state":"My card was charged twice for the same purchase.",
       "questions":{"department":{"type":"choice","instructions":"Which team should handle this?",
         "criteria":{"billing":null,"technical support":null,"sales":null}},
         "refund_requested":{"type":"noul","instructions":"Is a refund requested?"}}}'

The answer keeps the core fields (type, choice, probabilities, noul, score, legend) in the same POST /v1/systemone wire format.

Rebuild these files from the upstream adapter with scripts/build_linnaeus_gguf.sh (merge the rank-8 LoRA into Qwen/Qwen3.5-2B, export the head, then convert with --no-mtp; the base config declares an MTP layer but ships no MTP weights).

Downloads last month
208
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DreamBlooms/Linnaeus-0.1.0-2B-GGUF

Finetuned
Qwen/Qwen3.5-2B
Quantized
(5)
this model

Collection including DreamBlooms/Linnaeus-0.1.0-2B-GGUF