File size: 8,572 Bytes
8f27d8f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
---
library_name: dohnuts
base_model: Qwen/Qwen3.5-0.8B
base_model_relation: adapter
license: cc-by-nc-sa-4.0
tags:
- dohnuts
- decision-model
- multimodal
- rlcd
- lora
- research
---

# Dohnuts-0.1.0-0.8B

A multimodal decision model built on Qwen3.5-0.8B. It scores supplied candidates
from text and images, returning candidate probabilities, truth estimates, or
ordered scores. Independent questions about one input share computation in a
single forward pass. The interface returns decisions without generating reasoning
or free-form answers.

## Model details

| Property | Value |
| --- | --- |
| Base model | Qwen/Qwen3.5-0.8B |
| Training | Joint RLCD and auxiliary cross-entropy; language LoRA and a candidate scorer |
| Selection | Seed 42, update 3,600; highest development macro accuracy |
| Calibration | One temperature per decision type, fitted on an independent partition |
| Runtime | Merged LoRA, BF16, fused operations, shared input prefixes |
| Inputs | Text and one decoded image; 2–128 candidates per question |
| Context | 4,096 tokens per question at inference; 2,048 during training |
| Hardware used | One AMD Radeon RX 7900 XTX, 24 GB |

## Use

Dohnuts is intended for tasks with explicit candidate answers: routing requests,
classifying content, estimating whether a condition holds, rating relevance, or
answering visual questions. An application supplies the choices and decides how
to act on the returned probabilities.

After [setting up the runtime](https://github.com/PsiACE/dohnuts/blob/main/docs/inference.md), load the model from Hugging Face:

```python
from dohnuts.predictor import Predictor

model = Predictor.from_checkpoint("PsiACE/Dohnuts-0.1.0-0.8B")
```

The compact checkpoint contains LoRA and scorer weights. The loader downloads
and caches it with the pinned base model, merges LoRA, and applies calibration.
Weights are not bundled with the Python package. See the [inference guide](https://github.com/PsiACE/dohnuts/blob/main/docs/inference.md) for question
definitions, images, and response fields, or [Bub integration](https://github.com/PsiACE/dohnuts/blob/main/docs/bub-agent.md)
for agent use.

## Evaluation

The model achieves **78.21% macro accuracy** over 26 held-out dataset groups
containing 180,031 decisions. Each group contributes equally to this mean.
Development data selects weights; calibration data fits temperatures; test data
does neither. This is one training seed, with no estimate of variation across seeds.

### JevBench

Accuracy on the same 231 public tasks from JevBench v1.2.2:

| Model | Correct | Accuracy |
| --- | ---: | ---: |
| Dohnuts-0.1.0-0.8B | 152 / 231 | 65.80% |
| Jev 1.13.0 | 200 / 231 | 86.58% |
| Laya multilingual | 110 / 231 | 47.62% |
| Laya Vision | 111 / 231 | 48.05% |

Jev results use published per-task outcomes; Dohnuts and the two Laya checkpoints
were measured locally. Dohnuts uses a 4,096-token limit; the Laya runs use their
native 1,024-token limit and truncation. The other 303 leaderboard tasks are
unavailable, including the judge tier. These accuracies are not the official
four-axis leaderboard score.

### Laya task suites

Dohnuts leads the published Laya multilingual reference on the 51-language
MASSIVE intent suite, at **60.14%** versus **36.61%** language macro accuracy.
Laya multilingual leads on the 15-language XNLI suite, at **73.84%** versus
**70.91%**. The upstream question builders are preserved, but the historical
reference's input hashes are unavailable, so exact input identity cannot be verified.

Against Laya Vision on the same local image examples, Dohnuts scores higher on
A-OKVQA and VQAv2 yes/no; Laya Vision scores higher on ScienceQA and has lower
calibration error on ScienceQA and VQAv2. These runs use different precision and
have reference data-exposure limitations, documented in the
[comparison protocols](https://github.com/PsiACE/dohnuts/blob/main/docs/upstream-alignment.md).

The [comparison gallery](https://github.com/PsiACE/dohnuts/blob/main/docs/figures/README.md) includes every application suite,
language results, vision accuracy and calibration, and paired JevBench outcomes.

### Inference speed

Warm end-to-end median latency on one RX 7900 XTX, using BF16 and each model's
native API:

| Workload | Dohnuts | Laya multilingual | Laya Vision |
| --- | ---: | ---: | ---: |
| Text, 1 question | 15.07 ms | 9.77 ms | 11.00 ms |
| Text, 50 questions | 112.51 ms | 47.19 ms | 125.63 ms |
| Image, 1 question | 24.43 ms | — | 75.41 ms |
| Image, 3 questions | 35.84 ms | — | 77.89 ms |

Measurements use three warmups and 20 synchronized repetitions, including
preprocessing and transfers. They exclude model loading, network, and queueing.
Image timings use warm caches. They do not measure uncached image encoding or
service throughput under load. See the [full charts](https://github.com/PsiACE/dohnuts/blob/main/docs/figures/README.md#same-hardware-inference-latency)
and [benchmark protocol](https://github.com/PsiACE/dohnuts/blob/main/docs/local-benchmarks.md).

## Training

The mixture covers 26 text, language, and vision task groups, including public
classification, question-answering, retrieval, policy, mail, and visual datasets.
A fixed cap provides 143,238 eligible training rows; sampling is uniform over
groups with replacement. The 3,600 updates process 115,200 sampled examples.
Related documents and identical images are grouped to prevent cross-split leakage.
Dataset sources, exclusions, and terms are listed in the
[data reference](https://github.com/PsiACE/dohnuts/blob/main/docs/data-and-evaluation.md).

The base model and vision encoder are frozen. Training updates rank-8 language
LoRA adapters and a shared candidate scorer. The joint RLCD and cross-entropy
objective follows the pinned Laya and Laya Vision implementations. Its four
samples perturb decision logits; they are not generated trajectories. LoRA is
merged before temperature fitting and final evaluation. The
[RLCD specification](https://github.com/PsiACE/dohnuts/blob/main/docs/rlcd.md) gives the objective and fixed schedule.

## Limitations

- Quality depends on the task. Jev leads on the public JevBench tasks; Laya
  multilingual leads on several application suites and small text-batch latency.
- Global temperatures do not improve every dataset's calibration. The API's
  `confidence` field summarizes a distribution; it is not a measured probability
  of correctness. Check calibration on the intended workload.
- Candidate wording, order, and input length can affect decisions. Over-budget
  inputs are rejected. A long-document result covers only the eligible subset.
- Laya Vision has possible VQAv2 training-pool exposure and A-OKVQA selection
  exposure. Backbone pretraining exposure is unverified. These comparisons do
  not establish performance on unseen data for every reference.
- Bub acceptance exercises the decision tool. It does not measure autonomous
  planning quality.

## Artifact and provenance

Base revision: `2fc06364715b967f1860aea9cf38778875588b17`.

Selected weight SHA-256:
`196be33a0282537bcd821e2115643b352d0ad2a0bbaf7b242a1b7fe5bd96cfdf`.

The exported checkpoint records calibration, selection, and partition hashes in
`dohnuts.json`. The [results](https://github.com/PsiACE/dohnuts/blob/main/results/README.md) include loss, development accuracy,
held-out quality, calibration bins, latency samples, and resource measurements.
Their manifest identifies the selected weights and checksums each published table.
No training Git revision was recorded. [Chart values](https://github.com/PsiACE/dohnuts/blob/main/docs/figures/chart-data.csv)
and a [figure manifest](https://github.com/PsiACE/dohnuts/blob/main/docs/figures/manifest.json) accompany the comparisons.

## License

The code is licensed under [Apache-2.0](https://github.com/PsiACE/dohnuts/blob/main/LICENSE). The decision weights are provided
under [CC BY-NC-SA 4.0](https://huggingface.co/PsiACE/Dohnuts-0.1.0-0.8B/blob/main/LICENSE)
for non-commercial research. This grant covers the Dohnuts LoRA and decision-head
contributions; the base model and source data retain their own terms.

The Qwen3.5-0.8B base is Apache-2.0. ScienceQA's dataset terms include non-commercial
and share-alike restrictions; other sources have their own research-use terms.
The checkpoint is not offered as a commercially cleared model. See the
[data reference](https://github.com/PsiACE/dohnuts/blob/main/docs/data-and-evaluation.md#release-assets-and-terms) and
[attributions](https://github.com/PsiACE/dohnuts/blob/main/NOTICE).