File size: 3,180 Bytes
7fcc5ab
0f59e8c
 
7fcc5ab
0f59e8c
 
 
7fcc5ab
0f59e8c
 
 
 
 
 
7fcc5ab
 
0f59e8c
 
 
 
 
 
 
 
7fcc5ab
0f59e8c
 
 
 
 
7fcc5ab
0f59e8c
7fcc5ab
0f59e8c
 
 
 
 
7fcc5ab
0f59e8c
 
 
7fcc5ab
 
0f59e8c
 
 
 
 
7fcc5ab
0f59e8c
 
 
7fcc5ab
0f59e8c
7fcc5ab
 
0f59e8c
 
 
 
7fcc5ab
 
0f59e8c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7fcc5ab
 
 
0f59e8c
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
---
language: en
license: other
library_name: llama.cpp
base_model: togethercomputer/Tev1-0.8B-experimental
base_model_relation: quantized
pipeline_tag: text-classification
tags:
  - decision-model
  - multiple-choice
  - typesafe
  - qwen3.5
  - gguf
  - llama.cpp
---

# Tev1-0.8B

Tev1-0.8B is an experimental **decision model** from Together AI: one document
(the *state*), a question, and 2-24 labeled options in, a single option letter
out. It is a supervised fine-tune of `Qwen/Qwen3.5-0.8B` that keeps the language
model's standard next-token head and reads the answer letter after a short chat
decision prompt. It is a Jev-inspired experiment, not a non-autoregressive Jev
runtime.

This repository holds quantized GGUF conversions for CPU inference through
[dohnuts.cpp](https://github.com/DreamBlooms/dohnuts.cpp), a native C++ port on
[llama.cpp](https://github.com/ggml-org/llama.cpp). It serves the same
`POST /v1/systemone` wire format as the other decision models. No GPU or Python
runtime is needed.

## GGUF

| File | Contents | Size |
| --- | --- | ---: |
| `tev1-f16.gguf` | F16 language model | 1.4 GB |
| `tev1-Q8_0.gguf` | Q8_0 language model | 774 MB |
| `tev1.json` | profile config and fitted temperature | 46 B |

`tev1.json` is required alongside the GGUF. The conversion is a full fine-tune,
converted directly with `--no-mtp`; there is no adapter to merge and no pointer
head.

```sh
git clone https://github.com/DreamBlooms/dohnuts.cpp
cd dohnuts.cpp
git submodule update --init --depth 1
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON
cmake --build build -j --target dohnuts-cli

build/dohnuts-cli --server --port 8080 \
  --model tev1-Q8_0.gguf --metadata tev1.json
```

Ask one state several questions (the same request shape as the other profiles):

```sh
curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
  -d '{"state":"Returns are allowed within 30 days. This purchase was 12 days ago.",
       "questions":{"window":{"type":"choice","instructions":"Is this return within the allowed window?",
         "criteria":{"yes":"Yes.","no":"No.","unknown":"Not enough information."}}}}'
```

The answer keeps the core fields (`type`, `choice`, `probabilities`, `noul`,
`score`, `confidence`) and adds this model's `certainty` (with `legend` for
`score`) under `native`.

Rebuild from the upstream checkpoint with `scripts/build_tev1_gguf.sh`, which
converts the full fine-tune directly and quantizes to Q8_0. No retraining is
involved.

## Limits

- The upstream checkpoint is experimental. Calibration, multilingual behavior,
  and out-of-distribution robustness have not been evaluated here.
- Training mainly used 2-8 options; the interface allows up to 24, but the wider
  range is untested.
- Generic chat is not the intended interface and may produce prose.

## License

The base Qwen3.5-0.8B model is Apache-2.0. The upstream release license for the
fine-tuned weights is being finalized; see the
[upstream model card](https://huggingface.co/togethercomputer/Tev1-0.8B-experimental)
for the current terms. This is an independent, Jev-inspired release and does not
use Jev's answers as training labels.