--- language: en license: other library_name: llama.cpp base_model: togethercomputer/Tev1-0.8B-experimental base_model_relation: quantized pipeline_tag: text-classification tags: - decision-model - multiple-choice - typesafe - qwen3.5 - gguf - llama.cpp --- # Tev1-0.8B Tev1-0.8B is an experimental **decision model** from Together AI: one document (the *state*), a question, and 2-24 labeled options in, a single option letter out. It is a supervised fine-tune of `Qwen/Qwen3.5-0.8B` that keeps the language model's standard next-token head and reads the answer letter after a short chat decision prompt. It is a Jev-inspired experiment, not a non-autoregressive Jev runtime. This repository holds quantized GGUF conversions for CPU inference through [dohnuts.cpp](https://github.com/DreamBlooms/dohnuts.cpp), a native C++ port on [llama.cpp](https://github.com/ggml-org/llama.cpp). It serves the same `POST /v1/systemone` wire format as the other decision models. No GPU or Python runtime is needed. ## GGUF | File | Contents | Size | | --- | --- | ---: | | `tev1-f16.gguf` | F16 language model | 1.4 GB | | `tev1-Q8_0.gguf` | Q8_0 language model | 774 MB | | `tev1.json` | profile config and fitted temperature | 46 B | `tev1.json` is required alongside the GGUF. The conversion is a full fine-tune, converted directly with `--no-mtp`; there is no adapter to merge and no pointer head. ```sh git clone https://github.com/DreamBlooms/dohnuts.cpp cd dohnuts.cpp git submodule update --init --depth 1 cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON cmake --build build -j --target dohnuts-cli build/dohnuts-cli --server --port 8080 \ --model tev1-Q8_0.gguf --metadata tev1.json ``` Ask one state several questions (the same request shape as the other profiles): ```sh curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \ -d '{"state":"Returns are allowed within 30 days. This purchase was 12 days ago.", "questions":{"window":{"type":"choice","instructions":"Is this return within the allowed window?", "criteria":{"yes":"Yes.","no":"No.","unknown":"Not enough information."}}}}' ``` The answer keeps the core fields (`type`, `choice`, `probabilities`, `noul`, `score`, `confidence`) and adds this model's `certainty` (with `legend` for `score`) under `native`. Rebuild from the upstream checkpoint with `scripts/build_tev1_gguf.sh`, which converts the full fine-tune directly and quantizes to Q8_0. No retraining is involved. ## Limits - The upstream checkpoint is experimental. Calibration, multilingual behavior, and out-of-distribution robustness have not been evaluated here. - Training mainly used 2-8 options; the interface allows up to 24, but the wider range is untested. - Generic chat is not the intended interface and may produce prose. ## License The base Qwen3.5-0.8B model is Apache-2.0. The upstream release license for the fine-tuned weights is being finalized; see the [upstream model card](https://huggingface.co/togethercomputer/Tev1-0.8B-experimental) for the current terms. This is an independent, Jev-inspired release and does not use Jev's answers as training labels.