Tev1-0.8B

Tev1-0.8B is an experimental decision model from Together AI: one document (the state), a question, and 2-24 labeled options in, a single option letter out. It is a supervised fine-tune of Qwen/Qwen3.5-0.8B that keeps the language model's standard next-token head and reads the answer letter after a short chat decision prompt. It is a Jev-inspired experiment, not a non-autoregressive Jev runtime.

This repository holds quantized GGUF conversions for CPU inference through dohnuts.cpp, a native C++ port on llama.cpp. It serves the same POST /v1/systemone wire format as the other decision models. No GPU or Python runtime is needed.

GGUF

File Contents Size
tev1-f16.gguf F16 language model 1.4 GB
tev1-Q8_0.gguf Q8_0 language model 774 MB
tev1.json profile config and fitted temperature 46 B

tev1.json is required alongside the GGUF. The conversion is a full fine-tune, converted directly with --no-mtp; there is no adapter to merge and no pointer head.

git clone https://github.com/DreamBlooms/dohnuts.cpp
cd dohnuts.cpp
git submodule update --init --depth 1
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON
cmake --build build -j --target dohnuts-cli

build/dohnuts-cli --server --port 8080 \
  --model tev1-Q8_0.gguf --metadata tev1.json

Ask one state several questions (the same request shape as the other profiles):

curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
  -d '{"state":"Returns are allowed within 30 days. This purchase was 12 days ago.",
       "questions":{"window":{"type":"choice","instructions":"Is this return within the allowed window?",
         "criteria":{"yes":"Yes.","no":"No.","unknown":"Not enough information."}}}}'

The answer keeps the core fields (type, choice, probabilities, noul, score, confidence) and adds this model's certainty (with legend for score) under native.

Rebuild from the upstream checkpoint with scripts/build_tev1_gguf.sh, which converts the full fine-tune directly and quantizes to Q8_0. No retraining is involved.

Limits

  • The upstream checkpoint is experimental. Calibration, multilingual behavior, and out-of-distribution robustness have not been evaluated here.
  • Training mainly used 2-8 options; the interface allows up to 24, but the wider range is untested.
  • Generic chat is not the intended interface and may produce prose.

License

The base Qwen3.5-0.8B model is Apache-2.0. The upstream release license for the fine-tuned weights is being finalized; see the upstream model card for the current terms. This is an independent, Jev-inspired release and does not use Jev's answers as training labels.

Downloads last month
399
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DreamBlooms/Tev1-0.8B-experimental-GGUF

Quantized
(5)
this model

Collection including DreamBlooms/Tev1-0.8B-experimental-GGUF