Tev1-4B-experimental-GGUF

Tev1-4B-experimental is an experimental 4-billion-parameter decision model from Together AI, a supervised fine-tune of Qwen3.5-4B trained to select a single option from a structured state, question, and list of 2–24 labeled choices — a "Jev-inspired" experiment that retains Qwen's standard autoregressive next-token head rather than implementing a true non-autoregressive Jev runtime. Its intended interface takes a system instruction plus a JSON-structured decision payload (treating the state field strictly as data rather than instructions) and returns just a single option letter at temperature=0 with enable_thinking=False, making it usable via the Together API or compatible chat completion endpoints. On Together's internal development evaluation, it scored 88.0% (880/1,000) on the main decision set and 100% (300/300) on a synthetic policy-transfer set, with all outputs valid single letters — though the authors caution these are development results without an untuned-Qwen baseline for comparison, not an independent benchmark. It is explicitly not intended for generic chat use, should not serve as the sole authority for high-impact decisions, and has not been comprehensively evaluated for prompt injection resistance, multilingual behavior, or calibration; the base Qwen3.5-4B is Apache-2.0, but the release license for these fine-tuned weights was still being finalized at time of publication.

Model Files

File Name Quant Type File Size File Link Description
Tev1-4B-experimental.BF16.gguf BF16 8.42 GB Link Full BF16 weights. Highest quality, largest file size.
Tev1-4B-experimental.Q3_K_L.gguf Q3_K_L 2.42 GB Link Lower quality but usable, good for low RAM availability.
Tev1-4B-experimental.Q3_K_M.gguf Q3_K_M 2.26 GB Link Low quality.
Tev1-4B-experimental.Q4_K_M.gguf Q4_K_M 2.71 GB Link Good quality, default size for most use cases, recommended.
Tev1-4B-experimental.Q4_K_S.gguf Q4_K_S 2.56 GB Link Slightly lower quality with more space savings, recommended.
Tev1-4B-experimental.Q5_K_M.gguf Q5_K_M 3.07 GB Link High quality, recommended.
Tev1-4B-experimental.Q5_K_S.gguf Q5_K_S 2.99 GB Link High quality, recommended.
Tev1-4B-experimental.Q6_K.gguf Q6_K 3.46 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
937
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/Tev1-4B-experimental-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(6)
this model

Collection including prithivMLmods/Tev1-4B-experimental-GGUF