Text Generation
GGUF
jev-style
English
decision-model
classification
calibration
qwen3.5
single-prefill
conversational
Instructions to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- jev-style
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with jev-style:
pip install jev-style # GGUF builds score through llama.cpp: build the jev-score binary once hf download chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF build_jev_score.sh jev_score.cpp --local-dir jev-score export JEV_SCORE_BIN=$(sh jev-score/build_jev_score.sh /path/to/llama.cpp | tail -n 1)
from jev_style import JevStyle, noul, choice js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF") out = js.decide("I was charged twice for one order.", { "billing": noul("This message is about billing."), "team": choice("Which team should handle it?", ["billing", "shipping", "tech"]), }) print(out["answers"]["team"]["choice"]) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Ollama
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Ollama:
ollama run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Docker Model Runner:
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Lemonade
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Jev-Style-Qwen3.5-2B-Decision-v2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Rename GGUF files so the quant type ends the filename (Ollama :Q4_K_M tags)
Browse filesFixes discussion #1. Files are byte-identical; only names changed. README and SHA256SUMS.json updated.
- .gitattributes +3 -0
- Jev-Style-v2-BF16-Calibrated.calibration.json → Jev-Style-v2-Calibrated-BF16.calibration.json +0 -0
- Jev-Style-v2-BF16-Calibrated.gguf → Jev-Style-v2-Calibrated-BF16.gguf +0 -0
- Jev-Style-v2-Q4_K_M-Calibrated.calibration.json → Jev-Style-v2-Calibrated-Q4_K_M.calibration.json +0 -0
- Jev-Style-v2-Q4_K_M-Calibrated.gguf → Jev-Style-v2-Calibrated-Q4_K_M.gguf +0 -0
- Jev-Style-v2-Q8_0-Calibrated.calibration.json → Jev-Style-v2-Calibrated-Q8_0.calibration.json +0 -0
- Jev-Style-v2-Q8_0-Calibrated.gguf → Jev-Style-v2-Calibrated-Q8_0.gguf +0 -0
- README.md +5 -5
- SHA256SUMS.json +7 -7
.gitattributes
CHANGED
|
@@ -39,3 +39,6 @@ figures/calibration.png filter=lfs diff=lfs merge=lfs -text
|
|
| 39 |
figures/robustness.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
Jev-Style-v2-Q4_K_M-Calibrated.gguf filter=lfs diff=lfs merge=lfs -text
|
| 41 |
Jev-Style-v2-BF16-Calibrated.gguf filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
figures/robustness.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
Jev-Style-v2-Q4_K_M-Calibrated.gguf filter=lfs diff=lfs merge=lfs -text
|
| 41 |
Jev-Style-v2-BF16-Calibrated.gguf filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
Jev-Style-v2-Calibrated-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
Jev-Style-v2-Calibrated-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
| 44 |
+
Jev-Style-v2-Calibrated-BF16.gguf filter=lfs diff=lfs merge=lfs -text
|
Jev-Style-v2-BF16-Calibrated.calibration.json → Jev-Style-v2-Calibrated-BF16.calibration.json
RENAMED
|
File without changes
|
Jev-Style-v2-BF16-Calibrated.gguf → Jev-Style-v2-Calibrated-BF16.gguf
RENAMED
|
File without changes
|
Jev-Style-v2-Q4_K_M-Calibrated.calibration.json → Jev-Style-v2-Calibrated-Q4_K_M.calibration.json
RENAMED
|
File without changes
|
Jev-Style-v2-Q4_K_M-Calibrated.gguf → Jev-Style-v2-Calibrated-Q4_K_M.gguf
RENAMED
|
File without changes
|
Jev-Style-v2-Q8_0-Calibrated.calibration.json → Jev-Style-v2-Calibrated-Q8_0.calibration.json
RENAMED
|
File without changes
|
Jev-Style-v2-Q8_0-Calibrated.gguf → Jev-Style-v2-Calibrated-Q8_0.gguf
RENAMED
|
File without changes
|
README.md
CHANGED
|
@@ -30,9 +30,9 @@ A **Jev-style decision model** for classification, routing and typed choices. Gi
|
|
| 30 |
|
| 31 |
| Precision | File size | Choice agreement | Macro accuracy |
|
| 32 |
|---|---:|---:|---:|
|
| 33 |
-
| [Q4_K_M](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-
|
| 34 |
-
| [Q8_0](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-
|
| 35 |
-
| [BF16](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-
|
| 36 |
|
| 37 |
All three files include independently fitted calibration; use **runtime temperature 1.0**. Agreement is against CUDA merged BF16 on the same frozen 500-decision subset; accuracy is the task-macro average over its real-label examples. [Full precision comparison](evaluation/quantization_summary.json).
|
| 38 |
|
|
@@ -136,7 +136,7 @@ The benchmark figures describe the fixed CUDA reference comparison. Reliability
|
|
| 136 |
python -m pip install -U huggingface_hub
|
| 137 |
hf download chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF --local-dir jev-v2-gguf
|
| 138 |
cd jev-v2-gguf
|
| 139 |
-
llama-server -m Jev-Style-v2-
|
| 140 |
```
|
| 141 |
|
| 142 |
In another terminal, from the same directory:
|
|
@@ -148,7 +148,7 @@ python jev_decision_client.py --url http://127.0.0.1:8080 \
|
|
| 148 |
--options negative positive
|
| 149 |
```
|
| 150 |
|
| 151 |
-
The quick-start command selects Q8_0. To use Q4_K_M or BF16, replace its model filename with `Jev-Style-v2-
|
| 152 |
|
| 153 |
Use a llama.cpp build with Qwen3.5 support. Conversion and native evaluation used commit `b29c606e28a01b1bc8c1351026a0fa6e616bf6c4`. The client uses the native `/completion` endpoint, requests complete declared-option log-probabilities and increases the candidate count as needed. The supplied `gguf_logits.cpp` reads all declared-option logits directly through the C API.
|
| 154 |
|
|
|
|
| 30 |
|
| 31 |
| Precision | File size | Choice agreement | Macro accuracy |
|
| 32 |
|---|---:|---:|---:|
|
| 33 |
+
| [Q4_K_M](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-Calibrated-Q4_K_M.gguf?download=true) | 1.27 GB | 91.4% | 78.18% |
|
| 34 |
+
| [Q8_0](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-Calibrated-Q8_0.gguf?download=true) | 2.01 GB | 99.2% | 78.69% |
|
| 35 |
+
| [BF16](https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-Calibrated-BF16.gguf?download=true) | 3.78 GB | 99.6% | 79.36% |
|
| 36 |
|
| 37 |
All three files include independently fitted calibration; use **runtime temperature 1.0**. Agreement is against CUDA merged BF16 on the same frozen 500-decision subset; accuracy is the task-macro average over its real-label examples. [Full precision comparison](evaluation/quantization_summary.json).
|
| 38 |
|
|
|
|
| 136 |
python -m pip install -U huggingface_hub
|
| 137 |
hf download chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF --local-dir jev-v2-gguf
|
| 138 |
cd jev-v2-gguf
|
| 139 |
+
llama-server -m Jev-Style-v2-Calibrated-Q8_0.gguf -c 2048 -ngl 99 --port 8080
|
| 140 |
```
|
| 141 |
|
| 142 |
In another terminal, from the same directory:
|
|
|
|
| 148 |
--options negative positive
|
| 149 |
```
|
| 150 |
|
| 151 |
+
The quick-start command selects Q8_0. To use Q4_K_M or BF16, replace its model filename with `Jev-Style-v2-Calibrated-Q4_K_M.gguf` or `Jev-Style-v2-Calibrated-BF16.gguf`.
|
| 152 |
|
| 153 |
Use a llama.cpp build with Qwen3.5 support. Conversion and native evaluation used commit `b29c606e28a01b1bc8c1351026a0fa6e616bf6c4`. The client uses the native `/completion` endpoint, requests complete declared-option log-probabilities and increases the candidate count as needed. The supplied `gguf_logits.cpp` reads all declared-option logits directly through the C API.
|
| 154 |
|
SHA256SUMS.json
CHANGED
|
@@ -1,9 +1,9 @@
|
|
| 1 |
{
|
| 2 |
-
"Jev-Style-v2-
|
| 3 |
"bytes": 1807,
|
| 4 |
"sha256": "7f3ae96ebf0271483365b8941124f3be9521e9e3ed4c3d62240e51c9b416ebb4"
|
| 5 |
},
|
| 6 |
-
"Jev-Style-v2-
|
| 7 |
"bytes": 2012004256,
|
| 8 |
"sha256": "5c2aa0d35b24a27f03228b2c62ebaaebd9b5b785844634d4217278d822751494"
|
| 9 |
},
|
|
@@ -13,7 +13,7 @@
|
|
| 13 |
},
|
| 14 |
"README.md": {
|
| 15 |
"bytes": 12086,
|
| 16 |
-
"sha256": "
|
| 17 |
},
|
| 18 |
"evaluation/baseline_sensitivity.json": {
|
| 19 |
"bytes": 2015,
|
|
@@ -91,11 +91,11 @@
|
|
| 91 |
"bytes": 147,
|
| 92 |
"sha256": "1f819d59b9ebf6ffb79953832c3f7ae4564d7bae7fa68209b6583a9d92db2bf5"
|
| 93 |
},
|
| 94 |
-
"Jev-Style-v2-
|
| 95 |
"bytes": 1274388384,
|
| 96 |
"sha256": "c697d3b29d07fdd37b6ebeb5c98066f4c31632162adca6258db23f75184eb0c4"
|
| 97 |
},
|
| 98 |
-
"Jev-Style-v2-
|
| 99 |
"bytes": 1824,
|
| 100 |
"sha256": "4a5abb23d86ddd44b8b6b1054b48a0a9c40560cc0ee3edd24ae6be8e623ad5f1"
|
| 101 |
},
|
|
@@ -107,11 +107,11 @@
|
|
| 107 |
"bytes": 694,
|
| 108 |
"sha256": "5a0e2b0696e3aa918aff6e8ad2ebf1532091edf81a10815bf3add221c36932d3"
|
| 109 |
},
|
| 110 |
-
"Jev-Style-v2-
|
| 111 |
"bytes": 3775700896,
|
| 112 |
"sha256": "8baa111eec6e30a5c9e97d559b53127a19ef30171e8aed26673153b4f77adcbb"
|
| 113 |
},
|
| 114 |
-
"Jev-Style-v2-
|
| 115 |
"bytes": 1840,
|
| 116 |
"sha256": "e5eb54d405d213c751ba61b334a8370d03332c915d9c49725da3b7a6ee295109"
|
| 117 |
},
|
|
|
|
| 1 |
{
|
| 2 |
+
"Jev-Style-v2-Calibrated-Q8_0.calibration.json": {
|
| 3 |
"bytes": 1807,
|
| 4 |
"sha256": "7f3ae96ebf0271483365b8941124f3be9521e9e3ed4c3d62240e51c9b416ebb4"
|
| 5 |
},
|
| 6 |
+
"Jev-Style-v2-Calibrated-Q8_0.gguf": {
|
| 7 |
"bytes": 2012004256,
|
| 8 |
"sha256": "5c2aa0d35b24a27f03228b2c62ebaaebd9b5b785844634d4217278d822751494"
|
| 9 |
},
|
|
|
|
| 13 |
},
|
| 14 |
"README.md": {
|
| 15 |
"bytes": 12086,
|
| 16 |
+
"sha256": "5d8614413dc1583b450b81ebeee569907a023c4c65d5752dd318f542b9071577"
|
| 17 |
},
|
| 18 |
"evaluation/baseline_sensitivity.json": {
|
| 19 |
"bytes": 2015,
|
|
|
|
| 91 |
"bytes": 147,
|
| 92 |
"sha256": "1f819d59b9ebf6ffb79953832c3f7ae4564d7bae7fa68209b6583a9d92db2bf5"
|
| 93 |
},
|
| 94 |
+
"Jev-Style-v2-Calibrated-Q4_K_M.gguf": {
|
| 95 |
"bytes": 1274388384,
|
| 96 |
"sha256": "c697d3b29d07fdd37b6ebeb5c98066f4c31632162adca6258db23f75184eb0c4"
|
| 97 |
},
|
| 98 |
+
"Jev-Style-v2-Calibrated-Q4_K_M.calibration.json": {
|
| 99 |
"bytes": 1824,
|
| 100 |
"sha256": "4a5abb23d86ddd44b8b6b1054b48a0a9c40560cc0ee3edd24ae6be8e623ad5f1"
|
| 101 |
},
|
|
|
|
| 107 |
"bytes": 694,
|
| 108 |
"sha256": "5a0e2b0696e3aa918aff6e8ad2ebf1532091edf81a10815bf3add221c36932d3"
|
| 109 |
},
|
| 110 |
+
"Jev-Style-v2-Calibrated-BF16.gguf": {
|
| 111 |
"bytes": 3775700896,
|
| 112 |
"sha256": "8baa111eec6e30a5c9e97d559b53127a19ef30171e8aed26673153b4f77adcbb"
|
| 113 |
},
|
| 114 |
+
"Jev-Style-v2-Calibrated-BF16.calibration.json": {
|
| 115 |
"bytes": 1840,
|
| 116 |
"sha256": "e5eb54d405d213c751ba61b334a8370d03332c915d9c49725da3b7a6ee295109"
|
| 117 |
},
|