Text Generation
GGUF
jev-style
English
decision-model
classification
calibration
qwen3.5
single-prefill
conversational
Instructions to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- jev-style
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with jev-style:
pip install jev-style # GGUF builds score through llama.cpp: build the jev-score binary once hf download chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF build_jev_score.sh jev_score.cpp --local-dir jev-score export JEV_SCORE_BIN=$(sh jev-score/build_jev_score.sh /path/to/llama.cpp | tail -n 1)
from jev_style import JevStyle, noul, choice js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF") out = js.decide("I was charged twice for one order.", { "billing": noul("This message is about billing."), "team": choice("Which team should handle it?", ["billing", "shipping", "tech"]), }) print(out["answers"]["team"]["choice"]) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Ollama
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Ollama:
ollama run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Docker Model Runner:
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Lemonade
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Jev-Style-Qwen3.5-2B-Decision-v2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
README: add Ollama usage; SHA256SUMS: update README, add template/params
Browse files- README.md +8 -0
- SHA256SUMS.json +10 -2
README.md
CHANGED
|
@@ -154,6 +154,14 @@ Use a llama.cpp build with Qwen3.5 support. Conversion and native evaluation use
|
|
| 154 |
|
| 155 |
**Runtime calibration temperature is 1.0** for this file: its fitted temperature has already been incorporated. The accompanying calibration JSON records the exact settings and checksum. Serve the raw decision prompt shown below, with the full declared option list.
|
| 156 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 157 |
|
| 158 |
## Decision interface
|
| 159 |
|
|
|
|
| 154 |
|
| 155 |
**Runtime calibration temperature is 1.0** for this file: its fitted temperature has already been incorporated. The accompanying calibration JSON records the exact settings and checksum. Serve the raw decision prompt shown below, with the full declared option list.
|
| 156 |
|
| 157 |
+
### Ollama
|
| 158 |
+
|
| 159 |
+
```bash
|
| 160 |
+
ollama run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
|
| 161 |
+
```
|
| 162 |
+
|
| 163 |
+
Send the raw decision prompt shown below as the message; the model replies with the letter of the selected option (` B` for the example). Replace `Q4_K_M` with `Q8_0` or `BF16` for another precision. This repository supplies its own Ollama template, so run `ollama pull` again if you pulled the model before 2026-09-25.
|
| 164 |
+
|
| 165 |
|
| 166 |
## Decision interface
|
| 167 |
|
SHA256SUMS.json
CHANGED
|
@@ -12,8 +12,8 @@
|
|
| 12 |
"sha256": "50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32"
|
| 13 |
},
|
| 14 |
"README.md": {
|
| 15 |
-
"bytes":
|
| 16 |
-
"sha256": "
|
| 17 |
},
|
| 18 |
"evaluation/baseline_sensitivity.json": {
|
| 19 |
"bytes": 2015,
|
|
@@ -122,5 +122,13 @@
|
|
| 122 |
"evaluation/bf16_http_smoke.json": {
|
| 123 |
"bytes": 690,
|
| 124 |
"sha256": "d729e4dd7ab5d425e5eb031ca1dc92d74302a7971d9aeeba782c15e4904a5f59"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 125 |
}
|
| 126 |
}
|
|
|
|
| 12 |
"sha256": "50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32"
|
| 13 |
},
|
| 14 |
"README.md": {
|
| 15 |
+
"bytes": 12508,
|
| 16 |
+
"sha256": "406e98c0183ed8033b9c5db744454f06c5a45321b82aa024f3051949f65be2ae"
|
| 17 |
},
|
| 18 |
"evaluation/baseline_sensitivity.json": {
|
| 19 |
"bytes": 2015,
|
|
|
|
| 122 |
"evaluation/bf16_http_smoke.json": {
|
| 123 |
"bytes": 690,
|
| 124 |
"sha256": "d729e4dd7ab5d425e5eb031ca1dc92d74302a7971d9aeeba782c15e4904a5f59"
|
| 125 |
+
},
|
| 126 |
+
"template": {
|
| 127 |
+
"bytes": 91,
|
| 128 |
+
"sha256": "244d5cfdeb97383239af17d352fc32c2123257b2eecda8fd8285a7cd29770648"
|
| 129 |
+
},
|
| 130 |
+
"params": {
|
| 131 |
+
"bytes": 31,
|
| 132 |
+
"sha256": "208f455e3685ae841dea00e45a06d3b7fd226a0fbf88e6f2cd1077ff8e6b36af"
|
| 133 |
}
|
| 134 |
}
|