Text Generation
GGUF
jev-style
English
decision-model
classification
calibration
qwen3.5
single-prefill
conversational
Instructions to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- jev-style
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with jev-style:
pip install jev-style # GGUF builds score through llama.cpp: build the jev-score binary once hf download chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF build_jev_score.sh jev_score.cpp --local-dir jev-score export JEV_SCORE_BIN=$(sh jev-score/build_jev_score.sh /path/to/llama.cpp | tail -n 1)
from jev_style import JevStyle, noul, choice js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF") out = js.decide("I was charged twice for one order.", { "billing": noul("This message is about billing."), "team": choice("Which team should handle it?", ["billing", "shipping", "tech"]), }) print(out["answers"]["team"]["choice"]) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Ollama
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Ollama:
ollama run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Docker Model Runner:
docker model run hf.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
- Lemonade
How to use chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Jev-Style-Qwen3.5-2B-Decision-v2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Download Jev-Style-v2-Calibrated-BF16.calibration.json from chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 1.84 kB
-
https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-Calibrated-BF16.calibration.json
- Command line
-
hf download hf://chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/Jev-Style-v2-Calibrated-BF16.calibration.json
-
curl -L -o Jev-Style-v2-Calibrated-BF16.calibration.json https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2-GGUF/resolve/main/Jev-Style-v2-Calibrated-BF16.calibration.json
1.84 kB
| { | |
| "temperature": 1.0, | |
| "fitted_temperature_folded": 1.0408715111841746, | |
| "temperature_folded": true, | |
| "folded_tensor": "output_norm.weight", | |
| "source_calibration": { | |
| "temperature": 1.0408715111841746, | |
| "calibration_n": 3100, | |
| "objective": "sample_mean_soft_cross_entropy", | |
| "bounds": [ | |
| 0.05, | |
| 20 | |
| ], | |
| "nll_before": 0.5144028513612269, | |
| "nll_after": 0.5140996476734587, | |
| "backend": "gguf", | |
| "model": "source_project/h100_v2/results/Jev-Style-v2-BF16-uncalibrated.gguf", | |
| "source_sha256": "2fde7f45dce3440abfde145bb30ad61e2643b1f853866b5760b235685328dc1c" | |
| }, | |
| "input_sha256": "79fb883ffc5831f165c11ca0119c7702b3c5209c916d939f2555fb1c7d45413c", | |
| "output_sha256": "8baa111eec6e30a5c9e97d559b53127a19ef30171e8aed26673153b4f77adcbb", | |
| "validation_required": false, | |
| "validation": { | |
| "format": "BF16", | |
| "n": 500, | |
| "argmax_agreement": 0.996, | |
| "cuda_same_subset": { | |
| "accuracy": 0.7910177949703642, | |
| "macro_f1": 0.7760857891780776, | |
| "nll": 0.6000106706289864, | |
| "brier": 0.317876961372947, | |
| "ece": 0.16696329399611096 | |
| }, | |
| "gguf_same_subset": { | |
| "accuracy": 0.7936151975677668, | |
| "macro_f1": 0.77850284511292, | |
| "nll": 0.6007551256850553, | |
| "brier": 0.31813150017828207, | |
| "ece": 0.16639447171288832 | |
| }, | |
| "accuracy_difference": 0.0025974025974025983, | |
| "nll_difference": 0.0007444550560689045, | |
| "temperature_folded": true, | |
| "runtime_temperature": 1.0, | |
| "calibration_n": 3100, | |
| "criteria": { | |
| "agreement_min": 0.99, | |
| "accuracy_loss_max": 0.005, | |
| "nll_increase_max": 0.015 | |
| }, | |
| "calibration_backend": "llama.cpp batched, checked against native readout", | |
| "evaluation_backend": "llama.cpp native single-sequence exact option logits", | |
| "weight_bytes": 3775700896, | |
| "passed": true | |
| } | |
| } | |