StartLux-Decision-4B-Q8_0-GGUF

StartLux-Decision-4B in Q8_0 (8-bit) as a GGUF file for llama.cpp. The decision server in this repository turns it into the same typed decisions, with a probability for every option, as the original model.

Precisions

Repository Bits Size Same decision as the original JevBench public, of 231
StartLux-Decision-4B-Q8_0-GGUF (this repository) 8-bit 4.48 GB 100.0% 204 (88.3%) recommended
StartLux-Decision-4B-Q4_K_M-GGUF 4-bit 2.71 GB 98.3% 201 (87.0%) smallest
StartLux-Decision-4B-BF16-GGUF 16-bit 8.42 GB 100.0% 204 (88.3%) the original weights, unchanged
StartLux-Decision-4B (original weights) 204 (88.3%) for comparison

"Same decision" is the share of the 231 public JevBench items on which the file picks the same answer as the original weights run through the startlux_decision package, with the same prompts, readout and temperatures. Q8_0 and Q4_K_M are llama.cpp's standard quantizations of BF16, without an importance matrix.

Download and run

hf download startlux-models/StartLux-Decision-4B-Q8_0-GGUF --local-dir StartLux-Decision-4B-Q8_0-GGUF
cd StartLux-Decision-4B-Q8_0-GGUF
pip install -r requirements.txt                     # transformers and torch; a CPU build of torch is enough

# llama.cpp serves the weights; the decision server puts the prompt format, readout and calibration on top
llama-server -m StartLux-Decision-4B-Q8_0.gguf -ngl 99 -c 16384 --parallel 4 --port 8081 &
python -m startlux_decision.gguf_server --model-dir . --llama http://127.0.0.1:8081 --port 8090

-ngl 99 puts every layer on the GPU (CUDA or Metal); leave it out on a CPU-only machine. llama.cpp has to be recent enough for this model (build b10454 or newer).

Requests and responses use the TypeSafe /v1/systemone format:

curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"ticket": "I was charged twice for order #4411 and the app still shows it as unpaid."},
  "questions": {
    "team":   {"type": "choice", "instructions": "Which team should handle this ticket?",
               "criteria": {"billing": "Payments, refunds and invoices",
                            "shipping": "Delivery and tracking",
                            "technical": "App, login and account problems"}},
    "urgent": {"type": "noul", "instructions": "Should this ticket be answered today?"}
  }
}'

Plain chat with a GGUF file does not give these decisions: the option-letter readout and the per-type temperatures live in gguf_server.py, not in the weights. The file holds the text decoder only.

License

The model weights are released under CC BY-NC 4.0: free for research and other non-commercial use, with attribution. Commercial use requires a separate license from StartLux Labs; contact contact@startlux.com. The inference code in startlux_decision/ is Apache-2.0. See LICENSE and NOTICE.

Downloads last month
3,581
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for startlux-models/StartLux-Decision-4B-Q8_0-GGUF

Quantized
(4)
this model

Space using startlux-models/StartLux-Decision-4B-Q8_0-GGUF 1

Collection including startlux-models/StartLux-Decision-4B-Q8_0-GGUF