Instructions to use DreamBlooms/jet-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use DreamBlooms/jet-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf DreamBlooms/jet-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf DreamBlooms/jet-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf DreamBlooms/jet-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf DreamBlooms/jet-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf DreamBlooms/jet-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf DreamBlooms/jet-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf DreamBlooms/jet-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf DreamBlooms/jet-GGUF:Q8_0
Use Docker
docker model run hf.co/DreamBlooms/jet-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use DreamBlooms/jet-GGUF with Ollama:
ollama run hf.co/DreamBlooms/jet-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use DreamBlooms/jet-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DreamBlooms/jet-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "DreamBlooms/jet-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use DreamBlooms/jet-GGUF with Docker Model Runner:
docker model run hf.co/DreamBlooms/jet-GGUF:Q8_0
- Lemonade
How to use DreamBlooms/jet-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull DreamBlooms/jet-GGUF:Q8_0
Run and chat with the model
lemonade run user.jet-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use DreamBlooms/jet-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DreamBlooms/jet-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default DreamBlooms/jet-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use DreamBlooms/jet-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf DreamBlooms/jet-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "DreamBlooms/jet-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Jet
v6.2.0 · released 2026-09-26 · focused continuation, selected step 250.
Jet returns typed decisions and probabilities from supplied options without generating free-form answers.
This is the full merged BF16 Qwen3.5-4B text model, not an adapter. No separate
base-model download is required. It replaces v6.1 in michaljach/jet; previous
releases remain accessible by their version tags. The product name remains Jet.
Run
Linux + NVIDIA CUDA, Python 3.12. Verified with PyTorch 2.11.0+cu128, Transformers 5.17.0 and flash-linear-attention 0.5.2.
hf download michaljach/jet --revision v6.2.0 --local-dir jet
cd jet
python -m pip install -r requirements.txt
python jet.py <<'JSON'
{"state":"I was charged twice this month.","questions":{"topic":{"type":"choice","instructions":"What is the primary issue?","criteria":{"billing":"billing or payment problem","bug":"the product is broken"}}}}
JSON
from jet import Jet
model = Jet()
result = model.decide("The item arrived broken.", {
"damaged": {"type": "noul", "instructions": "Is the item damaged?"}
})
| Type | Input | Output |
|---|---|---|
choice |
2–255 named options | Selected key and probabilities |
score |
2–10 ordered levels | Expected zero-based score, selected level, probabilities |
noul |
Yes/no question | Probability of yes |
Use the included native prompt and restricted label-token readout. A generic text-generation pipeline does not implement this API. Questions are processed separately. The runtime now accepts complete prompts up to 16,384 tokens; inputs and options are never silently truncated. The longest reconstructed API-Bank input (11,495 tokens) passed the standalone runtime check, but full API-Bank accuracy has not been measured for this release. Context feasibility does not establish long-context quality across tasks.
FLA is enabled; convolution uses the PyTorch fallback. confidence is normalized
inverse entropy, not a calibrated correctness probability. Temperatures are
inherited from v6.1 and were not refitted or established as calibrated for v6.2.
Training and selection
The parent is full merged Jet v6.1 at
f446b82727be57da348bb46eccf211414294ab3e. Parent shard hashes and the exact
adapter hash are in merge-provenance.json. A fresh rank/alpha-16 correction LoRA
was trained on that parent; the model was not reset to the original Qwen weights.
- 4,000 examples: 800 banking intent, 800 entity-specific financial sentiment, 400 sarcasm/literal, and 2,000 broad retention examples.
- BF16 backbone / FP32 adapter, dropout .05, microbatch 1, accumulation 4, seed 260925; two trials of 1,000 updates at peak learning rates 3e-6 and 8e-6.
- Validation-selected winner: 3e-6 at step 250. The higher-rate trial failed the selection guards and is not included in this release.
- Selection used 915 examples and an objective of 70% mean focus metrics plus 30% retention accuracy; every family must stay within two percentage points of its initial metric. The winner was frozen before final holdout evaluation.
- Banking uses the training partition of BANKING77, financial supervision uses SEntFiN, and sarcasm uses author-labeled iSarcasm training examples. Financial headlines and tweet/rephrase groups stay in one split.
Full-model holdout results
The results below were re-measured on the full merged model, after BF16 rounding. Focus holdouts exclude the previous local input corpus listed in the data audit; retention is reused. Small sample sizes mean these modest differences do not establish statistical significance or broad benchmark improvement.
| Holdout / metric | Cases | Jet v6.1 | Jet v6.2 merged |
|---|---|---|---|
| Banking / accuracy | 154 | 74.68% | 74.68% |
| Entity sentiment / macro-F1 | 180 | 71.16% | 71.21% |
| Sarcasm / positive-class F1 | 80 | 46.81% | 50.00% |
| Retention / accuracy | 500 | 93.40% | 93.40% |
Financial sentiment is a SEntFiN transfer holdout, not the FinEntity benchmark. Banking and sarcasm rows above are local source holdouts, not their public test benchmark scores. Exact content/group exclusions do not establish semantic or pretraining decontamination. Benchmarks informed training focus.
Official overall Decision Index: not measured. The 25-benchmark comparison linked on the website belongs to the earlier v6.1 candidate, and must not be attributed to this release. No subset average is substituted for an overall score.
Merge verification
FP32 B@A corrections were added to the already merged parent and rounded to BF16. All 426 tensors loaded without missing, unexpected or mismatched keys; 248 modules received updates. Nine shards contain 8,411,510,272 parameter bytes.
On 157 fixed verification cases, 0 selected answers changed between the adapter and merged model. The largest probability difference was 4.76 percentage points.
These checks include choice, score, noul and a long API-Bank input. The full 914-case holdout passed the predeclared release guards. See evaluation.json, merge-validation.json and release-manifest.json. Merge equivalence is approximate.
Source and experiment records · Training protocol · Release validation · Historical benchmark charts
Apache-2.0 model/runtime; source datasets retain their own licenses. Dataset rows are not redistributed here. Quality outside the evaluated domains is not established.
GGUF
CPU inference is available through
dohnuts.cpp, a native C++ port on
llama.cpp. It runs the model on CPU,
reads the answer in one forward pass, and serves the same POST /v1/systemone
wire format as the upstream release. No GPU or Python runtime is needed. Each question is one forward pass; the next-token logits are restricted to the option labels (A, B, ... for choice, the level digits for score, no/yes for noul) and read after the chat decision prompt. jet fits one temperature per question type.
| File | Quantization | Size |
|---|---|---|
jet-4b-Q8_0.gguf |
Q8_0 | 4.5 GB |
jet.json |
profile config and per-type temperatures | 219 B |
jet.json is required alongside the GGUF. It carries the three calibrated temperatures (t_choice, t_score, t_noul) from the release's calibration.json.
git clone https://github.com/DreamBlooms/dohnuts.cpp
cd dohnuts.cpp
git submodule update --init --depth 1
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON
cmake --build build -j --target dohnuts-cli
build/dohnuts-cli --server --port 8080 \
--model jet-4b-Q8_0.gguf --metadata jet.json
Ask one state several questions:
curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' \
-d '{"state":"Refund policy: full refund within 30 days of purchase; 50% until day 60; none after. Order 1182 was bought on 3 March and returned on 20 April.",
"questions":{"refund":{"type":"choice","instructions":"What refund does order 1182 get?","criteria":{"full":"Full refund","half":"50% refund","none":"No refund"}}}}'
The answer keeps the core fields (type, choice, probabilities, noul,
score, confidence) and adds this model's own statistics under native. jet adds certainty (with legend for score).
Rebuild the file from the upstream weights with
scripts/build_jet_gguf.sh. No retraining is involved.
- Downloads last month
- 76
8-bit