Instructions to use cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF # Run inference directly in the terminal: llama cli -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF # Run inference directly in the terminal: llama cli -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF # Run inference directly in the terminal: ./llama-cli -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
Use Docker
docker model run hf.co/cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
- LM Studio
- Jan
- vLLM
How to use cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
- Ollama
How to use cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF with Ollama:
ollama run hf.co/cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
- Unsloth Desktop
- Docker Model Runner
How to use cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF with Docker Model Runner:
docker model run hf.co/cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
- Lemonade
How to use cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull cdtio33/Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF
Run and chat with the model
lemonade run user.Ornith-1.5-397B-IQ2_XXS-DFlash-Drafter-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
DFlash Drafter for Ornith 1.5 397B A17B IQ2_XXS (GGUF)
This is a drafter, not a language model. On its own it produces nothing useful. It is
loaded through --model-draft alongside Ornith-1.5-397B-A17B-IQ2_XXS and predicts that
model's next few tokens from its hidden states.
It runs on ik_llama.cpp only. Mainline llama.cpp and the poolside fork both reject it; see Compatibility before downloading.
Against the previous drafter for this target, on code prompts it was never trained on, it
decodes +6.5% faster at n_max=3 (95% CI +0.3% to +12.6%, four server boots,
counterbalanced). Acceptance rate rises from 0.47 to 0.52 at the same setting. Full numbers
and their weaknesses are in Evaluation.
Files
| File | Size | Precision | SHA-256 |
|---|---|---|---|
Ornith15-DFLASH-codering-q4.gguf |
704 MiB | mixed iq4_kss / iq3_ks |
5ee79024b8337367f222ceaa050739b03553d4669551db394aa21227e2f271f6 |
This is the measured artifact: every number below was produced with this exact file, and the SHA-256 is there so you can confirm you are reading results about the bytes you downloaded.
A pure q8_0 build of the same weights exists locally and is not published, because it was
never benchmarked and would only invite the assumption that a larger file is a better one.
Ask if you want it.
The q4 mix was chosen to match the incumbent drafter's, so that an A/B between the two
compares weights rather than quantization:
dflash_fc.weight=iq4_kss, ffn_down=iq4_kss, ffn_gate=iq3_ks, ffn_up=iq3_ks
Attention projections and norms stay at Q8_0.
Usage
llama-server \
--model Ornith-1.5-397B-IQ2_XXS-00001-of-00003.gguf \
--model-draft Ornith15-DFLASH-codering-q4.gguf \
--spec-type dflash:n_max=3,cross_ctx=512 \
--n-gpu-layers 99 --n-gpu-layers-draft 99 \
--draft-params "-wgt 17" -wgt 8 \
--flash-attn on --parallel 1
All three target shards must sit in one directory; naming shard one loads the rest.
--parallel 1 is a hard requirement, not a tuning choice. This ik_llama.cpp build refuses
to start with speculative decoding and more than one slot, because the accept path indexes
the batch absolutely. It does not degrade to serialized slots.
The two -wgt flags reserve worst-case graph space for the target and draft contexts. They
were set during measurement, so omitting them will not reproduce the numbers below.
cross_ctx=512 is the drafter's ring depth and is fixed by the architecture. It also
constrains capture: exporting training features with a prefill batch larger than 512 evicts
ring rows before they are written, which silently corrupts the pack.
Choosing n_max
Swept 2026-08-28 on one boot at N_MAX=6, four unseen ~32k-token prompts, two repetitions,
speculative.n_max overridden per request so every depth shares a single model load:
n_max |
tok/s | vs spec-off | acceptance |
|---|---|---|---|
| 0 | 15.00 | β | β |
| 2 | 19.07 | +27.3% | 0.562 |
| 3 | 19.18 | +28.0% | 0.462 |
| 4 | 19.82 | +32.4% | 0.431 |
| 5 | 17.72 | +18.5% | 0.334 |
The curve rises to 4 and turns over at 5, so the optimum is bracketed and depths of 6 and above need no test. Acceptance falls monotonically with depth, as expected.
Depths 2, 3 and 4 are not statistically separable on this data. Paired by prompt against
the shipped setting: n_max=2 is β0.0% (t=0.00), n_max=4 is +4.5% (t=0.75, 95% CI
[β14.5%, +23.4%]), n_max=5 is β6.7% (t=β1.30).
The fall-off at 5 needs stating precisely, because which depth you compare it to changes the
answer. Against n_max=4 it is β2.10 tok/s with t = 10.9 and every prompt agreeing, which is
about as clean as this rig produces. Against the shipped n_max=3 it is β1.46 tok/s with
t = 1.45 and one prompt going the other way β not separable. So 5 is reliably worse than the
peak, and only ambiguously worse than what is shipped.
n_max=3 ships unchanged. 4 is the point-estimate peak, but nothing here distinguishes it
from 3, and the previous drafter's optimum of 4 is a prior rather than evidence about this
one.
Worth knowing if you extend this: repetition-to-repetition spread within a cell had a median
of 0.06 tok/s across 20 cells (one outlier at 1.27), while prompt-to-prompt spread ran 2.7 to
5.8 tok/s depending on depth, and one prompt (fresh-03) sat 3β5 tok/s below the others at
every depth. The uncertainty here is prompt heterogeneity, not measurement noise, so more
prompts would sharpen this and more repetitions would not.
Verify cost on this target is roughly 16.2 ms per extra draft position, and Ornith 1.5 is a sparse MoE whose experts do not amortize across draft positions. Deeper drafting therefore stops paying well before acceptance reaches zero.
Architecture
1.29B parameters in 6 layers. The drafter reads the target's hidden states from 8 layers
(1, 9, 17, 25, 33, 41, 49, 57), concatenated to a 32768-wide input, and projects them
through dflash_fc into its own 4096-wide residual stream. A block of 16 mask slots seeded
with the anchor token is decoded in one forward pass, and the resulting hidden states are
pushed through the target's frozen embedding and output head.
| Layers | 6 (sliding_attention Γ5, full_attention Γ1) |
| Hidden / intermediate | 4096 / 12288 |
| Attention heads | 32 query, 8 KV, head_dim 128 |
| Sliding window | 4096 |
| RoPE theta | 1e7 |
| Block size | 16 (mask token 248077) |
| Vocab | 248320 |
RoPE theta must read 1e7 in the GGUF metadata. At the Qwen default of 1e6 acceptance
collapses, and the converter reads the top-level rope_theta key rather than the nested
rope_parameters block, so a config carrying only the nested value converts to a drafter
that loads cleanly and drafts badly.
The file contains no target weights. Its 69 tensors are the six blocks, dflash_fc,
dflash_hidden_norm and output_norm. There is no token_embd and no output.weight;
both come from the target at runtime. Redistributing this drafter therefore redistributes
only weights derived from the z-lab DFlash initialization.
Known trainer/runtime divergence
The training script forces all six layers to full attention, while the shipped config.json
correctly declares sliding attention on layers 0β4. The drafter was therefore trained able to
attend to later slots inside its block, and is served without that ability in five of six
layers.
This was measured rather than assumed. A fixture that reproduces both graphs on the same weights and anchors puts the difference at Β±0.8 pp on per-position match, in both directions, at every slot count tested. The window is 4096 while intra-block distances are at most 15, so it never binds. The divergence is real and worth knowing about; it is not worth correcting.
Training
Initialized from the Ornith-1.0 drafter (itself initialized from z-lab/Qwen3.5-397B-A17B-DFlash),
then fine-tuned on target features captured from the served IQ2_XXS model.
The motivating observation: three previous retrains had produced almost nothing at deployed settings, and every one of them trained on anchors from long prose generations off short prompts. Production's drafter ring holds code at absolute positions above 30,000. The untested variable was the data distribution, not the loss or the schedule.
| Capture | 12 prompts Γ ~128k chars of C/C++ from the ik_llama.cpp tree, 4096 generated tokens each |
| Pack | 457,790 rows, 30 GB, 12 segments, 0 malformed |
| Split | 9 segments train (348,025 rows) / 3 eval (109,765), split by captured request |
| Mixed with | the previous prose pack, pooling to 21,950 anchors |
| Schedule | 2 epochs, 2,742 steps at ~1.3 s/step, batch 16, lr 1e-4, stride 5 |
| Loss | position-weighted CE, gamma 1.5, weights normalized to sum 15 (wβ..wβ = 7.30 / 3.75 / 1.92 / 0.99 / 0.51) |
The weighting concentrates on positions 1β3 because tau@3 β 1 + pβ + pβpβ β position 1 caps
everything downstream, and the deployed configuration only ever requests 3 or 4 slots.
Best tau@15 on held-out code was 4.42 against 3.91 for the initialization.
Evaluation
Two things are worth separating: whether the drafter is better, and by how much at the setting you will actually serve.
Offline, on the deployed graph
Measured with the runtime's mask and slot count rather than the trainer's, on 250 anchors, against the previous drafter for this target. Only genuinely held-out segments are included β see Data contamination.
| Pack | Metric | Previous | This | Ξ |
|---|---|---|---|---|
| Held-out code | tau@3 (serving n_max=2) |
2.592 | 2.704 | +4.3% |
| Held-out code | tau@4 (serving n_max=3) |
3.100 | 3.268 | +5.4% |
| Held-out code | position-1 match | 0.908 | 0.928 | +2.0 pp |
| Prose (old distribution) | tau@3 / tau@4 |
2.452 / 2.916 | 2.504 / 2.976 | +2.1% |
The prose row matters more than its size suggests: this drafter was trained on code, and it did not get worse anywhere measured.
End-to-end throughput
Four server boots in A-B-B-A order, four unseen ~32k-token code prompts, three repetitions, 256 tokens generated per request. Counterbalancing is necessary because boot-to-boot drift on this hardware measured ~8%, larger than the effect. Deltas are paired within each half and clustered by prompt, so the unit of analysis is the prompt rather than the request.
n_max |
Ξ tok/s | 95% CI | t (df=3) |
|---|---|---|---|
| 0 (control) | +0.6% | [+0.3%, +0.9%] | 6.70 |
| 2 | +3.9% | [β0.4%, +8.2%] | 2.88 |
| 3 | +6.5% | [+0.3%, +12.6%] | 3.34 |
Every prompt improved at n_max=3 (+11.2 / +7.3 / +1.9 / +5.5%). Acceptance rate rose in
both boots independently β 0.472 / 0.474 to 0.508 / 0.531 β which is read from the server's
own counters rather than a wall clock.
The control row is the honest caveat. With speculation disabled the two arms should be
identical, and instead the candidate's boots ran 0.6% faster with a CI excluding zero. Some
small boot-level advantage was not fully cancelled. Subtracting it gives roughly +5.9% at
n_max=3.
n_max=2 is directionally positive and not separable from zero.
Data contamination
The first version of this evaluation was partly contaminated, and the correction is documented here rather than quietly fixed.
The prompt generator rotated a six-entry list of source files, so requesting 12 prompts produced 12 prompts with only 9 unique bodies. The four benchmark prompts were the first four of that same generator's output, and three of them were in the training split. One of the three offline evaluation segments duplicated a trained segment.
After rebuilding the evaluation on one distinct unseen source file per prompt, hash-guarded against the training corpus at build time:
- Offline recomputed on clean segments only: +6.9% / +8.6% fell to +4.3% / +5.4%. Part of the original gap was memorization.
- End-to-end on fresh prompts: +5.6% became +6.5%, with the CI excluding zero for the first time.
The contaminated online number was not inflated. That is the outcome, not the justification β the original measurement could not have told you either way.
What is not established
- Anything outside code. The code corpus is nine unique prompt bodies drawn from six C/C++ source files in one project. The evidence for general text is the +2.1% no-regression row above, which rests on two held-out prose segments.
- The optimal
n_max. See Choosing n_max. - Other quantizations of the target. The drafter is distilled against IQ2_XXS specifically, using that quantization's embedding and output head. Behaviour on IQ2_KS, Q4_K_M or bf16 is unmeasured.
- Long-horizon serving. All measurements are 256-token generations on a benchmark rig, not production traffic over days.
Four prompts and four boots is a small experiment. Treat +6.5% as the center of a wide interval.
Compatibility
| Runtime | Result |
|---|---|
ik_llama.cpp (HEAD f96eaddb or later) |
Works |
| mainline llama.cpp | unknown model architecture: 'dflash-draft' |
| poolside llama.cpp | unknown model architecture: 'dflash-draft' β its DFlash arch is registered as dflash |
The published q4 file additionally uses ik-native quantization types that mainline ggml
does not define, so it fails at tensor load before even reaching the architecture check. The
quantization is not the obstacle, though: a pure q8_0 build of these weights, using only
mainline types, still fails on the architecture name. The format is the wall, not the mix.
Both forks implement DFlash. Their drafter formats are not interchangeable, and running this drafter on poolside would require retraining and reconverting to that format.
Reproducing
Measured on an NVIDIA GB10 (DGX Spark), 121.6 GiB unified memory shared between CPU and GPU.
The target at IQ2_XXS with a 131,072-token context and q8_0 KV cache leaves a few GiB
free, so nothing else of size can run alongside it β including, on this hardware, an
evaluation script that expects to use "the CPU" as if that were a separate pool.
The full command sequence, the capture settings, the split, the training invocation and the
statistical analysis are in the source repository under kernel-lab/dflash-retrain/:
CODERING-REPRO-RUNBOOK.md for the commands in execution order,
CODERING-AB-STATISTICAL-NOTES.md for the A/B analysed with its weaknesses stated.
License
Apache 2.0, inherited from z-lab/Qwen3.5-397B-A17B-DFlash, of which these weights are a
fine-tune.
Training features were distilled from Ornith-1.5-397B outputs. If you intend to redistribute or use this commercially, check Ornith 1.5's own license terms on distillation β that question is not resolved by this card.
- Downloads last month
- 140
We're not able to determine the quantization variants.