Clef-Flash GGUF
This repository holds Cloudflare/clef-flash
in one GGUF file for bobcat. The file holds the
Qwen3.5-9B backbone in Q8_0 and Clef's joint schema head in its original bfloat16, under
tensor names that start with clef..
bobcat pull clef:flash
bobcat serve -m clef:flash
The server answers SystemOne requests at POST /v1/systemone with the decision head, and
bobcat decide -m clef:flash answers one request from the command line.
llama.cpp's converter wrote the backbone with --no-mtp --outtype q8_0, and bobcat's
tools/clef_gguf.py added the head and set the tokenizer's pre-tokenizer to Qwen2's, which
Clef's tokenizer.json uses. The file carries no vision encoder, so it answers text states.
bobcat's decisions match Cloudflare's reference code to within 0.005 on every probability of its
test request. llama.cpp rejects the head's extra tensors, so the file runs in bobcat only.
The model and its head come from Cloudflare under the Apache 2.0 license.