Clef-Flash GGUF

This repository holds Cloudflare/clef-flash in one GGUF file for bobcat. The file holds the Qwen3.5-9B backbone in Q8_0 and Clef's joint schema head in its original bfloat16, under tensor names that start with clef..

bobcat pull clef:flash
bobcat serve -m clef:flash

The server answers SystemOne requests at POST /v1/systemone with the decision head, and bobcat decide -m clef:flash answers one request from the command line.

llama.cpp's converter wrote the backbone with --no-mtp --outtype q8_0, and bobcat's tools/clef_gguf.py added the head and set the tokenizer's pre-tokenizer to Qwen2's, which Clef's tokenizer.json uses. The file carries no vision encoder, so it answers text states. bobcat's decisions match Cloudflare's reference code to within 0.005 on every probability of its test request. llama.cpp rejects the head's extra tensors, so the file runs in bobcat only.

The model and its head come from Cloudflare under the Apache 2.0 license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jadidbourbaki/clef-flash-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(30)
this model