Cloudflare Clef-Flash EXL3

EXL3 conversion of Cloudflare/clef-flash with a functional text/JSON Clef SystemOne adapter.

Source revision: 17f0b0ad64efb65d273590632833508766b2aae6

Quantization

  • Qwen3.5 decoder: EXL3 4.00 bpw
  • Vision tower: EXL3 6 bpw
  • LM head: FP16 (head_bits=16)
  • Cloudflare joint schema head: original BF16 weights retained in joint_head.safetensors; loaded as FP16 by the adapter
  • Tested with ExLlamaV3 1.5.2+cu128.torch2.10.0 on an RTX 4090

The LM head is intentionally kept at FP16 because Clef uses its output embedding vectors when scoring schema options.

Clef validation

The bundled clef_exl3.py bridges ExLlamaV3 final hidden states into Cloudflare's original JointSchemaHead and returns the same noul, choice, and score answer structures used by SystemOne.

Check BF16 reference EXL3
Invoice status: overdue 0.9732 0.9753
Invoice total > $1000 0.9732 0.9726
Outage routing: technical 0.9601 0.9572
Urgency expected score 1.7852 1.7966
Service outage: true 0.8339 0.8310

Across all numeric values in the two bundled validation cases, mean absolute delta was 0.004373 and maximum absolute delta was 0.0115.

clef-exl3-smoke.json, clef-bf16-reference.json, and VALIDATION.json contain the validation outputs and provenance summary.

A standard generation probe on the RTX 4090 measured 75.372 tok/s. This is a loader/generation smoke benchmark, not SystemOne decision throughput.

Current scope

TabbyAPI compatibility is not validated in this release and is not used as a publish gate; Clef relies on its custom SystemOne decision path rather than a standard chat-completions path.

Text and JSON state inputs are validated. The vision tower is included and quantized, but clef_exl3.py does not yet wire image/video inputs into the Clef decision path. Do not treat this release as validated multimodal SystemOne inference.

A preprocessor_config.json compatibility shim is included because this Cloudflare release stores the same image-processor metadata inside processor_config.json, while the tested ExLlamaV3 Qwen3.5 loader expects the standalone file.

Usage

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("ramgpt/clef-flash-EXL3")
sys.path.insert(0, path)
from clef_exl3 import ClefEXL3

model = ClefEXL3(path)
response = model.systemone({
    "model": "clef-flash-exl3",
    "state": {"invoice": {"total": 1250, "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Paid", "overdue": "Past due", "draft": "Not sent"}
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"}
    }
})
print(response["answers"])

Attribution

The base model, Clef joint schema head, and joint_schema_model.py originate from Cloudflare/clef-flash and are provided under the source model's Apache-2.0 license. This repository adds the EXL3 conversion, metadata compatibility shim, adapter, and validation artifacts.

Downloads last month
55
Safetensors
Model size
4B params
Tensor type
BF16
路
F16
路
I16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for ramgpt/clef-flash-EXL3

Finetuned
Qwen/Qwen3.5-9B
Quantized
(33)
this model