--- base_model: Cloudflare/clef-flash license: apache-2.0 library_name: exllamav3 tags: - exl3 - exllamav3 - clef - systemone - qwen3.5 --- # Cloudflare Clef-Flash EXL3 EXL3 conversion of `Cloudflare/clef-flash` with a functional text/JSON Clef SystemOne adapter. Source revision: `17f0b0ad64efb65d273590632833508766b2aae6` ## Quantization - Qwen3.5 decoder: EXL3 4.00 bpw - Vision tower: EXL3 6 bpw - LM head: FP16 (`head_bits=16`) - Cloudflare joint schema head: original BF16 weights retained in `joint_head.safetensors`; loaded as FP16 by the adapter - Tested with ExLlamaV3 `1.5.2+cu128.torch2.10.0` on an RTX 4090 The LM head is intentionally kept at FP16 because Clef uses its output embedding vectors when scoring schema options. ## Clef validation The bundled `clef_exl3.py` bridges ExLlamaV3 final hidden states into Cloudflare's original `JointSchemaHead` and returns the same `noul`, `choice`, and `score` answer structures used by SystemOne. | Check | BF16 reference | EXL3 | |---|---:|---:| | Invoice status: overdue | 0.9732 | 0.9753 | | Invoice total > $1000 | 0.9732 | 0.9726 | | Outage routing: technical | 0.9601 | 0.9572 | | Urgency expected score | 1.7852 | 1.7966 | | Service outage: true | 0.8339 | 0.8310 | Across all numeric values in the two bundled validation cases, mean absolute delta was `0.004373` and maximum absolute delta was `0.0115`. `clef-exl3-smoke.json`, `clef-bf16-reference.json`, and `VALIDATION.json` contain the validation outputs and provenance summary. A standard generation probe on the RTX 4090 measured `75.372 tok/s`. This is a loader/generation smoke benchmark, not SystemOne decision throughput. ## Current scope TabbyAPI compatibility is not validated in this release and is not used as a publish gate; Clef relies on its custom SystemOne decision path rather than a standard chat-completions path. Text and JSON state inputs are validated. The vision tower is included and quantized, but `clef_exl3.py` does not yet wire image/video inputs into the Clef decision path. Do not treat this release as validated multimodal SystemOne inference. A `preprocessor_config.json` compatibility shim is included because this Cloudflare release stores the same image-processor metadata inside `processor_config.json`, while the tested ExLlamaV3 Qwen3.5 loader expects the standalone file. ## Usage ```python import sys from huggingface_hub import snapshot_download path = snapshot_download("ramgpt/clef-flash-EXL3") sys.path.insert(0, path) from clef_exl3 import ClefEXL3 model = ClefEXL3(path) response = model.systemone({ "model": "clef-flash-exl3", "state": {"invoice": {"total": 1250, "status": "overdue"}}, "questions": { "status": { "type": "choice", "instructions": "What is the invoice status?", "criteria": {"paid": "Paid", "overdue": "Past due", "draft": "Not sent"} }, "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"} } }) print(response["answers"]) ``` ## Attribution The base model, Clef joint schema head, and `joint_schema_model.py` originate from `Cloudflare/clef-flash` and are provided under the source model's Apache-2.0 license. This repository adds the EXL3 conversion, metadata compatibility shim, adapter, and validation artifacts.