--- license: apache-2.0 library_name: transformers pipeline_tag: image-text-to-text base_model: Cloudflare/clef-flash base_model_relation: quantized tags: - clef - cloudflare - systemone - qwen3.5 - post-train - image-text-to-typed-output - multimodal - structured-output - classification - custom-code - bitsandbytes - 4-bit - nf4 --- # Clef-Flash (4-bit NF4 Quantized) This repository contains the **4-bit NF4 quantized** version of Cloudflare's **Clef-Flash** multimodal decision model. - **Base Model:** [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) - **Quantization:** 4-bit NormalFloat (NF4) with double quantization via `bitsandbytes` - **Compute Dtype:** `bfloat16` - **Backbone:** Qwen3.5-9B - **Joint Schema Head:** Unquantized BF16 precision for accurate scoring and routing - **Format:** Safetensors ## Quickstart / Usage ```python import sys import torch from huggingface_hub import snapshot_download path = snapshot_download("meossistant/clef-flash-4bit") sys.path.insert(0, path) from joint_schema_model import load_release_model, systemone model, processor = load_release_model(path, device="cuda") response = systemone(model, processor, { "model": "clef-flash", "state": "Our checkout started returning errors and orders are blocked.", "questions": { "department": { "type": "choice", "instructions": "Which team should handle the message?", "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}, }, "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]}, "outage": {"type": "noul", "instructions": "Is a service down?"}, }, }) print(response["answers"]) ``` ## Overview Clef-Flash is a 9B multimodal model that turns a state and a schema of typed questions into decisions in a single forward pass. This 4-bit quantized version reduces the VRAM requirement to ~6 GB, making it easily runnable on consumer GPUs.