--- license: apache-2.0 library_name: transformers pipeline_tag: image-text-to-text base_model: Cloudflare/clef base_model_relation: quantized datasets: - mit-han-lab/pile-val-backup tags: - clef - qwen3.8 - quark - quantized - mxfp4 - awq - amd - rocm - multimodal - structured-output - classification - custom-code --- # Clef MXFP4 This is an MXFP4 quantization of [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), a 27B multimodal decision model post-trained from Qwen3.8-27B. It was quantized with [AMD Quark](https://quark.docs.amd.com/) 0.13. The checkpoint is a standard Quark `hf_format` export. It loads in vLLM, and in Transformers with `amd-quark` installed. It is not tied to one GPU family: it runs natively on hardware with MXFP4 matrix units (for example MI350/MI355X) and emulated elsewhere. | | BF16 original | This repo | |---|---|---| | Backbone weights on disk | 54.7 GB | 18.9 GB | | Joint schema head | BF16 | BF16, byte-identical | | vLLM weight memory | 50.2 GiB | 17.0 GiB | ## Quantization | | | |---|---| | Format | OCP MXFP4: E2M1 elements with one E8M0 scale per 32 values along the input dimension | | Weights | MXFP4, static, `even` scale rounding | | Activations | MXFP4, dynamic per 32-value block | | Algorithm | AWQ (Quark's `qwen3_5` template) on the MLP projections | | Calibration | 128 samples × 512 tokens from `pileval` (`mit-han-lab/pile-val-backup`) | | Quantized | All language-model linear layers: full attention, Gated DeltaNet projections and MLP, in all 64 layers | | Kept in BF16 | Vision encoder and merger, `lm_head`, embeddings, norms, the `conv1d` layers, and the joint schema head | `lm_head` stays in BF16 deliberately. Clef's joint schema head reads the output-embedding matrix directly to embed answer options. The recipe is the same as AMD's own [amd/Qwen3.8-27B-Quark-AWQ-MXFP4](https://huggingface.co/amd/Qwen3.8-27B-Quark-AWQ-MXFP4) for the base model: ```bash python3 quantize_quark.py \ --model_dir Cloudflare/clef \ --output_dir ./clef-MXFP4 \ --quant_scheme mxfp4 \ --quant_algo awq \ --num_calib_data 128 \ --seq_len 512 \ --model_export hf_format \ --data_type auto \ --device cuda \ --skip_evaluation ``` `quantize_quark.py` is the LLM PTQ example from the [Quark v0.13 repository](https://github.com/amd/Quark/tree/v0.13/examples/torch/language_modeling/llm_ptq). After export, the backbone was resharded into five files with a safetensors index (bit-identical tensors), and the joint head files were copied unchanged. `algo_config` is set to `null` in `config.json` because AWQ is already folded into the weights and vLLM does not parse that field. The exported original is kept as `config.json.orig_with_algo_config`. ## Usage ### Clef decisions (Transformers) Use the original Clef code unchanged; only the repository name changes. Loading needs `amd-quark`: ```bash pip install amd-quark transformers pillow ``` ```python import sys import torch from huggingface_hub import snapshot_download path = snapshot_download("EliovpAI/clef-MXFP4") sys.path.insert(0, path) from joint_schema_model import load_release_model, systemone model, processor = load_release_model(path, device="cuda") response = systemone(model, processor, { "model": "clef", "state": "Our checkout started returning errors and orders are blocked.", "questions": { "department": { "type": "choice", "instructions": "Which team should handle the message?", "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}, }, "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]}, "outage": {"type": "noul", "instructions": "Is a service down?"}, }, }) print(response["answers"]) ``` See the [Clef model card](https://huggingface.co/Cloudflare/clef) for the input format, image and video inputs, and batching. In Transformers, Quark simulates the MXFP4 arithmetic. This is the reference path for accuracy, not for speed. ### Backbone serving (vLLM) vLLM loads the quantized backbone and runs native MXFP4 GEMMs where the hardware supports them. The Clef joint head is not part of vLLM; use the Transformers path above for decisions. ```bash vllm serve EliovpAI/clef-MXFP4 --max-model-len 16384 ``` ## Evaluation All runs were on one AMD Instinct MI355X (gfx950) with ROCm 7.2.3. BF16 is `Cloudflare/clef` at revision `2f3de3d`. Both models used the same code and inputs. ### Clef decisions: release code (Transformers 5.8.1, Quark 0.13) | Task | Records | BF16 accuracy | MXFP4 accuracy | Top-1 agreement | Mean KL (BF16 ‖ MXFP4) | |---|---|---|---|---|---| | BANKING77 test, 77-way choice | 300 | 94.0 % | 93.3 % | 98.0 % | 0.022 | | CLINC150 plus test, 151-way choice incl. out-of-scope | 300 | 96.7 % | 96.7 % | 98.7 % | 0.018 | | Receipt image, 2 × noul + 3-way choice | 3 questions | – | – | 100 % | 0.0008 | The records are a fixed random sample (seed 0). Each label is a choice option whose description is the humanized label name. These are our own prompts, not the Decision Index protocol, so the absolute numbers are not comparable to Cloudflare's published results. Use the BF16/MXFP4 difference. The SystemOne example from the Clef model card gives the same answers. `technical` is chosen at 0.918 (BF16 0.917), `urgency` peaks at "Today" with 0.920 (BF16 0.862), and `outage` is 0.841 (BF16 0.895). ### Backbone language modelling (vLLM 0.21.0) | | BF16 | MXFP4 | |---|---|---| | WikiText-2 perplexity (64 × 1,024 tokens) | 9.974 | 10.505 (+5.3 %) | | Greedy chat answers (4 prompts) | coherent | coherent, same content | ## Files | File | Purpose | |---|---| | `model-0000X-of-00005.safetensors`, `model.safetensors.index.json` | Quantized backbone (MXFP4 language model, BF16 vision encoder) | | `config.json` | Model and Quark quantization config | | `config.json.orig_with_algo_config` | Config as exported by Quark, including the AWQ settings | | `joint_head.safetensors`, `joint_head_config.json`, `joint_schema_model.py` | Clef joint schema head and code, unchanged from the original | | `tokenizer*`, `chat_template.jinja`, `processor_config.json`, `preprocessor_config.json`, `generation_config.json` | Tokenizer and processors | | `LICENSE` | Apache-2.0, from the original repository | | `SHA256SUMS` | Checksums of every file above | ## License and attribution Apache-2.0, the same as [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), which is post-trained from [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). All credit for the model goes to Cloudflare and the Qwen team. This repository changes only the weight format of the backbone, as described above. It is not affiliated with or endorsed by Cloudflare, Qwen or AMD.