Image-Text-to-Text
Transformers
Safetensors
qwen3_5
clef
qwen3.8
quark
quantized
mxfp4
awq
amd
rocm
multimodal
structured-output
classification
custom-code
conversational
8-bit precision
Instructions to use EliovpAI/clef-MXFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use EliovpAI/clef-MXFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="EliovpAI/clef-MXFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("EliovpAI/clef-MXFP4") model = AutoModelForMultimodalLM.from_pretrained("EliovpAI/clef-MXFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use EliovpAI/clef-MXFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EliovpAI/clef-MXFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EliovpAI/clef-MXFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/EliovpAI/clef-MXFP4
- SGLang
How to use EliovpAI/clef-MXFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "EliovpAI/clef-MXFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EliovpAI/clef-MXFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "EliovpAI/clef-MXFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EliovpAI/clef-MXFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use EliovpAI/clef-MXFP4 with Docker Model Runner:
docker model run hf.co/EliovpAI/clef-MXFP4
|
Download README.md from EliovpAI/clef-MXFP4: direct link, hf CLI and curl.
- Browser
- Download file 6.8 kB
-
https://huggingface.co/EliovpAI/clef-MXFP4/resolve/main/README.md
- Command line
-
hf download hf://EliovpAI/clef-MXFP4/README.md
-
curl -L -o README.md https://huggingface.co/EliovpAI/clef-MXFP4/resolve/main/README.md
6.8 kB
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: image-text-to-text | |
| base_model: Cloudflare/clef | |
| base_model_relation: quantized | |
| datasets: | |
| - mit-han-lab/pile-val-backup | |
| tags: | |
| - clef | |
| - qwen3.8 | |
| - quark | |
| - quantized | |
| - mxfp4 | |
| - awq | |
| - amd | |
| - rocm | |
| - multimodal | |
| - structured-output | |
| - classification | |
| - custom-code | |
| # Clef MXFP4 | |
| This is an MXFP4 quantization of [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), a 27B multimodal decision | |
| model post-trained from Qwen3.8-27B. It was quantized with [AMD Quark](https://quark.docs.amd.com/) 0.13. | |
| The checkpoint is a standard Quark `hf_format` export. It loads in vLLM, and in Transformers with `amd-quark` | |
| installed. It is not tied to one GPU family: it runs natively on hardware with MXFP4 matrix units (for example | |
| MI350/MI355X) and emulated elsewhere. | |
| | | BF16 original | This repo | | |
| |---|---|---| | |
| | Backbone weights on disk | 54.7 GB | 18.9 GB | | |
| | Joint schema head | BF16 | BF16, byte-identical | | |
| | vLLM weight memory | 50.2 GiB | 17.0 GiB | | |
| ## Quantization | |
| | | | | |
| |---|---| | |
| | Format | OCP MXFP4: E2M1 elements with one E8M0 scale per 32 values along the input dimension | | |
| | Weights | MXFP4, static, `even` scale rounding | | |
| | Activations | MXFP4, dynamic per 32-value block | | |
| | Algorithm | AWQ (Quark's `qwen3_5` template) on the MLP projections | | |
| | Calibration | 128 samples × 512 tokens from `pileval` (`mit-han-lab/pile-val-backup`) | | |
| | Quantized | All language-model linear layers: full attention, Gated DeltaNet projections and MLP, in all 64 layers | | |
| | Kept in BF16 | Vision encoder and merger, `lm_head`, embeddings, norms, the `conv1d` layers, and the joint schema head | | |
| `lm_head` stays in BF16 deliberately. Clef's joint schema head reads the output-embedding matrix directly to embed answer | |
| options. | |
| The recipe is the same as AMD's own [amd/Qwen3.8-27B-Quark-AWQ-MXFP4](https://huggingface.co/amd/Qwen3.8-27B-Quark-AWQ-MXFP4) | |
| for the base model: | |
| ```bash | |
| python3 quantize_quark.py \ | |
| --model_dir Cloudflare/clef \ | |
| --output_dir ./clef-MXFP4 \ | |
| --quant_scheme mxfp4 \ | |
| --quant_algo awq \ | |
| --num_calib_data 128 \ | |
| --seq_len 512 \ | |
| --model_export hf_format \ | |
| --data_type auto \ | |
| --device cuda \ | |
| --skip_evaluation | |
| ``` | |
| `quantize_quark.py` is the LLM PTQ example from the [Quark v0.13 repository](https://github.com/amd/Quark/tree/v0.13/examples/torch/language_modeling/llm_ptq). | |
| After export, the backbone was resharded into five files with a safetensors index (bit-identical tensors), and the | |
| joint head files were copied unchanged. `algo_config` is set to `null` in `config.json` because AWQ is already folded | |
| into the weights and vLLM does not parse that field. The exported original is kept as `config.json.orig_with_algo_config`. | |
| ## Usage | |
| ### Clef decisions (Transformers) | |
| Use the original Clef code unchanged; only the repository name changes. Loading needs `amd-quark`: | |
| ```bash | |
| pip install amd-quark transformers pillow | |
| ``` | |
| ```python | |
| import sys | |
| import torch | |
| from huggingface_hub import snapshot_download | |
| path = snapshot_download("EliovpAI/clef-MXFP4") | |
| sys.path.insert(0, path) | |
| from joint_schema_model import load_release_model, systemone | |
| model, processor = load_release_model(path, device="cuda") | |
| response = systemone(model, processor, { | |
| "model": "clef", | |
| "state": "Our checkout started returning errors and orders are blocked.", | |
| "questions": { | |
| "department": { | |
| "type": "choice", | |
| "instructions": "Which team should handle the message?", | |
| "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}, | |
| }, | |
| "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]}, | |
| "outage": {"type": "noul", "instructions": "Is a service down?"}, | |
| }, | |
| }) | |
| print(response["answers"]) | |
| ``` | |
| See the [Clef model card](https://huggingface.co/Cloudflare/clef) for the input format, image and video inputs, and | |
| batching. In Transformers, Quark simulates the MXFP4 arithmetic. This is the reference path for accuracy, not for | |
| speed. | |
| ### Backbone serving (vLLM) | |
| vLLM loads the quantized backbone and runs native MXFP4 GEMMs where the hardware supports them. The Clef joint head is | |
| not part of vLLM; use the Transformers path above for decisions. | |
| ```bash | |
| vllm serve EliovpAI/clef-MXFP4 --max-model-len 16384 | |
| ``` | |
| ## Evaluation | |
| All runs were on one AMD Instinct MI355X (gfx950) with ROCm 7.2.3. BF16 is `Cloudflare/clef` at revision `2f3de3d`. | |
| Both models used the same code and inputs. | |
| ### Clef decisions: release code (Transformers 5.8.1, Quark 0.13) | |
| | Task | Records | BF16 accuracy | MXFP4 accuracy | Top-1 agreement | Mean KL (BF16 ‖ MXFP4) | | |
| |---|---|---|---|---|---| | |
| | BANKING77 test, 77-way choice | 300 | 94.0 % | 93.3 % | 98.0 % | 0.022 | | |
| | CLINC150 plus test, 151-way choice incl. out-of-scope | 300 | 96.7 % | 96.7 % | 98.7 % | 0.018 | | |
| | Receipt image, 2 × noul + 3-way choice | 3 questions | – | – | 100 % | 0.0008 | | |
| The records are a fixed random sample (seed 0). Each label is a choice option whose description is the | |
| humanized label name. These are our own prompts, not the Decision Index protocol, so the absolute numbers are not | |
| comparable to Cloudflare's published results. Use the BF16/MXFP4 difference. | |
| The SystemOne example from the Clef model card gives the same answers. `technical` is chosen at 0.918 | |
| (BF16 0.917), `urgency` peaks at "Today" with 0.920 (BF16 0.862), and `outage` is 0.841 (BF16 0.895). | |
| ### Backbone language modelling (vLLM 0.21.0) | |
| | | BF16 | MXFP4 | | |
| |---|---|---| | |
| | WikiText-2 perplexity (64 × 1,024 tokens) | 9.974 | 10.505 (+5.3 %) | | |
| | Greedy chat answers (4 prompts) | coherent | coherent, same content | | |
| ## Files | |
| | File | Purpose | | |
| |---|---| | |
| | `model-0000X-of-00005.safetensors`, `model.safetensors.index.json` | Quantized backbone (MXFP4 language model, BF16 vision encoder) | | |
| | `config.json` | Model and Quark quantization config | | |
| | `config.json.orig_with_algo_config` | Config as exported by Quark, including the AWQ settings | | |
| | `joint_head.safetensors`, `joint_head_config.json`, `joint_schema_model.py` | Clef joint schema head and code, unchanged from the original | | |
| | `tokenizer*`, `chat_template.jinja`, `processor_config.json`, `preprocessor_config.json`, `generation_config.json` | Tokenizer and processors | | |
| | `LICENSE` | Apache-2.0, from the original repository | | |
| | `SHA256SUMS` | Checksums of every file above | | |
| ## License and attribution | |
| Apache-2.0, the same as [Cloudflare/clef](https://huggingface.co/Cloudflare/clef), which is post-trained from | |
| [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). All credit for the model goes to Cloudflare and the Qwen | |
| team. This repository changes only the weight format of the backbone, as described above. It is not affiliated with or | |
| endorsed by Cloudflare, Qwen or AMD. | |