Text Classification
MLX
Safetensors
qwen3_5
clef
decision-model
systemone
structured-output
classification
qwen3.5
4-bit precision
Instructions to use TrevorJS/clef-flash-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TrevorJS/clef-flash-mlx-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download TrevorJS/clef-flash-mlx-4bit --local-dir clef-flash-mlx-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
File size: 4,321 Bytes
6d4dc3f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 | ---
license: apache-2.0
library_name: mlx
base_model: Cloudflare/clef-flash
base_model_relation: quantized
pipeline_tag: text-classification
tags:
- mlx
- clef
- decision-model
- systemone
- structured-output
- classification
- qwen3.5
---
# clef-flash-mlx-4bit
[Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) converted to [MLX](https://github.com/ml-explore/mlx) for Apple Silicon,
with the backbone quantized to 4-bit. Clef-Flash is a decision model: given a state and a schema of typed questions
(`choice`, `score`, `noul`), it returns a probability for every allowed option of every question from one prefill pass,
with no text generation.
This is an unofficial conversion. It is not made or endorsed by Cloudflare. All credit for the model goes to the
Clef authors; see the [announcement](https://blog.cloudflare.com/clef-decision-models/).
## What is in this repo
| File | Contents |
|---|---|
| `model*.safetensors`, `config.json` | Qwen3.5-9B backbone from Clef-Flash, converted with `mlx_lm.convert` (affine, 4-bit, group size 64). Vision encoder dropped. |
| `joint_head.safetensors`, `joint_head_config.json` | Clef's joint schema head, unchanged from the original release (bf16) |
| `clef_mlx.py` | MLX port of the release's `joint_schema_model.py` (record encoding and joint schema head); the head runs in float32 |
| tokenizer, chat template, `LICENSE` | From the original release |
**Text only.** The vision encoder is not included, so image and video inputs are not supported.
## Usage
```bash
pip install mlx-lm huggingface_hub
```
```python
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("TrevorJS/clef-flash-mlx-4bit")
sys.path.insert(0, path)
from clef_mlx import load, decide
clef = load(path)
print(decide(clef, {
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
}))
# {'department': {'billing': 0.043, 'technical': 0.957}, 'urgency': {'0': 0.096, '1': 0.072, '2': 0.832},
# 'outage': {'true': 0.818, 'false': 0.182}} (4-bit output)
```
`decide` takes the same record shape as the original `encode_record` / `systemone` (state as text or JSON, questions keyed
by ID) and returns `{question_id: {option_id: probability}}`. All questions in a record are scored jointly.
## Verification
- **Head:** on identical inputs, `clef_mlx.JointSchemaHead` matches the release's torch `JointSchemaHead` to within 4e-6 on
the logits.
- **Encoding:** `clef_mlx.encode_record` produces the same token IDs, spans and option IDs as the release's
`encode_record` on 11 of 11 sampled records.
- **Backbone:** quantization is the only source of drift. No bf16 reference was run, so the end-to-end check is task
accuracy and the agreement between the two quantizations (below).
**JevBench public items** (231 items, pinned commit `bb05a335`), scored by this repo's code on an Apple M2 (24 GB):
| Variant | All | Easy | Standard | Hard | Peak memory | Median latency (M2) |
|---|---|---|---|---|---|---|
| 8-bit | 188/231 | 48/48 | 71/72 | 69/111 | 10.1 GB | 4.8 s |
| 4-bit | 182/231 | 48/48 | 68/72 | 66/111 | 5.9 GB | 2.9 s |
The two variants pick the same option on 216 of 231 items (median max per-option probability difference 0.011).
Latency is for one M2 and reflects that machine, not the model on a GPU server.
Other variant: [TrevorJS/clef-flash-mlx-8bit](https://huggingface.co/TrevorJS/clef-flash-mlx-8bit).
## Conversion
`mlx-lm 0.31.3`, `mlx 0.32.3`:
```bash
python -m mlx_lm convert --hf-path Cloudflare/clef-flash --mlx-path clef-flash-mlx-4bit -q --q-bits 4 --q-group-size 64
```
then `joint_head.safetensors`, `joint_head_config.json` and `LICENSE` copied from the original repo.
## License
Apache-2.0, following [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) and its base model
Qwen/Qwen3.5-9B. `clef_mlx.py` is a port of Cloudflare's Apache-2.0 `joint_schema_model.py`.
|