Instructions to use Vishva007/clef-flash-W4A16-AutoRound-GPTQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- vLLM
How to use Vishva007/clef-flash-W4A16-AutoRound-GPTQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Vishva007/clef-flash-W4A16-AutoRound-GPTQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vishva007/clef-flash-W4A16-AutoRound-GPTQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Vishva007/clef-flash-W4A16-AutoRound-GPTQ
- SGLang
How to use Vishva007/clef-flash-W4A16-AutoRound-GPTQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Vishva007/clef-flash-W4A16-AutoRound-GPTQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vishva007/clef-flash-W4A16-AutoRound-GPTQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Vishva007/clef-flash-W4A16-AutoRound-GPTQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vishva007/clef-flash-W4A16-AutoRound-GPTQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Vishva007/clef-flash-W4A16-AutoRound-GPTQ with Docker Model Runner:
docker model run hf.co/Vishva007/clef-flash-W4A16-AutoRound-GPTQ
Clef-Flash-W4A16-AutoRound-GPTQ
Quantized W4A16 (4-bit weights, 16-bit activations) release of Cloudflare/clef-flash using Intel AutoRound.
Clef-Flash is a 9B multimodal decision model post-trained from Qwen3.5 that evaluates structured schemas (text, JSON, image, video) and returns calibrated probability distributions across typed questions in a single forward pass without autoregressive token generation.
Quantization Details
- Method: Intel AutoRound W4A16 (
sym=True,group_size=32) - Vision Tower: Preserved in native BF16 (ensures zero OCR/visual degradation)
- Joint Schema Head (
joint_head.safetensors): Preserved in native BF16 - Linear Convolutions: Preserved in native BF16 to prevent drift in Qwen3.5 linear attention layers
- Accuracy: Zero loss observed across text, JSON state, and multimodal test suites compared to the base checkpoint.
Quickstart
Installation
pip install torch transformers huggingface_hub pillow
Usage (systemone API)
import sys
from huggingface_hub import snapshot_download
# Download repo and load custom joint schema model
repo_id = "Vishva007/clef-flash-W4A16-AutoRound" # or auto-gptq / llm-compressor variant
model_path = snapshot_download(repo_id)
sys.path.insert(0, model_path)
from joint_schema_model import load_release_model, systemone
model, processor = load_release_model(model_path, device="cuda")
# 1. Text / JSON Decision Example
response = systemone(model, processor, {
"model": "clef-flash",
"state": "Prod database latency spiked to 4,000ms. Checkout failing with 504 Gateway Timeouts.",
"questions": {
"severity": {
"type": "choice",
"instructions": "Determine incident severity level",
"criteria": {
"SEV_1": "Critical revenue outage",
"SEV_2": "Major feature degradation",
"SEV_3": "Minor issue"
}
},
"urgency": {
"type": "score",
"instructions": "Urgency rating",
"criteria": ["Low", "Medium", "Immediate page"]
},
"rollback": {
"type": "noul",
"instructions": "Should a rollback be initiated?"
}
}
})
print(response["answers"])
Multimodal (Image) Evaluation
from PIL import Image
response = systemone(model, processor, {
"model": "clef-flash",
"state": "Review the uploaded invoice receipt.",
"images": [Image.open("receipt.png")],
"questions": {
"legible": {"type": "noul", "instructions": "Is the receipt text clear and legible?"},
"amount_exceeds_1000": {"type": "noul", "instructions": "Is total > $1000 USD?"}
}
})
print(response["answers"])
Formats Available
| Repository | Format | Engine Target |
|---|---|---|
Vishva007/clef-flash-W4A16-AutoRound |
AutoRound / AutoGPTQ | Transformers / Native Python |
Vishva007/clef-flash-W4A16-AutoRound-GPTQ |
AutoGPTQ Standard | Transformers / ExLlama / AutoGPTQ |
Vishva007/clef-flash-W4A16-AutoRound-LLM-Compressor |
Compressed-Tensors | vLLM / SGLang |
🚀 Deploy on RunPod
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.
🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
PyTorch 2.14
PyTorch 2.13
PyTorch 2.12
Acknowledgements
- Base model by Cloudflare.
- Quantization powered by Intel AutoRound.
- Downloads last month
- 47