Qwen3-4B-Instruct-2507 โ ONNX
ONNX export of Qwen/Qwen3-4B-Instruct-2507, published by Liodon AI.
Exported with optimum (optimum.exporters.onnx.main_export,
task text-generation-with-past, so the graph exposes past-key-value inputs/outputs for KV-cached
autoregressive decoding).
Files
| File | Size | Notes |
|---|---|---|
model.onnx |
17.65 GB | FP32, full precision |
model_fp16.onnx |
9.58 GB | FP16, for GPU execution providers |
Quick Start
import onnxruntime as ort
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("liodon-ai/Qwen3-4B-Instruct-2507-ONNX")
sess = ort.InferenceSession("model_quantized.onnx", providers=["CPUExecutionProvider"])
# past_key_values.*.key / .value inputs must be supplied (zero-length tensors
# for the first forward pass) -- see optimum's ORTModelForCausalLM for a
# ready-made wrapper that handles KV-cache bookkeeping automatically:
# from optimum.onnxruntime import ORTModelForCausalLM
# model = ORTModelForCausalLM.from_pretrained("liodon-ai/Qwen3-4B-Instruct-2507-ONNX", file_name="model_quantized.onnx")
Source
- Model: Qwen/Qwen3-4B-Instruct-2507
- License: other
Citation
@misc{liodonai_qwen3_4b_instruct_2507_onnx,
title = {Qwen3-4B-Instruct-2507 โ ONNX},
author = {{Liodon AI}},
year = {2026},
howpublished = {\url{https://huggingface.co/liodon-ai/Qwen3-4B-Instruct-2507-ONNX}},
note = {ONNX export of Qwen/Qwen3-4B-Instruct-2507}
}
Exported by Liodon AI
- Downloads last month
- 202
Model tree for liodon-ai/Qwen3-4B-Instruct-2507-ONNX
Base model
Qwen/Qwen3-4B-Instruct-2507