File size: 11,585 Bytes
e1f719a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 | ---
base_model: microsoft/Phi-4-multimodal-instruct-onnx
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: onnx
tags:
- onnx
- onnxruntime-genai
- int8
- gptq
- weight-only-quantization
- vision-language
- speech
- phi4mm
license: mit
license_link: https://huggingface.co/microsoft/Phi-4-multimodal-instruct/resolve/main/LICENSE
---
# Phi-4-multimodal-instruct β ONNX INT8 decoder + GPTQ INT4 vision
An **INT8 language decoder** paired with a **GPTQ-calibrated INT4 vision tower**, rebuilt from
[`microsoft/Phi-4-multimodal-instruct-onnx`](https://huggingface.co/microsoft/Phi-4-multimodal-instruct-onnx)
for [`onnxruntime-genai`](https://github.com/microsoft/onnxruntime-genai) on the **CUDA** execution
provider. 5.6 B parameters: a 3.8 B Phi-4-mini language backbone, a ~0.45 B SigLIP vision encoder,
a ~0.46 B conformer speech encoder, and the modality LoRA adapters.
**This is the byte-exact build.** Its sibling
[`jarvis-pet/Phi-4-multimodal-instruct-gptq-int4-onnx`](https://huggingface.co/jarvis-pet/Phi-4-multimodal-instruct-gptq-int4-onnx)
is 1.99 GiB smaller and buys a bigger context window plus both modalities in one process on a 12 GB
card, but it drops the occasional character in short strings. This one does not: pick it when
byte-exact short-field extraction matters more than VRAM headroom.
**Why it is not simply the published artifact.** Microsoft's GPU release quantizes both towers with
**INT4 RTN**, and both are damaged by it. On the vision tower, `image_features` measured against
FP16 have a **relative L2 error of 0.94β1.09** β error equal to signal β so OCR either degenerates
into repetition or confabulates. The INT4 decoder separately breaks EOS, so page transcription never
terminates. `int4` was the whole problem, in both components; this build replaces the tower with
data-aware GPTQ and the decoder with 8-bit.
**The name states the DECODER's precision**, matching the convention already used for this account's
OpenVINO uploads (where `-int4-ov` likewise names the language model while other submodels sit at
INT8). The vision tower here is GPTQ INT4, as the table below records.
Released under the **MIT License**, inherited from the base model
[`microsoft/Phi-4-multimodal-instruct`](https://huggingface.co/microsoft/Phi-4-multimodal-instruct).
This is a derivative quantization only; all model rights and terms follow the base model.
## Files
19 files, **6.78 GiB** total:
| Component | Precision | Notes |
|---|---|---|
| `phi-4-mm-text.onnx{,.data}` | **INT8** | the decoder, **4.55 GB** β 673 `MatMulNBits`, `bits: 8` |
| `phi-4-mm-vision.onnx{,.data}` | **INT4 GPTQ**, block 128 | vision tower, 370 MB β 80 `MatMulNBits` |
| `phi-4-mm-speech.onnx{,.data}` | as published | conformer speech encoder, 520 MB |
| `phi-4-mm-embedding.onnx{,.data}` | as published | shared embedding graph, 1.23 GB |
| `phi-4-mm-vision.onnx_adapter`, `phi-4-mm-speech.onnx_adapter` | as published | **present but not referenced** β see below |
| `genai_config.json`, `vision_processor.json`, `speech_processor.json` | β | `genai_config.json` differs from the published one |
| `tokenizer.json`, `tokenizer_config.json`, `vocab.json`, `merges.txt`, `added_tokens.json`, `special_tokens_map.json` | β | unchanged |
**The adapter files are vestigial here, deliberately.** The locally built decoder bakes the LoRAs in
(896 `lora_*` initializers, 0 adapter inputs β `Phi4MMModel.make_layer` makes **vision** the baked
default), so `adapter_filename` has been **removed** from the `vision` and `speech` sections of
`genai_config.json`. Leaving it in makes the runtime bind tensors the graph does not have. Losing the
*speech* LoRA was expected to cost audio quality β measured, it does not (see below).
## Audio input β 16 kHz required, 8 kHz optional
The speech encoder consumes **80-bin log-mel features at a 25 ms window and a 10 ms hop**.
`speech_processor.json` declares `AudioDecoderEx.target_sample_rates: [8000, 16000]`, and
`Phi4AudioEmbed` carries a complete STFT/log-mel parameter set for **each** rate:
| | **16 kHz β required / primary** | **8 kHz β optional** |
|---|---|---|
| `n_fft` | 512 | 256 |
| `frame_length` | 400 β 25 ms | 200 β 25 ms |
| `hop_length` | 160 β 10 ms | 80 β 10 ms |
| `n_mel` | 80 | 80 |
Both paths emit the same feature geometry β 80 mel bins at 100 frames/second β which is what the
encoder is built for; `audio_compression_rate: 8` then decimates that 8Γ along time. What differs is
bandwidth, not shape.
- **Supply 16 kHz mono PCM16 WAV.** This is the primary path and what every audio measurement below
used.
- **8 kHz is natively supported**, not merely tolerated β a first-class second parameter set intended
for telephony-band input. Nyquist is 4 kHz there, so the 4β8 kHz band is absent from the features;
expect that to cost fricative and sibilant detail.
- **Any other rate is resampled by `AudioDecoderEx`**, not rejected. That path is untested here, so
do not rely on its fidelity β resample to 16 kHz yourself if you care about the transcript. This
differs from an OpenVINO serve of the same model, where nothing resamples and a wrong-rate waveform
transcribes confidently wrong.
## Measured results
Measured on an **RTX 3060 12 GB** (Windows 11, driver 610.88, no CUDA toolkit β pip wheels only),
`onnxruntime-genai-cuda 0.15.2` + `onnxruntime-gpu 1.29.0`, at **`max_length = 4096`**:
| Probe | Result |
|---|---|
| `Ocr.Tests.ps1` | **PASS=4 FAIL=0** β `"JARVIS OCR OCRPASS-4713"` read **byte-perfect** in 1.3 s |
| `OrtGenai.Tests.ps1` | **36/36** |
| Dense 850Γ1100 invoice, full transcription | `finish_reason=stop`, 186 tokens, 7.8 s, longest identical-line repeat **1**, **6/6** gold fields including `Subtotal 2,355.21` and routing `021000021` |
| Short-answer extraction | `Total Due` β `2,549.51`, 6 tokens, 1.3 s |
| Audio, 10 s 16 kHz clip | **byte-perfect** transcription *including* `"at sunrise"` (which the published INT4 build got wrong), plus content Q&A and correct French translation |
| VRAM | 8,849 MiB idle β **11,683 MiB after a page** |
Against the published INT4-RTN artifact on the same invoice: `finish_reason` `length` β **`stop`**,
identical-line repeats 10 β **1**, gold fields 5/6 β **6/6**, `Subtotal` wrong β **exact**, routing
number MISS β **HIT**, and severe character doubling β **none**.
### Cost: 12 GB serves one modality per process
Weights are 7.4 GB, and ORT's BFC arena grows and never releases, so:
- **`max_length = 8192` does not fit** a 12 GB card at all. 4096 works, for OCR-only or audio-only.
- **An audio request after a page request dies** with
`AllocateRawInternal Failed to allocate memory for requested buffer of size 1338028032` β the
vision activation buffer β with ~1.4 GB nominally free. A serve that must do both in one process
needs a bigger card.
Both limits are what the INT4-decoder sibling exists to lift: at 4.77 GB of weights it holds 8192
*and* takes audio after a page in the same process, on the same 12 GB card.
## How the quantization was produced
- **Decoder β INT8.** Upstream's `build_text()` hardcodes `precision = "int4"` and ignores
`--precision`, so `create_model` was called directly with `precision="int8"` and
`exclude_embeds=true`, `filename=phi-4-mm-text.onnx` to match what `genai_config.json` names.
- **Vision tower β GPTQ INT4, block 128.** `build_vision()` runs the nbits quantizer
*unconditionally* at INT4 RTN, so the FP16 tower was exported with that step patched out and
`torch.onnx.export(..., dynamo=False)` (torch β₯ 2.9's dynamo exporter fails to capture the SigLIP
tower and silently rewrites opset 14 β 18), then quantized with
`MatMulNBitsQuantizer(bits=4, block_size=128, is_symmetric=True)` under
`GPTQWeightOnlyQuantConfig(percdamp=0.1, actorder=False, mse=False, perchannel=True)`.
- **GPTQ is Python-API-only:** `--quant_method gptq` from the CLI builds the config with no
`calibration_data_reader`, so INC iterates `None` and dies with
`TypeError: 'NoneType' object is not iterable`.
Fidelity of the vision tower against FP16, five page fixtures, CUDA EP:
| tower | size | cosine | rel L2 |
|---|---|---|---|
| INT8 block 32 | 536 MB | 0.99990β0.99992 | 0.012 |
| **INT4 GPTQ block 128 (this build)** | **370 MB** | **0.950β0.961** | 0.29β0.33 |
| INT4 GPTQ block 32 | 383 MB | 0.926β0.944 | 0.33β0.38 |
| INT4 RTN block 32 (published) | 383 MB | **0.414β0.594** | **0.94β1.09** |
Cosine 0.95 passes every end-to-end probe and scored **identically** to an INT8 tower while being
166 MB smaller, which is why the INT4 GPTQ tower is used here rather than an INT8 one. Note finer
groups made GPTQ *worse*, the opposite of the usual expectation.
### Gotchas worth knowing
- **`device_id` must be 0.** ORT's CUDA EP creates its cudnn handle before `cudaSetDevice`, so any
non-zero `device_id` dies at model load with
`CUDNN failure 2007: CUDNN_STATUS_BAD_PARAM_STREAM_MISMATCH`. To serve on a different physical
card, mask it with `CUDA_VISIBLE_DEVICES=<n>` (plus `CUDA_DEVICE_ORDER=PCI_BUS_ID`) and keep
`device_id=0`.
- **`max_length` reserves the whole KV cache up front** β `genai_config.json` sets
`past_present_share_buffer: true`. The artifact declares `context_length: 131072`, which on this
geometry (32 layers Γ 8 KV heads Γ 128 head size Γ 2 Γ 2 B) would ask ~17 GB.
- **ORT's arena never releases.** Peak VRAM is set by the widest activation the process has *ever*
allocated, not by the current request β which is why modality order matters here.
- **`nvidia_awq` cannot quantize either half of this model.** `NVAWQWeightOnlyQuantConfig`
synthesises `input_ids`/`attention_mask`/`position_ids`/past-KV, but the vision tower takes
`pixel_values`/`image_attention_mask`/`image_sizes` and the decoder β built `exclude_embeds=true` β
takes `inputs_embeds` with no `input_ids`. Structural, not a tuning problem.
## Usage
```python
import onnxruntime as ort
import onnxruntime_genai as og
# The pip CUDA wheels place their DLLs outside the process search path.
ort.preload_dlls()
cfg = og.Config("Phi-4-multimodal-instruct-int8-onnx")
cfg.clear_providers()
cfg.append_provider("cuda")
cfg.set_provider_option("cuda", "device_id", "0") # must be 0 β see gotchas
model = og.Model(cfg)
tokenizer = og.Tokenizer(model)
params = og.GeneratorParams(model)
params.set_search_options(max_length=4096, do_sample=False) # 8192 does not fit 12 GB here
gen = og.Generator(model, params)
gen.append_tokens(tokenizer.encode(
"<|user|><|image_1|>\nTranscribe all text in this image exactly.<|end|><|assistant|>"))
```
Images and audio are passed as `og.Images.open_bytes(...)` / `og.Audios.open_bytes(...)` to
`model.create_multimodal_processor()`, and the prompt must carry the inline markers this
architecture expects β `<|image_1|>`, `<|audio_1|>`, numbered from 1 independently per modality.
**Use one modality per process on a 12 GB card**, per the arena limit above.
## Acknowledgements
Base model Β© Microsoft β
[`microsoft/Phi-4-multimodal-instruct`](https://huggingface.co/microsoft/Phi-4-multimodal-instruct)
and its ONNX release
[`microsoft/Phi-4-multimodal-instruct-onnx`](https://huggingface.co/microsoft/Phi-4-multimodal-instruct-onnx),
whose embedding and speech graphs, adapters and tokenizer ship here unchanged. Quantization tooling:
ONNX Runtime, its bundled Intel Neural Compressor GPTQ implementation, and onnxruntime-genai's model
builder.
|