Card: quantization and parity notes
Browse files
README.md
CHANGED
|
@@ -26,14 +26,12 @@ answered with calibrated probabilities, no text generation. Requests follow Type
|
|
| 26 |
- The backbone is exported with the onnxruntime-genai builder without the LM head and MTP layer: the graph
|
| 27 |
returns `hidden_states` plus the recurrent/conv/KV cache, so the state is encoded once and each question
|
| 28 |
runs as its own row on that cache (Kev's row form).
|
| 29 |
-
- Weights are int4 (RTN, block 32) with the Gated DeltaNet projections and their MLPs in int8; activations
|
| 30 |
-
|
| 31 |
-
WebGPU build (`onnxruntime-web/webgpu`) or its WASM backend.
|
| 32 |
- The pointer head is `head.bin` (fp32: `q.weight`, `q.bias`, `k.weight`, `k.bias`); `kev.json` holds the
|
| 33 |
calibration temperature, delimiter token ids and cache layout.
|
| 34 |
|
| 35 |
-
On 318 questions from Kev's development suites, probabilities differ from Kev's PyTorch fp32 path
|
| 36 |
-
average (the fp32 export matches to 2e-5); 16 answers change, none with a margin above 0.2.
|
| 37 |
Recipe: `training/kev-onnx` in the runonweb repo.
|
| 38 |
|
| 39 |
## Use
|
|
|
|
| 26 |
- The backbone is exported with the onnxruntime-genai builder without the LM head and MTP layer: the graph
|
| 27 |
returns `hidden_states` plus the recurrent/conv/KV cache, so the state is encoded once and each question
|
| 28 |
runs as its own row on that cache (Kev's row form).
|
| 29 |
+
- Weights are int4 (RTN, block 32) with the Gated DeltaNet projections and their MLPs in int8; activations fp32. ~735 MB. Needs the `LinearAttention` / `CausalConvWithState` contrib ops: ONNX Runtime Web's native WebGPU
|
| 30 |
+
build (`onnxruntime-web/webgpu`) on the GPU, `onnxruntime-web/wasm` on the CPU.
|
|
|
|
| 31 |
- The pointer head is `head.bin` (fp32: `q.weight`, `q.bias`, `k.weight`, `k.bias`); `kev.json` holds the
|
| 32 |
calibration temperature, delimiter token ids and cache layout.
|
| 33 |
|
| 34 |
+
On 318 questions from Kev's development suites, probabilities differ from the fp32 export by 0.028 on average; 16 answers change, none with a margin above 0.2 (accuracy 0.657 vs 0.664). The fp32 export of the same graph matches Kev's PyTorch fp32 path to 2e-5.
|
|
|
|
| 35 |
Recipe: `training/kev-onnx` in the runonweb repo.
|
| 36 |
|
| 37 |
## Use
|