midudev commited on
Commit
0c1d24f
·
verified ·
1 Parent(s): 22be876

Card: quantization and parity notes

Browse files
Files changed (1) hide show
  1. README.md +3 -5
README.md CHANGED
@@ -26,14 +26,12 @@ answered with calibrated probabilities, no text generation. Requests follow Type
26
  - The backbone is exported with the onnxruntime-genai builder without the LM head and MTP layer: the graph
27
  returns `hidden_states` plus the recurrent/conv/KV cache, so the state is encoded once and each question
28
  runs as its own row on that cache (Kev's row form).
29
- - Weights are int4 (RTN, block 32) with the Gated DeltaNet projections and their MLPs in int8; activations
30
- fp32. Needs `LinearAttention` / `CausalConvWithState` and 8-bit `MatMulNBits`: ONNX Runtime Web's native
31
- WebGPU build (`onnxruntime-web/webgpu`) or its WASM backend.
32
  - The pointer head is `head.bin` (fp32: `q.weight`, `q.bias`, `k.weight`, `k.bias`); `kev.json` holds the
33
  calibration temperature, delimiter token ids and cache layout.
34
 
35
- On 318 questions from Kev's development suites, probabilities differ from Kev's PyTorch fp32 path by 0.03 on
36
- average (the fp32 export matches to 2e-5); 16 answers change, none with a margin above 0.2.
37
  Recipe: `training/kev-onnx` in the runonweb repo.
38
 
39
  ## Use
 
26
  - The backbone is exported with the onnxruntime-genai builder without the LM head and MTP layer: the graph
27
  returns `hidden_states` plus the recurrent/conv/KV cache, so the state is encoded once and each question
28
  runs as its own row on that cache (Kev's row form).
29
+ - Weights are int4 (RTN, block 32) with the Gated DeltaNet projections and their MLPs in int8; activations fp32. ~735 MB. Needs the `LinearAttention` / `CausalConvWithState` contrib ops: ONNX Runtime Web's native WebGPU
30
+ build (`onnxruntime-web/webgpu`) on the GPU, `onnxruntime-web/wasm` on the CPU.
 
31
  - The pointer head is `head.bin` (fp32: `q.weight`, `q.bias`, `k.weight`, `k.bias`); `kev.json` holds the
32
  calibration temperature, delimiter token ids and cache layout.
33
 
34
+ On 318 questions from Kev's development suites, probabilities differ from the fp32 export by 0.028 on average; 16 answers change, none with a margin above 0.2 (accuracy 0.657 vs 0.664). The fp32 export of the same graph matches Kev's PyTorch fp32 path to 2e-5.
 
35
  Recipe: `training/kev-onnx` in the runonweb repo.
36
 
37
  ## Use