File size: 2,378 Bytes
22be876 0c1d24f 22be876 bcf4311 22be876 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 | ---
license: apache-2.0
base_model: jaredpalmer/kev-0.8b
library_name: onnxruntime
pipeline_tag: text-classification
tags:
- kev
- decision-model
- typesafe
- qwen3.5
- onnx
- webgpu
---
# kev-0.8b-ONNX
[kev-0.8b](https://huggingface.co/jaredpalmer/kev-0.8b) (revision `9a45d25eb2ab761841196625383fa1dff0e56c1e`) by Jared Palmer, packaged for the browser by
[runonweb](https://runonweb.ai/models/classify) (`runonweb/classify`).
Kev is a Jev-style decision model: typed questions about one text (`noul` yes/no, `choice`, `score`)
answered with calibrated probabilities, no text generation. Requests follow TypeSafe's System One API.
## What changed from the original
- The rank-16 LoRA is merged into `Qwen/Qwen3.5-0.8B-Base` (revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`) in fp32.
- The backbone is exported with the onnxruntime-genai builder without the LM head and MTP layer: the graph
returns `hidden_states` plus the recurrent/conv/KV cache, so the state is encoded once and each question
runs as its own row on that cache (Kev's row form).
- Weights are int4 (RTN, block 32) with the Gated DeltaNet projections and their MLPs in int8; activations fp32. ~735 MB. Needs the `LinearAttention` / `CausalConvWithState` contrib ops: ONNX Runtime Web's native WebGPU
build (`onnxruntime-web/webgpu`) on the GPU, `onnxruntime-web/wasm` on the CPU.
- The pointer head is `head.bin` (fp32: `q.weight`, `q.bias`, `k.weight`, `k.bias`); `kev.json` holds the
calibration temperature, delimiter token ids and cache layout.
On 318 questions from Kev's development suites, probabilities differ from the fp32 export by 0.028 on average; 16 answers change, none with a margin above 0.2 (accuracy 0.657 vs 0.664). The fp32 export of the same graph matches Kev's PyTorch fp32 path within 4e-5.
Recipe: `training/kev-onnx` in the runonweb repo.
## Use
```ts
import { Classifier } from 'runonweb/classify'
const classifier = new Classifier({ model: 'midudev/kev-0.8b-ONNX' })
const { answers } = await classifier.classify({
state: 'I was charged twice. Please fix this ASAP.',
questions: { billing: { type: 'noul', instructions: 'Is this ticket about billing?' } },
})
```
## License
Apache-2.0, like Kev and the Qwen3.5 base. Kev's training datasets have their own licenses; see the
[original model card](https://huggingface.co/jaredpalmer/kev-0.8b).
|