File size: 2,289 Bytes
22be876 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 | ---
license: apache-2.0
base_model: jaredpalmer/kev-0.8b
library_name: onnxruntime
pipeline_tag: text-classification
tags:
- kev
- decision-model
- typesafe
- qwen3.5
- onnx
- webgpu
---
# kev-0.8b-ONNX
[kev-0.8b](https://huggingface.co/jaredpalmer/kev-0.8b) (revision `9a45d25eb2ab761841196625383fa1dff0e56c1e`) by Jared Palmer, packaged for the browser by
[runonweb](https://runonweb.ai/models/classify) (`runonweb/classify`).
Kev is a Jev-style decision model: typed questions about one text (`noul` yes/no, `choice`, `score`)
answered with calibrated probabilities, no text generation. Requests follow TypeSafe's System One API.
## What changed from the original
- The rank-16 LoRA is merged into `Qwen/Qwen3.5-0.8B-Base` (revision `dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68`) in fp32.
- The backbone is exported with the onnxruntime-genai builder without the LM head and MTP layer: the graph
returns `hidden_states` plus the recurrent/conv/KV cache, so the state is encoded once and each question
runs as its own row on that cache (Kev's row form).
- Weights are int4 (RTN, block 32) with the Gated DeltaNet projections and their MLPs in int8; activations
fp32. Needs `LinearAttention` / `CausalConvWithState` and 8-bit `MatMulNBits`: ONNX Runtime Web's native
WebGPU build (`onnxruntime-web/webgpu`) or its WASM backend.
- The pointer head is `head.bin` (fp32: `q.weight`, `q.bias`, `k.weight`, `k.bias`); `kev.json` holds the
calibration temperature, delimiter token ids and cache layout.
On 318 questions from Kev's development suites, probabilities differ from Kev's PyTorch fp32 path by 0.03 on
average (the fp32 export matches to 2e-5); 16 answers change, none with a margin above 0.2.
Recipe: `training/kev-onnx` in the runonweb repo.
## Use
```ts
import { Classifier } from 'runonweb/classify'
const classifier = new Classifier({ model: 'midudev/kev-0.8b-ONNX' })
const { answers } = await classifier.classify({
state: 'I was charged twice. Please fix this ASAP.',
questions: { billing: { type: 'noul', instructions: 'Is this ticket about billing?' } },
})
```
## License
Apache-2.0, like Kev and the Qwen3.5 base. Kev's training datasets have their own licenses; see the
[original model card](https://huggingface.co/jaredpalmer/kev-0.8b).
|