kev-9b-merged: unofficial merged derivative of Kev
This is an unofficial derivative of jaredpalmer/kev-9b by Jared Palmer, prepared by Avartha. It is not endorsed by the Kev author. Kev and its weights are Apache-2.0, as are the Qwen3.5 (hybrid Gated DeltaNet) base weights; see LICENSE (Kev) and LICENSE-QWEN (base).
Modifications: the Kev LoRA adapter is merged into the base weights; the base's MTP tensors are removed; the pointer head is also provided as safetensors.
Sources (exact revisions)
| repo | revision | |
|---|---|---|
| Kev release | jaredpalmer/kev-9b | db029f08b290afd9fee4aa4bbcd9ae48602d1eb0 |
| Base | Qwen/Qwen3.5-9B-Base | 68c46c4b3498877f3ef123c856ecfde50c39f404 |
| Kev code (encoder, head, merge rule) | github.com/jaredpalmer/kev@fe64b1274ea7f80d4095866df90666abb03e9cf6 | Apache-2.0 |
What this repository is
- A Qwen3.5 (hybrid Gated DeltaNet) backbone in bf16 with the Kev fine-tune already applied. Kev never uses the LM head.
head.safetensors: the Kev pointer head (fp32).head.ptis the unchanged original, andkev_head.jsondescribes the contract:
| tensor | shape | dtype |
|---|---|---|
q.weight |
[256, 4096] | float32 |
q.bias |
[256] | float32 |
k.weight |
[256, 4096] | float32 |
k.bias |
[256] | float32 |
temperature |
[] | float32 |
- Readout:
logit_j = ((W_k h_opt_j + b_k) . (W_q h_decide + b_q)) / sqrt(256) / T, then a softmax over the question's options.his the backbone's final-norm hidden state. - Temperature T:
2.193649959389252fromhead.pt. - Token layout:
<|fim_prefix|>state,<|fim_middle|>question,<|box_start|>/<|box_end|>option,<|fim_suffix|>decide. There is no chat template; each question is a causal row continuing the state. - The tokenizer,
config.jsonand preprocessor files come from the base at the pinned revision. That is what Kev's loader uses: it always loads the tokenizer from the base.
Procedure
Merged with upstream/kev_merge.py (sha256 260f41d81bb9901b6991b36869dabe0930611c360ab3c0c1a4d7eb7de9331eed), streaming one base shard at a time on CPU:
- For each LoRA-adapted Linear:
W_bf16 = bf16( fp32(W) + (B @ A) * 2.0 ). Here2.0 = lora_alpha / r = 32 / 16comes from the adapter'sadapter_config.json(plain LoRA: no DoRA, no rsLoRA, no rank/alpha patterns). This is exactly kev-srckev/checkpoint.py:296-312: an fp32 base, PEFTmerge_and_unload, then a single cast to bf16. - Adapter keys
base_model.model.<module>.lora_{A,B}.weightmap to base keysmodel.language_model.<module>.weight. The prefix was chosen as the only candidate under which every adapted module exists in the base index. - 248 of 248 adapted modules were applied. The script fails if any adapted module is missing.
- Every other tensor is copied bit for bit, including tensors the base stores in fp32.
- Dropped: 15
mtp.*tensors. Kev builds its backbone asAutoModelForCausalLM(...).model, the text model only, so it never uses them.model.visual.*is kept unchanged, so the unchanged baseconfig.jsonstill describes the files. - Shard file names follow the base. Output: 760 tensors, 18,819,635,168 bytes.
Verification (CPU, before upload)
Run with upstream/kev_verify.py through kev-src's own encode(), rows_of(), DecisionModel.probs() and PointerHead. It used 5 short System One requests built with kev.api.to_record (token counts [49, 49, 43, 36, 63]). No GPU was used.
- Inventory vs base: 760 tensors. Names, shapes and dtypes equal the base's minus the 15 dropped tensors (ok=True).
- Bitwise vs Kev's own bf16 merge: all 426 backbone tensors of kev-src's PEFT fp32 merge, cast to bf16, equal this artifact bit for bit (ok=True). This covers all 248 adapted modules. 48 tensors are stored in fp32 and compared in fp32.
- Head conversion:
head.safetensorstensors equalhead.pt's (ok=True). The temperature equals the loader's, rounded to fp32 (ok=True).
| arm vs Kev reference (fp32, adapter unmerged, head.pt) | hidden max abs | hidden max rel L2 | probs max abs | argmax agree |
|---|---|---|---|---|
| fp32 merged (PEFT merge_and_unload) vs fp32 unmerged | 3.62e-05 | 2.66e-06 | 1.46e-06 | 7/7 |
| this artifact, bf16 weights, fp32 compute | 0.233 | 0.00847 | 0.0018 | 7/7 |
| this artifact, bf16 weights, bf16 compute (served form) | 0.476 | 0.0256 | 0.00929 | 7/7 |
The bf16 drift is inherent to bf16 serving. Kev's own cards report served bf16 within 0.017 of fp32 for the 4B. Not measured here: accuracy on Kev's evaluation suites, and GPU end-to-end serving.
Files
| file | bytes | sha256 |
|---|---|---|
LICENSE |
11,343 | |
LICENSE-QWEN |
11,343 | |
config.json |
3,126 | |
head.pt |
8,395,071 | 8e1dab2c8e3664f6fee843e0257d57946c25e0210c61901ea77e71761f4244e1 |
head.safetensors |
8,391,148 | 9bb5de2c1c6eb07508049e5040e54950bfad896c73c31b4ba7346929eb4ab46b |
kev_head.json |
2,793 | |
merge_manifest.json |
17,332 | |
merges.txt |
3,353,259 | |
model.safetensors-00001-of-00004.safetensors |
5,276,436,240 | 388d45e94da7a76afd4d5121682ad42d639a9675b2074281d7da220f5ce8d78f |
model.safetensors-00002-of-00004.safetensors |
5,335,161,608 | 1f2200def7573e1a0d75218ae999c21586ccbad0e9343605f3b74e79c01693d2 |
model.safetensors-00003-of-00004.safetensors |
4,932,509,248 | b2f98aec5b8a11d3765faaf153296bd5f775bf7336d463dba41be9394fa7f0aa |
model.safetensors-00004-of-00004.safetensors |
3,275,621,016 | 103cd9ff8f9f2f15543a52d2d1d5d7bd175288e7f82d91ebc9fce2a2152cbabb |
model.safetensors.index.json |
78,337 | |
preprocessor_config.json |
390 | |
tokenizer.json |
12,807,196 | |
tokenizer_config.json |
16,713 | |
upstream/adapter_config.json |
1,271 | |
upstream/kev_head.py |
4,878 | |
upstream/kev_merge.py |
7,700 | |
upstream/kev_model_card.md |
23,163 | |
upstream/kev_verify.py |
10,381 | |
upstream/provenance.json |
4,216 | |
upstream/training_config.json |
1,759 | |
upstream/training_metrics.json |
325 | |
verification.json |
2,148 | |
video_preprocessor_config.json |
386 | |
vocab.json |
6,722,759 | |
README.md |
(this file) |
- Downloads last month
- 17
Model tree for avartha/kev-9b-merged
Base model
Qwen/Qwen3.5-9B-Base