kev-8b-merged: unofficial merged derivative of Kev

This is an unofficial derivative of jaredpalmer/kev-8b by Jared Palmer, prepared by Avartha. It is not endorsed by the Kev author. Kev and its weights are Apache-2.0, as are the Qwen3 (attention-only) base weights; see LICENSE (Kev) and LICENSE-QWEN (base).

Modifications: the Kev LoRA adapter is merged into the base weights; the pointer head is also provided as safetensors.

Sources (exact revisions)

repo revision
Kev release jaredpalmer/kev-8b c80773da7f383f93c4dbff0c0b008e0463f9145a
Base Qwen/Qwen3-8B-Base 49e3418fbbbca6ecbdf9608b4d22e5a407081db4
Kev code (encoder, head, merge rule) github.com/jaredpalmer/kev@fe64b1274ea7f80d4095866df90666abb03e9cf6 Apache-2.0

What this repository is

  • A Qwen3 (attention-only) backbone in bf16 with the Kev fine-tune already applied. Kev never uses the LM head.
  • head.safetensors: the Kev pointer head (fp32). head.pt is the unchanged original, and kev_head.json describes the contract:
tensor shape dtype
q.weight [256, 4096] float32
q.bias [256] float32
k.weight [256, 4096] float32
k.bias [256] float32
temperature [] float32
  • Readout: logit_j = ((W_k h_opt_j + b_k) . (W_q h_decide + b_q)) / sqrt(256) / T, then a softmax over the question's options. h is the backbone's final-norm hidden state.
  • Temperature T: 1.0: head.pt carries no temperature, and kev-src's loader then applies 1.0.
  • Token layout: <|fim_prefix|> state, <|fim_middle|> question, <|box_start|>/<|box_end|> option, <|fim_suffix|> decide. There is no chat template; each question is a causal row continuing the state.
  • The tokenizer, config.json and preprocessor files come from the base at the pinned revision. That is what Kev's loader uses: it always loads the tokenizer from the base.

Procedure

Merged with upstream/kev_merge.py (sha256 260f41d81bb9901b6991b36869dabe0930611c360ab3c0c1a4d7eb7de9331eed), streaming one base shard at a time on CPU:

  • For each LoRA-adapted Linear: W_bf16 = bf16( fp32(W) + (B @ A) * 2.0 ). Here 2.0 = lora_alpha / r = 32 / 16 comes from the adapter's adapter_config.json (plain LoRA: no DoRA, no rsLoRA, no rank/alpha patterns). This is exactly kev-src kev/checkpoint.py:296-312: an fp32 base, PEFT merge_and_unload, then a single cast to bf16.
  • Adapter keys base_model.model.<module>.lora_{A,B}.weight map to base keys model.<module>.weight. The prefix was chosen as the only candidate under which every adapted module exists in the base index.
  • 252 of 252 adapted modules were applied. The script fails if any adapted module is missing.
  • Every other tensor is copied bit for bit, including tensors the base stores in fp32.
  • Nothing dropped: the base has no MTP or vision tensors.
  • Shard file names follow the base. Output: 399 tensors, 16,381,470,720 bytes.

Verification (CPU, before upload)

Run with upstream/kev_verify.py through kev-src's own encode(), rows_of(), DecisionModel.probs() and PointerHead. It used 5 short System One requests built with kev.api.to_record (token counts [48, 50, 43, 35, 64]). No GPU was used.

  • Inventory vs base: 399 tensors. Names, shapes and dtypes equal the base's minus the 0 dropped tensors (ok=True).
  • Bitwise vs Kev's own bf16 merge: all 398 backbone tensors of kev-src's PEFT fp32 merge, cast to bf16, equal this artifact bit for bit (ok=True). This covers all 252 adapted modules. 0 tensors are stored in fp32 and compared in fp32.
  • Head conversion: head.safetensors tensors equal head.pt's (ok=True). The temperature equals the loader's, rounded to fp32 (ok=True).
arm vs Kev reference (fp32, adapter unmerged, head.pt) hidden max abs hidden max rel L2 probs max abs argmax agree
fp32 merged (PEFT merge_and_unload) vs fp32 unmerged 0.000137 1.73e-06 1.19e-07 7/7
this artifact, bf16 weights, fp32 compute 1.7 0.024 0.00145 7/7
this artifact, bf16 weights, bf16 compute (served form) 11.6 0.0769 0.00426 7/7

The bf16 drift is inherent to bf16 serving. Kev's own cards report served bf16 within 0.017 of fp32 for the 4B. Not measured here: accuracy on Kev's evaluation suites, and GPU end-to-end serving.

Files

file bytes sha256
LICENSE 11,343
config.json 729
generation_config.json 138
head.pt 8,393,855 35ebe2a18c16ebd74f8c2090e7b2de8c2b007856fde96908c3a83459a2ca32bd
head.safetensors 8,391,132 785371819ae57f7812deeae51e8fb7460a12ed5cadfc21fabc6da64b4ad363e8
kev_head.json 2,776
merge_manifest.json 12,582
merges.txt 1,671,853
model-00001-of-00005.safetensors 3,996,250,744 216224c450fba853123a78674088ec59ea75198eb8e279d54f7deef498ac5e0e
model-00002-of-00005.safetensors 3,993,160,032 e0b570f79e6c9cb7625c82ec0407bc59c4aec86626d58f77dfca825646b89fde
model-00003-of-00005.safetensors 3,959,604,768 4c1ba54202e76e1d648dcb1386dc97f0243553b496bf4925480a1510bdbdb42e
model-00004-of-00005.safetensors 3,187,841,392 e2c7a666793d70fef877cb823671cacb944df767895a574a7e95a2fccdf90e3b
model-00005-of-00005.safetensors 1,244,659,840 fbf24915d47ea030bb68ab0b9488f4515a907185baa6dc26837c9c3f2326a550
model.safetensors.index.json 32,878
tokenizer.json 7,031,645
tokenizer_config.json 9,678
upstream/adapter_config.json 1,183
upstream/kev_head.py 4,878
upstream/kev_merge.py 7,700
upstream/kev_model_card.md 7,116
upstream/kev_verify.py 10,381
upstream/provenance.json 2,934
upstream/training_config.json 1,086
upstream/training_metrics.json 324
verification.json 2,108
vocab.json 2,776,833
README.md (this file)
Downloads last month
17
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for avartha/kev-8b-merged

Finetuned
(607)
this model