|
Download README.md from avartha/kev-0.6b-merged: direct link, hf CLI and curl.
- Browser
- Download file 5.77 kB
-
https://huggingface.co/avartha/kev-0.6b-merged/resolve/main/README.md
- Command line
-
hf download hf://avartha/kev-0.6b-merged/README.md
-
curl -L -o README.md https://huggingface.co/avartha/kev-0.6b-merged/resolve/main/README.md
5.77 kB
| license: apache-2.0 | |
| base_model: | |
| - Qwen/Qwen3-0.6B-Base | |
| - jaredpalmer/kev-0.6b | |
| tags: | |
| - kev | |
| - decision-model | |
| - pointer-head | |
| - unofficial | |
| # kev-0.6b-merged: unofficial merged derivative of Kev | |
| This is an **unofficial derivative** of [jaredpalmer/kev-0.6b](https://huggingface.co/jaredpalmer/kev-0.6b) by Jared Palmer, prepared by Avartha. It is not endorsed by the Kev author. Kev and its weights are Apache-2.0, as are the Qwen3 (attention-only) base weights; see `LICENSE` (Kev) and `LICENSE-QWEN` (base). | |
| **Modifications:** the Kev LoRA adapter is merged into the base weights; the pointer head is also provided as safetensors. | |
| ## Sources (exact revisions) | |
| | | repo | revision | | |
| |---|---|---| | |
| | Kev release | [jaredpalmer/kev-0.6b](https://huggingface.co/jaredpalmer/kev-0.6b) | `dece6dba8d43f0f7ded45e9f5b9df12474d90843` | | |
| | Base | [Qwen/Qwen3-0.6B-Base](https://huggingface.co/Qwen/Qwen3-0.6B-Base) | `da87bfb608c14b7cf20ba1ce41287e8de496c0cd` | | |
| | Kev code (encoder, head, merge rule) | github.com/jaredpalmer/kev@fe64b1274ea7f80d4095866df90666abb03e9cf6 | Apache-2.0 | | |
| ## What this repository is | |
| - A Qwen3 (attention-only) backbone in bf16 with the Kev fine-tune already applied. Kev never uses the LM head. | |
| - `head.safetensors`: the Kev pointer head (fp32). `head.pt` is the unchanged original, and `kev_head.json` describes the contract: | |
| | tensor | shape | dtype | | |
| |---|---|---| | |
| | `q.weight` | [256, 1024] | float32 | | |
| | `q.bias` | [256] | float32 | | |
| | `k.weight` | [256, 1024] | float32 | | |
| | `k.bias` | [256] | float32 | | |
| | `temperature` | [] | float32 | | |
| - **Readout:** `logit_j = ((W_k h_opt_j + b_k) . (W_q h_decide + b_q)) / sqrt(256) / T`, then a softmax over the question's options. `h` is the backbone's final-norm hidden state. | |
| - **Temperature T:** `1.0`: `head.pt` carries no temperature, and kev-src's loader then applies 1.0. | |
| - **Token layout:** `<|fim_prefix|>` state, `<|fim_middle|>` question, `<|box_start|>`/`<|box_end|>` option, `<|fim_suffix|>` decide. There is no chat template; each question is a causal row continuing the state. | |
| - The tokenizer, `config.json` and preprocessor files come from the base at the pinned revision. That is what Kev's loader uses: it always loads the tokenizer from the base. | |
| ## Procedure | |
| Merged with `upstream/kev_merge.py` (sha256 `260f41d81bb9901b6991b36869dabe0930611c360ab3c0c1a4d7eb7de9331eed`), streaming one base shard at a time on CPU: | |
| - For each LoRA-adapted Linear: `W_bf16 = bf16( fp32(W) + (B @ A) * 2.0 )`. Here `2.0 = lora_alpha / r = 32 / 16` comes from the adapter's `adapter_config.json` (plain LoRA: no DoRA, no rsLoRA, no rank/alpha patterns). This is exactly kev-src `kev/checkpoint.py:296-312`: an fp32 base, PEFT `merge_and_unload`, then a single cast to bf16. | |
| - Adapter keys `base_model.model.<module>.lora_{A,B}.weight` map to base keys `model.<module>.weight`. The prefix was chosen as the only candidate under which every adapted module exists in the base index. | |
| - 196 of 196 adapted modules were applied. The script fails if any adapted module is missing. | |
| - Every other tensor is copied bit for bit, including tensors the base stores in fp32. | |
| - Nothing dropped: the base has no MTP or vision tensors. | |
| - Shard file names follow the base. Output: 310 tensors, 1,192,099,840 bytes. | |
| ## Verification (CPU, before upload) | |
| Run with `upstream/kev_verify.py` through kev-src's own `encode()`, `rows_of()`, `DecisionModel.probs()` and `PointerHead`. It used 5 short System One requests built with `kev.api.to_record` (token counts [48, 50, 43, 35, 64]). No GPU was used. | |
| - **Inventory vs base:** 310 tensors. Names, shapes and dtypes equal the base's minus the 0 dropped tensors (ok=True). | |
| - **Bitwise vs Kev's own bf16 merge:** all 310 backbone tensors of kev-src's PEFT fp32 merge, cast to bf16, equal this artifact bit for bit (ok=True). This covers all 196 adapted modules. 0 tensors are stored in fp32 and compared in fp32. | |
| - **Head conversion:** `head.safetensors` tensors equal `head.pt`'s (ok=True). The temperature equals the loader's, rounded to fp32 (ok=True). | |
| | arm vs Kev reference (fp32, adapter unmerged, head.pt) | hidden max abs | hidden max rel L2 | probs max abs | argmax agree | | |
| |---|---|---|---|---| | |
| | fp32 merged (PEFT merge_and_unload) vs fp32 unmerged | 0.00021 | 1.42e-05 | 2.92e-06 | 7/7 | | |
| | this artifact, bf16 weights, fp32 compute | 0.663 | 0.0313 | 0.00504 | 7/7 | | |
| | this artifact, bf16 weights, bf16 compute (served form) | 2.38 | 0.0783 | 0.00625 | 7/7 | | |
| The bf16 drift is inherent to bf16 serving. Kev's own cards report served bf16 within 0.017 of fp32 for the 4B. Not measured here: accuracy on Kev's evaluation suites, and GPU end-to-end serving. | |
| ## Files | |
| | file | bytes | sha256 | | |
| |---|---|---| | |
| | `LICENSE` | 11,343 | | | |
| | `LICENSE-QWEN` | 11,343 | | | |
| | `config.json` | 727 | | | |
| | `generation_config.json` | 138 | | | |
| | `head.pt` | 2,102,399 | ce6cd9ffc54db41c179b65a33d60973dc8280c29e886218eb3ec07dc28d6f28b | | |
| | `head.safetensors` | 2,099,676 | fafed82bab4cd26cf863be15c8ce8bf92451062e0e51c33cd0a5de65849ceaa8 | | |
| | `kev_head.json` | 2,778 | | | |
| | `merge_manifest.json` | 9,590 | | | |
| | `merges.txt` | 1,671,853 | | | |
| | `model.safetensors` | 1,192,135,096 | 283ce6ee55ed651cb3c6b9e64c107076de263a7294785b10e66a037ab76a945e | | |
| | `tokenizer.json` | 7,031,645 | | | |
| | `tokenizer_config.json` | 9,678 | | | |
| | `upstream/adapter_config.json` | 1,185 | | | |
| | `upstream/kev_head.py` | 4,878 | | | |
| | `upstream/kev_merge.py` | 7,700 | | | |
| | `upstream/kev_model_card.md` | 6,439 | | | |
| | `upstream/kev_verify.py` | 10,381 | | | |
| | `upstream/provenance.json` | 2,848 | | | |
| | `upstream/training_config.json` | 1,047 | | | |
| | `upstream/training_metrics.json` | 323 | | | |
| | `verification.json` | 2,094 | | | |
| | `vocab.json` | 2,776,833 | | | |
| | `README.md` | (this file) | | | |