jaredpalmer commited on
Commit
d20697f
·
verified ·
1 Parent(s): 504b7c6

Card: Kev display names; Kev-0.5B scored on the same out-of-domain items (0.561)

Browse files
Files changed (1) hide show
  1. README.md +6 -6
README.md CHANGED
@@ -24,7 +24,7 @@ metrics:
24
  - expected_calibration_error
25
  - nll
26
  model-index:
27
- - name: kev-0.5b
28
  results:
29
  - task: { type: text-classification, name: typed decision (choice / noul / score) }
30
  dataset: { type: mixed, name: "held-out split of the six training sources (1,350 questions)" }
@@ -34,13 +34,13 @@ model-index:
34
  - { type: expected_calibration_error, value: 0.031, name: "ECE after temperature scaling (T=1.47)" }
35
  ---
36
 
37
- # kev-0.5b — prototype (superseded)
38
 
39
- `kev-0.5b` is a **decision model**. It takes one document (the *state*) and a set of typed questions, and returns a probability distribution for each question in one forward pass. It does not generate text.
40
 
41
  It is a LoRA adapter plus a small pointer head on top of `Qwen/Qwen2.5-0.5B`. It reproduces the architecture that Archer Hume inferred for TypeSafe's Jev in [*Jev's Architecture Unmasked*](https://archerhume.com/posts/jevs-architecture-unmasked), and it serves TypeSafe's public `/v1/systemone` API contract.
42
 
43
- This checkpoint is the **original prototype**, trained on a laptop in September 2026 to show the mechanism works. It is superseded by [`kev-0.6b`](kev-0.6b.md), [`kev-4b`](kev-4b.md) and [`kev-8b`](kev-8b.md), which use a Qwen3 base, frozen checksummed suites, and a recipe found through ~100 controlled trials; on the same out-of-domain items this model scores 0.575 against 0.620 / 0.790 / 0.796. It stays on the Hub for reference and reproducibility; use the current family for anything else.
44
 
45
  - Hub: [jaredpalmer/kev-0.5b](https://huggingface.co/jaredpalmer/kev-0.5b) (tag `v0.1`)
46
  - Hub: [jaredpalmer/kev-0.5b](https://huggingface.co/jaredpalmer/kev-0.5b) (this repo, run `kev`)
@@ -63,7 +63,7 @@ This checkpoint is the **original prototype**, trained on a laptop in September
63
  | Question types | `noul` (yes/no), `choice` (2–255 options), `score` (2–255 ordered levels) |
64
  | Language | English |
65
  | License | Apache-2.0 for the adapter and head. The base model is under the Qwen license (Apache-2.0 for Qwen2.5-0.5B). Datasets carry their own licenses. |
66
- | Version | `kev-0.5b` v0.1, trained 2026-09-17 |
67
 
68
  ## Intended use
69
 
@@ -134,7 +134,7 @@ Held-out **test / validation** splits of the same six sources, 150 records per s
134
 
135
  ### Accuracy and calibration
136
 
137
- | source | K | zero-shot base | zero-shot Instruct | **kev-0.5b** |
138
  |---|---|---|---|---|
139
  | | | acc / ECE | acc / ECE | acc / ECE / NLL |
140
  | banking77 | 77 | – | – | 0.860 / 0.057 / 0.56 |
 
24
  - expected_calibration_error
25
  - nll
26
  model-index:
27
+ - name: Kev-0.5B
28
  results:
29
  - task: { type: text-classification, name: typed decision (choice / noul / score) }
30
  dataset: { type: mixed, name: "held-out split of the six training sources (1,350 questions)" }
 
34
  - { type: expected_calibration_error, value: 0.031, name: "ECE after temperature scaling (T=1.47)" }
35
  ---
36
 
37
+ # Kev-0.5B — prototype (superseded)
38
 
39
+ Kev-0.5B is a **decision model**. It takes one document (the *state*) and a set of typed questions, and returns a probability distribution for each question in one forward pass. It does not generate text.
40
 
41
  It is a LoRA adapter plus a small pointer head on top of `Qwen/Qwen2.5-0.5B`. It reproduces the architecture that Archer Hume inferred for TypeSafe's Jev in [*Jev's Architecture Unmasked*](https://archerhume.com/posts/jevs-architecture-unmasked), and it serves TypeSafe's public `/v1/systemone` API contract.
42
 
43
+ This checkpoint is the **original prototype**, trained on a laptop in September 2026 to show the mechanism works. It is superseded by [Kev-0.6B](kev-0.6b.md), [Kev-4B](kev-4b.md) and [Kev-8B](kev-8b.md), which use a Qwen3 base, frozen checksummed suites, and a recipe found through ~100 controlled trials; on the same out-of-domain items (transfer-v4 dev) this model scores 0.561 against 0.620 / 0.790 / 0.796. It stays on the Hub for reference and reproducibility; use the current family for anything else.
44
 
45
  - Hub: [jaredpalmer/kev-0.5b](https://huggingface.co/jaredpalmer/kev-0.5b) (tag `v0.1`)
46
  - Hub: [jaredpalmer/kev-0.5b](https://huggingface.co/jaredpalmer/kev-0.5b) (this repo, run `kev`)
 
63
  | Question types | `noul` (yes/no), `choice` (2–255 options), `score` (2–255 ordered levels) |
64
  | Language | English |
65
  | License | Apache-2.0 for the adapter and head. The base model is under the Qwen license (Apache-2.0 for Qwen2.5-0.5B). Datasets carry their own licenses. |
66
+ | Version | Kev-0.5B v0.1, trained 2026-09-17 |
67
 
68
  ## Intended use
69
 
 
134
 
135
  ### Accuracy and calibration
136
 
137
+ | source | K | zero-shot base | zero-shot Instruct | **Kev-0.5B** |
138
  |---|---|---|---|---|
139
  | | | acc / ECE | acc / ECE | acc / ECE / NLL |
140
  | banking77 | 77 | – | – | 0.860 / 0.057 / 0.56 |