Text Classification
PEFT
Safetensors
English
decision-model
calibration
lora
multiple-choice
typesafe
qwen3.5
Eval Results (legacy)
Instructions to use jaredpalmer/kev-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use jaredpalmer/kev-4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Kev-4B: real-document delta (round 8; documents-v1 locked test 0.804 -> 0.904)
Browse files- README.md +49 -13
- adapter_model.safetensors +1 -1
- head.pt +2 -2
- provenance.json +46 -29
- result.json +0 -0
- train.log +94 -46
- training_config.json +15 -8
- training_metrics.json +8 -8
README.md
CHANGED
|
@@ -30,6 +30,11 @@ metrics:
|
|
| 30 |
model-index:
|
| 31 |
- name: Kev-4B
|
| 32 |
results:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
- task: { type: text-classification, name: typed decision (choice / noul / score) }
|
| 34 |
dataset: { type: mixed, name: "decision-v7 development (1,204 records; ten trained public sources + programmatic policy data)" }
|
| 35 |
metrics:
|
|
@@ -51,20 +56,51 @@ model-index:
|
|
| 51 |
|
| 52 |
Kev-4B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 33.8M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-4B-Base` (revision `1001bb4d`), serving TypeSafe's public `/v1/systemone` contract.
|
| 53 |
|
| 54 |
-
**
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
-
|
| 57 |
-
- Demo: [huggingface.co/spaces/jaredpalmer/kev](https://huggingface.co/spaces/jaredpalmer/kev) runs Kev-4B and Kev-0.8B on ZeroGPU with the same encoder and API code as `kev.serve`.
|
| 58 |
-
- Code, suites, every trial with hashes and paired bootstraps: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev) — `PLAN_Qwen35.md` (the port and this experiment), `PLAN.md`, `runs/leaderboard.md`
|
| 59 |
|
| 60 |
-
## Results (same frozen items for every row)
|
| 61 |
|
| 62 |
| | Kev-4B (Qwen3) | Kev-4B before the delta (`v7-base`) | **Kev-4B, raw logits** | **Kev-4B as served (T = 2.14)** | Jev |
|
| 63 |
|---|---|---|---|---|---|
|
| 64 |
| in-distribution accuracy (decision-v7 dev, 1,204 records) | 0.854 | 0.877 | 0.872 | 0.872 | 0.845 |
|
| 65 |
| out-of-domain accuracy (transfer-v4 dev, 764 records) | 0.790 | 0.794 | **0.797** | 0.797 | 0.857 |
|
| 66 |
| out-of-domain Brier | 0.328 | 0.316 | 0.299 | **0.264** | 0.211 |
|
| 67 |
-
| out-of-domain ECE | 0.102 | 0.130 | 0.122 | **0.
|
| 68 |
| confident errors out of domain (p ≥ 0.9 and wrong) | 8.2% | 8.2% | 6.9% | **2.6%** | 3.7% |
|
| 69 |
| coverage at ≤ 5% error (share of decisions automatable) | 0.31 | 0.54 | 0.54 | 0.57 | 0.70 |
|
| 70 |
| held-out policy structures, both siblings correct | 0.73 | 0.78 | 0.78 | 0.78 | 0.86 |
|
|
@@ -80,18 +116,18 @@ Per-source out-of-domain accuracy (Kev-4B / Jev): QNLI 0.91 / 0.93, SciQ 0.97 /
|
|
| 80 |
|
| 81 |
**What the delta cost.** MMLU-Pro fell 0.500 → 0.490 and scienthoon's ECE rose 0.086 → 0.116; coverage at ≤ 5% error was unchanged (0.54 development, 0.67 → 0.68 locked test) and confident errors fell (8.2% → 6.9%). The pre-registered criteria for the delta (`PLAN.md`, "Tonight's autoresearch") were met for dates and for the unknowable-confidence behaviour; the coverage criterion asked for +5 pp and got 0; the locked read decided promotion.
|
| 82 |
|
| 83 |
-
**Newer evaluation columns** (`transfer-v9` development, Kev-4B / Jev): MMLU-Pro (10-way) 0.490 / 0.840; state buried among unrelated records 0.
|
| 84 |
|
| 85 |
-
**External suites** (same items as their published Jev numbers): SemIf's authored 144 — 0.896 before the delta (live Jev 0.965; SemIf's untrained Qwen3.5-4B 0.813); scienthoon's 900 tickets — queue 0.918, angry 0.790, ECE 0.116 (Jev 0.897, 0.914, 0.105). On ekzhang's 1,000-question MMLU-Pro sample the
|
| 86 |
|
| 87 |
-
## How it was built
|
| 88 |
|
| 89 |
- **Base model**: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers. Because the recurrent layers cannot honour a block-causal mask, questions run as separate causal rows that continue from the shared state (`kev/model.py: forward_rows_batch`); isolation is exact by construction (together vs alone within 1e-5) and on attention-only models this form is bit-identical to the packed one.
|
| 90 |
-
- **Recipe**: `decision-v7`, two epochs, LoRA r=16 (attention, MLP and DeltaNet projections), lr 5e-5 — the same data and settings as every other Kev, so the Qwen3 → Qwen3.5 difference is the base (`
|
| 91 |
- **Delta**: `kev.train --init_from jaredpalmer/kev-4b@v7-base --data evals/night2/dates_unknowable.jsonl --replay 2000 --lr 2e-5 --epochs 1`. The 1,425 new records are generated (no public dataset): 900 date-bearing policy cases, a third rendered plainly, a third with a relational day-count sentence, a third with a `date_facts` field; 255 cases with the deciding sentence removed and a uniform soft target over the options, plus their 270 intact controls. Record hashes are in `evals/night2/manifest.json`; the source checkpoint's hashes are in `training_config.json`.
|
| 92 |
- Why a delta and not a retrain: it is a controlled change (one fixed checkpoint, one data addition, 9 minutes), and the results section shows exactly what it moved.
|
| 93 |
|
| 94 |
-
## Known limits
|
| 95 |
|
| 96 |
- Use [Kev-9B](kev-9b.md) when accuracy and calibration matter more than memory: 0.852 vs 0.837 out of domain on the locked test, Brier 0.237 vs 0.255.
|
| 97 |
|
|
@@ -102,11 +138,11 @@ Per-source out-of-domain accuracy (Kev-4B / Jev): QNLI 0.91 / 0.93, SciQ 0.97 /
|
|
| 102 |
- The raw logits are over-confident out of domain; the built-in temperature (T = 2.14) fixes most of it without changing any answer. `KEV_TEMPERATURE=1.0` gives the raw values. Coverage at a 5% error budget is 0.54–0.68 against Jev's 0.70.
|
| 103 |
- 4B bf16 needs ~9 GB of GPU memory for serving; training took 56 min on one H100 (peak 24.6 GB).
|
| 104 |
|
| 105 |
-
## Training
|
| 106 |
|
| 107 |
Frozen suite `evals/v7/decision-v7`: 10,000 public records (1,000 per source), 896 policy minimal-pair records over nine template families, 1,680 records from 60 randomly generated rule structures in four rendering styles. Two epochs, LoRA r=16 α=32 on `q/k/v/o_proj`, `gate/up/down_proj`, `in_proj_qkv/z/a/b`, `out_proj`; pointer head from scratch; cross-entropy on the option distribution; lr 5e-5 (OneCycle), effective batch 8, bf16 autocast with fp32 master weights, gradient checkpointing; option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records. Then the delta described above (one epoch, lr 2e-5, 3,937 records seen, 9 minutes on one H100). No Jev outputs were used for training.
|
| 108 |
|
| 109 |
-
## Evaluation protocol
|
| 110 |
|
| 111 |
Development partitions select models; the locked test partition is read at most once per candidate (`runs/locked/kev-4b-night2-du-ungated/`; the pre-delta read is `runs/locked/kev-4b-q35/`). Every number carries suite hash, code hashes and git commit in `result.json`. Untrained-base baselines use zero-shot letter logits on the same items (`scripts/base_mmlu_probe.py`).
|
| 112 |
|
|
|
|
| 30 |
model-index:
|
| 31 |
- name: Kev-4B
|
| 32 |
results:
|
| 33 |
+
- task: { type: text-classification, name: typed decision, real documents, locked test }
|
| 34 |
+
dataset: { type: mixed, name: "documents-v1 test (936 questions on CFPB complaint narratives; read once)" }
|
| 35 |
+
metrics:
|
| 36 |
+
- { type: accuracy, value: 0.904 }
|
| 37 |
+
- { type: brier_score, value: 0.156 }
|
| 38 |
- task: { type: text-classification, name: typed decision (choice / noul / score) }
|
| 39 |
dataset: { type: mixed, name: "decision-v7 development (1,204 records; ten trained public sources + programmatic policy data)" }
|
| 40 |
metrics:
|
|
|
|
| 56 |
|
| 57 |
Kev-4B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 33.8M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-4B-Base` (revision `1001bb4d`), serving TypeSafe's public `/v1/systemone` contract.
|
| 58 |
|
| 59 |
+
**This version (2026-09-24): real-document delta.** The previous Kev-4B plus one epoch (lr 2e-5) on `documents-v1` train: 5,219 real US consumer-finance complaint narratives (CFPB, 2015-2024, up to ~7k tokens) with 7,488 questions (which product, which main issue), labels kept only where two open-weight teachers agreed with the consumer's own filing, mixed with 2,000 replayed `decision-v7` records. On complaint narratives it has never seen, accuracy goes from 0.804 to **0.904** on the locked test (+9.9 pp [+7.5, +12.4], 936 questions) and from 0.811 to **0.891** on a private held-out set (`documents-v2`, 953 questions); on the development split it scores 0.895 against Jev's 0.868. Everything else is unchanged within noise: locked out-of-domain test 0.835 (previous 0.837), served Brier 0.233 (0.232).
|
| 60 |
+
|
| 61 |
+
**Read this before relying on the documents numbers.** The gain is measured **in distribution**: training and every documents suite share one source (CFPB complaints) and the same two question templates. It shows Kev-4B learns real long documents from a few thousand labelled examples; it does not show the same gain on other kinds of documents. Evaluation labels are AI-adjudicated (a unanimous three-model judge panel, or two agreeing adjudications) and human spot-checked (47/50 and 50/50).
|
| 62 |
+
|
| 63 |
+
- Hub: `jaredpalmer/kev-4b` (this repo; trial `r8-small/00-trial-0`, `PLAN.md` round 8). The previous version is at tag `night2-du-release`; the pre-delta v7 checkpoint at `v7-base`; the Qwen3 generation at `qwen3` ([its card](kev-4b-qwen3.md)).
|
| 64 |
+
- Code, suites, every trial with hashes and paired bootstraps: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev). The numbers below are in `runs/release/kev-4b-r8.json`.
|
| 65 |
+
|
| 66 |
+
## Results (as served: each checkpoint at its own fitted temperature)
|
| 67 |
+
|
| 68 |
+
| | **Kev-4B (this version, T = 2.96)** | previous Kev-4B (T = 2.14) | Jev |
|
| 69 |
+
|---|---|---|---|
|
| 70 |
+
| **real documents**, locked test (`documents-v1`, 936 questions) | **0.904** | 0.804 | – |
|
| 71 |
+
| real documents, private held-out (`documents-v2`, 953) | **0.891** | 0.811 | – |
|
| 72 |
+
| real documents, development (920) | **0.895** | 0.811 | 0.868 |
|
| 73 |
+
| real documents, Brier (locked test) | **0.156** | 0.286 | – |
|
| 74 |
+
| in-distribution accuracy (decision-v7 dev, 1,264 questions) | 0.873 | 0.872 | 0.845 |
|
| 75 |
+
| out-of-domain accuracy (transfer-v4 dev) | 0.802 | 0.797 | 0.857 |
|
| 76 |
+
| out-of-domain Brier / ECE | 0.265 / 0.043 | 0.264 / 0.040 | 0.211 / 0.049 |
|
| 77 |
+
| confident errors out of domain (p ≥ 0.9 and wrong) | 2.9% | 2.6% | 3.7% |
|
| 78 |
+
| coverage at ≤ 5% error | 0.552 | 0.573 | 0.70 |
|
| 79 |
+
| held-out policy structures, both siblings correct | 0.781 | 0.781 | 0.86 |
|
| 80 |
+
| unknowable items answered at ≥ 0.9 (transfer-v9) | 0.00 | 0.00 | 0.09 |
|
| 81 |
+
| MMLU-Pro (transfer-v9 dev, 10-way) | 0.515 | 0.490 | 0.840 |
|
| 82 |
+
| **locked test**, out-of-domain accuracy / Brier | **0.835 / 0.233** | 0.837 / 0.232 | – |
|
| 83 |
+
| **locked test**, in-distribution accuracy | 0.875 | 0.871 | – |
|
| 84 |
+
| SemIf (144 authored decisions) | 0.882 | 0.889 | – |
|
| 85 |
+
| scienthoon (873 support tickets) | 0.723 | 0.696 | – |
|
| 86 |
+
| WANLI-v2 (1,002 NLI pairs) | 0.691 | 0.699 | – |
|
| 87 |
+
| TypeSafe (89 answered rows) | 0.843 | 0.843 | – |
|
| 88 |
+
|
| 89 |
+
Paired against the previous version (record-clustered bootstrap, 95 %): documents dev +8.4 pp [+6.0, +10.6], locked test +9.9 [+7.5, +12.4], private held-out +8.0 [+5.6, +10.3]; scienthoon +2.6 [+0.8, +4.4]; SemIf −0.7 [−2.8, +1.4]; WANLI-v2 −0.8 [−2.0, +0.4]; TypeSafe identical on all 89 rows. The release was decided by a rule registered before any read (`PLAN.md`, round 8): documents development lower bound > 0, short-state and external guards sized to what each suite can resolve, then one read of the untouched documents test and one locked read. A first seed (round 7) gave the same documents gain and failed only per-suite lower bounds on the two smallest suites; it was not released.
|
| 90 |
+
|
| 91 |
+
**Calibration.** The delta sharpened the raw logits (fitted temperature 2.14 → 2.96; raw out-of-domain Brier 0.327, raw locked Brier 0.278). As served, calibration is unchanged. `KEV_TEMPERATURE=1.0` gives the raw values.
|
| 92 |
+
|
| 93 |
+
## Previous version: `night2-du` (2026-09-21), kept at tag `night2-du-release`
|
| 94 |
|
| 95 |
+
**At its release, the recommended Kev.** The best accuracy per byte: out of domain 0.797 on the development partition and **0.837 on the locked test**, Brier 0.255 on the test, held-out rule pairs 0.77–0.78. This checkpoint is the `decision-v7` recipe (trial `q35-4b-s23/00-trial-0`, seed 2, selected on development accuracy) followed by a 9-minute **delta fine-tune** (`--init_from`, lr 2e-5, one epoch) on 1,425 additional records — date-bearing policy cases rendered with explicit day counts, and evidence-free cases with uniform targets — mixed with 2,000 replayed training records. Against the pre-delta checkpoint on the locked test: +1.0 pp [−0.1, +2.1], Brier 0.266 → 0.255, `deadline` 0.65 → 0.75.
|
|
|
|
|
|
|
| 96 |
|
|
|
|
| 97 |
|
| 98 |
| | Kev-4B (Qwen3) | Kev-4B before the delta (`v7-base`) | **Kev-4B, raw logits** | **Kev-4B as served (T = 2.14)** | Jev |
|
| 99 |
|---|---|---|---|---|---|
|
| 100 |
| in-distribution accuracy (decision-v7 dev, 1,204 records) | 0.854 | 0.877 | 0.872 | 0.872 | 0.845 |
|
| 101 |
| out-of-domain accuracy (transfer-v4 dev, 764 records) | 0.790 | 0.794 | **0.797** | 0.797 | 0.857 |
|
| 102 |
| out-of-domain Brier | 0.328 | 0.316 | 0.299 | **0.264** | 0.211 |
|
| 103 |
+
| out-of-domain ECE | 0.102 | 0.130 | 0.122 | **0.040** | 0.049 |
|
| 104 |
| confident errors out of domain (p ≥ 0.9 and wrong) | 8.2% | 8.2% | 6.9% | **2.6%** | 3.7% |
|
| 105 |
| coverage at ≤ 5% error (share of decisions automatable) | 0.31 | 0.54 | 0.54 | 0.57 | 0.70 |
|
| 106 |
| held-out policy structures, both siblings correct | 0.73 | 0.78 | 0.78 | 0.78 | 0.86 |
|
|
|
|
| 116 |
|
| 117 |
**What the delta cost.** MMLU-Pro fell 0.500 → 0.490 and scienthoon's ECE rose 0.086 → 0.116; coverage at ≤ 5% error was unchanged (0.54 development, 0.67 → 0.68 locked test) and confident errors fell (8.2% → 6.9%). The pre-registered criteria for the delta (`PLAN.md`, "Tonight's autoresearch") were met for dates and for the unknowable-confidence behaviour; the coverage criterion asked for +5 pp and got 0; the locked read decided promotion.
|
| 118 |
|
| 119 |
+
**Newer evaluation columns** (`transfer-v9` development, Kev-4B / Jev): MMLU-Pro (10-way) 0.490 / 0.840; state buried among unrelated records 0.67 / 0.70; unknowable share at ≥ 0.9 confidence 0.00 / 0.09 (intact controls 0.94).
|
| 120 |
|
| 121 |
+
**External suites** (same items as their published Jev numbers): SemIf's authored 144 — 0.896 before the delta (live Jev 0.965; SemIf's untrained Qwen3.5-4B 0.813); scienthoon's 900 tickets — queue 0.918, angry 0.790, ECE 0.116 (Jev 0.897, 0.914, 0.105). On ekzhang's 1,000-question MMLU-Pro sample the shipped checkpoint scores 0.468 over all 1,000 questions (8 exceed the state limit and count as wrong; live Jev 0.835 on the same items, ekzhang reports 0.829). On SemIf's pinned third-party selections (`evals/external/{wanli,typesafe}-v1`): WANLI-256 accuracy 0.695 (live Jev 0.758); TypeSafe-102 equal-case agreement / total-variation distance 0.856 / 0.231 over the 89 rows within the 8,192-token serving context (13 rejected), 0.770 / 0.308 over all 102 with rejected rows scored as wrong (live Jev 0.891 / 0.125; published TypeSafe answers 0.883 / 0.127); plain accuracy on the answered rows 0.843, coverage at <= 5% error 0.02 (Jev 0.892, 0.84). The shipped temperature is fitted in distribution and does not transfer to every workload. On WANLI, a single temperature fitted on the workload's own labelled rows (`python -m kev.calibrate`, group-disjoint out-of-fold) lowers ECE from 0.166 as shipped to 0.052 (workload T 3.91 against the shipped 2.14). Accuracy is unchanged and coverage at <= 5% error does not improve. On TypeSafe the shipped temperature already fits and refitting does not help (ECE 0.158 as shipped, 0.175 out of fold).
|
| 122 |
|
| 123 |
+
### How it was built
|
| 124 |
|
| 125 |
- **Base model**: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers. Because the recurrent layers cannot honour a block-causal mask, questions run as separate causal rows that continue from the shared state (`kev/model.py: forward_rows_batch`); isolation is exact by construction (together vs alone within 1e-5) and on attention-only models this form is bit-identical to the packed one.
|
| 126 |
+
- **Recipe**: `decision-v7`, two epochs, LoRA r=16 (attention, MLP and DeltaNet projections), lr 5e-5 — the same data and settings as every other Kev, so the Qwen3 → Qwen3.5 difference is the base (`PLAN.md`, Qwen3.5 port §10: locked test +7.3 pp [+2.8, +11.7] over Kev-8B).
|
| 127 |
- **Delta**: `kev.train --init_from jaredpalmer/kev-4b@v7-base --data evals/night2/dates_unknowable.jsonl --replay 2000 --lr 2e-5 --epochs 1`. The 1,425 new records are generated (no public dataset): 900 date-bearing policy cases, a third rendered plainly, a third with a relational day-count sentence, a third with a `date_facts` field; 255 cases with the deciding sentence removed and a uniform soft target over the options, plus their 270 intact controls. Record hashes are in `evals/night2/manifest.json`; the source checkpoint's hashes are in `training_config.json`.
|
| 128 |
- Why a delta and not a retrain: it is a controlled change (one fixed checkpoint, one data addition, 9 minutes), and the results section shows exactly what it moved.
|
| 129 |
|
| 130 |
+
### Known limits
|
| 131 |
|
| 132 |
- Use [Kev-9B](kev-9b.md) when accuracy and calibration matter more than memory: 0.852 vs 0.837 out of domain on the locked test, Brier 0.237 vs 0.255.
|
| 133 |
|
|
|
|
| 138 |
- The raw logits are over-confident out of domain; the built-in temperature (T = 2.14) fixes most of it without changing any answer. `KEV_TEMPERATURE=1.0` gives the raw values. Coverage at a 5% error budget is 0.54–0.68 against Jev's 0.70.
|
| 139 |
- 4B bf16 needs ~9 GB of GPU memory for serving; training took 56 min on one H100 (peak 24.6 GB).
|
| 140 |
|
| 141 |
+
### Training
|
| 142 |
|
| 143 |
Frozen suite `evals/v7/decision-v7`: 10,000 public records (1,000 per source), 896 policy minimal-pair records over nine template families, 1,680 records from 60 randomly generated rule structures in four rendering styles. Two epochs, LoRA r=16 α=32 on `q/k/v/o_proj`, `gate/up/down_proj`, `in_proj_qkv/z/a/b`, `out_proj`; pointer head from scratch; cross-entropy on the option distribution; lr 5e-5 (OneCycle), effective batch 8, bf16 autocast with fp32 master weights, gradient checkpointing; option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records. Then the delta described above (one epoch, lr 2e-5, 3,937 records seen, 9 minutes on one H100). No Jev outputs were used for training.
|
| 144 |
|
| 145 |
+
### Evaluation protocol
|
| 146 |
|
| 147 |
Development partitions select models; the locked test partition is read at most once per candidate (`runs/locked/kev-4b-night2-du-ungated/`; the pre-delta read is `runs/locked/kev-4b-q35/`). Every number carries suite hash, code hashes and git commit in `result.json`. Untrained-base baselines use zero-shot letter logits on the same items (`scripts/base_mmlu_probe.py`).
|
| 148 |
|
adapter_model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 129924032
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c2c99077b58cd4a604304b743b7db5391187ef1d58e575f90c554c744aa23d65
|
| 3 |
size 129924032
|
head.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9a3051c3676f0be29d885935b6f50c287f7acf04e3cb29a6fed82ca3e5540b7e
|
| 3 |
+
size 5249727
|
provenance.json
CHANGED
|
@@ -1,11 +1,11 @@
|
|
| 1 |
{
|
| 2 |
"config": {
|
| 3 |
"epochs": 1,
|
| 4 |
-
"seed":
|
| 5 |
"lr": 2e-05,
|
| 6 |
"lora": 16,
|
| 7 |
-
"accum":
|
| 8 |
-
"batch":
|
| 9 |
"perm_kl": 0.0,
|
| 10 |
"perm_frac": 0.3,
|
| 11 |
"ord_w": 0.0,
|
|
@@ -18,45 +18,62 @@
|
|
| 18 |
"head_lr": 0.0,
|
| 19 |
"weight_decay": 0.01,
|
| 20 |
"anchor_w": 0.0,
|
|
|
|
|
|
|
|
|
|
| 21 |
"dtype": "bf16",
|
| 22 |
"checkpointing": 1,
|
| 23 |
"replay": 2000,
|
|
|
|
|
|
|
| 24 |
"base": "Qwen/Qwen3.5-4B-Base",
|
| 25 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 26 |
-
"init_from": "jaredpalmer/kev-4b"
|
| 27 |
-
"data": "evals/night2/dates_unknowable.jsonl"
|
| 28 |
},
|
| 29 |
-
"config_sha256": "
|
| 30 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 31 |
"source_hashes": {
|
| 32 |
"kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 33 |
-
"kev/anchors.py": "
|
| 34 |
-
"kev/api.py": "
|
| 35 |
-
"kev/autoresearch.py": "
|
| 36 |
-
"kev/benchmark.py": "
|
| 37 |
-
"kev/
|
|
|
|
|
|
|
| 38 |
"kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
|
| 39 |
"kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
|
| 40 |
-
"kev/data.py": "
|
| 41 |
-
"kev/
|
| 42 |
-
"kev/
|
| 43 |
-
"kev/
|
| 44 |
-
"kev/
|
| 45 |
-
"kev/
|
| 46 |
-
"kev/
|
| 47 |
-
"kev/
|
| 48 |
-
"kev/
|
| 49 |
-
"kev/
|
| 50 |
-
"kev/
|
| 51 |
-
"kev/
|
| 52 |
-
"
|
| 53 |
-
"
|
| 54 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
},
|
| 56 |
-
"git_commit": "
|
| 57 |
"platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
|
| 58 |
"torch": "2.8.0+cu128",
|
| 59 |
"device": "cuda",
|
| 60 |
-
"gpu": "NVIDIA
|
| 61 |
-
"legacy_checkpoint": false
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"config": {
|
| 3 |
"epochs": 1,
|
| 4 |
+
"seed": 2,
|
| 5 |
"lr": 2e-05,
|
| 6 |
"lora": 16,
|
| 7 |
+
"accum": 4,
|
| 8 |
+
"batch": 2,
|
| 9 |
"perm_kl": 0.0,
|
| 10 |
"perm_frac": 0.3,
|
| 11 |
"ord_w": 0.0,
|
|
|
|
| 18 |
"head_lr": 0.0,
|
| 19 |
"weight_decay": 0.01,
|
| 20 |
"anchor_w": 0.0,
|
| 21 |
+
"label_smoothing": 0.0,
|
| 22 |
+
"brier_w": 0.0,
|
| 23 |
+
"focal_gamma": 0.0,
|
| 24 |
"dtype": "bf16",
|
| 25 |
"checkpointing": 1,
|
| 26 |
"replay": 2000,
|
| 27 |
+
"max_state": 7552,
|
| 28 |
+
"data": "evals/documents-v1/train.jsonl",
|
| 29 |
"base": "Qwen/Qwen3.5-4B-Base",
|
| 30 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 31 |
+
"init_from": "jaredpalmer/kev-4b"
|
|
|
|
| 32 |
},
|
| 33 |
+
"config_sha256": "6fb0132214a793f590cce025ce669b85433bab47a63522c369c226f0fb34ff3c",
|
| 34 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 35 |
"source_hashes": {
|
| 36 |
"kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 37 |
+
"kev/anchors.py": "6963eafdb276db5a6c94d939eae675c761555553b8d0f963ffd803448beaabbb",
|
| 38 |
+
"kev/api.py": "7bffacfb762c626b8bc2f670f350295af5ccb0883dbe90239c7d2f8e5ef56582",
|
| 39 |
+
"kev/autoresearch.py": "f7d7fabb9c0df7065bee3fec4aa4bee7028c5d1565847aefe2d74e3d7d40ed8a",
|
| 40 |
+
"kev/benchmark.py": "9c442efd8b6b852c574a646192c61ee91afa1306e3b2978bb6391ae1b91a3fa0",
|
| 41 |
+
"kev/calibrate.py": "c1eff4744fabd349e8abca86777a7aa0cbea44904c8223d3b61fca3d43741519",
|
| 42 |
+
"kev/checkpoint.py": "7cd4fbedfcfd11c2997490df2bd88eaa12af60e5ad60da2e5390e6df1ac1f626",
|
| 43 |
+
"kev/compare.py": "3606dbf98bb305d66420158498cb837304d2edb8dab01cdaf2ec42964bef3435",
|
| 44 |
"kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
|
| 45 |
"kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
|
| 46 |
+
"kev/data.py": "ef1be396e7218ba5c94f0f67de2ceff23a39656b57199cb3aa3192cde53e8320",
|
| 47 |
+
"kev/device.py": "d1677fd98ec0979c7284546306e34e0d09298ca4042fc39eb2b32a74f4e975c4",
|
| 48 |
+
"kev/evaluate.py": "999264f2837dcfbf2ec93601aa4745e701698ab850a43ad89324672a6604965c",
|
| 49 |
+
"kev/experiment.py": "644adedf43fbd10566830fa6094e16dd001b8cad29d9b9da82f0e0f245ca06b1",
|
| 50 |
+
"kev/jev.py": "acd4cc3f1844e438cc83a8d409c15ef78a5ef64d39ea10f583d75a6c7646b243",
|
| 51 |
+
"kev/metrics.py": "dba8d90550999edcd642581fb4da0367bd4b39d08471815184fa732d11e137b9",
|
| 52 |
+
"kev/mlx_model.py": "f582428796faf6bf962b772a69a6227ac3259ccbd3215c0c8b55ce28dbbd0909",
|
| 53 |
+
"kev/model.py": "387f281bc0b72da1620a8d2fcf508996fc30b9d8e342211ecd67c6d930f9cbaa",
|
| 54 |
+
"kev/plot.py": "0874bcff2885d8155a1cceade0de8a275a7163d3c3fd6aa0c294e8ccf5ff7e02",
|
| 55 |
+
"kev/predictors.py": "b2a66dd9f0f0a8026bc94603383bacfdd622e75d1c935173997175606626a9b3",
|
| 56 |
+
"kev/publish.py": "8abc9bcf4a01365697b05cf3c5ad0013462bd92d005f5616954e40ab1100c7aa",
|
| 57 |
+
"kev/serve.py": "c210b1f4598e64b49c399f73603ed372f4314c165ceda50e1340c0466ec1c735",
|
| 58 |
+
"kev/study_v3.py": "7891150aa623e4479c7c789d4dc4662f186132b37b7bac1e1f072b9e08ee13e3",
|
| 59 |
+
"kev/suite.py": "d524663fe2656d873c548a6216a2ddfebaf7299650f07ba31ae24e4c453398f1",
|
| 60 |
+
"kev/train.py": "68428ed43b3e362e52d4650dc5296ab88767061f7998a525715e1b306eeef474",
|
| 61 |
+
"kev/transfer_v9.py": "0902409742151250a28af1fd8f42258b70fcc161f3deed3ee35fe3f473f3c763",
|
| 62 |
+
"modal_app.py": "28318ef72dbeba7855f3c6b05339ee8c0f2e57562d21ede72352fba98d0e6ae3",
|
| 63 |
+
"pyproject.toml": "7c17fbe9efcadda3eb488b59adf1b6db5938cd33af4d3bd389ce4359b73fcc7d",
|
| 64 |
+
"uv.lock": "a9922dbb89acdef78299fd2b4a8c3f7f0fa1b2bc08b55595b6926fa785a9c466"
|
| 65 |
},
|
| 66 |
+
"git_commit": "ea271bf9dd0f2a280b1ef096a7d237388484620c",
|
| 67 |
"platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
|
| 68 |
"torch": "2.8.0+cu128",
|
| 69 |
"device": "cuda",
|
| 70 |
+
"gpu": "NVIDIA H200",
|
| 71 |
+
"legacy_checkpoint": false,
|
| 72 |
+
"measured_checkpoint": {
|
| 73 |
+
"requested": "/runs/r8-small/00-trial-0/checkpoint",
|
| 74 |
+
"resolved": "/runs/r8-small/00-trial-0/checkpoint",
|
| 75 |
+
"head_sha256": "19fdcdfb65380b51e90458edeb7e8dd7554532f357379c4cda44152cbbcd4195",
|
| 76 |
+
"adapter_sha256": "c2c99077b58cd4a604304b743b7db5391187ef1d58e575f90c554c744aa23d65",
|
| 77 |
+
"inference_temperature": 1.0
|
| 78 |
+
}
|
| 79 |
}
|
result.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
train.log
CHANGED
|
@@ -1,49 +1,97 @@
|
|
| 1 |
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
|
| 2 |
-
delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/
|
| 3 |
device=cuda trainable params=33.8M
|
| 4 |
-
replay: 2000 of 12576 suite training records mixed with
|
| 5 |
-
|
| 6 |
[transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
|
| 7 |
-
ep0 step 10/
|
| 8 |
-
ep0 step 20/
|
| 9 |
-
ep0 step 30/
|
| 10 |
-
ep0 step 40/
|
| 11 |
-
ep0 step 50/
|
| 12 |
-
ep0 step 60/
|
| 13 |
-
ep0 step 70/
|
| 14 |
-
ep0 step 80/
|
| 15 |
-
ep0 step 90/
|
| 16 |
-
ep0 step 100/
|
| 17 |
-
ep0 step 110/
|
| 18 |
-
ep0 step 120/
|
| 19 |
-
ep0 step 130/
|
| 20 |
-
ep0 step 140/
|
| 21 |
-
ep0 step 150/
|
| 22 |
-
ep0 step 160/
|
| 23 |
-
ep0 step 170/
|
| 24 |
-
ep0 step 180/
|
| 25 |
-
ep0 step 190/
|
| 26 |
-
ep0 step 200/
|
| 27 |
-
ep0 step 210/
|
| 28 |
-
ep0 step 220/
|
| 29 |
-
ep0 step 230/
|
| 30 |
-
ep0 step 240/
|
| 31 |
-
ep0 step 250/
|
| 32 |
-
ep0 step 260/
|
| 33 |
-
ep0 step 270/
|
| 34 |
-
ep0 step 280/
|
| 35 |
-
ep0 step 290/
|
| 36 |
-
ep0 step 300/
|
| 37 |
-
ep0 step 310/
|
| 38 |
-
ep0 step 320/
|
| 39 |
-
ep0 step 330/
|
| 40 |
-
ep0 step 340/
|
| 41 |
-
ep0 step 350/
|
| 42 |
-
ep0 step 360/
|
| 43 |
-
ep0 step 370/
|
| 44 |
-
ep0 step 380/
|
| 45 |
-
ep0 step 390/
|
| 46 |
-
ep0 step 400/
|
| 47 |
-
ep0 step 410/
|
| 48 |
-
ep0 step 420/
|
| 49 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
|
| 2 |
+
delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/485ace8703592fcf405488b262449990824cfed1: 496 adapter tensors and the pointer head loaded
|
| 3 |
device=cuda trainable params=33.8M
|
| 4 |
+
replay: 2000 of 12576 suite training records mixed with 5219 from evals/documents-v1/train.jsonl
|
| 5 |
+
7219 training requests (holdout=[]), questions by type {'choice': 8557, 'score': 580, 'noul': 837}
|
| 6 |
[transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
|
| 7 |
+
ep0 step 10/903 loss 0.574 kl 0.000 anchor 0.000 2.168s/rec
|
| 8 |
+
ep0 step 20/903 loss 0.461 kl 0.000 anchor 0.000 1.289s/rec
|
| 9 |
+
ep0 step 30/903 loss 0.478 kl 0.000 anchor 0.000 1.091s/rec
|
| 10 |
+
ep0 step 40/903 loss 0.427 kl 0.000 anchor 0.000 0.865s/rec
|
| 11 |
+
ep0 step 50/903 loss 0.417 kl 0.000 anchor 0.000 0.726s/rec
|
| 12 |
+
ep0 step 60/903 loss 0.334 kl 0.000 anchor 0.000 0.640s/rec
|
| 13 |
+
ep0 step 70/903 loss 0.464 kl 0.000 anchor 0.000 0.570s/rec
|
| 14 |
+
ep0 step 80/903 loss 0.365 kl 0.000 anchor 0.000 0.550s/rec
|
| 15 |
+
ep0 step 90/903 loss 0.221 kl 0.000 anchor 0.000 0.510s/rec
|
| 16 |
+
ep0 step 100/903 loss 0.299 kl 0.000 anchor 0.000 0.480s/rec
|
| 17 |
+
ep0 step 110/903 loss 0.673 kl 0.000 anchor 0.000 0.453s/rec
|
| 18 |
+
ep0 step 120/903 loss 0.232 kl 0.000 anchor 0.000 0.435s/rec
|
| 19 |
+
ep0 step 130/903 loss 0.207 kl 0.000 anchor 0.000 0.417s/rec
|
| 20 |
+
ep0 step 140/903 loss 0.257 kl 0.000 anchor 0.000 0.406s/rec
|
| 21 |
+
ep0 step 150/903 loss 0.363 kl 0.000 anchor 0.000 0.393s/rec
|
| 22 |
+
ep0 step 160/903 loss 0.418 kl 0.000 anchor 0.000 0.381s/rec
|
| 23 |
+
ep0 step 170/903 loss 0.426 kl 0.000 anchor 0.000 0.371s/rec
|
| 24 |
+
ep0 step 180/903 loss 0.244 kl 0.000 anchor 0.000 0.365s/rec
|
| 25 |
+
ep0 step 190/903 loss 0.142 kl 0.000 anchor 0.000 0.357s/rec
|
| 26 |
+
ep0 step 200/903 loss 0.196 kl 0.000 anchor 0.000 0.349s/rec
|
| 27 |
+
ep0 step 210/903 loss 0.224 kl 0.000 anchor 0.000 0.343s/rec
|
| 28 |
+
ep0 step 220/903 loss 0.290 kl 0.000 anchor 0.000 0.337s/rec
|
| 29 |
+
ep0 step 230/903 loss 0.248 kl 0.000 anchor 0.000 0.331s/rec
|
| 30 |
+
ep0 step 240/903 loss 0.176 kl 0.000 anchor 0.000 0.325s/rec
|
| 31 |
+
ep0 step 250/903 loss 0.135 kl 0.000 anchor 0.000 0.321s/rec
|
| 32 |
+
ep0 step 260/903 loss 0.328 kl 0.000 anchor 0.000 0.317s/rec
|
| 33 |
+
ep0 step 270/903 loss 0.398 kl 0.000 anchor 0.000 0.314s/rec
|
| 34 |
+
ep0 step 280/903 loss 0.311 kl 0.000 anchor 0.000 0.310s/rec
|
| 35 |
+
ep0 step 290/903 loss 0.075 kl 0.000 anchor 0.000 0.308s/rec
|
| 36 |
+
ep0 step 300/903 loss 0.155 kl 0.000 anchor 0.000 0.304s/rec
|
| 37 |
+
ep0 step 310/903 loss 0.367 kl 0.000 anchor 0.000 0.303s/rec
|
| 38 |
+
ep0 step 320/903 loss 0.358 kl 0.000 anchor 0.000 0.300s/rec
|
| 39 |
+
ep0 step 330/903 loss 0.253 kl 0.000 anchor 0.000 0.298s/rec
|
| 40 |
+
ep0 step 340/903 loss 0.400 kl 0.000 anchor 0.000 0.295s/rec
|
| 41 |
+
ep0 step 350/903 loss 0.087 kl 0.000 anchor 0.000 0.293s/rec
|
| 42 |
+
ep0 step 360/903 loss 0.247 kl 0.000 anchor 0.000 0.290s/rec
|
| 43 |
+
ep0 step 370/903 loss 0.225 kl 0.000 anchor 0.000 0.289s/rec
|
| 44 |
+
ep0 step 380/903 loss 0.171 kl 0.000 anchor 0.000 0.286s/rec
|
| 45 |
+
ep0 step 390/903 loss 0.402 kl 0.000 anchor 0.000 0.284s/rec
|
| 46 |
+
ep0 step 400/903 loss 0.151 kl 0.000 anchor 0.000 0.283s/rec
|
| 47 |
+
ep0 step 410/903 loss 0.175 kl 0.000 anchor 0.000 0.281s/rec
|
| 48 |
+
ep0 step 420/903 loss 0.200 kl 0.000 anchor 0.000 0.278s/rec
|
| 49 |
+
ep0 step 430/903 loss 0.460 kl 0.000 anchor 0.000 0.276s/rec
|
| 50 |
+
ep0 step 440/903 loss 0.176 kl 0.000 anchor 0.000 0.274s/rec
|
| 51 |
+
ep0 step 450/903 loss 0.124 kl 0.000 anchor 0.000 0.273s/rec
|
| 52 |
+
ep0 step 460/903 loss 0.139 kl 0.000 anchor 0.000 0.272s/rec
|
| 53 |
+
ep0 step 470/903 loss 0.285 kl 0.000 anchor 0.000 0.271s/rec
|
| 54 |
+
ep0 step 480/903 loss 0.172 kl 0.000 anchor 0.000 0.270s/rec
|
| 55 |
+
ep0 step 490/903 loss 0.296 kl 0.000 anchor 0.000 0.269s/rec
|
| 56 |
+
ep0 step 500/903 loss 0.254 kl 0.000 anchor 0.000 0.268s/rec
|
| 57 |
+
ep0 step 510/903 loss 0.147 kl 0.000 anchor 0.000 0.266s/rec
|
| 58 |
+
ep0 step 520/903 loss 0.168 kl 0.000 anchor 0.000 0.266s/rec
|
| 59 |
+
ep0 step 530/903 loss 0.166 kl 0.000 anchor 0.000 0.265s/rec
|
| 60 |
+
ep0 step 540/903 loss 0.222 kl 0.000 anchor 0.000 0.264s/rec
|
| 61 |
+
ep0 step 550/903 loss 0.178 kl 0.000 anchor 0.000 0.267s/rec
|
| 62 |
+
ep0 step 560/903 loss 0.183 kl 0.000 anchor 0.000 0.266s/rec
|
| 63 |
+
ep0 step 570/903 loss 0.383 kl 0.000 anchor 0.000 0.266s/rec
|
| 64 |
+
ep0 step 580/903 loss 0.316 kl 0.000 anchor 0.000 0.265s/rec
|
| 65 |
+
ep0 step 590/903 loss 0.082 kl 0.000 anchor 0.000 0.264s/rec
|
| 66 |
+
ep0 step 600/903 loss 0.176 kl 0.000 anchor 0.000 0.263s/rec
|
| 67 |
+
ep0 step 610/903 loss 0.064 kl 0.000 anchor 0.000 0.263s/rec
|
| 68 |
+
ep0 step 620/903 loss 0.108 kl 0.000 anchor 0.000 0.263s/rec
|
| 69 |
+
ep0 step 630/903 loss 0.181 kl 0.000 anchor 0.000 0.263s/rec
|
| 70 |
+
ep0 step 640/903 loss 0.073 kl 0.000 anchor 0.000 0.262s/rec
|
| 71 |
+
ep0 step 650/903 loss 0.152 kl 0.000 anchor 0.000 0.261s/rec
|
| 72 |
+
ep0 step 660/903 loss 0.142 kl 0.000 anchor 0.000 0.260s/rec
|
| 73 |
+
ep0 step 670/903 loss 0.452 kl 0.000 anchor 0.000 0.259s/rec
|
| 74 |
+
ep0 step 680/903 loss 0.165 kl 0.000 anchor 0.000 0.259s/rec
|
| 75 |
+
ep0 step 690/903 loss 0.126 kl 0.000 anchor 0.000 0.258s/rec
|
| 76 |
+
ep0 step 700/903 loss 0.177 kl 0.000 anchor 0.000 0.257s/rec
|
| 77 |
+
ep0 step 710/903 loss 0.200 kl 0.000 anchor 0.000 0.259s/rec
|
| 78 |
+
ep0 step 720/903 loss 0.218 kl 0.000 anchor 0.000 0.258s/rec
|
| 79 |
+
ep0 step 730/903 loss 0.173 kl 0.000 anchor 0.000 0.257s/rec
|
| 80 |
+
ep0 step 740/903 loss 0.233 kl 0.000 anchor 0.000 0.257s/rec
|
| 81 |
+
ep0 step 750/903 loss 0.123 kl 0.000 anchor 0.000 0.256s/rec
|
| 82 |
+
ep0 step 760/903 loss 0.070 kl 0.000 anchor 0.000 0.256s/rec
|
| 83 |
+
ep0 step 770/903 loss 0.339 kl 0.000 anchor 0.000 0.255s/rec
|
| 84 |
+
ep0 step 780/903 loss 0.338 kl 0.000 anchor 0.000 0.255s/rec
|
| 85 |
+
ep0 step 790/903 loss 0.131 kl 0.000 anchor 0.000 0.255s/rec
|
| 86 |
+
ep0 step 800/903 loss 0.110 kl 0.000 anchor 0.000 0.254s/rec
|
| 87 |
+
ep0 step 810/903 loss 0.110 kl 0.000 anchor 0.000 0.256s/rec
|
| 88 |
+
ep0 step 820/903 loss 0.100 kl 0.000 anchor 0.000 0.255s/rec
|
| 89 |
+
ep0 step 830/903 loss 0.390 kl 0.000 anchor 0.000 0.255s/rec
|
| 90 |
+
ep0 step 840/903 loss 0.119 kl 0.000 anchor 0.000 0.255s/rec
|
| 91 |
+
ep0 step 850/903 loss 0.233 kl 0.000 anchor 0.000 0.255s/rec
|
| 92 |
+
ep0 step 860/903 loss 0.226 kl 0.000 anchor 0.000 0.254s/rec
|
| 93 |
+
ep0 step 870/903 loss 0.121 kl 0.000 anchor 0.000 0.253s/rec
|
| 94 |
+
ep0 step 880/903 loss 0.125 kl 0.000 anchor 0.000 0.254s/rec
|
| 95 |
+
ep0 step 890/903 loss 0.604 kl 0.000 anchor 0.000 0.253s/rec
|
| 96 |
+
ep0 step 900/903 loss 0.379 kl 0.000 anchor 0.000 0.253s/rec
|
| 97 |
+
saved /runs/r8-small/00-trial-0/checkpoint
|
training_config.json
CHANGED
|
@@ -7,21 +7,26 @@
|
|
| 7 |
"head_lr": 0.0,
|
| 8 |
"weight_decay": 0.01,
|
| 9 |
"lora": 16,
|
| 10 |
-
"accum":
|
| 11 |
"holdout": "",
|
| 12 |
"perm_kl": 0.0,
|
| 13 |
"perm_frac": 0.3,
|
| 14 |
"ord_w": 0.0,
|
|
|
|
|
|
|
|
|
|
| 15 |
"suite": "/root/evals/v7/decision-v7",
|
| 16 |
"train_sources": "",
|
| 17 |
"device": "cuda",
|
| 18 |
-
"batch":
|
| 19 |
"dtype": "bf16",
|
|
|
|
| 20 |
"checkpointing": 1,
|
| 21 |
"option_isolation": 0,
|
| 22 |
"special_embeddings": 0,
|
| 23 |
"head_dim": 256,
|
| 24 |
"lora_targets": "all",
|
|
|
|
| 25 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 26 |
"p_none": 0.1,
|
| 27 |
"p_none_distract": 0.12,
|
|
@@ -32,19 +37,21 @@
|
|
| 32 |
"anchor": "",
|
| 33 |
"anchor_w": 0.0,
|
| 34 |
"anchor_sources": "",
|
| 35 |
-
"out": "/runs/
|
| 36 |
-
"data": "evals/
|
|
|
|
| 37 |
"replay": 2000,
|
| 38 |
"init_from": "jaredpalmer/kev-4b",
|
| 39 |
-
"seed":
|
| 40 |
},
|
| 41 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 42 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 43 |
"init_source": {
|
| 44 |
"init_from": "jaredpalmer/kev-4b",
|
| 45 |
-
"resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/
|
| 46 |
-
"adapter_sha256": "
|
| 47 |
-
"head_sha256": "
|
|
|
|
| 48 |
},
|
| 49 |
"ordinal_objective": "ranked_probability_score",
|
| 50 |
"holdout": []
|
|
|
|
| 7 |
"head_lr": 0.0,
|
| 8 |
"weight_decay": 0.01,
|
| 9 |
"lora": 16,
|
| 10 |
+
"accum": 4,
|
| 11 |
"holdout": "",
|
| 12 |
"perm_kl": 0.0,
|
| 13 |
"perm_frac": 0.3,
|
| 14 |
"ord_w": 0.0,
|
| 15 |
+
"label_smoothing": 0.0,
|
| 16 |
+
"brier_w": 0.0,
|
| 17 |
+
"focal_gamma": 0.0,
|
| 18 |
"suite": "/root/evals/v7/decision-v7",
|
| 19 |
"train_sources": "",
|
| 20 |
"device": "cuda",
|
| 21 |
+
"batch": 2,
|
| 22 |
"dtype": "bf16",
|
| 23 |
+
"weights_dtype": "fp32",
|
| 24 |
"checkpointing": 1,
|
| 25 |
"option_isolation": 0,
|
| 26 |
"special_embeddings": 0,
|
| 27 |
"head_dim": 256,
|
| 28 |
"lora_targets": "all",
|
| 29 |
+
"lora_placement": "full",
|
| 30 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 31 |
"p_none": 0.1,
|
| 32 |
"p_none_distract": 0.12,
|
|
|
|
| 37 |
"anchor": "",
|
| 38 |
"anchor_w": 0.0,
|
| 39 |
"anchor_sources": "",
|
| 40 |
+
"out": "/runs/r8-small/00-trial-0/checkpoint",
|
| 41 |
+
"data": "evals/documents-v1/train.jsonl",
|
| 42 |
+
"max_state": 7552,
|
| 43 |
"replay": 2000,
|
| 44 |
"init_from": "jaredpalmer/kev-4b",
|
| 45 |
+
"seed": 2
|
| 46 |
},
|
| 47 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 48 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 49 |
"init_source": {
|
| 50 |
"init_from": "jaredpalmer/kev-4b",
|
| 51 |
+
"resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/485ace8703592fcf405488b262449990824cfed1",
|
| 52 |
+
"adapter_sha256": "9797de69a42188e411b17b7b4fcb66a23374dcebc21d71a7a66f836b5d34df2b",
|
| 53 |
+
"head_sha256": "d8f796da36ff7bd7c0fb9496b452139bb7851af4fc82b07b500b682d3f721d6a",
|
| 54 |
+
"adapter_tensors": 496
|
| 55 |
},
|
| 56 |
"ordinal_objective": "ranked_probability_score",
|
| 57 |
"holdout": []
|
training_metrics.json
CHANGED
|
@@ -1,14 +1,14 @@
|
|
| 1 |
{
|
| 2 |
-
"wall_seconds":
|
| 3 |
-
"records_seen":
|
| 4 |
-
"requested_records":
|
| 5 |
"truncated_records": 0,
|
| 6 |
"rejected_records": 0,
|
| 7 |
-
"optimizer_steps":
|
| 8 |
-
"forward_tokens":
|
| 9 |
-
"peak_device_bytes":
|
| 10 |
"device": "cuda",
|
| 11 |
"dtype": "bf16",
|
| 12 |
-
"batch":
|
| 13 |
-
"peak_rss_bytes":
|
| 14 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"wall_seconds": 2574.7955305576324,
|
| 3 |
+
"records_seen": 10179,
|
| 4 |
+
"requested_records": 7219,
|
| 5 |
"truncated_records": 0,
|
| 6 |
"rejected_records": 0,
|
| 7 |
+
"optimizer_steps": 903,
|
| 8 |
+
"forward_tokens": 6749150,
|
| 9 |
+
"peak_device_bytes": 39839947776,
|
| 10 |
"device": "cuda",
|
| 11 |
"dtype": "bf16",
|
| 12 |
+
"batch": 2,
|
| 13 |
+
"peak_rss_bytes": 31563161600
|
| 14 |
}
|