Text Classification
PEFT
Safetensors
English
decision-model
calibration
lora
multiple-choice
typesafe
qwen3.5
Eval Results (legacy)
Instructions to use jaredpalmer/kev-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use jaredpalmer/kev-4b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Kev-4B round 10: skills delta (hard-v1 + devtools-v1) on top of round 8
Browse files- README.md +73 -19
- adapter_model.safetensors +1 -1
- head.pt +2 -2
- provenance.json +19 -18
- result.json +0 -0
- train.log +195 -94
- training_config.json +9 -9
- training_metrics.json +7 -7
README.md
CHANGED
|
@@ -30,42 +30,96 @@ metrics:
|
|
| 30 |
model-index:
|
| 31 |
- name: Kev-4B
|
| 32 |
results:
|
| 33 |
-
- task: { type: text-classification, name: typed decision,
|
| 34 |
-
dataset: { type: mixed, name: "
|
| 35 |
metrics:
|
| 36 |
-
- { type: accuracy, value: 0.
|
| 37 |
-
- { type: brier_score, value: 0.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
- task: { type: text-classification, name: typed decision (choice / noul / score) }
|
| 39 |
-
dataset: { type: mixed, name: "decision-v7 development (1,
|
| 40 |
metrics:
|
| 41 |
-
- { type: accuracy, value: 0.
|
| 42 |
-
- { type: expected_calibration_error, value: 0.
|
| 43 |
- task: { type: text-classification, name: typed decision, out-of-domain }
|
| 44 |
-
dataset: { type: mixed, name: "transfer-v4 development (
|
| 45 |
metrics:
|
| 46 |
-
- { type: accuracy, value: 0.
|
| 47 |
-
- { type: brier_score, value: 0.
|
| 48 |
- task: { type: text-classification, name: typed decision, out-of-domain, locked test }
|
| 49 |
dataset: { type: mixed, name: "transfer-v4 test (read once)" }
|
| 50 |
metrics:
|
| 51 |
-
- { type: accuracy, value: 0.
|
| 52 |
-
- { type: brier_score, value: 0.
|
| 53 |
---
|
| 54 |
|
| 55 |
# Kev-4B
|
| 56 |
|
| 57 |
Kev-4B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 33.8M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-4B-Base` (revision `1001bb4d`), serving TypeSafe's public `/v1/systemone` contract.
|
| 58 |
|
| 59 |
-
**This version (2026-09-24):
|
| 60 |
|
| 61 |
-
**Read this before relying on the
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
-
- Hub: `jaredpalmer/kev-4b` (this repo; trial `
|
| 64 |
-
- Code, suites, every trial with hashes and paired bootstraps: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev). The numbers below are in `runs/release/kev-4b-
|
| 65 |
|
| 66 |
## Results (as served: each checkpoint at its own fitted temperature)
|
| 67 |
|
| 68 |
-
| | **Kev-4B (this version, T = 2.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
|---|---|---|---|
|
| 70 |
| **real documents**, locked test (`documents-v1`, 936 questions) | **0.904** | 0.804 | – |
|
| 71 |
| real documents, private held-out (`documents-v2`, 953) | **0.891** | 0.811 | – |
|
|
@@ -86,11 +140,11 @@ Kev-4B is a **decision model**: one document (the *state*) and a set of typed qu
|
|
| 86 |
| WANLI-v2 (1,002 NLI pairs) | 0.691 | 0.699 | – |
|
| 87 |
| TypeSafe (89 answered rows) | 0.843 | 0.843 | – |
|
| 88 |
|
| 89 |
-
Paired against the
|
| 90 |
|
| 91 |
**Calibration.** The delta sharpened the raw logits (fitted temperature 2.14 → 2.96; raw out-of-domain Brier 0.327, raw locked Brier 0.278). As served, calibration is unchanged. `KEV_TEMPERATURE=1.0` gives the raw values.
|
| 92 |
|
| 93 |
-
##
|
| 94 |
|
| 95 |
**At its release, the recommended Kev.** The best accuracy per byte: out of domain 0.797 on the development partition and **0.837 on the locked test**, Brier 0.255 on the test, held-out rule pairs 0.77–0.78. This checkpoint is the `decision-v7` recipe (trial `q35-4b-s23/00-trial-0`, seed 2, selected on development accuracy) followed by a 9-minute **delta fine-tune** (`--init_from`, lr 2e-5, one epoch) on 1,425 additional records — date-bearing policy cases rendered with explicit day counts, and evidence-free cases with uniform targets — mixed with 2,000 replayed training records. Against the pre-delta checkpoint on the locked test: +1.0 pp [−0.1, +2.1], Brier 0.266 → 0.255, `deadline` 0.65 → 0.75.
|
| 96 |
|
|
|
|
| 30 |
model-index:
|
| 31 |
- name: Kev-4B
|
| 32 |
results:
|
| 33 |
+
- task: { type: text-classification, name: typed decision, skill records, locked test }
|
| 34 |
+
dataset: { type: mixed, name: "hard-v1 test (1,088 questions; programmatic labels, held-out templates; read once)" }
|
| 35 |
metrics:
|
| 36 |
+
- { type: accuracy, value: 0.803 }
|
| 37 |
+
- { type: brier_score, value: 0.278 }
|
| 38 |
+
- task: { type: text-classification, name: typed decision, developer tooling, locked test }
|
| 39 |
+
dataset: { type: mixed, name: "devtools-v1 test (1,071 questions; six public developer-tooling sources; read once)" }
|
| 40 |
+
metrics:
|
| 41 |
+
- { type: accuracy, value: 0.756 }
|
| 42 |
+
- { type: brier_score, value: 0.342 }
|
| 43 |
- task: { type: text-classification, name: typed decision (choice / noul / score) }
|
| 44 |
+
dataset: { type: mixed, name: "decision-v7 development (1,264 questions; ten trained public sources + programmatic policy data)" }
|
| 45 |
metrics:
|
| 46 |
+
- { type: accuracy, value: 0.873 }
|
| 47 |
+
- { type: expected_calibration_error, value: 0.013, name: "ECE, as served" }
|
| 48 |
- task: { type: text-classification, name: typed decision, out-of-domain }
|
| 49 |
+
dataset: { type: mixed, name: "transfer-v4 development (656 questions; six never-trained sources + held-out policy structures)" }
|
| 50 |
metrics:
|
| 51 |
+
- { type: accuracy, value: 0.817 }
|
| 52 |
+
- { type: brier_score, value: 0.243 }
|
| 53 |
- task: { type: text-classification, name: typed decision, out-of-domain, locked test }
|
| 54 |
dataset: { type: mixed, name: "transfer-v4 test (read once)" }
|
| 55 |
metrics:
|
| 56 |
+
- { type: accuracy, value: 0.838 }
|
| 57 |
+
- { type: brier_score, value: 0.224 }
|
| 58 |
---
|
| 59 |
|
| 60 |
# Kev-4B
|
| 61 |
|
| 62 |
Kev-4B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 33.8M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-4B-Base` (revision `1001bb4d`), serving TypeSafe's public `/v1/systemone` contract.
|
| 63 |
|
| 64 |
+
**This version (2026-09-24, second update): skills delta.** The round-8 Kev-4B (below) plus one epoch (lr 2e-5) on 11,320 new training records mixed with 4,000 replayed `decision-v7` records. 6,000 come from `hard-v1`, our programmatically labelled suite of the skills Kev was worst at: long policy documents with exceptions and sublimits, trade-offs under stated priorities, probability and expected value, multi-hop over several facts, dates and arithmetic, judging a proposed answer, and abstaining when a fact is missing. 5,320 come from `devtools-v1`, developer-tooling decisions from four licence-checked public datasets (CodeReviewer, CommitPackFT, FlakeFlagger, Aegis). On the held-out test splits, read once, accuracy goes from 0.540 to **0.803** on `hard-v1` (+26.3 pp [+23.3, +29.5], 1,088 questions) and from 0.623 to **0.756** on `devtools-v1` (+13.4 pp [+10.1, +16.1], 1,071 questions). On the development splits it scores 0.786 on hard-v1 against Jev's 0.777 and 0.739 on devtools-v1 against Jev's 0.713. The locked out-of-domain test is unchanged within noise (0.835 → 0.838, +0.3 pp [−1.8, +2.3]; served Brier 0.233 → 0.224). On JevBench's public items, which no training or selection step saw, the hard tier goes from 0.450 to **0.541**.
|
| 65 |
|
| 66 |
+
**Read this before relying on the hard-v1 and devtools-v1 numbers.**
|
| 67 |
+
|
| 68 |
+
- **Both gains are measured in distribution.** `hard-v1` is generated: every label is computed by its family's solver, and the splits hold out *templates* (0-3 train, 4 development, 5 test) of the same seven generators. A held-out template is a new wording of a skill the model was trained on, not a new skill. JevBench's public hard tier is the out-of-distribution check, and there the gain is about a third as large (+9.0 pp, below).
|
| 69 |
+
- **devtools-v1 labels are the public datasets' own, not adjudicated for this suite.** Some are human (CodeReviewer: whether a reviewer commented on the hunk; Aegis: human safety labels), some heuristic or by construction (CommitPackFT: the commit type is the first verb of the subject; FlakeFlagger: the test both passed and failed over reruns; When2Call: built by NVIDIA's pipeline). Before this delta, every model we scored was near chance on two of the binary sources, Jev included: on development, CodeReviewer 0.473 to 0.553 and FlakeFlagger 0.500 to 0.520 across Kev-0.8B, Kev-4B, Kev-9B, Kev-27B and Jev. This version reaches 0.633 and 0.693 on them after training on the same sources, which may be the labelling proxy being learned rather than the decision. When2Call and the prompt-injection source are never trained on: When2Call rises 0.573 → 0.660, prompt injection stays at 0.753 (Jev 0.893).
|
| 70 |
+
- **Two reads went down.** The locked in-distribution test moved 0.875 → 0.865, and TypeSafe's 89 answered rows moved 0.843 → 0.798. That is 4 questions, but the paired interval (−4.5 pp [−9.3, −1.0]) excludes zero. TypeSafe is too small to gate on, so the rule only counts it inside the pooled external guard. Real documents are unchanged (development 0.895 → 0.891, −0.3 pp [−1.5, +0.9]).
|
| 71 |
|
| 72 |
+
- Hub: `jaredpalmer/kev-4b` (this repo; trial `r10-skills/00-trial-0`; the registration and every read are in `PLAN.md` round 10 on the `research/overnight-r6` branch). The previous (round-8) version is at tag `r8-documents-release`; the `night2-du` version at `night2-du-release`; the pre-delta v7 checkpoint at `v7-base`; the Qwen3 generation at `qwen3` ([its card](kev-4b-qwen3.md)).
|
| 73 |
+
- Code, suites, every trial with hashes and paired bootstraps: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev). The numbers below are in `runs/release/kev-4b-r10.json`; JevBench in `runs/jevbench-public/kev-4b-r10/`.
|
| 74 |
|
| 75 |
## Results (as served: each checkpoint at its own fitted temperature)
|
| 76 |
|
| 77 |
+
| | **Kev-4B (this version, T = 2.41)** | round-8 Kev-4B (T = 2.96) | Jev |
|
| 78 |
+
|---|---|---|---|
|
| 79 |
+
| **hard-v1**, test (1,088 questions, read once) | **0.803** | 0.540 | – |
|
| 80 |
+
| hard-v1, development (1,083) | **0.786** | 0.503 | 0.777 |
|
| 81 |
+
| hard-v1 ECE, test / development | 0.084 / 0.095 | 0.112 / 0.137 | – / 0.035 |
|
| 82 |
+
| **devtools-v1**, test (1,071, read once) | **0.756** | 0.623 | – |
|
| 83 |
+
| devtools-v1, development (1,072) | **0.739** | 0.605 | 0.713 |
|
| 84 |
+
| real documents, development (`documents-v1`, 920) | 0.891 | 0.895 | 0.868 |
|
| 85 |
+
| in-distribution accuracy (decision-v7 dev, 1,264 questions) | 0.873 | 0.873 | 0.845 |
|
| 86 |
+
| out-of-domain accuracy (transfer-v4 dev, 656) | 0.817 | 0.802 | 0.857 |
|
| 87 |
+
| out-of-domain Brier / ECE | 0.243 / 0.042 | 0.265 / 0.043 | 0.211 / 0.049 |
|
| 88 |
+
| confident errors out of domain (p ≥ 0.9 and wrong) | 0.9% | 2.9% | 3.7% |
|
| 89 |
+
| coverage at ≤ 5% error | 0.620 | 0.552 | 0.70 |
|
| 90 |
+
| held-out policy structures, both siblings correct | 0.812 | 0.781 | 0.86 |
|
| 91 |
+
| unknowable items answered at ≥ 0.9 (transfer-v9) | 0.00 | 0.00 | 0.09 |
|
| 92 |
+
| MMLU-Pro (transfer-v9 dev, 10-way) | 0.565 | 0.515 | 0.840 |
|
| 93 |
+
| **locked test**, out-of-domain accuracy / Brier | **0.838 / 0.224** | 0.835 / 0.233 | – |
|
| 94 |
+
| **locked test**, in-distribution accuracy | 0.865 | 0.875 | – |
|
| 95 |
+
| SemIf (144 authored decisions) | 0.889 | 0.882 | – |
|
| 96 |
+
| scienthoon (873 support tickets) | 0.723 | 0.723 | – |
|
| 97 |
+
| WANLI-v2 (1,002 NLI pairs) | 0.693 | 0.691 | – |
|
| 98 |
+
| TypeSafe (89 answered rows) | 0.798 | 0.843 | ��� |
|
| 99 |
+
| JevBench public items, all 231 (report only) | 0.758 | 0.714 | – |
|
| 100 |
+
| JevBench public, hard tier (111): accuracy / ECE | 0.541 / 0.112 | 0.450 / 0.263 | – |
|
| 101 |
+
|
| 102 |
+
Jev's devtools-v1 figure is over all 1,074 development questions. Kev's rows drop the one CodeReviewer id that the builder reused for two different records (2 questions; `PLAN.md` round-10 amendment), because paired bootstraps need unique ids.
|
| 103 |
+
|
| 104 |
+
Paired against the round-8 version (record-clustered bootstrap, 95 %): hard-v1 development +28.3 pp [+25.1, +31.4], test +26.3 [+23.3, +29.5]; devtools-v1 development +13.3 [+11.3, +15.4], test +13.4 [+10.1, +16.1]; the two tests pooled +19.9 [+17.8, +21.8]; documents development −0.3 [−1.5, +0.9]; SemIf +0.7 [−2.8, +4.2]; scienthoon 0.0 [−1.4, +1.5]; WANLI-v2 +0.2 [−1.8, +2.2]; TypeSafe −4.5 [−9.3, −1.0]; locked out-of-domain test +0.3 [−1.8, +2.3].
|
| 105 |
+
|
| 106 |
+
**How the release was decided.** By a rule registered before the training data was built (`PLAN.md` round 10, `research/overnight-r6` branch). The primary criterion is the pooled hard-v1 + devtools-v1 development accuracy, with a lower bound above zero; this arm scored +20.8 pp [+18.8, +22.8]. Guards on short states, documents, WANLI-v2, scienthoon, the pooled externals and unknowable confidence are each sized to what the suite can resolve (short states +1.5 pp [−0.3, +3.5], pooled externals 0.0 [−1.2, +1.1]). Hard-set ECE may be no worse than the parent's plus 0.01. Then come one read of the two untouched test splits (pooled lower bound above zero) and one locked read. Three arms were trained: both sets, hard-v1 only, and devtools-v1 only. The two single-set arms failed external guards; this arm is the only one that passed.
|
| 107 |
+
|
| 108 |
+
**JevBench public items, report only.** `runs/jevbench-public/kev-4b-r10` holds JevBench's unchanged harness run against this checkpoint, served by the `kev-deploy` template. Across all 231 public items, accuracy goes from 0.714 (round-8 version) to 0.758. On the hard tier it goes from 0.450 to 0.541: paired over the 111 hard items that is +9.0 pp [+2.7, +15.3], with 12 items newly right and 2 newly wrong (exact McNemar p = 0.013). Hard-tier ECE falls 0.263 → 0.112. No JevBench item was used for training or selection. `evals/hard-v1/overlap.json` checks all 7,400 hard-v1 records against the 231 public items (8-gram Jaccard, threshold 0.2) and finds none above the threshold. JevBench's sealed half has not been read.
|
| 109 |
+
|
| 110 |
+
**Calibration.** This delta softened the raw logits (fitted temperature 2.96 → 2.41; raw out-of-domain Brier 0.269 on development, 0.242 on the locked test), the reverse of round 8. As served, out-of-domain Brier improved from 0.265 to 0.243 on development and confident errors from 2.9% to 0.9%. `KEV_TEMPERATURE=1.0` gives the raw values.
|
| 111 |
+
|
| 112 |
+
## Previous version: round-8 real-document delta (2026-09-24) at tag `r8-documents-release`
|
| 113 |
+
|
| 114 |
+
**Real-document delta.** The `night2-du` Kev-4B plus one epoch (lr 2e-5) on `documents-v1` train: 5,219 real US consumer-finance complaint narratives (CFPB, 2015-2024, up to ~7k tokens) with 7,488 questions (which product, which main issue), labels kept only where two open-weight teachers agreed with the consumer's own filing, mixed with 2,000 replayed `decision-v7` records. On complaint narratives it has never seen, accuracy goes from 0.804 to **0.904** on the locked test (+9.9 pp [+7.5, +12.4], 936 questions) and from 0.811 to **0.891** on a private held-out set (`documents-v2`, 953 questions); on the development split it scores 0.895 against Jev's 0.868. Everything else is unchanged within noise: locked out-of-domain test 0.835 (previous 0.837), served Brier 0.233 (0.232).
|
| 115 |
+
|
| 116 |
+
**Read this before relying on the documents numbers.** The gain is measured **in distribution**: training and every documents suite share one source (CFPB complaints) and the same two question templates. It shows Kev-4B learns real long documents from a few thousand labelled examples; it does not show the same gain on other kinds of documents. Evaluation labels are AI-adjudicated (a unanimous three-model judge panel, or two agreeing adjudications) and human spot-checked (47/50 and 50/50).
|
| 117 |
+
|
| 118 |
+
- Trial `r8-small/00-trial-0` (Hub revision `957b91e7`); the registration and every read are in `PLAN.md` round 8 on the `research/overnight-r6` branch. The numbers below are in `runs/release/kev-4b-r8.json`.
|
| 119 |
+
|
| 120 |
+
### Results (as served: each checkpoint at its own fitted temperature)
|
| 121 |
+
|
| 122 |
+
| | **round-8 Kev-4B (T = 2.96)** | `night2-du` Kev-4B (T = 2.14) | Jev |
|
| 123 |
|---|---|---|---|
|
| 124 |
| **real documents**, locked test (`documents-v1`, 936 questions) | **0.904** | 0.804 | – |
|
| 125 |
| real documents, private held-out (`documents-v2`, 953) | **0.891** | 0.811 | – |
|
|
|
|
| 140 |
| WANLI-v2 (1,002 NLI pairs) | 0.691 | 0.699 | – |
|
| 141 |
| TypeSafe (89 answered rows) | 0.843 | 0.843 | – |
|
| 142 |
|
| 143 |
+
Paired against the `night2-du` version (record-clustered bootstrap, 95 %): documents dev +8.4 pp [+6.0, +10.6], locked test +9.9 [+7.5, +12.4], private held-out +8.0 [+5.6, +10.3]; scienthoon +2.6 [+0.8, +4.4]; SemIf −0.7 [−2.8, +1.4]; WANLI-v2 −0.8 [−2.0, +0.4]; TypeSafe identical on all 89 rows. The release was decided by a rule registered before any read (`PLAN.md` round 8, `research/overnight-r6` branch): documents development lower bound > 0, short-state and external guards sized to what each suite can resolve, then one read of the untouched documents test and one locked read. A first seed (round 7) gave the same documents gain and failed only per-suite lower bounds on the two smallest suites; it was not released.
|
| 144 |
|
| 145 |
**Calibration.** The delta sharpened the raw logits (fitted temperature 2.14 → 2.96; raw out-of-domain Brier 0.327, raw locked Brier 0.278). As served, calibration is unchanged. `KEV_TEMPERATURE=1.0` gives the raw values.
|
| 146 |
|
| 147 |
+
## Earlier version: `night2-du` (2026-09-21), kept at tag `night2-du-release`
|
| 148 |
|
| 149 |
**At its release, the recommended Kev.** The best accuracy per byte: out of domain 0.797 on the development partition and **0.837 on the locked test**, Brier 0.255 on the test, held-out rule pairs 0.77–0.78. This checkpoint is the `decision-v7` recipe (trial `q35-4b-s23/00-trial-0`, seed 2, selected on development accuracy) followed by a 9-minute **delta fine-tune** (`--init_from`, lr 2e-5, one epoch) on 1,425 additional records — date-bearing policy cases rendered with explicit day counts, and evidence-free cases with uniform targets — mixed with 2,000 replayed training records. Against the pre-delta checkpoint on the locked test: +1.0 pp [−0.1, +2.1], Brier 0.266 → 0.255, `deadline` 0.65 → 0.75.
|
| 150 |
|
adapter_model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 129924032
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:90e817356246e7f18bfa7ca3d31794cd4fbeb3332a66a84cb51d9ceae925f2b2
|
| 3 |
size 129924032
|
head.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:dd633435998ecc751ac538717a3742e32149500fabf7d7276287dbf0693f347c
|
| 3 |
+
size 5249791
|
provenance.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
| 1 |
{
|
| 2 |
"config": {
|
| 3 |
"epochs": 1,
|
| 4 |
-
"seed":
|
| 5 |
"lr": 2e-05,
|
| 6 |
"lora": 16,
|
| 7 |
"accum": 4,
|
|
@@ -23,57 +23,58 @@
|
|
| 23 |
"focal_gamma": 0.0,
|
| 24 |
"dtype": "bf16",
|
| 25 |
"checkpointing": 1,
|
| 26 |
-
"replay": 2000,
|
| 27 |
"max_state": 7552,
|
| 28 |
-
"data": "evals/documents-v1/train.jsonl",
|
| 29 |
"base": "Qwen/Qwen3.5-4B-Base",
|
| 30 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 31 |
-
"init_from": "jaredpalmer/kev-4b"
|
|
|
|
|
|
|
| 32 |
},
|
| 33 |
-
"config_sha256": "
|
| 34 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 35 |
"source_hashes": {
|
| 36 |
"kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 37 |
"kev/anchors.py": "6963eafdb276db5a6c94d939eae675c761555553b8d0f963ffd803448beaabbb",
|
| 38 |
"kev/api.py": "7bffacfb762c626b8bc2f670f350295af5ccb0883dbe90239c7d2f8e5ef56582",
|
| 39 |
"kev/autoresearch.py": "f7d7fabb9c0df7065bee3fec4aa4bee7028c5d1565847aefe2d74e3d7d40ed8a",
|
| 40 |
-
"kev/benchmark.py": "
|
| 41 |
"kev/calibrate.py": "c1eff4744fabd349e8abca86777a7aa0cbea44904c8223d3b61fca3d43741519",
|
| 42 |
-
"kev/checkpoint.py": "
|
| 43 |
"kev/compare.py": "3606dbf98bb305d66420158498cb837304d2edb8dab01cdaf2ec42964bef3435",
|
| 44 |
"kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
|
| 45 |
"kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
|
| 46 |
-
"kev/
|
|
|
|
| 47 |
"kev/device.py": "d1677fd98ec0979c7284546306e34e0d09298ca4042fc39eb2b32a74f4e975c4",
|
| 48 |
"kev/evaluate.py": "999264f2837dcfbf2ec93601aa4745e701698ab850a43ad89324672a6604965c",
|
| 49 |
"kev/experiment.py": "644adedf43fbd10566830fa6094e16dd001b8cad29d9b9da82f0e0f245ca06b1",
|
| 50 |
"kev/jev.py": "acd4cc3f1844e438cc83a8d409c15ef78a5ef64d39ea10f583d75a6c7646b243",
|
| 51 |
"kev/metrics.py": "dba8d90550999edcd642581fb4da0367bd4b39d08471815184fa732d11e137b9",
|
| 52 |
"kev/mlx_model.py": "f582428796faf6bf962b772a69a6227ac3259ccbd3215c0c8b55ce28dbbd0909",
|
| 53 |
-
"kev/model.py": "
|
| 54 |
"kev/plot.py": "0874bcff2885d8155a1cceade0de8a275a7163d3c3fd6aa0c294e8ccf5ff7e02",
|
| 55 |
-
"kev/predictors.py": "
|
| 56 |
"kev/publish.py": "8abc9bcf4a01365697b05cf3c5ad0013462bd92d005f5616954e40ab1100c7aa",
|
| 57 |
-
"kev/serve.py": "
|
| 58 |
"kev/study_v3.py": "7891150aa623e4479c7c789d4dc4662f186132b37b7bac1e1f072b9e08ee13e3",
|
| 59 |
-
"kev/suite.py": "
|
| 60 |
"kev/train.py": "68428ed43b3e362e52d4650dc5296ab88767061f7998a525715e1b306eeef474",
|
| 61 |
"kev/transfer_v9.py": "0902409742151250a28af1fd8f42258b70fcc161f3deed3ee35fe3f473f3c763",
|
| 62 |
-
"modal_app.py": "
|
| 63 |
"pyproject.toml": "7c17fbe9efcadda3eb488b59adf1b6db5938cd33af4d3bd389ce4359b73fcc7d",
|
| 64 |
"uv.lock": "a9922dbb89acdef78299fd2b4a8c3f7f0fa1b2bc08b55595b6926fa785a9c466"
|
| 65 |
},
|
| 66 |
-
"git_commit": "
|
| 67 |
"platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
|
| 68 |
"torch": "2.8.0+cu128",
|
| 69 |
"device": "cuda",
|
| 70 |
"gpu": "NVIDIA H200",
|
| 71 |
"legacy_checkpoint": false,
|
| 72 |
"measured_checkpoint": {
|
| 73 |
-
"requested": "/runs/
|
| 74 |
-
"resolved": "/runs/
|
| 75 |
-
"head_sha256": "
|
| 76 |
-
"adapter_sha256": "
|
| 77 |
"inference_temperature": 1.0
|
| 78 |
}
|
| 79 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"config": {
|
| 3 |
"epochs": 1,
|
| 4 |
+
"seed": 1,
|
| 5 |
"lr": 2e-05,
|
| 6 |
"lora": 16,
|
| 7 |
"accum": 4,
|
|
|
|
| 23 |
"focal_gamma": 0.0,
|
| 24 |
"dtype": "bf16",
|
| 25 |
"checkpointing": 1,
|
|
|
|
| 26 |
"max_state": 7552,
|
|
|
|
| 27 |
"base": "Qwen/Qwen3.5-4B-Base",
|
| 28 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 29 |
+
"init_from": "jaredpalmer/kev-4b@957b91e762e883935830246eeb02381f9d2694b6",
|
| 30 |
+
"data": "evals/round10/skills/train.jsonl",
|
| 31 |
+
"replay": 4000
|
| 32 |
},
|
| 33 |
+
"config_sha256": "08aa7634eef0409340edd9df72b4b347bdda02e87b25c920e205dc727911525d",
|
| 34 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 35 |
"source_hashes": {
|
| 36 |
"kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 37 |
"kev/anchors.py": "6963eafdb276db5a6c94d939eae675c761555553b8d0f963ffd803448beaabbb",
|
| 38 |
"kev/api.py": "7bffacfb762c626b8bc2f670f350295af5ccb0883dbe90239c7d2f8e5ef56582",
|
| 39 |
"kev/autoresearch.py": "f7d7fabb9c0df7065bee3fec4aa4bee7028c5d1565847aefe2d74e3d7d40ed8a",
|
| 40 |
+
"kev/benchmark.py": "10a892d76cb4c200e94d0a68ee74054aa1870d7f484c0530a531cacfa75032ed",
|
| 41 |
"kev/calibrate.py": "c1eff4744fabd349e8abca86777a7aa0cbea44904c8223d3b61fca3d43741519",
|
| 42 |
+
"kev/checkpoint.py": "f3edb4d159c0aa12d7006fc599fe6ebc3f2c288b3b69890c289f7dac451dd03d",
|
| 43 |
"kev/compare.py": "3606dbf98bb305d66420158498cb837304d2edb8dab01cdaf2ec42964bef3435",
|
| 44 |
"kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
|
| 45 |
"kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
|
| 46 |
+
"kev/cuda_graphs.py": "12f1953fea8f77c3c9dff29468202fbb37545145f93d3618194565a89949fd55",
|
| 47 |
+
"kev/data.py": "e77a2fffeb8ee7e05b118893be8e38b8358aef97c3ca2e143cdce208201b92f9",
|
| 48 |
"kev/device.py": "d1677fd98ec0979c7284546306e34e0d09298ca4042fc39eb2b32a74f4e975c4",
|
| 49 |
"kev/evaluate.py": "999264f2837dcfbf2ec93601aa4745e701698ab850a43ad89324672a6604965c",
|
| 50 |
"kev/experiment.py": "644adedf43fbd10566830fa6094e16dd001b8cad29d9b9da82f0e0f245ca06b1",
|
| 51 |
"kev/jev.py": "acd4cc3f1844e438cc83a8d409c15ef78a5ef64d39ea10f583d75a6c7646b243",
|
| 52 |
"kev/metrics.py": "dba8d90550999edcd642581fb4da0367bd4b39d08471815184fa732d11e137b9",
|
| 53 |
"kev/mlx_model.py": "f582428796faf6bf962b772a69a6227ac3259ccbd3215c0c8b55ce28dbbd0909",
|
| 54 |
+
"kev/model.py": "2634ffe7747d69473fb6596942c248e5df2586df3610bd1adb47e7e9acd99f96",
|
| 55 |
"kev/plot.py": "0874bcff2885d8155a1cceade0de8a275a7163d3c3fd6aa0c294e8ccf5ff7e02",
|
| 56 |
+
"kev/predictors.py": "4d02906c285d78bcbd91908f4fc86577ccc95c320d704817f6444aabbc4fcd4f",
|
| 57 |
"kev/publish.py": "8abc9bcf4a01365697b05cf3c5ad0013462bd92d005f5616954e40ab1100c7aa",
|
| 58 |
+
"kev/serve.py": "57d379bbcdeb3e4d4c3fdde5471e4dddeaf7710b8ed61ecad3466f3b67e6c610",
|
| 59 |
"kev/study_v3.py": "7891150aa623e4479c7c789d4dc4662f186132b37b7bac1e1f072b9e08ee13e3",
|
| 60 |
+
"kev/suite.py": "32b2882d1ea0fe5d03ba4d676a6180d580545f157e36fa715e7e89169ec99130",
|
| 61 |
"kev/train.py": "68428ed43b3e362e52d4650dc5296ab88767061f7998a525715e1b306eeef474",
|
| 62 |
"kev/transfer_v9.py": "0902409742151250a28af1fd8f42258b70fcc161f3deed3ee35fe3f473f3c763",
|
| 63 |
+
"modal_app.py": "7c6bc859f7d3e7011539a8f5e9e90390288cc72b5d80f911c8841c068f4d1136",
|
| 64 |
"pyproject.toml": "7c17fbe9efcadda3eb488b59adf1b6db5938cd33af4d3bd389ce4359b73fcc7d",
|
| 65 |
"uv.lock": "a9922dbb89acdef78299fd2b4a8c3f7f0fa1b2bc08b55595b6926fa785a9c466"
|
| 66 |
},
|
| 67 |
+
"git_commit": "6d02f5d066cd34958dfd15ffa5d2f6f0f4c21a63",
|
| 68 |
"platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
|
| 69 |
"torch": "2.8.0+cu128",
|
| 70 |
"device": "cuda",
|
| 71 |
"gpu": "NVIDIA H200",
|
| 72 |
"legacy_checkpoint": false,
|
| 73 |
"measured_checkpoint": {
|
| 74 |
+
"requested": "/runs/r10-skills/00-trial-0/checkpoint",
|
| 75 |
+
"resolved": "/runs/r10-skills/00-trial-0/checkpoint",
|
| 76 |
+
"head_sha256": "979aaa60135274973a6dc57fe24dfc4f1ce88e65df46c73d86b2047c94d897df",
|
| 77 |
+
"adapter_sha256": "90e817356246e7f18bfa7ca3d31794cd4fbeb3332a66a84cb51d9ceae925f2b2",
|
| 78 |
"inference_temperature": 1.0
|
| 79 |
}
|
| 80 |
}
|
result.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
train.log
CHANGED
|
@@ -1,97 +1,198 @@
|
|
| 1 |
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
|
| 2 |
-
delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/
|
| 3 |
device=cuda trainable params=33.8M
|
| 4 |
-
replay:
|
| 5 |
-
|
| 6 |
[transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
|
| 7 |
-
ep0 step 10/
|
| 8 |
-
ep0 step 20/
|
| 9 |
-
ep0 step 30/
|
| 10 |
-
ep0 step 40/
|
| 11 |
-
ep0 step 50/
|
| 12 |
-
ep0 step 60/
|
| 13 |
-
ep0 step 70/
|
| 14 |
-
ep0 step 80/
|
| 15 |
-
ep0 step 90/
|
| 16 |
-
ep0 step 100/
|
| 17 |
-
ep0 step 110/
|
| 18 |
-
ep0 step 120/
|
| 19 |
-
ep0 step 130/
|
| 20 |
-
ep0 step 140/
|
| 21 |
-
ep0 step 150/
|
| 22 |
-
ep0 step 160/
|
| 23 |
-
ep0 step 170/
|
| 24 |
-
ep0 step 180/
|
| 25 |
-
ep0 step 190/
|
| 26 |
-
ep0 step 200/
|
| 27 |
-
ep0 step 210/
|
| 28 |
-
ep0 step 220/
|
| 29 |
-
ep0 step 230/
|
| 30 |
-
ep0 step 240/
|
| 31 |
-
ep0 step 250/
|
| 32 |
-
ep0 step 260/
|
| 33 |
-
ep0 step 270/
|
| 34 |
-
ep0 step 280/
|
| 35 |
-
ep0 step 290/
|
| 36 |
-
ep0 step 300/
|
| 37 |
-
ep0 step 310/
|
| 38 |
-
ep0 step 320/
|
| 39 |
-
ep0 step 330/
|
| 40 |
-
ep0 step 340/
|
| 41 |
-
ep0 step 350/
|
| 42 |
-
ep0 step 360/
|
| 43 |
-
ep0 step 370/
|
| 44 |
-
ep0 step 380/
|
| 45 |
-
ep0 step 390/
|
| 46 |
-
ep0 step 400/
|
| 47 |
-
ep0 step 410/
|
| 48 |
-
ep0 step 420/
|
| 49 |
-
ep0 step 430/
|
| 50 |
-
ep0 step 440/
|
| 51 |
-
ep0 step 450/
|
| 52 |
-
ep0 step 460/
|
| 53 |
-
ep0 step 470/
|
| 54 |
-
ep0 step 480/
|
| 55 |
-
ep0 step 490/
|
| 56 |
-
ep0 step 500/
|
| 57 |
-
ep0 step 510/
|
| 58 |
-
ep0 step 520/
|
| 59 |
-
ep0 step 530/
|
| 60 |
-
ep0 step 540/
|
| 61 |
-
ep0 step 550/
|
| 62 |
-
ep0 step 560/
|
| 63 |
-
ep0 step 570/
|
| 64 |
-
ep0 step 580/
|
| 65 |
-
ep0 step 590/
|
| 66 |
-
ep0 step 600/
|
| 67 |
-
ep0 step 610/
|
| 68 |
-
ep0 step 620/
|
| 69 |
-
ep0 step 630/
|
| 70 |
-
ep0 step 640/
|
| 71 |
-
ep0 step 650/
|
| 72 |
-
ep0 step 660/
|
| 73 |
-
ep0 step 670/
|
| 74 |
-
ep0 step 680/
|
| 75 |
-
ep0 step 690/
|
| 76 |
-
ep0 step 700/
|
| 77 |
-
ep0 step 710/
|
| 78 |
-
ep0 step 720/
|
| 79 |
-
ep0 step 730/
|
| 80 |
-
ep0 step 740/
|
| 81 |
-
ep0 step 750/
|
| 82 |
-
ep0 step 760/
|
| 83 |
-
ep0 step 770/
|
| 84 |
-
ep0 step 780/
|
| 85 |
-
ep0 step 790/
|
| 86 |
-
ep0 step 800/
|
| 87 |
-
ep0 step 810/
|
| 88 |
-
ep0 step 820/
|
| 89 |
-
ep0 step 830/
|
| 90 |
-
ep0 step 840/
|
| 91 |
-
ep0 step 850/
|
| 92 |
-
ep0 step 860/
|
| 93 |
-
ep0 step 870/
|
| 94 |
-
ep0 step 880/
|
| 95 |
-
ep0 step 890/
|
| 96 |
-
ep0 step 900/
|
| 97 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
|
| 2 |
+
delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/957b91e762e883935830246eeb02381f9d2694b6: 496 adapter tensors and the pointer head loaded
|
| 3 |
device=cuda trainable params=33.8M
|
| 4 |
+
replay: 4000 of 12576 suite training records mixed with 11320 from evals/round10/skills/train.jsonl
|
| 5 |
+
15320 training requests (holdout=[]), questions by type {'choice': 10917, 'noul': 8981, 'score': 1377}
|
| 6 |
[transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
|
| 7 |
+
ep0 step 10/1915 loss 1.451 kl 0.000 anchor 0.000 0.441s/rec
|
| 8 |
+
ep0 step 20/1915 loss 1.678 kl 0.000 anchor 0.000 0.340s/rec
|
| 9 |
+
ep0 step 30/1915 loss 1.520 kl 0.000 anchor 0.000 0.308s/rec
|
| 10 |
+
ep0 step 40/1915 loss 1.340 kl 0.000 anchor 0.000 0.343s/rec
|
| 11 |
+
ep0 step 50/1915 loss 1.612 kl 0.000 anchor 0.000 0.312s/rec
|
| 12 |
+
ep0 step 60/1915 loss 0.875 kl 0.000 anchor 0.000 0.304s/rec
|
| 13 |
+
ep0 step 70/1915 loss 0.809 kl 0.000 anchor 0.000 0.295s/rec
|
| 14 |
+
ep0 step 80/1915 loss 1.151 kl 0.000 anchor 0.000 0.295s/rec
|
| 15 |
+
ep0 step 90/1915 loss 0.824 kl 0.000 anchor 0.000 0.292s/rec
|
| 16 |
+
ep0 step 100/1915 loss 0.846 kl 0.000 anchor 0.000 0.311s/rec
|
| 17 |
+
ep0 step 110/1915 loss 0.823 kl 0.000 anchor 0.000 0.300s/rec
|
| 18 |
+
ep0 step 120/1915 loss 0.807 kl 0.000 anchor 0.000 0.299s/rec
|
| 19 |
+
ep0 step 130/1915 loss 0.850 kl 0.000 anchor 0.000 0.296s/rec
|
| 20 |
+
ep0 step 140/1915 loss 0.655 kl 0.000 anchor 0.000 0.293s/rec
|
| 21 |
+
ep0 step 150/1915 loss 0.886 kl 0.000 anchor 0.000 0.291s/rec
|
| 22 |
+
ep0 step 160/1915 loss 0.604 kl 0.000 anchor 0.000 0.288s/rec
|
| 23 |
+
ep0 step 170/1915 loss 0.811 kl 0.000 anchor 0.000 0.290s/rec
|
| 24 |
+
ep0 step 180/1915 loss 0.616 kl 0.000 anchor 0.000 0.287s/rec
|
| 25 |
+
ep0 step 190/1915 loss 0.784 kl 0.000 anchor 0.000 0.284s/rec
|
| 26 |
+
ep0 step 200/1915 loss 0.731 kl 0.000 anchor 0.000 0.293s/rec
|
| 27 |
+
ep0 step 210/1915 loss 0.741 kl 0.000 anchor 0.000 0.288s/rec
|
| 28 |
+
ep0 step 220/1915 loss 0.638 kl 0.000 anchor 0.000 0.285s/rec
|
| 29 |
+
ep0 step 230/1915 loss 0.950 kl 0.000 anchor 0.000 0.284s/rec
|
| 30 |
+
ep0 step 240/1915 loss 0.637 kl 0.000 anchor 0.000 0.283s/rec
|
| 31 |
+
ep0 step 250/1915 loss 0.727 kl 0.000 anchor 0.000 0.280s/rec
|
| 32 |
+
ep0 step 260/1915 loss 0.724 kl 0.000 anchor 0.000 0.279s/rec
|
| 33 |
+
ep0 step 270/1915 loss 0.598 kl 0.000 anchor 0.000 0.281s/rec
|
| 34 |
+
ep0 step 280/1915 loss 0.507 kl 0.000 anchor 0.000 0.279s/rec
|
| 35 |
+
ep0 step 290/1915 loss 0.581 kl 0.000 anchor 0.000 0.276s/rec
|
| 36 |
+
ep0 step 300/1915 loss 0.649 kl 0.000 anchor 0.000 0.277s/rec
|
| 37 |
+
ep0 step 310/1915 loss 0.608 kl 0.000 anchor 0.000 0.280s/rec
|
| 38 |
+
ep0 step 320/1915 loss 0.773 kl 0.000 anchor 0.000 0.279s/rec
|
| 39 |
+
ep0 step 330/1915 loss 0.695 kl 0.000 anchor 0.000 0.277s/rec
|
| 40 |
+
ep0 step 340/1915 loss 0.689 kl 0.000 anchor 0.000 0.277s/rec
|
| 41 |
+
ep0 step 350/1915 loss 0.615 kl 0.000 anchor 0.000 0.276s/rec
|
| 42 |
+
ep0 step 360/1915 loss 0.618 kl 0.000 anchor 0.000 0.274s/rec
|
| 43 |
+
ep0 step 370/1915 loss 0.754 kl 0.000 anchor 0.000 0.274s/rec
|
| 44 |
+
ep0 step 380/1915 loss 0.514 kl 0.000 anchor 0.000 0.275s/rec
|
| 45 |
+
ep0 step 390/1915 loss 0.569 kl 0.000 anchor 0.000 0.273s/rec
|
| 46 |
+
ep0 step 400/1915 loss 0.574 kl 0.000 anchor 0.000 0.273s/rec
|
| 47 |
+
ep0 step 410/1915 loss 0.591 kl 0.000 anchor 0.000 0.274s/rec
|
| 48 |
+
ep0 step 420/1915 loss 0.690 kl 0.000 anchor 0.000 0.272s/rec
|
| 49 |
+
ep0 step 430/1915 loss 0.643 kl 0.000 anchor 0.000 0.272s/rec
|
| 50 |
+
ep0 step 440/1915 loss 0.616 kl 0.000 anchor 0.000 0.272s/rec
|
| 51 |
+
ep0 step 450/1915 loss 0.595 kl 0.000 anchor 0.000 0.270s/rec
|
| 52 |
+
ep0 step 460/1915 loss 0.643 kl 0.000 anchor 0.000 0.270s/rec
|
| 53 |
+
ep0 step 470/1915 loss 0.646 kl 0.000 anchor 0.000 0.270s/rec
|
| 54 |
+
ep0 step 480/1915 loss 0.608 kl 0.000 anchor 0.000 0.269s/rec
|
| 55 |
+
ep0 step 490/1915 loss 0.488 kl 0.000 anchor 0.000 0.268s/rec
|
| 56 |
+
ep0 step 500/1915 loss 0.586 kl 0.000 anchor 0.000 0.267s/rec
|
| 57 |
+
ep0 step 510/1915 loss 0.407 kl 0.000 anchor 0.000 0.265s/rec
|
| 58 |
+
ep0 step 520/1915 loss 0.483 kl 0.000 anchor 0.000 0.264s/rec
|
| 59 |
+
ep0 step 530/1915 loss 0.484 kl 0.000 anchor 0.000 0.264s/rec
|
| 60 |
+
ep0 step 540/1915 loss 0.771 kl 0.000 anchor 0.000 0.265s/rec
|
| 61 |
+
ep0 step 550/1915 loss 0.588 kl 0.000 anchor 0.000 0.264s/rec
|
| 62 |
+
ep0 step 560/1915 loss 0.443 kl 0.000 anchor 0.000 0.264s/rec
|
| 63 |
+
ep0 step 570/1915 loss 0.489 kl 0.000 anchor 0.000 0.263s/rec
|
| 64 |
+
ep0 step 580/1915 loss 0.473 kl 0.000 anchor 0.000 0.262s/rec
|
| 65 |
+
ep0 step 590/1915 loss 0.718 kl 0.000 anchor 0.000 0.262s/rec
|
| 66 |
+
ep0 step 600/1915 loss 0.634 kl 0.000 anchor 0.000 0.263s/rec
|
| 67 |
+
ep0 step 610/1915 loss 0.480 kl 0.000 anchor 0.000 0.262s/rec
|
| 68 |
+
ep0 step 620/1915 loss 0.579 kl 0.000 anchor 0.000 0.261s/rec
|
| 69 |
+
ep0 step 630/1915 loss 0.386 kl 0.000 anchor 0.000 0.261s/rec
|
| 70 |
+
ep0 step 640/1915 loss 0.660 kl 0.000 anchor 0.000 0.260s/rec
|
| 71 |
+
ep0 step 650/1915 loss 0.761 kl 0.000 anchor 0.000 0.259s/rec
|
| 72 |
+
ep0 step 660/1915 loss 0.703 kl 0.000 anchor 0.000 0.259s/rec
|
| 73 |
+
ep0 step 670/1915 loss 0.393 kl 0.000 anchor 0.000 0.259s/rec
|
| 74 |
+
ep0 step 680/1915 loss 0.502 kl 0.000 anchor 0.000 0.258s/rec
|
| 75 |
+
ep0 step 690/1915 loss 0.718 kl 0.000 anchor 0.000 0.258s/rec
|
| 76 |
+
ep0 step 700/1915 loss 0.582 kl 0.000 anchor 0.000 0.258s/rec
|
| 77 |
+
ep0 step 710/1915 loss 0.458 kl 0.000 anchor 0.000 0.258s/rec
|
| 78 |
+
ep0 step 720/1915 loss 0.496 kl 0.000 anchor 0.000 0.257s/rec
|
| 79 |
+
ep0 step 730/1915 loss 0.458 kl 0.000 anchor 0.000 0.257s/rec
|
| 80 |
+
ep0 step 740/1915 loss 0.533 kl 0.000 anchor 0.000 0.257s/rec
|
| 81 |
+
ep0 step 750/1915 loss 0.559 kl 0.000 anchor 0.000 0.257s/rec
|
| 82 |
+
ep0 step 760/1915 loss 0.540 kl 0.000 anchor 0.000 0.256s/rec
|
| 83 |
+
ep0 step 770/1915 loss 0.455 kl 0.000 anchor 0.000 0.256s/rec
|
| 84 |
+
ep0 step 780/1915 loss 0.595 kl 0.000 anchor 0.000 0.255s/rec
|
| 85 |
+
ep0 step 790/1915 loss 0.337 kl 0.000 anchor 0.000 0.255s/rec
|
| 86 |
+
ep0 step 800/1915 loss 0.603 kl 0.000 anchor 0.000 0.255s/rec
|
| 87 |
+
ep0 step 810/1915 loss 0.524 kl 0.000 anchor 0.000 0.254s/rec
|
| 88 |
+
ep0 step 820/1915 loss 0.654 kl 0.000 anchor 0.000 0.254s/rec
|
| 89 |
+
ep0 step 830/1915 loss 0.543 kl 0.000 anchor 0.000 0.254s/rec
|
| 90 |
+
ep0 step 840/1915 loss 0.531 kl 0.000 anchor 0.000 0.254s/rec
|
| 91 |
+
ep0 step 850/1915 loss 0.569 kl 0.000 anchor 0.000 0.253s/rec
|
| 92 |
+
ep0 step 860/1915 loss 0.371 kl 0.000 anchor 0.000 0.253s/rec
|
| 93 |
+
ep0 step 870/1915 loss 0.442 kl 0.000 anchor 0.000 0.253s/rec
|
| 94 |
+
ep0 step 880/1915 loss 0.486 kl 0.000 anchor 0.000 0.253s/rec
|
| 95 |
+
ep0 step 890/1915 loss 0.513 kl 0.000 anchor 0.000 0.252s/rec
|
| 96 |
+
ep0 step 900/1915 loss 0.537 kl 0.000 anchor 0.000 0.252s/rec
|
| 97 |
+
ep0 step 910/1915 loss 0.571 kl 0.000 anchor 0.000 0.252s/rec
|
| 98 |
+
ep0 step 920/1915 loss 0.574 kl 0.000 anchor 0.000 0.252s/rec
|
| 99 |
+
ep0 step 930/1915 loss 0.512 kl 0.000 anchor 0.000 0.252s/rec
|
| 100 |
+
ep0 step 940/1915 loss 0.634 kl 0.000 anchor 0.000 0.251s/rec
|
| 101 |
+
ep0 step 950/1915 loss 0.476 kl 0.000 anchor 0.000 0.252s/rec
|
| 102 |
+
ep0 step 960/1915 loss 0.529 kl 0.000 anchor 0.000 0.251s/rec
|
| 103 |
+
ep0 step 970/1915 loss 0.621 kl 0.000 anchor 0.000 0.251s/rec
|
| 104 |
+
ep0 step 980/1915 loss 0.529 kl 0.000 anchor 0.000 0.251s/rec
|
| 105 |
+
ep0 step 990/1915 loss 0.565 kl 0.000 anchor 0.000 0.250s/rec
|
| 106 |
+
ep0 step 1000/1915 loss 0.638 kl 0.000 anchor 0.000 0.250s/rec
|
| 107 |
+
ep0 step 1010/1915 loss 0.600 kl 0.000 anchor 0.000 0.250s/rec
|
| 108 |
+
ep0 step 1020/1915 loss 0.581 kl 0.000 anchor 0.000 0.250s/rec
|
| 109 |
+
ep0 step 1030/1915 loss 0.594 kl 0.000 anchor 0.000 0.250s/rec
|
| 110 |
+
ep0 step 1040/1915 loss 0.539 kl 0.000 anchor 0.000 0.250s/rec
|
| 111 |
+
ep0 step 1050/1915 loss 0.547 kl 0.000 anchor 0.000 0.250s/rec
|
| 112 |
+
ep0 step 1060/1915 loss 0.714 kl 0.000 anchor 0.000 0.250s/rec
|
| 113 |
+
ep0 step 1070/1915 loss 0.378 kl 0.000 anchor 0.000 0.249s/rec
|
| 114 |
+
ep0 step 1080/1915 loss 0.657 kl 0.000 anchor 0.000 0.251s/rec
|
| 115 |
+
ep0 step 1090/1915 loss 0.463 kl 0.000 anchor 0.000 0.251s/rec
|
| 116 |
+
ep0 step 1100/1915 loss 0.585 kl 0.000 anchor 0.000 0.252s/rec
|
| 117 |
+
ep0 step 1110/1915 loss 0.312 kl 0.000 anchor 0.000 0.252s/rec
|
| 118 |
+
ep0 step 1120/1915 loss 0.539 kl 0.000 anchor 0.000 0.252s/rec
|
| 119 |
+
ep0 step 1130/1915 loss 0.541 kl 0.000 anchor 0.000 0.252s/rec
|
| 120 |
+
ep0 step 1140/1915 loss 0.784 kl 0.000 anchor 0.000 0.252s/rec
|
| 121 |
+
ep0 step 1150/1915 loss 0.558 kl 0.000 anchor 0.000 0.252s/rec
|
| 122 |
+
ep0 step 1160/1915 loss 0.600 kl 0.000 anchor 0.000 0.252s/rec
|
| 123 |
+
ep0 step 1170/1915 loss 0.562 kl 0.000 anchor 0.000 0.252s/rec
|
| 124 |
+
ep0 step 1180/1915 loss 0.723 kl 0.000 anchor 0.000 0.252s/rec
|
| 125 |
+
ep0 step 1190/1915 loss 0.394 kl 0.000 anchor 0.000 0.254s/rec
|
| 126 |
+
ep0 step 1200/1915 loss 0.394 kl 0.000 anchor 0.000 0.254s/rec
|
| 127 |
+
ep0 step 1210/1915 loss 0.331 kl 0.000 anchor 0.000 0.254s/rec
|
| 128 |
+
ep0 step 1220/1915 loss 0.334 kl 0.000 anchor 0.000 0.253s/rec
|
| 129 |
+
ep0 step 1230/1915 loss 0.447 kl 0.000 anchor 0.000 0.253s/rec
|
| 130 |
+
ep0 step 1240/1915 loss 0.551 kl 0.000 anchor 0.000 0.253s/rec
|
| 131 |
+
ep0 step 1250/1915 loss 0.509 kl 0.000 anchor 0.000 0.252s/rec
|
| 132 |
+
ep0 step 1260/1915 loss 0.694 kl 0.000 anchor 0.000 0.253s/rec
|
| 133 |
+
ep0 step 1270/1915 loss 0.532 kl 0.000 anchor 0.000 0.252s/rec
|
| 134 |
+
ep0 step 1280/1915 loss 0.496 kl 0.000 anchor 0.000 0.252s/rec
|
| 135 |
+
ep0 step 1290/1915 loss 0.556 kl 0.000 anchor 0.000 0.252s/rec
|
| 136 |
+
ep0 step 1300/1915 loss 0.403 kl 0.000 anchor 0.000 0.252s/rec
|
| 137 |
+
ep0 step 1310/1915 loss 0.537 kl 0.000 anchor 0.000 0.252s/rec
|
| 138 |
+
ep0 step 1320/1915 loss 0.517 kl 0.000 anchor 0.000 0.252s/rec
|
| 139 |
+
ep0 step 1330/1915 loss 0.633 kl 0.000 anchor 0.000 0.252s/rec
|
| 140 |
+
ep0 step 1340/1915 loss 0.491 kl 0.000 anchor 0.000 0.252s/rec
|
| 141 |
+
ep0 step 1350/1915 loss 0.512 kl 0.000 anchor 0.000 0.253s/rec
|
| 142 |
+
ep0 step 1360/1915 loss 0.398 kl 0.000 anchor 0.000 0.253s/rec
|
| 143 |
+
ep0 step 1370/1915 loss 0.396 kl 0.000 anchor 0.000 0.253s/rec
|
| 144 |
+
ep0 step 1380/1915 loss 0.484 kl 0.000 anchor 0.000 0.253s/rec
|
| 145 |
+
ep0 step 1390/1915 loss 0.492 kl 0.000 anchor 0.000 0.253s/rec
|
| 146 |
+
ep0 step 1400/1915 loss 0.487 kl 0.000 anchor 0.000 0.253s/rec
|
| 147 |
+
ep0 step 1410/1915 loss 0.410 kl 0.000 anchor 0.000 0.253s/rec
|
| 148 |
+
ep0 step 1420/1915 loss 0.405 kl 0.000 anchor 0.000 0.253s/rec
|
| 149 |
+
ep0 step 1430/1915 loss 0.391 kl 0.000 anchor 0.000 0.253s/rec
|
| 150 |
+
ep0 step 1440/1915 loss 0.408 kl 0.000 anchor 0.000 0.253s/rec
|
| 151 |
+
ep0 step 1450/1915 loss 0.394 kl 0.000 anchor 0.000 0.253s/rec
|
| 152 |
+
ep0 step 1460/1915 loss 0.461 kl 0.000 anchor 0.000 0.253s/rec
|
| 153 |
+
ep0 step 1470/1915 loss 0.485 kl 0.000 anchor 0.000 0.253s/rec
|
| 154 |
+
ep0 step 1480/1915 loss 0.480 kl 0.000 anchor 0.000 0.252s/rec
|
| 155 |
+
ep0 step 1490/1915 loss 0.539 kl 0.000 anchor 0.000 0.252s/rec
|
| 156 |
+
ep0 step 1500/1915 loss 0.541 kl 0.000 anchor 0.000 0.254s/rec
|
| 157 |
+
ep0 step 1510/1915 loss 0.554 kl 0.000 anchor 0.000 0.254s/rec
|
| 158 |
+
ep0 step 1520/1915 loss 0.513 kl 0.000 anchor 0.000 0.254s/rec
|
| 159 |
+
ep0 step 1530/1915 loss 0.423 kl 0.000 anchor 0.000 0.254s/rec
|
| 160 |
+
ep0 step 1540/1915 loss 0.471 kl 0.000 anchor 0.000 0.254s/rec
|
| 161 |
+
ep0 step 1550/1915 loss 0.342 kl 0.000 anchor 0.000 0.254s/rec
|
| 162 |
+
ep0 step 1560/1915 loss 0.547 kl 0.000 anchor 0.000 0.253s/rec
|
| 163 |
+
ep0 step 1570/1915 loss 0.472 kl 0.000 anchor 0.000 0.253s/rec
|
| 164 |
+
ep0 step 1580/1915 loss 0.477 kl 0.000 anchor 0.000 0.253s/rec
|
| 165 |
+
ep0 step 1590/1915 loss 0.593 kl 0.000 anchor 0.000 0.253s/rec
|
| 166 |
+
ep0 step 1600/1915 loss 0.995 kl 0.000 anchor 0.000 0.253s/rec
|
| 167 |
+
ep0 step 1610/1915 loss 0.459 kl 0.000 anchor 0.000 0.253s/rec
|
| 168 |
+
ep0 step 1620/1915 loss 0.609 kl 0.000 anchor 0.000 0.253s/rec
|
| 169 |
+
ep0 step 1630/1915 loss 0.386 kl 0.000 anchor 0.000 0.253s/rec
|
| 170 |
+
ep0 step 1640/1915 loss 0.555 kl 0.000 anchor 0.000 0.253s/rec
|
| 171 |
+
ep0 step 1650/1915 loss 0.517 kl 0.000 anchor 0.000 0.253s/rec
|
| 172 |
+
ep0 step 1660/1915 loss 0.390 kl 0.000 anchor 0.000 0.252s/rec
|
| 173 |
+
ep0 step 1670/1915 loss 0.503 kl 0.000 anchor 0.000 0.253s/rec
|
| 174 |
+
ep0 step 1680/1915 loss 0.870 kl 0.000 anchor 0.000 0.252s/rec
|
| 175 |
+
ep0 step 1690/1915 loss 0.419 kl 0.000 anchor 0.000 0.253s/rec
|
| 176 |
+
ep0 step 1700/1915 loss 0.429 kl 0.000 anchor 0.000 0.253s/rec
|
| 177 |
+
ep0 step 1710/1915 loss 0.552 kl 0.000 anchor 0.000 0.253s/rec
|
| 178 |
+
ep0 step 1720/1915 loss 0.504 kl 0.000 anchor 0.000 0.253s/rec
|
| 179 |
+
ep0 step 1730/1915 loss 0.402 kl 0.000 anchor 0.000 0.253s/rec
|
| 180 |
+
ep0 step 1740/1915 loss 0.383 kl 0.000 anchor 0.000 0.253s/rec
|
| 181 |
+
ep0 step 1750/1915 loss 0.631 kl 0.000 anchor 0.000 0.253s/rec
|
| 182 |
+
ep0 step 1760/1915 loss 0.618 kl 0.000 anchor 0.000 0.252s/rec
|
| 183 |
+
ep0 step 1770/1915 loss 0.335 kl 0.000 anchor 0.000 0.253s/rec
|
| 184 |
+
ep0 step 1780/1915 loss 0.250 kl 0.000 anchor 0.000 0.253s/rec
|
| 185 |
+
ep0 step 1790/1915 loss 0.499 kl 0.000 anchor 0.000 0.253s/rec
|
| 186 |
+
ep0 step 1800/1915 loss 0.351 kl 0.000 anchor 0.000 0.252s/rec
|
| 187 |
+
ep0 step 1810/1915 loss 0.613 kl 0.000 anchor 0.000 0.252s/rec
|
| 188 |
+
ep0 step 1820/1915 loss 0.355 kl 0.000 anchor 0.000 0.252s/rec
|
| 189 |
+
ep0 step 1830/1915 loss 0.556 kl 0.000 anchor 0.000 0.252s/rec
|
| 190 |
+
ep0 step 1840/1915 loss 0.328 kl 0.000 anchor 0.000 0.252s/rec
|
| 191 |
+
ep0 step 1850/1915 loss 0.863 kl 0.000 anchor 0.000 0.252s/rec
|
| 192 |
+
ep0 step 1860/1915 loss 0.441 kl 0.000 anchor 0.000 0.252s/rec
|
| 193 |
+
ep0 step 1870/1915 loss 0.385 kl 0.000 anchor 0.000 0.253s/rec
|
| 194 |
+
ep0 step 1880/1915 loss 0.398 kl 0.000 anchor 0.000 0.252s/rec
|
| 195 |
+
ep0 step 1890/1915 loss 0.507 kl 0.000 anchor 0.000 0.252s/rec
|
| 196 |
+
ep0 step 1900/1915 loss 0.602 kl 0.000 anchor 0.000 0.252s/rec
|
| 197 |
+
ep0 step 1910/1915 loss 0.570 kl 0.000 anchor 0.000 0.251s/rec
|
| 198 |
+
saved /runs/r10-skills/00-trial-0/checkpoint
|
training_config.json
CHANGED
|
@@ -37,20 +37,20 @@
|
|
| 37 |
"anchor": "",
|
| 38 |
"anchor_w": 0.0,
|
| 39 |
"anchor_sources": "",
|
| 40 |
-
"out": "/runs/
|
| 41 |
-
"data": "evals/
|
| 42 |
"max_state": 7552,
|
| 43 |
-
"replay":
|
| 44 |
-
"init_from": "jaredpalmer/kev-4b",
|
| 45 |
-
"seed":
|
| 46 |
},
|
| 47 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 48 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 49 |
"init_source": {
|
| 50 |
-
"init_from": "jaredpalmer/kev-4b",
|
| 51 |
-
"resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/
|
| 52 |
-
"adapter_sha256": "
|
| 53 |
-
"head_sha256": "
|
| 54 |
"adapter_tensors": 496
|
| 55 |
},
|
| 56 |
"ordinal_objective": "ranked_probability_score",
|
|
|
|
| 37 |
"anchor": "",
|
| 38 |
"anchor_w": 0.0,
|
| 39 |
"anchor_sources": "",
|
| 40 |
+
"out": "/runs/r10-skills/00-trial-0/checkpoint",
|
| 41 |
+
"data": "evals/round10/skills/train.jsonl",
|
| 42 |
"max_state": 7552,
|
| 43 |
+
"replay": 4000,
|
| 44 |
+
"init_from": "jaredpalmer/kev-4b@957b91e762e883935830246eeb02381f9d2694b6",
|
| 45 |
+
"seed": 1
|
| 46 |
},
|
| 47 |
"suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
|
| 48 |
"base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
|
| 49 |
"init_source": {
|
| 50 |
+
"init_from": "jaredpalmer/kev-4b@957b91e762e883935830246eeb02381f9d2694b6",
|
| 51 |
+
"resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/957b91e762e883935830246eeb02381f9d2694b6",
|
| 52 |
+
"adapter_sha256": "c2c99077b58cd4a604304b743b7db5391187ef1d58e575f90c554c744aa23d65",
|
| 53 |
+
"head_sha256": "9a3051c3676f0be29d885935b6f50c287f7acf04e3cb29a6fed82ca3e5540b7e",
|
| 54 |
"adapter_tensors": 496
|
| 55 |
},
|
| 56 |
"ordinal_objective": "ranked_probability_score",
|
training_metrics.json
CHANGED
|
@@ -1,14 +1,14 @@
|
|
| 1 |
{
|
| 2 |
-
"wall_seconds":
|
| 3 |
-
"records_seen":
|
| 4 |
-
"requested_records":
|
| 5 |
"truncated_records": 0,
|
| 6 |
"rejected_records": 0,
|
| 7 |
-
"optimizer_steps":
|
| 8 |
-
"forward_tokens":
|
| 9 |
-
"peak_device_bytes":
|
| 10 |
"device": "cuda",
|
| 11 |
"dtype": "bf16",
|
| 12 |
"batch": 2,
|
| 13 |
-
"peak_rss_bytes":
|
| 14 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"wall_seconds": 5031.164765119553,
|
| 3 |
+
"records_seen": 19994,
|
| 4 |
+
"requested_records": 15320,
|
| 5 |
"truncated_records": 0,
|
| 6 |
"rejected_records": 0,
|
| 7 |
+
"optimizer_steps": 1915,
|
| 8 |
+
"forward_tokens": 8738926,
|
| 9 |
+
"peak_device_bytes": 47737451520,
|
| 10 |
"device": "cuda",
|
| 11 |
"dtype": "bf16",
|
| 12 |
"batch": 2,
|
| 13 |
+
"peak_rss_bytes": 31552528384
|
| 14 |
}
|