Publish quantization/data reproduction source and correct recipe documentation
Browse filesRecovered source, runbooks, recipe/data/environment provenance, CPU verification and path relocation. Correct ARVQ-v2 corpus, objective, schedule, boundary weighting and FP16 reference decode; preserve model weights and original layer reports. Explicitly document historical reproduction gaps.
This view is limited to 50 files because it contains too many changes. See raw diff
- METHODOLOGY.md +4 -339
- README.md +16 -22
- arvq88_reference.py +87 -9
- calibration_corpus.json +16 -23
- pv_progress.json +18 -101
- reproduce/AUDIT.md +13 -0
- reproduce/README.md +88 -0
- reproduce/RUNBOOK.md +29 -0
- reproduce/THIRD_PARTY.md +1 -0
- reproduce/checkpoint_before_audit.json +100 -0
- reproduce/data_artifacts.json +333 -0
- reproduce/environment_observed.json +16 -0
- reproduce/external_sources.json +18 -0
- reproduce/historical_metadata/METHODOLOGY.md +342 -0
- reproduce/historical_metadata/backbone_sources.json +0 -0
- reproduce/historical_metadata/build_provenance.json +178 -0
- reproduce/historical_metadata/calibration_corpus.json +24 -0
- reproduce/historical_metadata/cold_assignment.json +0 -0
- reproduce/historical_metadata/pv_progress.json +187 -0
- reproduce/historical_metadata/source.json +14 -0
- reproduce/layer_recipe_summary.json +1727 -0
- reproduce/package_manifest.json +247 -0
- reproduce/prepare_workspace.py +26 -0
- reproduce/recipe.json +116 -0
- reproduce/requirements.in +10 -0
- reproduce/source/btx53/arvq88/__init__.py +1 -0
- reproduce/source/btx53/arvq88/activation.py +44 -0
- reproduce/source/btx53/arvq88/alternating.py +64 -0
- reproduce/source/btx53/arvq88/arvq_reap.py +174 -0
- reproduce/source/btx53/arvq88/backbone.py +44 -0
- reproduce/source/btx53/arvq88/checkpoint.py +168 -0
- reproduce/source/btx53/arvq88/cli.py +154 -0
- reproduce/source/btx53/arvq88/docs/expert_experiments.md +76 -0
- reproduce/source/btx53/arvq88/docs/gradient_pv_recipe.md +180 -0
- reproduce/source/btx53/arvq88/docs/per_expert_handoff.md +78 -0
- reproduce/source/btx53/arvq88/docs/sequential_pilot.md +65 -0
- reproduce/source/btx53/arvq88/encoder.py +321 -0
- reproduce/source/btx53/arvq88/experiment.py +129 -0
- reproduce/source/btx53/arvq88/final_publish.py +69 -0
- reproduce/source/btx53/arvq88/fit.py +48 -0
- reproduce/source/btx53/arvq88/gate_worker.py +5 -0
- reproduce/source/btx53/arvq88/gradient_indices.py +120 -0
- reproduce/source/btx53/arvq88/incremental_publish.py +122 -0
- reproduce/source/btx53/arvq88/inputs.py +198 -0
- reproduce/source/btx53/arvq88/jobs.py +28 -0
- reproduce/source/btx53/arvq88/local_source.py +43 -0
- reproduce/source/btx53/arvq88/pack.py +142 -0
- reproduce/source/btx53/arvq88/perf/FP16_BLOCK_SCALE_HANDOFF.md +19 -0
- reproduce/source/btx53/arvq88/perf/FULL_CORPUS_PIPELINE.md +80 -0
- reproduce/source/btx53/arvq88/perf/README.md +145 -0
METHODOLOGY.md
CHANGED
|
@@ -1,342 +1,7 @@
|
|
| 1 |
-
# ARVQ
|
| 2 |
|
| 3 |
-
|
| 4 |
-
(`GLM-5.3-Vision-NVFP4-ARVQ-hybrid`) were quantized and tuned. This document
|
| 5 |
-
reconstructs the actual pipeline and tuning decisions from the build transcripts
|
| 6 |
-
and the fitting/serving code; numbers are cited from those sources. It is a
|
| 7 |
-
methodology record, not a quality claim: full-model task quality and native
|
| 8 |
-
SM120 execution were **not** evaluated for this checkpoint.
|
| 9 |
|
| 10 |
-
---
|
| 11 |
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
GLM-5.3 is a large MoE. Each MoE block routes each token to its top-8 of 256
|
| 15 |
-
routed experts (plus a shared expert). This checkpoint is a **hybrid** in which
|
| 16 |
-
each expert is stored in one of two ways:
|
| 17 |
-
|
| 18 |
-
- **Hot experts** — kept in the NVFP4 donor format (unchanged).
|
| 19 |
-
- **Cold experts** — re-quantized to ~2 bits with **ARVQ** (additive residual
|
| 20 |
-
vector quantization), the subject of this document.
|
| 21 |
-
|
| 22 |
-
The hot/cold split is fixed by a REAP allocation that keeps the **5,750** most
|
| 23 |
-
important experts hot; the remaining **13,450** experts across the model are
|
| 24 |
-
cold (per the 09-14 build transcript, which verified "13,450 cold-expert source
|
| 25 |
-
files"). The allocation is 75% text-REAP / 25% multimodal-salience weighted and
|
| 26 |
-
is **not** modified by this campaign.
|
| 27 |
-
|
| 28 |
-
Only the **75 MoE blocks, layers 3–77**, are touched. Layers 0–2 are dense and
|
| 29 |
-
stay frozen. Everything except the cold experts is inherited unchanged:
|
| 30 |
-
attention, backbone, shared experts, hot NVFP4 experts, BF16 MTP, and the
|
| 31 |
-
vision components.
|
| 32 |
-
|
| 33 |
-
Cold experts are quantized per (layer, projection). The two projections are the
|
| 34 |
-
fused gate/up `w13` (N=4096, K=6144) and the down projection `w2` (N=6144,
|
| 35 |
-
K=2048).
|
| 36 |
-
|
| 37 |
-
---
|
| 38 |
-
|
| 39 |
-
## 2. ARVQ representation and bit budget
|
| 40 |
-
|
| 41 |
-
Each cold expert weight group of **8 contiguous columns** is represented as the
|
| 42 |
-
sum of two codebook atoms times a scale:
|
| 43 |
-
|
| 44 |
-
```
|
| 45 |
-
w_group = global * block_scale * (c0[a] + c1[b])
|
| 46 |
-
```
|
| 47 |
-
|
| 48 |
-
- **Two codebooks** `c0`, `c1`, each **256 entries × 8 dims** (512 codewords per
|
| 49 |
-
pair). This is the "rvq256_256x8" format. Codebooks are **per-expert** in this
|
| 50 |
-
v3 checkpoint (`codebook_scope = expert`, packed format
|
| 51 |
-
`rvq256_256x8_expert`, version 3).
|
| 52 |
-
- **Indices** `a`, `b` are one uint8 each per 8-weight group → **16 index bits
|
| 53 |
-
per group = 2.0 bits/weight**.
|
| 54 |
-
- **Codebook atoms are constrained to the FP4 grid** `{0, ±0.5, ±1, ±1.5, ±2,
|
| 55 |
-
±3, ±4, ±6}` (project-to-FP4 after every update). The FP4 grid on the atoms is
|
| 56 |
-
why no incoherence/Hadamard rotation is used: the target space cannot be
|
| 57 |
-
rescaled without leaving the grid (see `encoder.py` docstring).
|
| 58 |
-
- **Block scales** are stored as **`float8_e4m3fn`**, one scale per **[row,
|
| 59 |
-
128-column block]** (48 blocks/row for `w13`, 16 blocks/row for `w2`). At use
|
| 60 |
-
time each block scale is `repeat_interleave(16)` to cover the 16 eight-wide
|
| 61 |
-
groups in a 128-column block. Adding the FP8 scale (8 bits / 128 weights =
|
| 62 |
-
0.0625 bit/weight) gives the reported **2.0625 bits/weight** plus the
|
| 63 |
-
amortized per-expert codebooks.
|
| 64 |
-
- A single fp32 `global` scalar per (layer, projection) sits above the block
|
| 65 |
-
scales so the index/scale layout is unchanged from earlier versions.
|
| 66 |
-
|
| 67 |
-
Packing (`pack.py`) writes MMA-fragment index tiles (64 uint32 words/tile) and
|
| 68 |
-
verifies, per layer, that every expert's indices round-trip exactly and that a
|
| 69 |
-
sample of decoded weights matches bit-for-bit before publishing.
|
| 70 |
-
|
| 71 |
-
---
|
| 72 |
-
|
| 73 |
-
## 3. Initial fit (Hessian-aware, no PV yet)
|
| 74 |
-
|
| 75 |
-
Before any gradient tuning, each (layer, projection) gets an
|
| 76 |
-
activation-Hessian-tuned initialization (`encoder.py`, `fit.py`):
|
| 77 |
-
|
| 78 |
-
1. **Per-expert raw-basis Hessian.** `w13` uses `H = E[x xᵀ]` over routed
|
| 79 |
-
tokens; `w2` uses `H = E[m mᵀ]` where `m = silu(x·Wgᵀ)·(x·Wuᵀ)` is the BF16
|
| 80 |
-
SwiGLU intermediate. Experts with <32 routed tokens fall back to all captured
|
| 81 |
-
tokens. Activations are stored/computed in **float32**; Hessians are half.
|
| 82 |
-
2. **Escalating-damp Cholesky** of `H⁻¹`: damping starts at `1e-3 · mean(diag H)`
|
| 83 |
-
and multiplies by 4× (up to 12 attempts) until the factorization succeeds.
|
| 84 |
-
3. **Codebook EM** on a Hessian-importance-weighted subsample (20,000 groups/
|
| 85 |
-
expert) of normalized 8-dim groups: k-means warm start → 8 alternating
|
| 86 |
-
refine iterations, FP4-projecting the atoms after every M-step.
|
| 87 |
-
4. **Per-expert LDLQ/GPTQ error-feedback column sweep** with the fixed
|
| 88 |
-
codebooks (col_block=128, 2 sweep passes, 1 inner refine): plain-L2 inner
|
| 89 |
-
assignment, off-diagonal Hessian carries cross-group error feedback.
|
| 90 |
-
5. **Least-squares refit of the E4M3 block scales** given the chosen codes.
|
| 91 |
-
|
| 92 |
-
An earlier decision, recorded 09-14: a **tuned Hadamard rotation was tested and
|
| 93 |
-
rejected** — a matched 48-expert down-projection refit gave 0.372 (H128 rotated)
|
| 94 |
-
versus **0.191 unrotated**, so the unrotated raw-basis fit was adopted.
|
| 95 |
-
|
| 96 |
-
Nonfinite guards are raised as errors throughout (the fit/tuner refuse to
|
| 97 |
-
proceed on NaN/Inf rather than silently scrubbing).
|
| 98 |
-
|
| 99 |
-
---
|
| 100 |
-
|
| 101 |
-
## 4. The PV objective: `--target reference`
|
| 102 |
-
|
| 103 |
-
The cold experts are then output-tuned per layer. The objective for this
|
| 104 |
-
checkpoint is **`--target reference`** (`sequential_pv_full_corpus.py`).
|
| 105 |
-
|
| 106 |
-
For each captured token position the tuner forms:
|
| 107 |
-
|
| 108 |
-
- `frozen` — the block output from everything that is **not** a cold expert on
|
| 109 |
-
the **student** input (retained tuned upstream layers → attention → shared
|
| 110 |
-
expert → hot NVFP4 experts). This is fixed.
|
| 111 |
-
- `reference` — the **original FP8 reference block output**: the unchanged donor
|
| 112 |
-
backbone with the **original FP8 routed experts** (all 256), evaluated on the
|
| 113 |
-
**original reference trajectory** for that token position.
|
| 114 |
-
- `required = reference − frozen` — the residual the cold experts must supply.
|
| 115 |
-
|
| 116 |
-
The loss is the relative squared error of the cold-expert sum against `required`:
|
| 117 |
-
|
| 118 |
-
```
|
| 119 |
-
loss = || student_cold(x_student) − required ||² / (denominator · Σw)
|
| 120 |
-
```
|
| 121 |
-
|
| 122 |
-
where `denominator` is the per-row mean energy of `reference − frozen` over the
|
| 123 |
-
corpus (so the scale is comparable across layers). Selection metric is
|
| 124 |
-
`reference_rel = ||(frozen + cold) − reference|| / ||reference||`, measured over
|
| 125 |
-
the **whole transformer block including the residual**, not the cold-experts sum
|
| 126 |
-
alone.
|
| 127 |
-
|
| 128 |
-
**Why reference (and its known weakness).** The transcripts weighed
|
| 129 |
-
`reference` against a `same_input` target (match the block's output on the
|
| 130 |
-
*student's own* drifted input). The rationale for reference (09-16): "Matching
|
| 131 |
-
the original reference trajectory is more directly aligned with preserving the
|
| 132 |
-
model's behavior." The acknowledged risk: because each layer receives a
|
| 133 |
-
**different (student) input** than the original, the cold experts may lack the
|
| 134 |
-
capacity to reproduce the reference output and could "learn corrections that
|
| 135 |
-
work on training examples but fail on unseen ones" — i.e. reference-target is
|
| 136 |
-
the more out-of-distribution objective. It was chosen only after a matched
|
| 137 |
-
layer-4 A/B:
|
| 138 |
-
|
| 139 |
-
| Layer-4 metric | reference vs same_input |
|
| 140 |
-
|---|---|
|
| 141 |
-
| validation reference error | **0.21% better** with reference |
|
| 142 |
-
| development-audit error | 0.08% **worse** with reference |
|
| 143 |
-
|
| 144 |
-
A small mixed difference; reference was retained for behavioral alignment. Two
|
| 145 |
-
important design consequences: the target is the **complete block output**
|
| 146 |
-
(including the residual), so matching only the MoE contribution cannot leave the
|
| 147 |
-
incoming residual error untouched; and the loss compares the **routing-weighted
|
| 148 |
-
sum of all cold experts** jointly (gate/up and down together).
|
| 149 |
-
|
| 150 |
-
---
|
| 151 |
-
|
| 152 |
-
## 5. Sequential per-layer pipeline with trajectory coupling
|
| 153 |
-
|
| 154 |
-
Layers are processed **sequentially, 3 → 77** (`full_reference.py` /
|
| 155 |
-
`sequential_capture.py`):
|
| 156 |
-
|
| 157 |
-
1. **Capture** (8-way sequence-parallel). For layer L, the student input is the
|
| 158 |
-
**retained tuned output of block L-1** (`student_outputs.pt`), while the
|
| 159 |
-
reference target re-runs the **donor backbone + original FP8 experts** on the
|
| 160 |
-
original reference hidden state. Both trajectories are carried forward in a
|
| 161 |
-
rolling cache; layer 3 is the control (no preceding ARVQ layer, so student =
|
| 162 |
-
reference input). This cross-layer coupling means each layer is tuned against
|
| 163 |
-
the exact drifted inputs it will see at serving time.
|
| 164 |
-
2. **Fit / PV tune** the cold experts (Section 6).
|
| 165 |
-
3. **Gate** (Section 7).
|
| 166 |
-
4. **Export + serialized replay** — decode to the packed format and require the
|
| 167 |
-
reloaded forward to match the tuned forward to <2e-6 before use.
|
| 168 |
-
5. **Per-layer HF publish** — `publish_reference.py` re-assembles cold slots
|
| 169 |
-
across the 8 expert shards, checks slot coverage/allocation/global metadata,
|
| 170 |
-
runs the exact index round-trip and sampled decoded-weight parity, then
|
| 171 |
-
**atomically replaces that layer's two tensor files + reports in one commit**
|
| 172 |
-
with parent-commit protection; remote hashes are re-verified.
|
| 173 |
-
|
| 174 |
-
Parallelism: cold experts are **sharded across 8 GPUs** (each rank owns distinct
|
| 175 |
-
experts); routed outputs are summed with an all-reduce so a single coupled-output
|
| 176 |
-
objective is optimized without a dense-model replica.
|
| 177 |
-
|
| 178 |
-
---
|
| 179 |
-
|
| 180 |
-
## 6. PV tuning numerics (the published full-corpus campaign)
|
| 181 |
-
|
| 182 |
-
Per the model card, all 75 layers were replaced by a **full-corpus sequential
|
| 183 |
-
PV campaign** drawing sequentially from **18,001,846 training tokens** at
|
| 184 |
-
context **1024** (53,248 legacy tokens were removed around 17 exact matches to
|
| 185 |
-
held-out prompts). Fixed **validation** and **development-audit** sets each hold
|
| 186 |
-
**16,384 tokens**.
|
| 187 |
-
|
| 188 |
-
- **Optimizer:** Adam. Two trainable parameter groups — FP4-constrained
|
| 189 |
-
per-expert **codebooks** and FP8-constrained per-block **scales**. (The
|
| 190 |
-
faster fork trains per-row scale deltas instead; the published full-corpus
|
| 191 |
-
fork trains the per-128-block log-scales.)
|
| 192 |
-
- **Effective batch:** 262,144 tokens, accumulated as **four 65,536-token
|
| 193 |
-
microbatch passes** (256 whole 1024-token sequences), one shuffled
|
| 194 |
-
no-replacement pass over the corpus.
|
| 195 |
-
- **Learning-rate schedule:** book/scale LR **0.048 / 0.032** through update 45,
|
| 196 |
-
then dropped by ×0.25 to **0.012 / 0.008** (`lr_decay_after=45`,
|
| 197 |
-
`lr_decay_factor=0.25`). (Note: these campaign LRs are higher than the
|
| 198 |
-
in-repo code defaults of 0.003/0.002, consistent with the much larger
|
| 199 |
-
262k-token batch.)
|
| 200 |
-
- **Update budget:** layers 4–26 use a fixed **69 updates**. From layer 27, 69
|
| 201 |
-
is the maximum with an **adaptive early stop**: three validation checks more
|
| 202 |
-
than 0.1% worse than best trigger an earlier LR reduction; 15 updates without
|
| 203 |
-
improvement after that reduction permit stopping.
|
| 204 |
-
- **Validation every 5 updates**, retaining the best checkpoint.
|
| 205 |
-
- **Index reassignment every 20 updates** (Section 6.1).
|
| 206 |
-
- **Regularization:** codebook drift-from-anchor penalty + scale drift penalty,
|
| 207 |
-
weighted 0.01; grad-norm clip 1.0; codebook atoms clamped to [-6, 6];
|
| 208 |
-
log-scales clamped to within ±0.35 of their initial value.
|
| 209 |
-
- **Coverage gate:** experts with fewer than **256 routed training rows** are
|
| 210 |
-
frozen (gradient masked) and keep their initial books/scales.
|
| 211 |
-
|
| 212 |
-
### 6.1 Output-gradient index reassignment
|
| 213 |
-
|
| 214 |
-
Continuous Adam cannot move the discrete indices, so every 20 updates a discrete
|
| 215 |
-
proposal pass runs (`gradient_indices.py`, method
|
| 216 |
-
`expert_parallel_output_gradient_prefix_backtracking_v1`):
|
| 217 |
-
|
| 218 |
-
1. Backprop the **output** loss to the reconstructed weight of one expert
|
| 219 |
-
projection.
|
| 220 |
-
2. Propose alternative `(a,b)` index pairs near a small gradient step, capped at
|
| 221 |
-
`max_fraction=0.001` of groups, `trust_ratio=0.01`, `target_ratio=0.03`; the
|
| 222 |
-
current pair is always in the candidate set (changing nothing stays legal).
|
| 223 |
-
3. **Prefix backtracking acceptance:** try the top-k proposals with
|
| 224 |
-
k∈{full, ¼, 1/16, 1}, accept only if the **training** SSE strictly drops
|
| 225 |
-
**and** a separate **check batch** SSE does not increase (beyond 1e-7 slack);
|
| 226 |
-
otherwise revert. Proposal deltas are broadcast sparsely (only routed rows).
|
| 227 |
-
|
| 228 |
-
The transcripts flag this as the method's main theoretical weakness versus
|
| 229 |
-
published PV-Tuning: codebooks are tuned for output accuracy while index
|
| 230 |
-
proposals originally came from weight-reconstruction Hessians — the
|
| 231 |
-
output-gradient proposal above was added to close that gap.
|
| 232 |
-
|
| 233 |
-
### 6.2 Cold arithmetic emulation
|
| 234 |
-
|
| 235 |
-
Both training and evaluation emulate the serving numerics (`activation.py`,
|
| 236 |
-
pinned to serving revision `b1380cf7…`): **four FP4 activation planes**
|
| 237 |
-
(`fp4_planes4_fp16_boundaries_v1`) and **FP16 SwiGLU boundaries** (FP32 GEMM,
|
| 238 |
-
FP16 SiLU/product), with straight-through estimators for gradients. This
|
| 239 |
-
emulates the quantization boundaries only — **native SM120 MMA accumulation and
|
| 240 |
-
TP reduction are not bit-exact** and were not qualified.
|
| 241 |
-
|
| 242 |
-
---
|
| 243 |
-
|
| 244 |
-
## 7. Acceptance gates
|
| 245 |
-
|
| 246 |
-
Publication of a layer requires all of:
|
| 247 |
-
|
| 248 |
-
- **Validation non-regression:** `0 ≤ final_reference_rel ≤ initial·(1+1e-6)`.
|
| 249 |
-
- **Development-audit non-regression:** the held-out audit split (never used for
|
| 250 |
-
updates or checkpoint selection) must satisfy
|
| 251 |
-
`audit_reference_rel ≤ initial_audit·(1+1e-6)`; otherwise the layer keeps its
|
| 252 |
-
initialization. There is no absolute error floor — the rule is purely
|
| 253 |
-
final ≤ initial.
|
| 254 |
-
- **Serialized-replay parity:** decoded/reloaded forward matches the tuned
|
| 255 |
-
forward to <2e-6, and remote file hashes are re-verified after upload.
|
| 256 |
-
|
| 257 |
-
The audit is honestly labeled **development data, not an untouched final test**;
|
| 258 |
-
it is a split of the same calibration tokens the initialization already saw. Two
|
| 259 |
-
notable per-layer decisions recorded in the card: **layer 30** retains its
|
| 260 |
-
audit-qualified update-15 checkpoint after its validation-best update-20 failed
|
| 261 |
-
the development audit; **layer 3** retains a separately qualified lower-LR
|
| 262 |
-
refinement.
|
| 263 |
-
|
| 264 |
-
---
|
| 265 |
-
|
| 266 |
-
## 8. Calibration corpus
|
| 267 |
-
|
| 268 |
-
The text corpus (`build_calib_v31.py` + `reasoning_slice.py`) is deterministic
|
| 269 |
-
(seed 42), GLM-tokenized, with the following domain mix by tokens:
|
| 270 |
-
|
| 271 |
-
| Share | Domain | Sources |
|
| 272 |
-
|---|---|---|
|
| 273 |
-
| ~32% | code | local vLLM sources, m-a-p/CodeFeedback, jtatman/python-code-500k |
|
| 274 |
-
| ~20% | tool-calling / agentic | generated GLM chat-template sessions |
|
| 275 |
-
| ~15% | reasoning | OpenR1-Math-220k + dolphin-r1 `<think>…</think>` → answer |
|
| 276 |
-
| ~13% | instruction chat | tatsu-lab/alpaca + coding chat |
|
| 277 |
-
| ~10% | medical | MedQA textbook continuation + medical Q&A |
|
| 278 |
-
| ~10% | prose | vLLM docs markdown + databricks-dolly-15k |
|
| 279 |
-
|
| 280 |
-
The **reasoning slice was added specifically** because the earlier calib-v3 mix
|
| 281 |
-
starved the cold 2-bit experts of reasoning-termination behavior: 87% of its
|
| 282 |
-
`<think>` blocks were empty. The slice emits multi-turn conversations where an
|
| 283 |
-
**earlier** assistant turn carries a real `<think>…</think>` that terminates and
|
| 284 |
-
hands off to an answer, followed by a trailing user turn, so the think-close
|
| 285 |
-
token **`</think>` (id 154842)** renders in-stream rather than as a trailing
|
| 286 |
-
generation prompt. Held-out prompts are explicitly excluded (last 25 vLLM docs,
|
| 287 |
-
last 3 MedQA files, vLLM code beyond index 400 of the seed-42 shuffle), and
|
| 288 |
-
17 exact-match prompts were purged from the training set.
|
| 289 |
-
|
| 290 |
-
The capture/training code also supports **up-weighting boundary rows** (rows
|
| 291 |
-
whose next token is a boundary such as `</think>` / `<|endoftext|>`) via a
|
| 292 |
-
`row_weight` term whose sum normalizes the loss, and an optional matched-mixed
|
| 293 |
-
validation set; the published card describes reasoning-based validation.
|
| 294 |
-
|
| 295 |
-
---
|
| 296 |
-
|
| 297 |
-
## 9. Results captured during the build
|
| 298 |
-
|
| 299 |
-
All errors below are **held-out relative L2**, either over the whole transformer
|
| 300 |
-
block (including residual) or over the cold-experts sum, as noted. They are
|
| 301 |
-
local reconstruction errors, **not** token-accuracy or perplexity.
|
| 302 |
-
|
| 303 |
-
- **Pilot smoke test** (layer 3, one step): full-block error 0.005497 → 0.005456
|
| 304 |
-
(~0.75%), 304/388 expert-projection proposals accepted; exported weights
|
| 305 |
-
reproduced the retained result.
|
| 306 |
-
- **Layer 3** (200-step pilot): full-block reference error **0.005497 →
|
| 307 |
-
0.005271 (−4.12%)**, best at step 200.
|
| 308 |
-
- **Layer 4** (200-step pilot): full-block reference error **0.015760 →
|
| 309 |
-
0.015596 (−1.05%)**.
|
| 310 |
-
- **reference vs same_input** (layer 4, matched): reference 0.21% better on
|
| 311 |
-
validation, 0.08% worse on development audit.
|
| 312 |
-
- **Faster fork parity:** the optimized fork (larger microbatch, specialized
|
| 313 |
-
embedding-lookup backward, sparse proposal broadcast) matched the reference
|
| 314 |
-
fork's validation and audit **exactly** on layers 3 and 4, with all propagated
|
| 315 |
-
BF16 outputs equal; training loops fell from ~1525 s → ~337 s (layer 3) and
|
| 316 |
-
~1226 s → ~303 s (layer 4).
|
| 317 |
-
- **Early-layer stability example** (layer 18, full-corpus): initial
|
| 318 |
-
0.017205, step-5 0.017176, step-69 0.017183 — differences ≤0.17%, treated as
|
| 319 |
-
a signal to watch LR/noise rather than proof of convergence.
|
| 320 |
-
|
| 321 |
-
Older cold-only checkpoints reported ~0.2211→0.2070 (layer 3) and 0.2364→0.2320,
|
| 322 |
-
but those measured **cold-expert output only on different captures** and are
|
| 323 |
-
**not comparable** to the full-block errors above.
|
| 324 |
-
|
| 325 |
-
**No comparison against the FP8 donor or an AQLM variant on perplexity / KLD /
|
| 326 |
-
top-1 / wikitext was recorded** in the reviewed transcripts, and no full-model
|
| 327 |
-
task evaluation was performed. Those remain open.
|
| 328 |
-
|
| 329 |
-
---
|
| 330 |
-
|
| 331 |
-
## 10. Honest limitations
|
| 332 |
-
|
| 333 |
-
- Full-model quality and native **SM120** execution were **not** evaluated; the
|
| 334 |
-
arithmetic emulation covers FP4/FP16 boundaries only, not MMA/TP bit-exactness.
|
| 335 |
-
- The development audit is a split of calibration data, not an independent test.
|
| 336 |
-
- The reference target is the more OOD objective; its generalization advantage
|
| 337 |
-
over same_input was small and mixed on the one matched layer tested.
|
| 338 |
-
- Gates enforce local non-regression, not any absolute quality bar.
|
| 339 |
-
|
| 340 |
-
*Prepared from the GLM-5.3 ARVQ build transcripts (2026-09-14 and 2026-09-16)
|
| 341 |
-
and the `btx53/arvq88` fitting/serving code. Where a number could not be sourced
|
| 342 |
-
it is stated as unknown rather than estimated.*
|
|
|
|
| 1 |
+
# ARVQ-v2 methodology correction
|
| 2 |
|
| 3 |
+
The previous file was copied from v1 and described the wrong target, scale dtype, corpus and schedule. The current methodology is documented in [reproduce/README.md](reproduce/README.md), with exact stage order in [RUNBOOK.md](reproduce/RUNBOOK.md) and settings in [recipe.json](reproduce/recipe.json).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
+
Key facts: same-input target; original FP8 expert teacher; FP4 per-expert books; FP16 block scales;15,007,754-token cleaned corpus;50× next-boundary loss weighting; reassignment every10; decay by25 or earlier; patience10. Validation selection is unweighted target-relative L2. Most layers stop early. All75 layer receipts remain available. Historical contradictory metadata is archived under reproduce/historical_metadata.
|
| 6 |
|
| 7 |
+
The bundled reference packer/decoder now handles both uint8 E4M3 and actual FP16 scales, tested on CPU. No model weights were changed by this correction.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
README.md
CHANGED
|
@@ -9,26 +9,15 @@ Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation.
|
|
| 9 |
Other layers retain their previously published weights; see pv_progress.json.
|
| 10 |
Each layer's two tensor files and reports are replaced together in one commit.
|
| 11 |
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
retains its separately qualified lower-LR refinement. Output-gradient index
|
| 22 |
-
reassignment occurs every 20 updates; validation every five updates retains
|
| 23 |
-
the best checkpoint. Audit non-regression and export replay gate publication.
|
| 24 |
-
|
| 25 |
-
Student inputs reflect retained tuned upstream layers. Fixed reference targets
|
| 26 |
-
use original FP8-source routed experts and the donor backbone. A rolling cache
|
| 27 |
-
carries both trajectories. Cold arithmetic emulates FP4 activation planes and
|
| 28 |
-
FP16 boundaries; native SM120 kernel execution has not yet been smoke-tested
|
| 29 |
-
(the evaluation below decodes the published weights to BF16 in software).
|
| 30 |
-
The audit set is historical development data, not an untouched final test.
|
| 31 |
-
Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
|
| 32 |
|
| 33 |
## Full-model evaluation (2026-09-22)
|
| 34 |
|
|
@@ -46,8 +35,7 @@ four held-out corpora):
|
|
| 46 |
In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
|
| 47 |
ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
|
| 48 |
shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
|
| 49 |
-
calibration-domain bias
|
| 50 |
-
regression; treat general-English quality accordingly.
|
| 51 |
|
| 52 |
**Repo history note (2026-09-22):** the hot tier, `aqlm_layer_books` and
|
| 53 |
`cold_assignment.json` were re-exported to match the cold_manifests allocation.
|
|
@@ -78,3 +66,9 @@ format version 4 to match. Earlier branches cannot load this repo:
|
|
| 78 |
and predates the v4 format string; the `glm52-sm120` main branch and the v3
|
| 79 |
`arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
|
| 80 |
is for the v2r repo (quarter-scale residual decode), not this one.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
Other layers retain their previously published weights; see pv_progress.json.
|
| 10 |
Each layer's two tensor files and reports are replaced together in one commit.
|
| 11 |
|
| 12 |
+
The production recipe uses **same-input targets**, **FP16 block scales**, and a **50× boundary loss weight**. It differs from v1; previous versions of this README accidentally retained the v1 training description.
|
| 13 |
+
|
| 14 |
+
The cleaned v3.2 training stream contains **15,007,754 tokens**, context1024, effective batch262,144 and microbatch65,536. It removes empty think blocks and frames documents with canonical GLM delimiters. Reasoning validation and development audit captures contain16,384 tokens each; audit is WikiText/OOD and was used during development. The maximum one-pass budget is58updates, with adaptive early stopping: LR .048/.032,4× decay by update25 or earlier after three worsening validation checks, then patience10. Reassignment runs every10updates and validation every5. Published receipts show67/75 layers selecting update5; a full corpus is available, but most layers do **not** consume an entire pass.
|
| 15 |
+
|
| 16 |
+
Boundary weighting is50 on the **position preceding** token154842 (`</think>`) or154820 (`<|endoftext|>`),1 elsewhere. It multiplies squared output residuals once and normalizes by summed weights; it is not2500×. Window-final positions have weight1. Validation metrics and discrete reassignment checks remain unweighted. Books stay on the FP4 grid; learned FP16 block scales serialize directly. The cold index rate is2bpw, or2.125bpw including block scales, plus codebook/global overhead.
|
| 17 |
+
|
| 18 |
+
Student inputs are propagated through the selected upstream hybrid layers. The target is the original FP8-source routed experts evaluated on **those same student inputs**, with frozen hot contribution subtracted. This is not v1's independent original-reference trajectory objective. Two trajectories may still be carried for diagnostic comparison. Reference labels in some historical reports are generic/stale; the reports' `target: same_input` and fitting code determine the actual objective.
|
| 19 |
+
|
| 20 |
+
Hot experts, backbone and vision components are inherited; the original bootstrap hot-allocation mismatch was corrected as noted below. Native runtime correctness is a separate qualification from software decoding.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
|
| 22 |
## Full-model evaluation (2026-09-22)
|
| 23 |
|
|
|
|
| 35 |
In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
|
| 36 |
ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
|
| 37 |
shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
|
| 38 |
+
possible calibration-domain mismatch. These results do not isolate corpus bias from quantization format, optimization, or trajectory effects; treat general-English quality accordingly.
|
|
|
|
| 39 |
|
| 40 |
**Repo history note (2026-09-22):** the hot tier, `aqlm_layer_books` and
|
| 41 |
`cold_assignment.json` were re-exported to match the cold_manifests allocation.
|
|
|
|
| 66 |
and predates the v4 format string; the `glm52-sm120` main branch and the v3
|
| 67 |
`arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
|
| 68 |
is for the v2r repo (quarter-scale residual decode), not this one.
|
| 69 |
+
|
| 70 |
+
## Reproduce this release
|
| 71 |
+
|
| 72 |
+
The [reproducibility guide](reproduce/README.md), [stage-by-stage runbook](reproduce/RUNBOOK.md), [recipe](reproduce/recipe.json), [data artifact hashes](reproduce/data_artifacts.json), and [source tree](reproduce/source) now accompany this checkpoint. The package contains202 recovered source files spanning corpus preparation, capture, fitting, PV, export and evaluation. Run `python reproduce/verify.py` for CPU checks.
|
| 73 |
+
|
| 74 |
+
**Scope:** the2026-09-26 audit adds code/documentation and corrects metadata; it does not change model weights. Exact historical retraining is not yet guaranteed: some original local corpus inputs, streamed-dataset revisions and environment locks were not preserved. Recovered artifacts, missing inputs and changed-source qualifications are explicitly documented rather than concealed.
|
arvq88_reference.py
CHANGED
|
@@ -20,7 +20,66 @@ def pack_codebooks(c0,c1):
|
|
| 20 |
levels=torch.tensor(LEVELS,device=cb.device);n=(cb[...,None]-levels).abs().argmin(-1)
|
| 21 |
if not torch.equal(levels[n],cb):raise ValueError('Codebooks must be exactly on the FP4 grid')
|
| 22 |
return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
|
| 23 |
-
def
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
|
| 25 |
default=(4096,6144) if proj=='w13' else (6144,2048)
|
| 26 |
N,K=default if N is None else (N,K)
|
|
@@ -28,16 +87,27 @@ def decode(t,layer,proj,e,N=None,K=None):
|
|
| 28 |
if cb.ndim==2:cb=cb[e]
|
| 29 |
if cb.shape!=(512,):raise ValueError('Invalid packed codebook shape')
|
| 30 |
values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
|
| 31 |
-
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
def sha(p):
|
| 34 |
h=hashlib.sha256()
|
| 35 |
with open(p,'rb') as f:
|
| 36 |
for chunk in iter(lambda:f.read(8<<20),b''):h.update(chunk)
|
| 37 |
return h.hexdigest()
|
| 38 |
-
def export_layer(source,dest,L):
|
| 39 |
from safetensors.torch import save_file,load_file
|
| 40 |
source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
original=json.loads((source/'arvq-manifest.json').read_text());files={}
|
| 42 |
for proj,tag in [('w13','gateup'),('w2','down')]:
|
| 43 |
d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
|
|
@@ -45,20 +115,28 @@ def export_layer(source,dest,L):
|
|
| 45 |
if expert and d['c0'].shape!=(E,256,8):raise ValueError('Codebook expert count mismatch')
|
| 46 |
if proj=='w13':scope='expert' if expert else 'layer'
|
| 47 |
elif scope!=('expert' if expert else 'layer'):raise ValueError('Mixed projection codebook scopes')
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
|
| 49 |
pre+'codebooks':pack_codebooks(d['c0'],d['c1']),
|
| 50 |
-
pre+'scales':
|
| 51 |
pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
|
| 52 |
for e in range(E):
|
| 53 |
a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
|
| 54 |
for e in sorted({0,E//2,E-1}):
|
| 55 |
c0,c1=(d['c0'][e],d['c1'][e]) if expert else (d['c0'],d['c1'])
|
| 56 |
-
ref=((c0[d['a'][e].long()]+c1[d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
|
| 57 |
-
assert torch.equal(ref,decode(t,L,proj,e,N,K))
|
|
|
|
|
|
|
| 58 |
name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
|
| 59 |
-
save_file(t,str(tmp),metadata={'format':'pt','arvq_format':'
|
| 60 |
tmp.replace(dest/name);del t,d
|
| 61 |
loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64);del loaded
|
| 62 |
files[proj]={'file':name,'sha256':sha(dest/name)}
|
| 63 |
-
|
|
|
|
| 64 |
(dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
|
|
|
|
| 20 |
levels=torch.tensor(LEVELS,device=cb.device);n=(cb[...,None]-levels).abs().argmin(-1)
|
| 21 |
if not torch.equal(levels[n],cb):raise ValueError('Codebooks must be exactly on the FP4 grid')
|
| 22 |
return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
|
| 23 |
+
def pack_codebooks_mb(cb,factors):
|
| 24 |
+
"""mcbook16: cb [E,256+B*256,8] in EFFECTIVE units. Rows are stored as plain-grid
|
| 25 |
+
nibbles; decode multiplies residual book m by factors[m] (exact powers of two)."""
|
| 26 |
+
B=len(factors)
|
| 27 |
+
if cb.ndim!=3 or cb.shape[1]!=256+B*256 or cb.shape[2]!=8:raise ValueError('Invalid mcbook codebook shape')
|
| 28 |
+
f=torch.cat([torch.ones(256),torch.tensor(factors).float().repeat_interleave(256)])[None,:,None]
|
| 29 |
+
grid=cb/f
|
| 30 |
+
levels=torch.tensor(LEVELS,device=cb.device);n=(grid[...,None]-levels).abs().argmin(-1)
|
| 31 |
+
if not torch.equal(levels[n]*f,cb):raise ValueError('Codebooks must be exactly on their per-book grids')
|
| 32 |
+
return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
|
| 33 |
+
def decode_mb(t,layer,proj,e,N,K):
|
| 34 |
+
pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
|
| 35 |
+
a,b=unpack(t[pre+'packed'][e],N,K)
|
| 36 |
+
cb=t[pre+'codebooks'][e].long();factors=t[pre+'book_factors'].float();B=len(factors)
|
| 37 |
+
if cb.shape!=(256+B*256,):raise ValueError('Invalid mcbook packed codebook shape')
|
| 38 |
+
values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
|
| 39 |
+
values=torch.cat([values[:256],values[256:]*factors.repeat_interleave(256)[:,None]])
|
| 40 |
+
sel=t[pre+'selectors'][e].long()
|
| 41 |
+
m=sel.repeat_interleave(16,0).repeat_interleave(8,1)
|
| 42 |
+
stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
|
| 43 |
+
if stored.dtype==torch.float16:scales=stored.float()
|
| 44 |
+
elif stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
|
| 45 |
+
else:raise ValueError('Unsupported packed block scale dtype')
|
| 46 |
+
return ((values[a.long()]+values[256+m*256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
|
| 47 |
+
def export_layer_mb(source,dest,L):
|
| 48 |
+
"""mcbook16 export (format version 5): 16 residual books, per-tile 4-bit selectors,
|
| 49 |
+
fp16 block scales; packed index stream and scales unchanged from v4."""
|
| 50 |
+
from safetensors.torch import save_file,load_file
|
| 51 |
+
source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
|
| 52 |
+
original=json.loads((source/'arvq-manifest.json').read_text());files={};first_factors=None
|
| 53 |
+
for proj,tag in [('w13','gateup'),('w2','down')]:
|
| 54 |
+
d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
|
| 55 |
+
if d.get('scale_dtype')!='fp16':raise ValueError('mcbook export requires fp16 block scales')
|
| 56 |
+
if not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
|
| 57 |
+
factors=[float(f) for f in d['book_factor']]
|
| 58 |
+
if first_factors is None:first_factors=factors
|
| 59 |
+
elif factors!=first_factors:raise ValueError('Mixed projection book factors')
|
| 60 |
+
sel=d['selector']
|
| 61 |
+
if sel.shape!=(E,N//16,K//64) or int(sel.max())>=len(factors):raise ValueError('Invalid selector tensor')
|
| 62 |
+
t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
|
| 63 |
+
pre+'codebooks':pack_codebooks_mb(d['cb'],factors),
|
| 64 |
+
pre+'selectors':sel.to(torch.uint8).contiguous(),
|
| 65 |
+
pre+'book_factors':torch.tensor(factors,dtype=torch.float32),
|
| 66 |
+
pre+'scales':d['s'].half().reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
|
| 67 |
+
pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
|
| 68 |
+
for e in range(E):
|
| 69 |
+
a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
|
| 70 |
+
for e in sorted({0,E//2,E-1}):
|
| 71 |
+
m=d['selector'][e].long().repeat_interleave(16,0).repeat_interleave(8,1)
|
| 72 |
+
cb=d['cb'][e]
|
| 73 |
+
ref=((cb[d['a'][e].long()]+cb[256+m*256+d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
|
| 74 |
+
assert torch.equal(ref,decode_mb(t,L,proj,e,N,K))
|
| 75 |
+
name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
|
| 76 |
+
save_file(t,str(tmp),metadata={'format':'pt','arvq_format':'rvq256_mb16_256x8_expert_fp16block','codebook_scope':'expert','residual_books':str(len(factors)),'book_factors':json.dumps(factors),'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires mcbook16 v5 loader/kernel (vllm-glm52-sm120 branch experiment/arvq-mcbook16)'})
|
| 77 |
+
tmp.replace(dest/name)
|
| 78 |
+
loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64) and loaded[pre+'selectors'].shape==(E,N//16,K//64);del loaded
|
| 79 |
+
files[proj]={'file':name,'sha256':sha(dest/name)};del t,d
|
| 80 |
+
m={**original,'format':'rvq256_mb16_256x8_expert_fp16block','version':5,'residual_scale_shift':0,'block_scale_dtype':'float16','codebook_scope':'expert','bits':2.125+4/1024,'codebook_sizes':[256,16*256],'index_bits':[8,8],'selector_bits_per_tile':4,'words_per_tile':64,'book_factors':first_factors,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True,'requires_mcbook16_loader':True}
|
| 81 |
+
(dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
|
| 82 |
+
def decode(t,layer,proj,e,N=None,K=None,residual_scale_shift=0):
|
| 83 |
pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
|
| 84 |
default=(4096,6144) if proj=='w13' else (6144,2048)
|
| 85 |
N,K=default if N is None else (N,K)
|
|
|
|
| 87 |
if cb.ndim==2:cb=cb[e]
|
| 88 |
if cb.shape!=(512,):raise ValueError('Invalid packed codebook shape')
|
| 89 |
values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
|
| 90 |
+
stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
|
| 91 |
+
if stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
|
| 92 |
+
elif stored.dtype==torch.float16:scales=stored.float()
|
| 93 |
+
else:raise ValueError('Unsupported packed block scale dtype')
|
| 94 |
+
f=2.0**(-int(residual_scale_shift))
|
| 95 |
+
return ((values[a.long()]+f*values[256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
|
| 96 |
def sha(p):
|
| 97 |
h=hashlib.sha256()
|
| 98 |
with open(p,'rb') as f:
|
| 99 |
for chunk in iter(lambda:f.read(8<<20),b''):h.update(chunk)
|
| 100 |
return h.hexdigest()
|
| 101 |
+
def export_layer(source,dest,L,residual_scale_shift=0):
|
| 102 |
from safetensors.torch import save_file,load_file
|
| 103 |
source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
|
| 104 |
+
probe=torch.load(source/'w13.pt',map_location='cpu',weights_only=True,mmap=True)['w13']
|
| 105 |
+
if 'selector' in probe:
|
| 106 |
+
del probe
|
| 107 |
+
if int(residual_scale_shift):raise ValueError('mcbook exports are shift-0 (per-book factors)')
|
| 108 |
+
return export_layer_mb(source,dest,L)
|
| 109 |
+
del probe
|
| 110 |
+
rss=int(residual_scale_shift);f=2.0**(-rss)
|
| 111 |
original=json.loads((source/'arvq-manifest.json').read_text());files={}
|
| 112 |
for proj,tag in [('w13','gateup'),('w2','down')]:
|
| 113 |
d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
|
|
|
|
| 115 |
if expert and d['c0'].shape!=(E,256,8):raise ValueError('Codebook expert count mismatch')
|
| 116 |
if proj=='w13':scope='expert' if expert else 'layer'
|
| 117 |
elif scope!=('expert' if expert else 'layer'):raise ValueError('Mixed projection codebook scopes')
|
| 118 |
+
fp16=d.get('scale_dtype','fp8_e4m3')=='fp16'
|
| 119 |
+
if proj=='w13':fp16_blocks=fp16
|
| 120 |
+
elif fp16_blocks!=fp16:raise ValueError('Mixed projection scale precision')
|
| 121 |
+
if fp16 and not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
|
| 122 |
+
stored_scales=d['s'].half() if fp16 else d['s'].to(torch.float8_e4m3fn).view(torch.uint8)
|
| 123 |
t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
|
| 124 |
pre+'codebooks':pack_codebooks(d['c0'],d['c1']),
|
| 125 |
+
pre+'scales':stored_scales.reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
|
| 126 |
pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
|
| 127 |
for e in range(E):
|
| 128 |
a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
|
| 129 |
for e in sorted({0,E//2,E-1}):
|
| 130 |
c0,c1=(d['c0'][e],d['c1'][e]) if expert else (d['c0'],d['c1'])
|
| 131 |
+
ref=((c0[d['a'][e].long()]+f*c1[d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
|
| 132 |
+
assert torch.equal(ref,decode(t,L,proj,e,N,K,rss))
|
| 133 |
+
base_fmt='rvq256_256x8_expert_fp16block' if fp16 else ('rvq256_256x8_expert' if expert else 'rvq256_256x8')
|
| 134 |
+
arvq_fmt=base_fmt+(f'_rs{int(2**rss)}' if rss else '')
|
| 135 |
name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
|
| 136 |
+
save_file(t,str(tmp),metadata={'format':'pt','arvq_format':arvq_fmt,'residual_scale_shift':str(rss),'codebook_scope':scope,'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires FP16 block-scale v4 loader/kernel' if fp16 else ('requires per-expert v3 loader/kernel' if expert else 'requires 8+8 loader/kernel')})
|
| 137 |
tmp.replace(dest/name);del t,d
|
| 138 |
loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64);del loaded
|
| 139 |
files[proj]={'file':name,'sha256':sha(dest/name)}
|
| 140 |
+
base_mfmt='rvq256_256x8_expert_fp16block' if fp16_blocks else ('rvq256_256x8_expert' if scope=='expert' else 'rvq256_256x8')
|
| 141 |
+
m={**original,'format':base_mfmt+(f'_rs{int(2**rss)}' if rss else ''),'version':4 if fp16_blocks else (3 if scope=='expert' else 2),'residual_scale_shift':rss,'block_scale_dtype':'float16' if fp16_blocks else 'float8_e4m3fn','codebook_scope':scope,'bits':2.125 if fp16_blocks else 2.0625,'codebook_sizes':[256,256],'index_bits':[8,8],'words_per_tile':64,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True}
|
| 142 |
(dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
|
calibration_corpus.json
CHANGED
|
@@ -1,24 +1,17 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
-
"
|
| 4 |
-
"
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
"partitions": "disjoint prompt hashes; separately captured",
|
| 19 |
-
"limitations": "Text-only captures in 2048-token windows; source datasets overlap earlier calibration. Teacher answers are not correctness-verified.",
|
| 20 |
-
"initial_fit_partition": "train only",
|
| 21 |
-
"capture_teacher": "RadixArk/GLM-5.3-NVFP4",
|
| 22 |
-
"rollout_teacher": "unc NVFP4 teacher",
|
| 23 |
-
"same_token_corpus_as": "/tmp/glm53-kernel-aware-pv/full"
|
| 24 |
-
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"stage": "current recipe provenance correction",
|
| 3 |
+
"variant": "ARVQ-v2",
|
| 4 |
+
"full_pv_training_tokens": 15007754,
|
| 5 |
+
"sequence_length": 1024,
|
| 6 |
+
"validation_rows": 16384,
|
| 7 |
+
"audit_rows": 16384,
|
| 8 |
+
"boundary_boost": 50,
|
| 9 |
+
"boundary_token_ids": [
|
| 10 |
+
154842,
|
| 11 |
+
154820
|
| 12 |
+
],
|
| 13 |
+
"data_pipeline": "reproduce/README.md",
|
| 14 |
+
"data_hashes": "reproduce/data_artifacts.json",
|
| 15 |
+
"historical_component_metadata": "reproduce/historical_metadata/calibration_corpus.json",
|
| 16 |
+
"exact_retraining_guaranteed": false
|
| 17 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
pv_progress.json
CHANGED
|
@@ -77,111 +77,28 @@
|
|
| 77 |
77
|
| 78 |
],
|
| 79 |
"total_layers": 75,
|
| 80 |
-
"recipe": "
|
| 81 |
-
"training_tokens":
|
| 82 |
"batch_tokens": 262144,
|
| 83 |
"microbatch_tokens": 65536,
|
| 84 |
-
"max_updates_per_layer":
|
| 85 |
"adaptive_schedule": {
|
| 86 |
-
"from_layer": 27,
|
| 87 |
"relative_worsening": 0.001,
|
| 88 |
"checks": 3,
|
| 89 |
-
"patience_updates":
|
| 90 |
-
"max_updates": 69,
|
| 91 |
-
"selection": "existing fixed reasoning validation; retain best checkpoint",
|
| 92 |
-
"description": "Drop LR4x after three consecutive validation checks >0.1% worse than best, or after45 at latest. After reduction stop following15 updates without a new best."
|
| 93 |
},
|
| 94 |
-
"lr_decay_after":
|
| 95 |
"lr_decay_factor": 0.25,
|
| 96 |
-
"layer3_exception": "
|
| 97 |
-
"
|
| 98 |
-
"
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
"legacy_pv_layers": [
|
| 110 |
-
9,
|
| 111 |
-
10,
|
| 112 |
-
11,
|
| 113 |
-
12,
|
| 114 |
-
13,
|
| 115 |
-
14
|
| 116 |
-
],
|
| 117 |
-
"initial_fit_layers": [
|
| 118 |
-
15,
|
| 119 |
-
16,
|
| 120 |
-
17,
|
| 121 |
-
18,
|
| 122 |
-
19,
|
| 123 |
-
20,
|
| 124 |
-
21,
|
| 125 |
-
22,
|
| 126 |
-
23,
|
| 127 |
-
24,
|
| 128 |
-
25,
|
| 129 |
-
26,
|
| 130 |
-
27,
|
| 131 |
-
28,
|
| 132 |
-
29,
|
| 133 |
-
30,
|
| 134 |
-
31,
|
| 135 |
-
32,
|
| 136 |
-
33,
|
| 137 |
-
34,
|
| 138 |
-
35,
|
| 139 |
-
36,
|
| 140 |
-
37,
|
| 141 |
-
38,
|
| 142 |
-
39,
|
| 143 |
-
40,
|
| 144 |
-
41,
|
| 145 |
-
42,
|
| 146 |
-
43,
|
| 147 |
-
44,
|
| 148 |
-
45,
|
| 149 |
-
46,
|
| 150 |
-
47,
|
| 151 |
-
48,
|
| 152 |
-
49,
|
| 153 |
-
50,
|
| 154 |
-
51,
|
| 155 |
-
52,
|
| 156 |
-
53,
|
| 157 |
-
54,
|
| 158 |
-
55,
|
| 159 |
-
56,
|
| 160 |
-
57,
|
| 161 |
-
58,
|
| 162 |
-
59,
|
| 163 |
-
60,
|
| 164 |
-
61,
|
| 165 |
-
62,
|
| 166 |
-
63,
|
| 167 |
-
64,
|
| 168 |
-
65,
|
| 169 |
-
66,
|
| 170 |
-
67,
|
| 171 |
-
68,
|
| 172 |
-
69,
|
| 173 |
-
70,
|
| 174 |
-
71,
|
| 175 |
-
72,
|
| 176 |
-
73,
|
| 177 |
-
74,
|
| 178 |
-
75,
|
| 179 |
-
76,
|
| 180 |
-
77
|
| 181 |
-
],
|
| 182 |
-
"format": "rvq256_256x8_expert",
|
| 183 |
-
"full_model_quality": "pending"
|
| 184 |
-
},
|
| 185 |
-
"format": "rvq256_256x8_expert",
|
| 186 |
-
"full_model_quality": "pending"
|
| 187 |
-
}
|
|
|
|
| 77 |
77
|
| 78 |
],
|
| 79 |
"total_layers": 75,
|
| 80 |
+
"recipe": "full_corpus_sequential_same_input_fp16block",
|
| 81 |
+
"training_tokens": 15007754,
|
| 82 |
"batch_tokens": 262144,
|
| 83 |
"microbatch_tokens": 65536,
|
| 84 |
+
"max_updates_per_layer": 58,
|
| 85 |
"adaptive_schedule": {
|
|
|
|
| 86 |
"relative_worsening": 0.001,
|
| 87 |
"checks": 3,
|
| 88 |
+
"patience_updates": 10
|
|
|
|
|
|
|
|
|
|
| 89 |
},
|
| 90 |
+
"lr_decay_after": 25,
|
| 91 |
"lr_decay_factor": 0.25,
|
| 92 |
+
"layer3_exception": "v2 layer3 uses same LR recipe; selected step50 after58 updates, per its published report",
|
| 93 |
+
"format": "rvq256_256x8_expert_fp16block",
|
| 94 |
+
"full_model_quality": "See README full-model evaluation; software-dequant evaluation, not a guarantee of native runtime parity",
|
| 95 |
+
"target": "same_input",
|
| 96 |
+
"boundary_boost": 50.0,
|
| 97 |
+
"boundary_token_ids": [
|
| 98 |
+
154842,
|
| 99 |
+
154820
|
| 100 |
+
],
|
| 101 |
+
"reassign_every": 10,
|
| 102 |
+
"block_scale_dtype": "float16",
|
| 103 |
+
"metadata_corrected_on": "2026-09-26"
|
| 104 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
reproduce/AUDIT.md
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Audit evidence and corrections, 2026-09-26
|
| 2 |
+
|
| 3 |
+
Evidence precedence: actual checkpoint tensor headers and75 per-layer PV reports; saved campaign config; executable capture/loss/packing source; local RUNLOG53 chronology and Claude build conversation; model-card prose. Private transcripts are not uploaded.
|
| 4 |
+
|
| 5 |
+
Confirmed v2 defects corrected: copied v1 README and METHODOLOGY; incorrect18,001,846-token claim (actual15,007,754); FP8 vs actual FP16 scales; reference vs same-input target; decay45/patience15/reassign20 vs25/10/10; missing50× boundary loss rule; stale pv_progress; root reference decoder reinterpreting FP16 bytes as E4M3. Header samples show U8 in v1 and F16 in v2. CPU fixtures verify corrected decode in both formats.
|
| 6 |
+
|
| 7 |
+
v1 methodology now distinguishes ARVQ output-benefit allocation from the old AQLM allocation. Both ARVQ corpus metadata files now distinguish earlier/component provenance from final PV streams. AQLM notes clarify sampled activation PV, global75/25 row mixture,65536-entry shared layer/projection books and absence of boundary weighting.
|
| 8 |
+
|
| 9 |
+
v2 WikiText degradation is measured; attributing it solely to corpus bias was too strong and has been corrected. Quantization and optimization effects were not isolated by a controlled ablation.
|
| 10 |
+
|
| 11 |
+
Historical per-layer `checkpoint_selection` strings can be stale even when `target` and recorded metrics are correct; original reports are retained, not retroactively rewritten. Old publisher scripts are archived and can regenerate stale prose; use the corrected root documentation and recipe files as the current documentation contract.
|
| 12 |
+
|
| 13 |
+
Missing provenance is explicitly listed. This package is a practical reconstruction/runbook and source release, not certification of byte-identical historical retraining.
|
reproduce/README.md
ADDED
|
@@ -0,0 +1,88 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Reproducing the GLM-5.3 hybrid releases
|
| 2 |
+
|
| 3 |
+
This package was reconstructed on 2026-09-26 from the local quantization source, saved configurations, per-layer reports, checkpoint tensor headers, and Claude build history. The original private conversations are not redistributed. `checkpoint_before_audit.json` pins the weights/reports audited; the present documentation-only revision does not retrain or replace weights.
|
| 4 |
+
|
| 5 |
+
## Reproducibility status
|
| 6 |
+
|
| 7 |
+
The fitting, capture, data-building, evaluation, packing and verification source is included in `source/` with SHA-256 checksums. This is a **recovered working-tree snapshot**, containing subsequent fixes and experiments, not a falsely reconstructed original Git commit. `recipe.json` selects the intended recipe; do not infer it from experimental filenames. Data hashes and current environment versions are included. Existing model weights and original per-layer reports remain the reference artifacts.
|
| 8 |
+
|
| 9 |
+
A bit-for-bit re-training claim is **not justified**: original streamed dataset revisions, local corpus inputs and the complete original dependency lockfile were not all retained. Some training paths were overwritten by later campaigns. The original AQLM `/data` activation and sample archives are no longer present at their recorded paths; the v1 18,001,846-token flat stream has not been recovered. These gaps must be resolved or a rerun must be labeled a new reproduction, with measured differences. Neither a fixed random seed nor the current package versions restores missing historical inputs.
|
| 10 |
+
|
| 11 |
+
## Package layout
|
| 12 |
+
|
| 13 |
+
- `source/tools/`: text corpus builders, source loading/NVFP4 emulation, capture, AQLM fitting, PV and validation.
|
| 14 |
+
- `source/tools/capture53/`: GLM5.3 wrappers, multimodal corpus/capture, allocation, checkpoint assembly, MTP conversion, vision graft and full-model PPL/KL evaluation.
|
| 15 |
+
- `source/btx53/`: FP8 source ingest, raw-basis ARVQ initializer, activation emulation, coupled sequential PV, capture/propagation, REAP allocation and export tooling; `arvq88/perf/` contains the full-corpus controller and its gates.
|
| 16 |
+
- `recipe.json`, `layer_recipe_summary.json`: recovered variant-specific settings and original layer exceptions.
|
| 17 |
+
- `data_artifacts.json`: SHA-256, sizes/shapes and token-boundary counts for recovered token files; paths are historical identities, not downloadable links.
|
| 18 |
+
- `historical_metadata/`: original conflicting metadata retained as historical evidence, not the current recipe.
|
| 19 |
+
- `prepare_workspace.py`: creates a new relocatable copy; does not overwrite the archive or execute a training job.
|
| 20 |
+
- `verify.py`: CPU-only source checksum, boundary-weight semantics and FP8/FP16 reference-decoder checks.
|
| 21 |
+
|
| 22 |
+
## Environment and paths
|
| 23 |
+
|
| 24 |
+
Training/capture used eight A100s with software dequantization/emulation. This is not proof of native SM120 kernel equivalence. `environment_observed.json` records the current audit environment, **not** the original training environment. `requirements.in` lists main Python dependencies, not an exact historical lock. Use a CUDA-compatible PyTorch build and a Transformers version supporting the donor tokenizer; install additional optional dependencies demanded by the chosen multimodal backend. Data-source scripts record their imports. Native serving requires the custom serving fork described on the model card.
|
| 25 |
+
|
| 26 |
+
```bash
|
| 27 |
+
python -m venv .venv
|
| 28 |
+
. .venv/bin/activate
|
| 29 |
+
pip install -r reproduce/requirements.in
|
| 30 |
+
python reproduce/verify.py
|
| 31 |
+
python reproduce/prepare_workspace.py --dest /absolute/path/glm53-reproduction
|
| 32 |
+
cd /absolute/path/glm53-reproduction
|
| 33 |
+
# Use the same activated interpreter for all following commands.
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
The preparer copies source and maps legacy `/data` and `/tmp` paths into the destination, and the old checkout path into this source tree. It writes the transformed source hashes and path map. Review `recipe.json` and scripts before GPU work: donor, tokenizer, corpus, source-cache and baseline locations must exist. Historical scripts that spawn `ROOT/.venv/bin/python` need that interpreter available (the preparer creates the link). Old publication/watchdog scripts are archival; **do not run them against the original public repos**. Export to a new directory and use your own destination repo after verification.
|
| 37 |
+
|
| 38 |
+
Acquire donor/base/vision snapshots using pinned revisions in `historical_metadata/source.json`, `backbone_sources.json`, available model-index provenance, and the audited checkpoint metadata. Where a revision is missing or marked `legacy_cache_revision_unverified`, record that gap rather than substituting today's HEAD as if it were historical. The source FP8 checkpoint and NVFP4 donor are distinct teachers. AQLM uses the NVFP4 donor; ARVQ cold fitting uses original FP8 expert weights. `btx53/tools/prefetch_cold_fp8.py`, `arvq88/inputs.py`, `local_source.py` and `tools/aqlm_quantize.py` implement source fetching/dequantization.
|
| 39 |
+
|
| 40 |
+
## Data pipelines and non-portable inputs
|
| 41 |
+
|
| 42 |
+
### AQLM text calibration (v3)
|
| 43 |
+
|
| 44 |
+
`tools/build_calib_v3.py`: seed42, target15M tokens, approximately40% code /25% agentic /15% instruction /10% medical /10% prose. Streams CodeFeedback, python-code-dataset-500k, Alpaca and Dolly; also uses local vLLM sources/docs, local tools, synthetic shell/agent sessions and MedQA material. See the builder for exact generator order and per-category stopping. It excludes reserved vLLM/MedQA files, but many dataset revisions were not pinned. The original `collect_expert_stats_v2.py` also reads host logs, diffs and shell outputs: these **host inputs must be frozen and reviewed** to reproduce the original corpus; they are not included here. Do not publish freshly collected private logs merely to recreate this pipeline.
|
| 45 |
+
|
| 46 |
+
The initial capture uses nonoverlapping2048-token windows. `stream_capture53.py stats` records routing/salience and sampled activations; it is not a claim that later AQLM PV trains on every15M token. AQLM convergence and PV use the retained activation sample.
|
| 47 |
+
|
| 48 |
+
### Multimodal branch
|
| 49 |
+
|
| 50 |
+
`capture53/build_calib_mm.py`: seed42,1200 image/caption samples,300 each from medical, natural images, OCR/documents and screenshots. Primary candidates: eltorio/ROCOv2-radiology, lmms-lab/COCO-Caption2017, naver-clova-ix/synthdog-en and HuggingFaceM4/WebSight v0.2; explicit fallbacks live in the script. Retain generated `samples.jsonl`, selected dataset/config/split, dataset revision, sample IDs and checksums. Images are resized448×448 and processed with the checkpoint KimiK25 processor into256 vision tokens. Source licenses differ (ROCOv2 includes a noncommercial/share-alike restriction); model MIT metadata does not relicense the datasets. No image or caption archive is redistributed here.
|
| 51 |
+
|
| 52 |
+
`capture_mm53.py` splices the GLM5.2V tower/projector outputs and masks padding. `solve_r3.py` allocates using .75*per-layer-normalized text REAP + .25*per-layer-normalized multimodal salience. Mixed activation files concatenate32768 text rows and10922 multimodal rows. Thus75/25 is a **global row mixture**, not a guarantee that each expert's routed Hessian contains exactly25% visual energy/rows.
|
| 53 |
+
|
| 54 |
+
### ARVQ v1 full-corpus stream
|
| 55 |
+
|
| 56 |
+
The published75 layer reports describe18,001,846 tokens (17,580 windows including the short tail), context1024. Historical metadata's3,048,596-token reasoning corpus identifies a component/earlier capture, not the complete final stream. The exact original concatenated stream is not recovered. Do not replace it silently with the current15M v32 file or label v32 a faithful v1 data reproduction. Capture reads `train_token_file`; supply a recovered checksum-verified v1 stream to reproduce that campaign. Use original layer reports for early-stop and recovery exceptions.
|
| 57 |
+
|
| 58 |
+
### ARVQ v2 v3.2 data rebuild
|
| 59 |
+
|
| 60 |
+
`tools/build_calib_v32.py` actual executable budget is22% code /15% agentic /30% reasoning /13% instruction /10% medical /10% prose, seed42, target15M; actual complete documents produce15,007,754 tokens. Its old opening docstring predates the30% reasoning increase; the budget constants and this runbook are authoritative. Medical data switches to MedRAG textbooks when old local MedQA is unavailable. Reasoning sources include OpenR1-Math-220k and dolphin-r1. Reasoning is shortened to3500 characters, answers1200. Dataset availability, streaming order and local files can affect exact regeneration.
|
| 61 |
+
|
| 62 |
+
`clean_think` removes empty/whitespace-only think blocks, trailing generation-only prompts and unmatched think tags while retaining content. `frame_ids` adds canonical `[gMASK]<sop>` prefix and `<|endoftext|>` suffix only when absent. Use donor tokenization; IDs154822/154824/154820 correspond to those delimiters. Concatenate `shard_*.npy` in sorted order into the configured `train.npy`, then compare length and SHA to `data_artifacts.json`.
|
| 63 |
+
|
| 64 |
+
`tools/build_trace_refit_v32.py` builds three eval-capture sources: training reference from the first1M calibration tokens; held-out reasoning with RNG2025 skipping9000 reasoning documents; and OOD WikiText test text. The capturer uses RNG1234 to select64/16/16 nonoverlapping1024-token sequences, yielding65536 training-probe and16384 validation/audit rows. The procedural skip is not an independent content-deduplication guarantee; the audit split was repeatedly inspected during development.
|
| 65 |
+
|
| 66 |
+
## Boundary weighting: exact semantics
|
| 67 |
+
|
| 68 |
+
For v2, within each1024-token sequence, set `w[t]=50` when the **next** token is154842 (`</think>`) or154820 (`<|endoftext|>`); otherwise1. The last row is1 because the implementation does not look across the window boundary. The same rule is in standalone capture and fused handoff. It weights the hidden output responsible for predicting a boundary, not the boundary token's own output.
|
| 69 |
+
|
| 70 |
+
The PV data term is `sum_t w[t] * ||cold_prediction[t] - required[t]||² / (D * sum_t w[t])`, with D the captured target-energy normalizer. The residual is squared **before** multiplying by50; the multiplier is50, not2500. The small anchor regularizer is separate. Reported validation/audit relative-L2 and boundary/bulk summaries are **unweighted**; checkpoint selection uses validation `target_rel`. Discrete index-reassignment proposal/check objectives are also unweighted in the recovered implementation. The50× multiplier is therefore a continuous PV-loss rule, not a blanket multiplier on all quantization/evaluation stages.
|
| 71 |
+
|
| 72 |
+
AQLM and ARVQ v1 have no identified50× boundary term in their production recipes. Do not retroactively attribute the v2 change to them. Both the corpus cleanup and weighting were motivated by boundary preservation; no causal guarantee of shorter generation follows from this alone.
|
| 73 |
+
|
| 74 |
+
## Verification and acceptance
|
| 75 |
+
|
| 76 |
+
Before fitting: verify corpus hashes/counts, donor identity, tokenizer special IDs, cold/hot partition coverage and target dimensions. Freeze the assignment. Initial-fit output must cover every cold expert. After PV: retain the selected checkpoint, check finite metrics and audit/non-regression policy, then pack and reload for index and decoded-weight parity. Propagate the selected weights and recompute the next layer's inputs; do not mix independently repaired layers without re-evaluation. Save scripts/config hashes, training receipts, all per-layer metrics and environment lockfiles for the new run.
|
| 77 |
+
|
| 78 |
+
`tools/capture53/eval3_kld.py` / `eval_hybrid53.py` evaluate teacher-forced PPL/KL. `btx53/arvq88/pack.py` verifies packed indices and sample decoded weights; the root `arvq88_reference.py` supports actual uint8-E4M3 and float16 scale storage. v3 cold storage is2+8/128=2.0625 bpw before codebook overhead; v4 is2+16/128=2.125 bpw before overhead. AQLM's2bits likewise describes index rate, excluding its65536×8 codebook and scales. Verify actual bytes if comparing compression budgets.
|
| 79 |
+
|
| 80 |
+
This audit ran lightweight CPU verification; it did not rerun a multi-day quantization or certify native GPU serving. Historical PPL/task results on the model card retain their original scope.
|
| 81 |
+
|
| 82 |
+
## Variant-specific commands
|
| 83 |
+
|
| 84 |
+
Read [RUNBOOK.md](RUNBOOK.md) for this repository’s stage order and recipe.
|
| 85 |
+
|
| 86 |
+
## External code and optional modules
|
| 87 |
+
|
| 88 |
+
`external_sources.json` records observed upstream commits. Install b12x at its recorded revision only for the optional BTX comparison modules; these are not part of the ARVQ production recipe. The vLLM checkout is also an input to local code/prose generation and must be placed at `source/vllm` (or the relocated workspace’s `vllm`). Its current commit is not proof of the original dirty training checkout. Dataset IDs/licenses appear in the builders; preserve the resolved revisions and generated sample manifests on a new run. See THIRD_PARTY.md.
|
reproduce/RUNBOOK.md
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# ARVQ-v2 runbook
|
| 2 |
+
|
| 3 |
+
Run the source from the prepared workspace. Preserve all75 original layer reports and the original cold allocation. `recipe.json` is the recovered v2 configuration.
|
| 4 |
+
|
| 5 |
+
1. Prepare source experts from `zai-org/GLM-5.3` and the `RadixArk/GLM-5.3-NVFP4` donor. The published `source.json` pins a source revision but also marks legacy-cache revision verification unavailable; retain that qualification. Source dequantization uses block128×128 E4M3. Prepare the vision extension and capture inputs using the included capture tools.
|
| 6 |
+
2. Run `tools/build_calib_v32.py`, concatenate sorted shards into `train.npy`, and run `tools/build_trace_refit_v32.py`. Confirm15,007,754 training tokens and the recorded hashes.
|
| 7 |
+
3. Generate initial ARVQ fits and the ARVQ output-benefit allocation. Historical v2 launch (replace paths):
|
| 8 |
+
|
| 9 |
+
```bash
|
| 10 |
+
python btx53/build_baseline.py --pv-steps 0 --workdir BASELINE \
|
| 11 |
+
--calib-tokens CALIB_SHARDS --calib-token-count 15000000 \
|
| 12 |
+
--src-cache FP8_EXPERT_CACHE --nvfp4-donor DONOR \
|
| 13 |
+
--dst baseline-only --gpus 0,1,2,3,4,5,6,7 --codebook-scope expert
|
| 14 |
+
```
|
| 15 |
+
|
| 16 |
+
`--codebook-scope expert` is required. `arvq88/encoder.py` and `fit.py` implement raw-basis Hessian/LDLQ fitting; `arvq_reap.py` blends text/MM error-reduction benefit75/25 and keeps5750 hot experts. To reproduce the published assignment rather than resolve it anew, use the published `cold_assignment.json` and verify against every cold manifest. Never substitute the original AQLM allocation merely because the hot count matches.
|
| 17 |
+
|
| 18 |
+
4. Place `recipe.json` as `WORK/config.json`, configure source/cache/corpus/baseline paths, and launch:
|
| 19 |
+
|
| 20 |
+
```bash
|
| 21 |
+
python btx53/arvq88/perf/full_pipeline.py --work WORK --first 3 --last 77
|
| 22 |
+
```
|
| 23 |
+
|
| 24 |
+
The controller captures full training rows plus fixed eval rows, runs coupled expert-parallel PV, gates retained outputs, and propagates the chosen student state. Fused handoff requires its bitwise replay qualification. The source includes later optimizations; preserve receipts/identity hashes and a fresh WORK if changing the recipe.
|
| 25 |
+
|
| 26 |
+
5. v2: same-input source target, FP16 block scales,50× boundary weighting, LR .048/.032 with4× decay by update25 or earlier, patience10, index reassignment every10, evaluation every5. 58updates would cover the full15,007,754-token corpus; adaptive stopping often uses fewer. The published reports show67layers selecting update5. Do not describe this as one complete pass for every layer.
|
| 27 |
+
6. Export through `arvq88/pack.py` / checkpoint assembly to a **new** destination. Run packed-index and decoded-weight parity, hot/cold consistency, full-model PPL/KL and native serving tests. The historical publishers reference public production repos and are provided for provenance, not as a safe rerun destination. Fixed metadata in this audit is the reference; historical publisher templates can reintroduce stale descriptions.
|
| 28 |
+
|
| 29 |
+
A v2 model-wide v4 marker requires FP16-scale cold layers. Do not deploy a partially replaced v3/v4 mixture. The old bootstrap also had a hot-tier mismatch, corrected at f53a8dcc; use a later audited snapshot.
|
reproduce/THIRD_PARTY.md
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
Source files retain their original notices. External checkouts and dataset material have their own licenses; the model card MIT tag does not override those. See external_sources.json and the dataset builder source for origins. The b12x/quant-toolkit/vLLM repositories are referenced rather than bulk vendored. No new license is asserted over third-party code or datasets by this audit.
|
reproduce/checkpoint_before_audit.json
ADDED
|
@@ -0,0 +1,100 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"repo": "jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid",
|
| 3 |
+
"revision": "cbde0aed09e981e68bfefc827cb22d5fbdb81cfb",
|
| 4 |
+
"files": [
|
| 5 |
+
"METHODOLOGY.md",
|
| 6 |
+
"README.md",
|
| 7 |
+
"arvq88_reference.py",
|
| 8 |
+
"backbone_sources.json",
|
| 9 |
+
"btx_hybrid_manifest.json",
|
| 10 |
+
"build_provenance.json",
|
| 11 |
+
"calibration_corpus.json",
|
| 12 |
+
"cold_assignment.json",
|
| 13 |
+
"config.json",
|
| 14 |
+
"gates_report.json",
|
| 15 |
+
"generation_config.json",
|
| 16 |
+
"kimi_k25_processor.py",
|
| 17 |
+
"kimi_k25_vision_processing.py",
|
| 18 |
+
"media_utils.py",
|
| 19 |
+
"model.safetensors.index.json",
|
| 20 |
+
"preprocessor_config.json",
|
| 21 |
+
"pv_layers/layer_003.json",
|
| 22 |
+
"pv_layers/layer_004.json",
|
| 23 |
+
"pv_layers/layer_005.json",
|
| 24 |
+
"pv_layers/layer_006.json",
|
| 25 |
+
"pv_layers/layer_007.json",
|
| 26 |
+
"pv_layers/layer_008.json",
|
| 27 |
+
"pv_layers/layer_009.json",
|
| 28 |
+
"pv_layers/layer_010.json",
|
| 29 |
+
"pv_layers/layer_011.json",
|
| 30 |
+
"pv_layers/layer_012.json",
|
| 31 |
+
"pv_layers/layer_013.json",
|
| 32 |
+
"pv_layers/layer_014.json",
|
| 33 |
+
"pv_layers/layer_015.json",
|
| 34 |
+
"pv_layers/layer_016.json",
|
| 35 |
+
"pv_layers/layer_017.json",
|
| 36 |
+
"pv_layers/layer_018.json",
|
| 37 |
+
"pv_layers/layer_019.json",
|
| 38 |
+
"pv_layers/layer_020.json",
|
| 39 |
+
"pv_layers/layer_021.json",
|
| 40 |
+
"pv_layers/layer_022.json",
|
| 41 |
+
"pv_layers/layer_023.json",
|
| 42 |
+
"pv_layers/layer_024.json",
|
| 43 |
+
"pv_layers/layer_025.json",
|
| 44 |
+
"pv_layers/layer_026.json",
|
| 45 |
+
"pv_layers/layer_027.json",
|
| 46 |
+
"pv_layers/layer_028.json",
|
| 47 |
+
"pv_layers/layer_029.json",
|
| 48 |
+
"pv_layers/layer_030.json",
|
| 49 |
+
"pv_layers/layer_031.json",
|
| 50 |
+
"pv_layers/layer_032.json",
|
| 51 |
+
"pv_layers/layer_033.json",
|
| 52 |
+
"pv_layers/layer_034.json",
|
| 53 |
+
"pv_layers/layer_035.json",
|
| 54 |
+
"pv_layers/layer_036.json",
|
| 55 |
+
"pv_layers/layer_037.json",
|
| 56 |
+
"pv_layers/layer_038.json",
|
| 57 |
+
"pv_layers/layer_039.json",
|
| 58 |
+
"pv_layers/layer_040.json",
|
| 59 |
+
"pv_layers/layer_041.json",
|
| 60 |
+
"pv_layers/layer_042.json",
|
| 61 |
+
"pv_layers/layer_043.json",
|
| 62 |
+
"pv_layers/layer_044.json",
|
| 63 |
+
"pv_layers/layer_045.json",
|
| 64 |
+
"pv_layers/layer_046.json",
|
| 65 |
+
"pv_layers/layer_047.json",
|
| 66 |
+
"pv_layers/layer_048.json",
|
| 67 |
+
"pv_layers/layer_049.json",
|
| 68 |
+
"pv_layers/layer_050.json",
|
| 69 |
+
"pv_layers/layer_051.json",
|
| 70 |
+
"pv_layers/layer_052.json",
|
| 71 |
+
"pv_layers/layer_053.json",
|
| 72 |
+
"pv_layers/layer_054.json",
|
| 73 |
+
"pv_layers/layer_055.json",
|
| 74 |
+
"pv_layers/layer_056.json",
|
| 75 |
+
"pv_layers/layer_057.json",
|
| 76 |
+
"pv_layers/layer_058.json",
|
| 77 |
+
"pv_layers/layer_059.json",
|
| 78 |
+
"pv_layers/layer_060.json",
|
| 79 |
+
"pv_layers/layer_061.json",
|
| 80 |
+
"pv_layers/layer_062.json",
|
| 81 |
+
"pv_layers/layer_063.json",
|
| 82 |
+
"pv_layers/layer_064.json",
|
| 83 |
+
"pv_layers/layer_065.json",
|
| 84 |
+
"pv_layers/layer_066.json",
|
| 85 |
+
"pv_layers/layer_067.json",
|
| 86 |
+
"pv_layers/layer_068.json",
|
| 87 |
+
"pv_layers/layer_069.json",
|
| 88 |
+
"pv_layers/layer_070.json",
|
| 89 |
+
"pv_layers/layer_071.json",
|
| 90 |
+
"pv_layers/layer_072.json",
|
| 91 |
+
"pv_layers/layer_073.json",
|
| 92 |
+
"pv_layers/layer_074.json",
|
| 93 |
+
"pv_layers/layer_075.json",
|
| 94 |
+
"pv_layers/layer_076.json",
|
| 95 |
+
"pv_layers/layer_077.json",
|
| 96 |
+
"pv_progress.json",
|
| 97 |
+
"source.json",
|
| 98 |
+
"tokenizer_config.json"
|
| 99 |
+
]
|
| 100 |
+
}
|
reproduce/data_artifacts.json
ADDED
|
@@ -0,0 +1,333 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"note": "Recovered local artifact hashes; current paths may have been reused. v32 hashes belong to v2. No raw texts, local logs, private chats or image bytes are included. v1 original 18,001,846-token flat stream is not recovered by this manifest.",
|
| 3 |
+
"files": [
|
| 4 |
+
{
|
| 5 |
+
"path": "/tmp/glm52-calib-v3/shard_00000.npy",
|
| 6 |
+
"sha256": "84dba15d4fec9e222ab68d2196a926dfd60123a7741530b3cbc02a84e1f20cb8",
|
| 7 |
+
"bytes": 4000128,
|
| 8 |
+
"shape": [
|
| 9 |
+
1000000
|
| 10 |
+
],
|
| 11 |
+
"dtype": "uint32"
|
| 12 |
+
},
|
| 13 |
+
{
|
| 14 |
+
"path": "/tmp/glm52-calib-v3/shard_00001.npy",
|
| 15 |
+
"sha256": "25310686c241a8ddfd5463fa240bf0a9549885772cc0d0ff75d15138f00c1ea0",
|
| 16 |
+
"bytes": 4000128,
|
| 17 |
+
"shape": [
|
| 18 |
+
1000000
|
| 19 |
+
],
|
| 20 |
+
"dtype": "uint32"
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"path": "/tmp/glm52-calib-v3/shard_00002.npy",
|
| 24 |
+
"sha256": "cbcfa8c283d8e894e83559198c52e65a95455ceee9bb5061c9cce90ec53478b9",
|
| 25 |
+
"bytes": 4000128,
|
| 26 |
+
"shape": [
|
| 27 |
+
1000000
|
| 28 |
+
],
|
| 29 |
+
"dtype": "uint32"
|
| 30 |
+
},
|
| 31 |
+
{
|
| 32 |
+
"path": "/tmp/glm52-calib-v3/shard_00003.npy",
|
| 33 |
+
"sha256": "b64418fda40e375922baa5755b8c01ef7b4b14b8d40c1fcd2ebd9c80e1cec381",
|
| 34 |
+
"bytes": 4000128,
|
| 35 |
+
"shape": [
|
| 36 |
+
1000000
|
| 37 |
+
],
|
| 38 |
+
"dtype": "uint32"
|
| 39 |
+
},
|
| 40 |
+
{
|
| 41 |
+
"path": "/tmp/glm52-calib-v3/shard_00004.npy",
|
| 42 |
+
"sha256": "b42bd5638c244923766898100cb0e75458919a3a127d42789c925d81ebe59699",
|
| 43 |
+
"bytes": 4000128,
|
| 44 |
+
"shape": [
|
| 45 |
+
1000000
|
| 46 |
+
],
|
| 47 |
+
"dtype": "uint32"
|
| 48 |
+
},
|
| 49 |
+
{
|
| 50 |
+
"path": "/tmp/glm52-calib-v3/shard_00005.npy",
|
| 51 |
+
"sha256": "a672f7431ddefb1b5233fc7e53e6f0fc09a36395231c5c2c8d8a7d4056fe74eb",
|
| 52 |
+
"bytes": 4000128,
|
| 53 |
+
"shape": [
|
| 54 |
+
1000000
|
| 55 |
+
],
|
| 56 |
+
"dtype": "uint32"
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"path": "/tmp/glm52-calib-v3/shard_00006.npy",
|
| 60 |
+
"sha256": "e6dcafae0e4018eb7d99b7517366f05375570d934476ad0456b4e4448c890719",
|
| 61 |
+
"bytes": 4000128,
|
| 62 |
+
"shape": [
|
| 63 |
+
1000000
|
| 64 |
+
],
|
| 65 |
+
"dtype": "uint32"
|
| 66 |
+
},
|
| 67 |
+
{
|
| 68 |
+
"path": "/tmp/glm52-calib-v3/shard_00007.npy",
|
| 69 |
+
"sha256": "eb22593102c11b48a182852989f17399ca32b3bc77196b73b22cb3e3d58190f9",
|
| 70 |
+
"bytes": 4000128,
|
| 71 |
+
"shape": [
|
| 72 |
+
1000000
|
| 73 |
+
],
|
| 74 |
+
"dtype": "uint32"
|
| 75 |
+
},
|
| 76 |
+
{
|
| 77 |
+
"path": "/tmp/glm52-calib-v3/shard_00008.npy",
|
| 78 |
+
"sha256": "610bfc9c5dcaf030eabfeca2c4150f409391102634f9dff003a58ca32bb5b0ed",
|
| 79 |
+
"bytes": 4000128,
|
| 80 |
+
"shape": [
|
| 81 |
+
1000000
|
| 82 |
+
],
|
| 83 |
+
"dtype": "uint32"
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"path": "/tmp/glm52-calib-v3/shard_00009.npy",
|
| 87 |
+
"sha256": "0464899366bb7429800be3f61486af04b871abde381a28c644a44b14793ff9a8",
|
| 88 |
+
"bytes": 4000128,
|
| 89 |
+
"shape": [
|
| 90 |
+
1000000
|
| 91 |
+
],
|
| 92 |
+
"dtype": "uint32"
|
| 93 |
+
},
|
| 94 |
+
{
|
| 95 |
+
"path": "/tmp/glm52-calib-v3/shard_00010.npy",
|
| 96 |
+
"sha256": "9691d0cfce1ec93e5ace299c4045e9c1bb2e53adcf713adccbd1336ebe9287a7",
|
| 97 |
+
"bytes": 4000128,
|
| 98 |
+
"shape": [
|
| 99 |
+
1000000
|
| 100 |
+
],
|
| 101 |
+
"dtype": "uint32"
|
| 102 |
+
},
|
| 103 |
+
{
|
| 104 |
+
"path": "/tmp/glm52-calib-v3/shard_00011.npy",
|
| 105 |
+
"sha256": "d48e3552bcbf2251ecbd98febc390167d15514f265d0cf4bab2169a173266438",
|
| 106 |
+
"bytes": 4000128,
|
| 107 |
+
"shape": [
|
| 108 |
+
1000000
|
| 109 |
+
],
|
| 110 |
+
"dtype": "uint32"
|
| 111 |
+
},
|
| 112 |
+
{
|
| 113 |
+
"path": "/tmp/glm52-calib-v3/shard_00012.npy",
|
| 114 |
+
"sha256": "255fa27435f0948f0af499e474810aa877951c68c84655afa5a7376c47592c04",
|
| 115 |
+
"bytes": 4000128,
|
| 116 |
+
"shape": [
|
| 117 |
+
1000000
|
| 118 |
+
],
|
| 119 |
+
"dtype": "uint32"
|
| 120 |
+
},
|
| 121 |
+
{
|
| 122 |
+
"path": "/tmp/glm52-calib-v3/shard_00013.npy",
|
| 123 |
+
"sha256": "94220e009467d6b645902d2bdd429aa29544a4eb9fa5d4f9b24bff823b8e0496",
|
| 124 |
+
"bytes": 4000128,
|
| 125 |
+
"shape": [
|
| 126 |
+
1000000
|
| 127 |
+
],
|
| 128 |
+
"dtype": "uint32"
|
| 129 |
+
},
|
| 130 |
+
{
|
| 131 |
+
"path": "/tmp/glm52-calib-v3/shard_00014.npy",
|
| 132 |
+
"sha256": "37d7b9604e62a12eef2c293ed899f6bb2a8032418e5c5b7c578c82ff29e5bf0c",
|
| 133 |
+
"bytes": 4000128,
|
| 134 |
+
"shape": [
|
| 135 |
+
1000000
|
| 136 |
+
],
|
| 137 |
+
"dtype": "uint32"
|
| 138 |
+
},
|
| 139 |
+
{
|
| 140 |
+
"path": "/tmp/glm52-calib-v3/shard_00015.npy",
|
| 141 |
+
"sha256": "a7a3a29377ef4436d6ec5185d5f45f2c7e615fbb093499e59a3651d78ef0181d",
|
| 142 |
+
"bytes": 26120,
|
| 143 |
+
"shape": [
|
| 144 |
+
6498
|
| 145 |
+
],
|
| 146 |
+
"dtype": "uint32"
|
| 147 |
+
},
|
| 148 |
+
{
|
| 149 |
+
"path": "/tmp/glm52-calib-v32/shard_00000.npy",
|
| 150 |
+
"sha256": "587d01f48f9d39719661004f51d288bed7ccba5808c2a86068ad83a72e8e9130",
|
| 151 |
+
"bytes": 4000128,
|
| 152 |
+
"shape": [
|
| 153 |
+
1000000
|
| 154 |
+
],
|
| 155 |
+
"dtype": "uint32"
|
| 156 |
+
},
|
| 157 |
+
{
|
| 158 |
+
"path": "/tmp/glm52-calib-v32/shard_00001.npy",
|
| 159 |
+
"sha256": "b34946701d15f586a7c7513ab37aa43c7bdbf1e9a51448a90ddea81850df8011",
|
| 160 |
+
"bytes": 4000128,
|
| 161 |
+
"shape": [
|
| 162 |
+
1000000
|
| 163 |
+
],
|
| 164 |
+
"dtype": "uint32"
|
| 165 |
+
},
|
| 166 |
+
{
|
| 167 |
+
"path": "/tmp/glm52-calib-v32/shard_00002.npy",
|
| 168 |
+
"sha256": "42b98201ecbfd0830fc333c6691ec67daaa6b82287b8d33a8ac2670dca14d842",
|
| 169 |
+
"bytes": 4000128,
|
| 170 |
+
"shape": [
|
| 171 |
+
1000000
|
| 172 |
+
],
|
| 173 |
+
"dtype": "uint32"
|
| 174 |
+
},
|
| 175 |
+
{
|
| 176 |
+
"path": "/tmp/glm52-calib-v32/shard_00003.npy",
|
| 177 |
+
"sha256": "87c619b34e746d40ab1b32827317aca92a5583d7a9a010c6199ccb63aeb52f8f",
|
| 178 |
+
"bytes": 4000128,
|
| 179 |
+
"shape": [
|
| 180 |
+
1000000
|
| 181 |
+
],
|
| 182 |
+
"dtype": "uint32"
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"path": "/tmp/glm52-calib-v32/shard_00004.npy",
|
| 186 |
+
"sha256": "9ea10b0107c7afbb82256c5e00ddc001f5b0288d9f35e69712e88ffd20d70b28",
|
| 187 |
+
"bytes": 4000128,
|
| 188 |
+
"shape": [
|
| 189 |
+
1000000
|
| 190 |
+
],
|
| 191 |
+
"dtype": "uint32"
|
| 192 |
+
},
|
| 193 |
+
{
|
| 194 |
+
"path": "/tmp/glm52-calib-v32/shard_00005.npy",
|
| 195 |
+
"sha256": "5941cd2282b9c9eeeffcb500044a20c8bf5cb27d03ca3e0783491d956141af56",
|
| 196 |
+
"bytes": 4000128,
|
| 197 |
+
"shape": [
|
| 198 |
+
1000000
|
| 199 |
+
],
|
| 200 |
+
"dtype": "uint32"
|
| 201 |
+
},
|
| 202 |
+
{
|
| 203 |
+
"path": "/tmp/glm52-calib-v32/shard_00006.npy",
|
| 204 |
+
"sha256": "3a83b4ae5266f229e4431193594462f38f82357f0ecb20bb5bf2b79c1703e1ab",
|
| 205 |
+
"bytes": 4000128,
|
| 206 |
+
"shape": [
|
| 207 |
+
1000000
|
| 208 |
+
],
|
| 209 |
+
"dtype": "uint32"
|
| 210 |
+
},
|
| 211 |
+
{
|
| 212 |
+
"path": "/tmp/glm52-calib-v32/shard_00007.npy",
|
| 213 |
+
"sha256": "564afecdb45368874c7b371ed4627188b14d7a84ab3a8e2f26c26433fa228548",
|
| 214 |
+
"bytes": 4000128,
|
| 215 |
+
"shape": [
|
| 216 |
+
1000000
|
| 217 |
+
],
|
| 218 |
+
"dtype": "uint32"
|
| 219 |
+
},
|
| 220 |
+
{
|
| 221 |
+
"path": "/tmp/glm52-calib-v32/shard_00008.npy",
|
| 222 |
+
"sha256": "0a56d27bcab2ad9149e138c85b513016c97cc0950d8911ae49ed83e12c6f6c8c",
|
| 223 |
+
"bytes": 4000128,
|
| 224 |
+
"shape": [
|
| 225 |
+
1000000
|
| 226 |
+
],
|
| 227 |
+
"dtype": "uint32"
|
| 228 |
+
},
|
| 229 |
+
{
|
| 230 |
+
"path": "/tmp/glm52-calib-v32/shard_00009.npy",
|
| 231 |
+
"sha256": "7fe80d4b868c2d7f1d0a162a7fd80782f7d69f8f50ec6976f6a75c1a24dc2aca",
|
| 232 |
+
"bytes": 4000128,
|
| 233 |
+
"shape": [
|
| 234 |
+
1000000
|
| 235 |
+
],
|
| 236 |
+
"dtype": "uint32"
|
| 237 |
+
},
|
| 238 |
+
{
|
| 239 |
+
"path": "/tmp/glm52-calib-v32/shard_00010.npy",
|
| 240 |
+
"sha256": "61d886db2533ec343ca240ccdd859118476bd37dd59ba37cfcb0dfb56c6d5f57",
|
| 241 |
+
"bytes": 4000128,
|
| 242 |
+
"shape": [
|
| 243 |
+
1000000
|
| 244 |
+
],
|
| 245 |
+
"dtype": "uint32"
|
| 246 |
+
},
|
| 247 |
+
{
|
| 248 |
+
"path": "/tmp/glm52-calib-v32/shard_00011.npy",
|
| 249 |
+
"sha256": "d7e4af049ce7bf3f0186973d73f66990403f365deac2b1ba1c26bd04be917bfb",
|
| 250 |
+
"bytes": 4000128,
|
| 251 |
+
"shape": [
|
| 252 |
+
1000000
|
| 253 |
+
],
|
| 254 |
+
"dtype": "uint32"
|
| 255 |
+
},
|
| 256 |
+
{
|
| 257 |
+
"path": "/tmp/glm52-calib-v32/shard_00012.npy",
|
| 258 |
+
"sha256": "23bbeb34c4ac0f753192f1d8d17d3def7ba5cf6449149549bfe1ff9c0ef404e7",
|
| 259 |
+
"bytes": 4000128,
|
| 260 |
+
"shape": [
|
| 261 |
+
1000000
|
| 262 |
+
],
|
| 263 |
+
"dtype": "uint32"
|
| 264 |
+
},
|
| 265 |
+
{
|
| 266 |
+
"path": "/tmp/glm52-calib-v32/shard_00013.npy",
|
| 267 |
+
"sha256": "36f1aba2e26747a9227f7d60edde31d4e18b64f402578b42ccd1bbd16258414a",
|
| 268 |
+
"bytes": 4000128,
|
| 269 |
+
"shape": [
|
| 270 |
+
1000000
|
| 271 |
+
],
|
| 272 |
+
"dtype": "uint32"
|
| 273 |
+
},
|
| 274 |
+
{
|
| 275 |
+
"path": "/tmp/glm52-calib-v32/shard_00014.npy",
|
| 276 |
+
"sha256": "4d86d75d69d69b86fae7eba6aaf31153d9a043620bb94daee285925b4ce8a353",
|
| 277 |
+
"bytes": 4000128,
|
| 278 |
+
"shape": [
|
| 279 |
+
1000000
|
| 280 |
+
],
|
| 281 |
+
"dtype": "uint32"
|
| 282 |
+
},
|
| 283 |
+
{
|
| 284 |
+
"path": "/tmp/glm52-calib-v32/shard_00015.npy",
|
| 285 |
+
"sha256": "6a5fc10154611f997bc9957ced9c43f37c14372289a5c89d850a5f405358c5b9",
|
| 286 |
+
"bytes": 31144,
|
| 287 |
+
"shape": [
|
| 288 |
+
7754
|
| 289 |
+
],
|
| 290 |
+
"dtype": "uint32"
|
| 291 |
+
},
|
| 292 |
+
{
|
| 293 |
+
"path": "/tmp/glm53-vision-trace-refit/audit/tokens/shard_000.npy",
|
| 294 |
+
"sha256": "3f9fead556891373590e7017016833d7395674bdcf60c052dc663fb3190932bc",
|
| 295 |
+
"bytes": 1160816,
|
| 296 |
+
"shape": [
|
| 297 |
+
290172
|
| 298 |
+
],
|
| 299 |
+
"dtype": "uint32"
|
| 300 |
+
},
|
| 301 |
+
{
|
| 302 |
+
"path": "/tmp/glm53-vision-trace-refit/train/tokens/shard_000.npy",
|
| 303 |
+
"sha256": "587d01f48f9d39719661004f51d288bed7ccba5808c2a86068ad83a72e8e9130",
|
| 304 |
+
"bytes": 4000128,
|
| 305 |
+
"shape": [
|
| 306 |
+
1000000
|
| 307 |
+
],
|
| 308 |
+
"dtype": "uint32"
|
| 309 |
+
},
|
| 310 |
+
{
|
| 311 |
+
"path": "/tmp/glm53-vision-trace-refit/validation/tokens/shard_000.npy",
|
| 312 |
+
"sha256": "9f943a73f4a9eed1f8f817b34653f242f212b5df108d66dfeea4e7f0a5357f47",
|
| 313 |
+
"bytes": 1605812,
|
| 314 |
+
"shape": [
|
| 315 |
+
401421
|
| 316 |
+
],
|
| 317 |
+
"dtype": "uint32"
|
| 318 |
+
},
|
| 319 |
+
{
|
| 320 |
+
"path": "/tmp/glm53-layer3-full-corpus/train.npy",
|
| 321 |
+
"sha256": "26812c90e966c66bbe33b0e0c8bf9572906deb65eda7ce022f73c58f04eda29a",
|
| 322 |
+
"shape": [
|
| 323 |
+
15007754
|
| 324 |
+
],
|
| 325 |
+
"dtype": "uint32",
|
| 326 |
+
"boundary_counts": {
|
| 327 |
+
"154820": 18854,
|
| 328 |
+
"154841": 3099,
|
| 329 |
+
"154842": 3099
|
| 330 |
+
}
|
| 331 |
+
}
|
| 332 |
+
]
|
| 333 |
+
}
|
reproduce/environment_observed.json
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"captured_on": "2026-09-26",
|
| 3 |
+
"versions": {
|
| 4 |
+
"torch": "2.11.0+cu130",
|
| 5 |
+
"numpy": "2.2.6",
|
| 6 |
+
"transformers": "5.13.0",
|
| 7 |
+
"datasets": "5.0.0",
|
| 8 |
+
"safetensors": "0.8.0",
|
| 9 |
+
"huggingface_hub": "1.22.0",
|
| 10 |
+
"Pillow": "12.3.0",
|
| 11 |
+
"scipy": "1.18.0",
|
| 12 |
+
"triton": "3.6.0"
|
| 13 |
+
},
|
| 14 |
+
"historical_environment_exactly_recovered": false,
|
| 15 |
+
"note": "Current audit environment, not an original training lockfile. Pin a tested CUDA/PyTorch/Transformers stack before rerunning."
|
| 16 |
+
}
|
reproduce/external_sources.json
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"scope": "Observed checkout commits, not proof of original dataset/training source state",
|
| 3 |
+
"vllm": {
|
| 4 |
+
"url": "https://github.com/vllm-project/vllm.git",
|
| 5 |
+
"revision": "f32e2837ffc41ad369fff3ba8b05461ce5b53b55",
|
| 6 |
+
"use": "optional serving/code-corpus checkout; local dirty diff not reproduced"
|
| 7 |
+
},
|
| 8 |
+
"b12x": {
|
| 9 |
+
"url": "https://github.com/local-inference-lab/b12x",
|
| 10 |
+
"revision": "9043b448622764a598969518d413b3fd8b3c0c07",
|
| 11 |
+
"use": "optional BTX comparison modules, not ARVQ production codec"
|
| 12 |
+
},
|
| 13 |
+
"quant-toolkit": {
|
| 14 |
+
"url": "https://github.com/local-inference-lab/quant-toolkit",
|
| 15 |
+
"revision": "8bdb1016e52dd15f91e42d35c09096c0f31325f3",
|
| 16 |
+
"use": "historical external toolkit; not vendored"
|
| 17 |
+
}
|
| 18 |
+
}
|
reproduce/historical_metadata/METHODOLOGY.md
ADDED
|
@@ -0,0 +1,342 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# ARVQ cold-expert quantization: methodology
|
| 2 |
+
|
| 3 |
+
How the cold Mixture-of-Experts weights in this repository
|
| 4 |
+
(`GLM-5.3-Vision-NVFP4-ARVQ-hybrid`) were quantized and tuned. This document
|
| 5 |
+
reconstructs the actual pipeline and tuning decisions from the build transcripts
|
| 6 |
+
and the fitting/serving code; numbers are cited from those sources. It is a
|
| 7 |
+
methodology record, not a quality claim: full-model task quality and native
|
| 8 |
+
SM120 execution were **not** evaluated for this checkpoint.
|
| 9 |
+
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
## 1. What this model is (the hybrid)
|
| 13 |
+
|
| 14 |
+
GLM-5.3 is a large MoE. Each MoE block routes each token to its top-8 of 256
|
| 15 |
+
routed experts (plus a shared expert). This checkpoint is a **hybrid** in which
|
| 16 |
+
each expert is stored in one of two ways:
|
| 17 |
+
|
| 18 |
+
- **Hot experts** — kept in the NVFP4 donor format (unchanged).
|
| 19 |
+
- **Cold experts** — re-quantized to ~2 bits with **ARVQ** (additive residual
|
| 20 |
+
vector quantization), the subject of this document.
|
| 21 |
+
|
| 22 |
+
The hot/cold split is fixed by a REAP allocation that keeps the **5,750** most
|
| 23 |
+
important experts hot; the remaining **13,450** experts across the model are
|
| 24 |
+
cold (per the 09-14 build transcript, which verified "13,450 cold-expert source
|
| 25 |
+
files"). The allocation is 75% text-REAP / 25% multimodal-salience weighted and
|
| 26 |
+
is **not** modified by this campaign.
|
| 27 |
+
|
| 28 |
+
Only the **75 MoE blocks, layers 3–77**, are touched. Layers 0–2 are dense and
|
| 29 |
+
stay frozen. Everything except the cold experts is inherited unchanged:
|
| 30 |
+
attention, backbone, shared experts, hot NVFP4 experts, BF16 MTP, and the
|
| 31 |
+
vision components.
|
| 32 |
+
|
| 33 |
+
Cold experts are quantized per (layer, projection). The two projections are the
|
| 34 |
+
fused gate/up `w13` (N=4096, K=6144) and the down projection `w2` (N=6144,
|
| 35 |
+
K=2048).
|
| 36 |
+
|
| 37 |
+
---
|
| 38 |
+
|
| 39 |
+
## 2. ARVQ representation and bit budget
|
| 40 |
+
|
| 41 |
+
Each cold expert weight group of **8 contiguous columns** is represented as the
|
| 42 |
+
sum of two codebook atoms times a scale:
|
| 43 |
+
|
| 44 |
+
```
|
| 45 |
+
w_group = global * block_scale * (c0[a] + c1[b])
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
- **Two codebooks** `c0`, `c1`, each **256 entries × 8 dims** (512 codewords per
|
| 49 |
+
pair). This is the "rvq256_256x8" format. Codebooks are **per-expert** in this
|
| 50 |
+
v3 checkpoint (`codebook_scope = expert`, packed format
|
| 51 |
+
`rvq256_256x8_expert`, version 3).
|
| 52 |
+
- **Indices** `a`, `b` are one uint8 each per 8-weight group → **16 index bits
|
| 53 |
+
per group = 2.0 bits/weight**.
|
| 54 |
+
- **Codebook atoms are constrained to the FP4 grid** `{0, ±0.5, ±1, ±1.5, ±2,
|
| 55 |
+
±3, ±4, ±6}` (project-to-FP4 after every update). The FP4 grid on the atoms is
|
| 56 |
+
why no incoherence/Hadamard rotation is used: the target space cannot be
|
| 57 |
+
rescaled without leaving the grid (see `encoder.py` docstring).
|
| 58 |
+
- **Block scales** are stored as **`float8_e4m3fn`**, one scale per **[row,
|
| 59 |
+
128-column block]** (48 blocks/row for `w13`, 16 blocks/row for `w2`). At use
|
| 60 |
+
time each block scale is `repeat_interleave(16)` to cover the 16 eight-wide
|
| 61 |
+
groups in a 128-column block. Adding the FP8 scale (8 bits / 128 weights =
|
| 62 |
+
0.0625 bit/weight) gives the reported **2.0625 bits/weight** plus the
|
| 63 |
+
amortized per-expert codebooks.
|
| 64 |
+
- A single fp32 `global` scalar per (layer, projection) sits above the block
|
| 65 |
+
scales so the index/scale layout is unchanged from earlier versions.
|
| 66 |
+
|
| 67 |
+
Packing (`pack.py`) writes MMA-fragment index tiles (64 uint32 words/tile) and
|
| 68 |
+
verifies, per layer, that every expert's indices round-trip exactly and that a
|
| 69 |
+
sample of decoded weights matches bit-for-bit before publishing.
|
| 70 |
+
|
| 71 |
+
---
|
| 72 |
+
|
| 73 |
+
## 3. Initial fit (Hessian-aware, no PV yet)
|
| 74 |
+
|
| 75 |
+
Before any gradient tuning, each (layer, projection) gets an
|
| 76 |
+
activation-Hessian-tuned initialization (`encoder.py`, `fit.py`):
|
| 77 |
+
|
| 78 |
+
1. **Per-expert raw-basis Hessian.** `w13` uses `H = E[x xᵀ]` over routed
|
| 79 |
+
tokens; `w2` uses `H = E[m mᵀ]` where `m = silu(x·Wgᵀ)·(x·Wuᵀ)` is the BF16
|
| 80 |
+
SwiGLU intermediate. Experts with <32 routed tokens fall back to all captured
|
| 81 |
+
tokens. Activations are stored/computed in **float32**; Hessians are half.
|
| 82 |
+
2. **Escalating-damp Cholesky** of `H⁻¹`: damping starts at `1e-3 · mean(diag H)`
|
| 83 |
+
and multiplies by 4× (up to 12 attempts) until the factorization succeeds.
|
| 84 |
+
3. **Codebook EM** on a Hessian-importance-weighted subsample (20,000 groups/
|
| 85 |
+
expert) of normalized 8-dim groups: k-means warm start → 8 alternating
|
| 86 |
+
refine iterations, FP4-projecting the atoms after every M-step.
|
| 87 |
+
4. **Per-expert LDLQ/GPTQ error-feedback column sweep** with the fixed
|
| 88 |
+
codebooks (col_block=128, 2 sweep passes, 1 inner refine): plain-L2 inner
|
| 89 |
+
assignment, off-diagonal Hessian carries cross-group error feedback.
|
| 90 |
+
5. **Least-squares refit of the E4M3 block scales** given the chosen codes.
|
| 91 |
+
|
| 92 |
+
An earlier decision, recorded 09-14: a **tuned Hadamard rotation was tested and
|
| 93 |
+
rejected** — a matched 48-expert down-projection refit gave 0.372 (H128 rotated)
|
| 94 |
+
versus **0.191 unrotated**, so the unrotated raw-basis fit was adopted.
|
| 95 |
+
|
| 96 |
+
Nonfinite guards are raised as errors throughout (the fit/tuner refuse to
|
| 97 |
+
proceed on NaN/Inf rather than silently scrubbing).
|
| 98 |
+
|
| 99 |
+
---
|
| 100 |
+
|
| 101 |
+
## 4. The PV objective: `--target reference`
|
| 102 |
+
|
| 103 |
+
The cold experts are then output-tuned per layer. The objective for this
|
| 104 |
+
checkpoint is **`--target reference`** (`sequential_pv_full_corpus.py`).
|
| 105 |
+
|
| 106 |
+
For each captured token position the tuner forms:
|
| 107 |
+
|
| 108 |
+
- `frozen` — the block output from everything that is **not** a cold expert on
|
| 109 |
+
the **student** input (retained tuned upstream layers → attention → shared
|
| 110 |
+
expert → hot NVFP4 experts). This is fixed.
|
| 111 |
+
- `reference` — the **original FP8 reference block output**: the unchanged donor
|
| 112 |
+
backbone with the **original FP8 routed experts** (all 256), evaluated on the
|
| 113 |
+
**original reference trajectory** for that token position.
|
| 114 |
+
- `required = reference − frozen` — the residual the cold experts must supply.
|
| 115 |
+
|
| 116 |
+
The loss is the relative squared error of the cold-expert sum against `required`:
|
| 117 |
+
|
| 118 |
+
```
|
| 119 |
+
loss = || student_cold(x_student) − required ||² / (denominator · Σw)
|
| 120 |
+
```
|
| 121 |
+
|
| 122 |
+
where `denominator` is the per-row mean energy of `reference − frozen` over the
|
| 123 |
+
corpus (so the scale is comparable across layers). Selection metric is
|
| 124 |
+
`reference_rel = ||(frozen + cold) − reference|| / ||reference||`, measured over
|
| 125 |
+
the **whole transformer block including the residual**, not the cold-experts sum
|
| 126 |
+
alone.
|
| 127 |
+
|
| 128 |
+
**Why reference (and its known weakness).** The transcripts weighed
|
| 129 |
+
`reference` against a `same_input` target (match the block's output on the
|
| 130 |
+
*student's own* drifted input). The rationale for reference (09-16): "Matching
|
| 131 |
+
the original reference trajectory is more directly aligned with preserving the
|
| 132 |
+
model's behavior." The acknowledged risk: because each layer receives a
|
| 133 |
+
**different (student) input** than the original, the cold experts may lack the
|
| 134 |
+
capacity to reproduce the reference output and could "learn corrections that
|
| 135 |
+
work on training examples but fail on unseen ones" — i.e. reference-target is
|
| 136 |
+
the more out-of-distribution objective. It was chosen only after a matched
|
| 137 |
+
layer-4 A/B:
|
| 138 |
+
|
| 139 |
+
| Layer-4 metric | reference vs same_input |
|
| 140 |
+
|---|---|
|
| 141 |
+
| validation reference error | **0.21% better** with reference |
|
| 142 |
+
| development-audit error | 0.08% **worse** with reference |
|
| 143 |
+
|
| 144 |
+
A small mixed difference; reference was retained for behavioral alignment. Two
|
| 145 |
+
important design consequences: the target is the **complete block output**
|
| 146 |
+
(including the residual), so matching only the MoE contribution cannot leave the
|
| 147 |
+
incoming residual error untouched; and the loss compares the **routing-weighted
|
| 148 |
+
sum of all cold experts** jointly (gate/up and down together).
|
| 149 |
+
|
| 150 |
+
---
|
| 151 |
+
|
| 152 |
+
## 5. Sequential per-layer pipeline with trajectory coupling
|
| 153 |
+
|
| 154 |
+
Layers are processed **sequentially, 3 → 77** (`full_reference.py` /
|
| 155 |
+
`sequential_capture.py`):
|
| 156 |
+
|
| 157 |
+
1. **Capture** (8-way sequence-parallel). For layer L, the student input is the
|
| 158 |
+
**retained tuned output of block L-1** (`student_outputs.pt`), while the
|
| 159 |
+
reference target re-runs the **donor backbone + original FP8 experts** on the
|
| 160 |
+
original reference hidden state. Both trajectories are carried forward in a
|
| 161 |
+
rolling cache; layer 3 is the control (no preceding ARVQ layer, so student =
|
| 162 |
+
reference input). This cross-layer coupling means each layer is tuned against
|
| 163 |
+
the exact drifted inputs it will see at serving time.
|
| 164 |
+
2. **Fit / PV tune** the cold experts (Section 6).
|
| 165 |
+
3. **Gate** (Section 7).
|
| 166 |
+
4. **Export + serialized replay** — decode to the packed format and require the
|
| 167 |
+
reloaded forward to match the tuned forward to <2e-6 before use.
|
| 168 |
+
5. **Per-layer HF publish** — `publish_reference.py` re-assembles cold slots
|
| 169 |
+
across the 8 expert shards, checks slot coverage/allocation/global metadata,
|
| 170 |
+
runs the exact index round-trip and sampled decoded-weight parity, then
|
| 171 |
+
**atomically replaces that layer's two tensor files + reports in one commit**
|
| 172 |
+
with parent-commit protection; remote hashes are re-verified.
|
| 173 |
+
|
| 174 |
+
Parallelism: cold experts are **sharded across 8 GPUs** (each rank owns distinct
|
| 175 |
+
experts); routed outputs are summed with an all-reduce so a single coupled-output
|
| 176 |
+
objective is optimized without a dense-model replica.
|
| 177 |
+
|
| 178 |
+
---
|
| 179 |
+
|
| 180 |
+
## 6. PV tuning numerics (the published full-corpus campaign)
|
| 181 |
+
|
| 182 |
+
Per the model card, all 75 layers were replaced by a **full-corpus sequential
|
| 183 |
+
PV campaign** drawing sequentially from **18,001,846 training tokens** at
|
| 184 |
+
context **1024** (53,248 legacy tokens were removed around 17 exact matches to
|
| 185 |
+
held-out prompts). Fixed **validation** and **development-audit** sets each hold
|
| 186 |
+
**16,384 tokens**.
|
| 187 |
+
|
| 188 |
+
- **Optimizer:** Adam. Two trainable parameter groups — FP4-constrained
|
| 189 |
+
per-expert **codebooks** and FP8-constrained per-block **scales**. (The
|
| 190 |
+
faster fork trains per-row scale deltas instead; the published full-corpus
|
| 191 |
+
fork trains the per-128-block log-scales.)
|
| 192 |
+
- **Effective batch:** 262,144 tokens, accumulated as **four 65,536-token
|
| 193 |
+
microbatch passes** (256 whole 1024-token sequences), one shuffled
|
| 194 |
+
no-replacement pass over the corpus.
|
| 195 |
+
- **Learning-rate schedule:** book/scale LR **0.048 / 0.032** through update 45,
|
| 196 |
+
then dropped by ×0.25 to **0.012 / 0.008** (`lr_decay_after=45`,
|
| 197 |
+
`lr_decay_factor=0.25`). (Note: these campaign LRs are higher than the
|
| 198 |
+
in-repo code defaults of 0.003/0.002, consistent with the much larger
|
| 199 |
+
262k-token batch.)
|
| 200 |
+
- **Update budget:** layers 4–26 use a fixed **69 updates**. From layer 27, 69
|
| 201 |
+
is the maximum with an **adaptive early stop**: three validation checks more
|
| 202 |
+
than 0.1% worse than best trigger an earlier LR reduction; 15 updates without
|
| 203 |
+
improvement after that reduction permit stopping.
|
| 204 |
+
- **Validation every 5 updates**, retaining the best checkpoint.
|
| 205 |
+
- **Index reassignment every 20 updates** (Section 6.1).
|
| 206 |
+
- **Regularization:** codebook drift-from-anchor penalty + scale drift penalty,
|
| 207 |
+
weighted 0.01; grad-norm clip 1.0; codebook atoms clamped to [-6, 6];
|
| 208 |
+
log-scales clamped to within ±0.35 of their initial value.
|
| 209 |
+
- **Coverage gate:** experts with fewer than **256 routed training rows** are
|
| 210 |
+
frozen (gradient masked) and keep their initial books/scales.
|
| 211 |
+
|
| 212 |
+
### 6.1 Output-gradient index reassignment
|
| 213 |
+
|
| 214 |
+
Continuous Adam cannot move the discrete indices, so every 20 updates a discrete
|
| 215 |
+
proposal pass runs (`gradient_indices.py`, method
|
| 216 |
+
`expert_parallel_output_gradient_prefix_backtracking_v1`):
|
| 217 |
+
|
| 218 |
+
1. Backprop the **output** loss to the reconstructed weight of one expert
|
| 219 |
+
projection.
|
| 220 |
+
2. Propose alternative `(a,b)` index pairs near a small gradient step, capped at
|
| 221 |
+
`max_fraction=0.001` of groups, `trust_ratio=0.01`, `target_ratio=0.03`; the
|
| 222 |
+
current pair is always in the candidate set (changing nothing stays legal).
|
| 223 |
+
3. **Prefix backtracking acceptance:** try the top-k proposals with
|
| 224 |
+
k∈{full, ¼, 1/16, 1}, accept only if the **training** SSE strictly drops
|
| 225 |
+
**and** a separate **check batch** SSE does not increase (beyond 1e-7 slack);
|
| 226 |
+
otherwise revert. Proposal deltas are broadcast sparsely (only routed rows).
|
| 227 |
+
|
| 228 |
+
The transcripts flag this as the method's main theoretical weakness versus
|
| 229 |
+
published PV-Tuning: codebooks are tuned for output accuracy while index
|
| 230 |
+
proposals originally came from weight-reconstruction Hessians — the
|
| 231 |
+
output-gradient proposal above was added to close that gap.
|
| 232 |
+
|
| 233 |
+
### 6.2 Cold arithmetic emulation
|
| 234 |
+
|
| 235 |
+
Both training and evaluation emulate the serving numerics (`activation.py`,
|
| 236 |
+
pinned to serving revision `b1380cf7…`): **four FP4 activation planes**
|
| 237 |
+
(`fp4_planes4_fp16_boundaries_v1`) and **FP16 SwiGLU boundaries** (FP32 GEMM,
|
| 238 |
+
FP16 SiLU/product), with straight-through estimators for gradients. This
|
| 239 |
+
emulates the quantization boundaries only — **native SM120 MMA accumulation and
|
| 240 |
+
TP reduction are not bit-exact** and were not qualified.
|
| 241 |
+
|
| 242 |
+
---
|
| 243 |
+
|
| 244 |
+
## 7. Acceptance gates
|
| 245 |
+
|
| 246 |
+
Publication of a layer requires all of:
|
| 247 |
+
|
| 248 |
+
- **Validation non-regression:** `0 ≤ final_reference_rel ≤ initial·(1+1e-6)`.
|
| 249 |
+
- **Development-audit non-regression:** the held-out audit split (never used for
|
| 250 |
+
updates or checkpoint selection) must satisfy
|
| 251 |
+
`audit_reference_rel ≤ initial_audit·(1+1e-6)`; otherwise the layer keeps its
|
| 252 |
+
initialization. There is no absolute error floor — the rule is purely
|
| 253 |
+
final ≤ initial.
|
| 254 |
+
- **Serialized-replay parity:** decoded/reloaded forward matches the tuned
|
| 255 |
+
forward to <2e-6, and remote file hashes are re-verified after upload.
|
| 256 |
+
|
| 257 |
+
The audit is honestly labeled **development data, not an untouched final test**;
|
| 258 |
+
it is a split of the same calibration tokens the initialization already saw. Two
|
| 259 |
+
notable per-layer decisions recorded in the card: **layer 30** retains its
|
| 260 |
+
audit-qualified update-15 checkpoint after its validation-best update-20 failed
|
| 261 |
+
the development audit; **layer 3** retains a separately qualified lower-LR
|
| 262 |
+
refinement.
|
| 263 |
+
|
| 264 |
+
---
|
| 265 |
+
|
| 266 |
+
## 8. Calibration corpus
|
| 267 |
+
|
| 268 |
+
The text corpus (`build_calib_v31.py` + `reasoning_slice.py`) is deterministic
|
| 269 |
+
(seed 42), GLM-tokenized, with the following domain mix by tokens:
|
| 270 |
+
|
| 271 |
+
| Share | Domain | Sources |
|
| 272 |
+
|---|---|---|
|
| 273 |
+
| ~32% | code | local vLLM sources, m-a-p/CodeFeedback, jtatman/python-code-500k |
|
| 274 |
+
| ~20% | tool-calling / agentic | generated GLM chat-template sessions |
|
| 275 |
+
| ~15% | reasoning | OpenR1-Math-220k + dolphin-r1 `<think>…</think>` → answer |
|
| 276 |
+
| ~13% | instruction chat | tatsu-lab/alpaca + coding chat |
|
| 277 |
+
| ~10% | medical | MedQA textbook continuation + medical Q&A |
|
| 278 |
+
| ~10% | prose | vLLM docs markdown + databricks-dolly-15k |
|
| 279 |
+
|
| 280 |
+
The **reasoning slice was added specifically** because the earlier calib-v3 mix
|
| 281 |
+
starved the cold 2-bit experts of reasoning-termination behavior: 87% of its
|
| 282 |
+
`<think>` blocks were empty. The slice emits multi-turn conversations where an
|
| 283 |
+
**earlier** assistant turn carries a real `<think>…</think>` that terminates and
|
| 284 |
+
hands off to an answer, followed by a trailing user turn, so the think-close
|
| 285 |
+
token **`</think>` (id 154842)** renders in-stream rather than as a trailing
|
| 286 |
+
generation prompt. Held-out prompts are explicitly excluded (last 25 vLLM docs,
|
| 287 |
+
last 3 MedQA files, vLLM code beyond index 400 of the seed-42 shuffle), and
|
| 288 |
+
17 exact-match prompts were purged from the training set.
|
| 289 |
+
|
| 290 |
+
The capture/training code also supports **up-weighting boundary rows** (rows
|
| 291 |
+
whose next token is a boundary such as `</think>` / `<|endoftext|>`) via a
|
| 292 |
+
`row_weight` term whose sum normalizes the loss, and an optional matched-mixed
|
| 293 |
+
validation set; the published card describes reasoning-based validation.
|
| 294 |
+
|
| 295 |
+
---
|
| 296 |
+
|
| 297 |
+
## 9. Results captured during the build
|
| 298 |
+
|
| 299 |
+
All errors below are **held-out relative L2**, either over the whole transformer
|
| 300 |
+
block (including residual) or over the cold-experts sum, as noted. They are
|
| 301 |
+
local reconstruction errors, **not** token-accuracy or perplexity.
|
| 302 |
+
|
| 303 |
+
- **Pilot smoke test** (layer 3, one step): full-block error 0.005497 → 0.005456
|
| 304 |
+
(~0.75%), 304/388 expert-projection proposals accepted; exported weights
|
| 305 |
+
reproduced the retained result.
|
| 306 |
+
- **Layer 3** (200-step pilot): full-block reference error **0.005497 →
|
| 307 |
+
0.005271 (−4.12%)**, best at step 200.
|
| 308 |
+
- **Layer 4** (200-step pilot): full-block reference error **0.015760 →
|
| 309 |
+
0.015596 (−1.05%)**.
|
| 310 |
+
- **reference vs same_input** (layer 4, matched): reference 0.21% better on
|
| 311 |
+
validation, 0.08% worse on development audit.
|
| 312 |
+
- **Faster fork parity:** the optimized fork (larger microbatch, specialized
|
| 313 |
+
embedding-lookup backward, sparse proposal broadcast) matched the reference
|
| 314 |
+
fork's validation and audit **exactly** on layers 3 and 4, with all propagated
|
| 315 |
+
BF16 outputs equal; training loops fell from ~1525 s → ~337 s (layer 3) and
|
| 316 |
+
~1226 s → ~303 s (layer 4).
|
| 317 |
+
- **Early-layer stability example** (layer 18, full-corpus): initial
|
| 318 |
+
0.017205, step-5 0.017176, step-69 0.017183 — differences ≤0.17%, treated as
|
| 319 |
+
a signal to watch LR/noise rather than proof of convergence.
|
| 320 |
+
|
| 321 |
+
Older cold-only checkpoints reported ~0.2211→0.2070 (layer 3) and 0.2364→0.2320,
|
| 322 |
+
but those measured **cold-expert output only on different captures** and are
|
| 323 |
+
**not comparable** to the full-block errors above.
|
| 324 |
+
|
| 325 |
+
**No comparison against the FP8 donor or an AQLM variant on perplexity / KLD /
|
| 326 |
+
top-1 / wikitext was recorded** in the reviewed transcripts, and no full-model
|
| 327 |
+
task evaluation was performed. Those remain open.
|
| 328 |
+
|
| 329 |
+
---
|
| 330 |
+
|
| 331 |
+
## 10. Honest limitations
|
| 332 |
+
|
| 333 |
+
- Full-model quality and native **SM120** execution were **not** evaluated; the
|
| 334 |
+
arithmetic emulation covers FP4/FP16 boundaries only, not MMA/TP bit-exactness.
|
| 335 |
+
- The development audit is a split of calibration data, not an independent test.
|
| 336 |
+
- The reference target is the more OOD objective; its generalization advantage
|
| 337 |
+
over same_input was small and mixed on the one matched layer tested.
|
| 338 |
+
- Gates enforce local non-regression, not any absolute quality bar.
|
| 339 |
+
|
| 340 |
+
*Prepared from the GLM-5.3 ARVQ build transcripts (2026-09-14 and 2026-09-16)
|
| 341 |
+
and the `btx53/arvq88` fitting/serving code. Where a number could not be sourced
|
| 342 |
+
it is stated as unknown rather than estimated.*
|
reproduce/historical_metadata/backbone_sources.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
reproduce/historical_metadata/build_provenance.json
ADDED
|
@@ -0,0 +1,178 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"stage": "per-expert initial fit; alternating PV in progress",
|
| 3 |
+
"allocation": {
|
| 4 |
+
"mode": "ARVQ output-benefit allocation",
|
| 5 |
+
"text_weight": 0.75,
|
| 6 |
+
"mm_weight": 0.25,
|
| 7 |
+
"hot_count": 5750,
|
| 8 |
+
"floor": 8,
|
| 9 |
+
"cap": 176,
|
| 10 |
+
"text_partition": "train only",
|
| 11 |
+
"metric": "positive routing-weighted squared-error reduction from retaining NVFP4 over fitted ARVQ",
|
| 12 |
+
"scores_sha256": {
|
| 13 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_3.npz": "99e689edad7a7f026759c9c66e682b499ca3263e6d19ad54f59c1f1dccdb3ea9",
|
| 14 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_4.npz": "7813ddafdb23f7c4802bf431a2d063f7afc9f3463356079bc97ef8289e30d823",
|
| 15 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_5.npz": "551940cef83bd52f2e032bd738a88fe293e74efc3e918cfdabf025e111c0d586",
|
| 16 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_6.npz": "1949e2540a0d92169ae9e5374789443d026b0868860149e785342eee19b521b5",
|
| 17 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_7.npz": "cb1d4f6597057c32d2d8136700ea357b86a12ccb881d4bb4a613033c6ea0a34e",
|
| 18 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_8.npz": "c2d4cfb7641a3a9fe96abffa76ab59774eb37cc629b058b3f9bebb1f98bbb09e",
|
| 19 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_9.npz": "c82bff242306c5b27acac0bdd8530e9ad603c1e00642855face4a9aefd03f97a",
|
| 20 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_10.npz": "59fcb8922894f0155704577296cd8752323d1f880a51e801a5b5b1c969647aa0",
|
| 21 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_11.npz": "45b94e85267b83b94203eddca795a21de68fb359a55dd0e304082597a8cb2cee",
|
| 22 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_12.npz": "a277261ee1c2b60dd45de874e0e7e4819356480af132fcb23dfff411a2775077",
|
| 23 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_13.npz": "911c1ba89ea9807704759cad0522532127329259b19ba3ade47ce95c1a5c4e49",
|
| 24 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_14.npz": "e672d24e525c3c077ff2a9a2617b6435e0d3d97398cd2c901f5298bd4a0b7197",
|
| 25 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_15.npz": "e163205f5eb51d0c8796df2070fde3bf80caf3bdf166311336669f53a695c437",
|
| 26 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_16.npz": "9cda2a0e36f882a5daa485c698b9d1eb5ff3aba7ffeac81f7706bed09214f6c6",
|
| 27 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_17.npz": "0f085320e0e1f09bc52183d619a8d91535bdb33b5d56c3de05f212b4337bc2a7",
|
| 28 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_18.npz": "627d542b236da70a5dd2ed098d21cee75dc4cb38db6105369f947361dbe7448a",
|
| 29 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_19.npz": "3613a6fa1e852c110ee05ad901a4c89e1ba659c4782b88988bc935a3a69406bf",
|
| 30 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_20.npz": "7be6a86b425bb811436bdf0f1e4a9a55a0d64aa8a2a7df163923a1898e88e3d9",
|
| 31 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_21.npz": "4abd00a72ff0e76af4cf708caace639e945bce378168ee47892db2645af0de81",
|
| 32 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_22.npz": "92deb3e93534084dd0f43a0f3ab6a11bde3156b5a6218128f03ad58c31e5da08",
|
| 33 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_23.npz": "6866eee649260c556ad663a9f2b198109d1a3aa65dff69a26ef9a33ebfe90b17",
|
| 34 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_24.npz": "9d2a097d7039b8016e50fe48ef7db9175bd6af7652fb8f02d2845e0608c085c8",
|
| 35 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_25.npz": "9675f20e29ebc295606eff124e7eea8e512123d3a8ec37774fac5db06797a5c6",
|
| 36 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_26.npz": "db25e41b5b130d39bcb6d508f719eaa76164b71813fa2bbdba61a249653f9fe3",
|
| 37 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_27.npz": "07d874eae8001ca5fc0e976ba7199afa6e83f54925606b0a8e2ea3836920dbcd",
|
| 38 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_28.npz": "b83eafbc8d8cf70d62c805c37cc3b6b6a2554f2236cd858a84d44696a45a526b",
|
| 39 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_29.npz": "084317cbf70cdb24996a55bda7423b32eb2ee807209d73a0baba7fcad1307cd7",
|
| 40 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_30.npz": "68a6b57cae5acdb0d01fd498ff324ae4e3b01a3c9e3121d7814b1d93fc53cb78",
|
| 41 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_31.npz": "b2f421260f518195814fd1f42b10dbf59060dfd08ad726e1afd0de2f6b5f7442",
|
| 42 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_32.npz": "88f73b9073cf9b4532aac041fad881f9df4392fea7b367101e361e643183a01a",
|
| 43 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_33.npz": "cce07637bfafcc225057a4c7e3c304c167965ca28b551e7da0f0d13f41b00505",
|
| 44 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_34.npz": "0251a6abee66ab96080a23ff1b545ad933841525a8782d3de3dac5b472a681bb",
|
| 45 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_35.npz": "1dc1d919c3fb1de2ae8ec4765e5576e53de6049d43608d3f2e2c5eb687b01b0b",
|
| 46 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_36.npz": "43815d9f766ec63f51268edd90433bf6331327766888a1a089cf7bf31c9a28b6",
|
| 47 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_37.npz": "dc7a1f5b296bc6105a8ca969f558e2f1261238889abfc73326b50ef83b2c0216",
|
| 48 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_38.npz": "2fb9251355a449817aef26486305cc41edace820a91d6113a2f3a0b8f05258a8",
|
| 49 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_39.npz": "f69d1d2da6fd0b3a277e0d3810be2eff1e95050742e26013a30226cb2ba86b50",
|
| 50 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_40.npz": "c59aeceb845ee137ebf890bf5789d181198272838d1f794bf2c96ae31502e940",
|
| 51 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_41.npz": "d962c915b2b672f58bd5f1bfa65e806f09dd5a842a5986c1ca4ac03422620130",
|
| 52 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_42.npz": "acbdc6f83401fac52ce2a5a5e4d174be609842ffba5f1a3e219e9faae913c55a",
|
| 53 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_43.npz": "7ffa833b873294a9f6f1f005644b1101fda459b10ecdb5375eb4c85f7cf6ad5b",
|
| 54 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_44.npz": "906e81150c5f462b68750f74966df5bb93527a30f3b2cb0ecc410da5c9859840",
|
| 55 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_45.npz": "df9f6938632d32af48ceef8013694d0a92bd7208f4335fb06dc74ee2cd38ba12",
|
| 56 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_46.npz": "9fd39f469414d71da762d10b7671e5e0bdcbf087808804e2c868b73193702065",
|
| 57 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_47.npz": "9d2eb792b6553fa078c5cecac1ac53db461e5f2b7bc3b971246f6b8bc7846477",
|
| 58 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_48.npz": "49b108964ddb44d54b4d2375e82772d14275e94ab43da17855953717367ff583",
|
| 59 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_49.npz": "6775fdc746a2b11a443645140a86d1a11dcb389df71289fbef3fe300dec787e2",
|
| 60 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_50.npz": "f6d604b98a1c161b023f42205abf6d1eeecc9119871a8e920623908af0784fc9",
|
| 61 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_51.npz": "0aac9c4e66efaf5c0924107749e82a1e19e70a0b8e31db46434acf38353b8312",
|
| 62 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_52.npz": "7601958d27deb50c0bbf21a6ebb79c20859ca08c962b6f751f972b0f3453bbe6",
|
| 63 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_53.npz": "7bf959678dfc05ec4ef9f7b063d74a3de80e3cf757a199ecf083988ca55d41b7",
|
| 64 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_54.npz": "53097dbffee6e800d22551ed5ea068fe5bae3bad836813e2288f0667612bb6e6",
|
| 65 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_55.npz": "1ec4c9979cc55150e3e3163f8bf52f5d4042e351c15429bd11acc8bc7a77d273",
|
| 66 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_56.npz": "1675552ccfc8929bd29d7eed04e18e387222858beadf170e9652e17b27966766",
|
| 67 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_57.npz": "7e7029ec0de486aad04720d1059b69d74358c59432a33aa16e526041f4506609",
|
| 68 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_58.npz": "813756873a97c71e40fdff33313809723d599dbef2b2ebf36af0d56118b4e62d",
|
| 69 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_59.npz": "fcaadf941258b398bd6bf6a96f7865fccfe9008bf4ac1b93df8896065f761383",
|
| 70 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_60.npz": "e73c14eb000ec4c2fdadb4078011f12ab44bfd470bf690c22c3e0ffdbe89d11a",
|
| 71 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_61.npz": "a4c10b888ee7879054c08e4de97540e786f6c44c8823f3b7988dfb0fbe1b8927",
|
| 72 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_62.npz": "b73347956fed5d783082fe8dd76027372ec50db12683447b180e9a053f9abe61",
|
| 73 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_63.npz": "fe9211a06376cee25e813131672c6101c4564a7c27b5820eafe54b3efa0623e1",
|
| 74 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_64.npz": "348ecb20769be5b1bec0cb31157f513b92f911ef36a8439bef7d45b31321f4d9",
|
| 75 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_65.npz": "2d37ba213e5e4c5e598aedb928c7f50f16205665253d9fc04cad569aa01384db",
|
| 76 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_66.npz": "28615fbdc3c2f0b2458d38a52f05e041dd02501b3d18448ba449b4626a334cd6",
|
| 77 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_67.npz": "9e5a6100b7acb907c99f80ce1cf89f2e46e392c7ca48bb000eda2c126ae1413a",
|
| 78 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_68.npz": "dc0fdd42f3332b73508943c6f9ed0c4166ca0065c2b444823eccbd37a451be6f",
|
| 79 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_69.npz": "4a2b3139ba9f6cd9327285f6ca2417eeacbca49c6e96d8dc525f80e1a7d6eed0",
|
| 80 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_70.npz": "6c03c2284504c6b89cd7ea6af120b6236691dee58ecb2ab1680be19e379068e0",
|
| 81 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_71.npz": "c51e64b3d7d9347bb8fa91a6f4b74631983b7ded564f291b44ba496ff198622c",
|
| 82 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_72.npz": "298c23118c8377d56d7a025cfc59ca782246c1269f0a0b47cdc44f8f53778dc6",
|
| 83 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_73.npz": "0e2aaa5468da7aa0f48a0007eb7fdbed01abe6fdcca5f62ccc8c0d0febccd582",
|
| 84 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_74.npz": "7c33c7cb1a136dd9f438ab36b83007554227ecbb2488fdf883d545d1c680cc94",
|
| 85 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_75.npz": "0c71e68caa217700531d8611731f67fb5842e53df15778bbb480472028d564f6",
|
| 86 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_76.npz": "9dc6ec2ca2dd15912b5505a62ea3b296be8a174bc3c69eeeaa0d6ba316dcaf67",
|
| 87 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_77.npz": "a2c9a43c8566867ede5e47f37b72c2623adc075edb44dd43508037d9a137bd06",
|
| 88 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_3.npz": "2e102e3892277e8808818ffa8373389b9e02e3ce3e1b280c8511dbf0e47353be",
|
| 89 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_4.npz": "c007140d4edbd361da7e15c6bd74bb49e10786ccddc49f0c343592c3939d5519",
|
| 90 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_5.npz": "99e460fca675b7c78122c054e9293b5bfc745d43a4cd57625c4b53dd6fb7dc09",
|
| 91 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_6.npz": "f1d5673cb87462244ddf87c79feee9f38887bcf3443092335b3b38c7ee350e0d",
|
| 92 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_7.npz": "b7e29194ef43c192fedd848312f2fd33687075be7271fbd5e8e882c600a5e4b7",
|
| 93 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_8.npz": "25b9e46d9c422178b825f2d20cf23a1adb09470dea7454e6b70000b927b58f98",
|
| 94 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_9.npz": "3757ebe69d33ab85557c141c3a37e5c800fe2c2ee2c966d2cdef3a6e689a1420",
|
| 95 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_10.npz": "07ab61f067a9fff3d4a1c9bb5039c745417d5bf305d824be85ddd6b114b9829f",
|
| 96 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_11.npz": "269c82abbd5c60440a169962d051b285670e86beef6f60f84a23ddb260e2419f",
|
| 97 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_12.npz": "0e2326f560e5e0b3b7441d8e6584720fefc42d6209bbd0ffc5c4af3679b5f542",
|
| 98 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_13.npz": "4e8c4f5e081cbe193d0468ad98842f4904ddceef140d61363eadce54bb9d8ce9",
|
| 99 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_14.npz": "16ccc9716e778239850f0d8c7f827ab3eefbd4673b775d1bdd17673141550973",
|
| 100 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_15.npz": "199256730bd069357ace4ff4a9528f4a4a2b316e5d299071d2b8751a05a0f026",
|
| 101 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_16.npz": "4631ef874d195778a08af4f0d03a2e6a6a2d1ed6c830b7016c758a0687a1a235",
|
| 102 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_17.npz": "66fbd376505e0f6f921e3f483253f8e733378f8df909733c8fc65194b6eeb092",
|
| 103 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_18.npz": "ee3134b30d8287684cdf95d1eb44da997b90c3ef4c681ac998a770109b35adbf",
|
| 104 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_19.npz": "611b57f18f927ae09479a29d77425fafba33514d0f15ff69dae366733c91e1d3",
|
| 105 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_20.npz": "b3d431eae9b9c3aa5a36c60975378bbec4112377cc858b5ccf0d0946778a6079",
|
| 106 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_21.npz": "cefc699b376b5b43c14769884cc481c22f4a284e168c1fa88feeb9614588f505",
|
| 107 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_22.npz": "7966b47efd7a838d8f87c2f5697320648c6f0bcdd5ffe7762cef03ff2ffbfa95",
|
| 108 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_23.npz": "9b55935db844c788fc59417a1c053722fd5685ad6ec4060a192a0a1fc1582e4f",
|
| 109 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_24.npz": "65c1134ad6a7a38dabea537c0b19899126aa8de849306c1385c14a3c09a22d37",
|
| 110 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_25.npz": "038eab98b035a9ee810e1c7bb08ad89d227bcb87012bf88164d3dea0ad3bc1cd",
|
| 111 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_26.npz": "8e31e623cf17a166d7ee056ba20b5c6f19a693e039c9910a851c0dc3a9cc983b",
|
| 112 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_27.npz": "1306ea155b6746f73a4463cb8be7474bb7f80b3870daead9930e37d0a8fa12cf",
|
| 113 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_28.npz": "3cd4ce023f5f11199b423dda6484bbb7377617b086e2e0bf5dc57afd29e7f566",
|
| 114 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_29.npz": "c7ee73af65346bc85e3c908825198faf4acc4987778f204df0821f4b642c7990",
|
| 115 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_30.npz": "81607884d312e7062274e230fec71dffc2a8a9f073ad8479d3ba19ea3df66899",
|
| 116 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_31.npz": "514aa850fc3099da75b19b74be73e944d7b58cb3577b4a221f58bef4169d9f75",
|
| 117 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_32.npz": "63b2e2a7e689a5efe641536527594ce26be5c807a32acc09a81aa6709150ccdb",
|
| 118 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_33.npz": "acdf16dc4552aef9184c5b687a9b7e0a4ea5edda0fa9ae496eb5f4f6bec12db1",
|
| 119 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_34.npz": "9c70b6174cf4a7a02a6300b14bd730fc51f6d1520dff274b8a3d6588368624e8",
|
| 120 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_35.npz": "25a82c26d0da013eb14be7d9b9f1811eeb9768fa45687e4a71c96c9b9752d9e6",
|
| 121 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_36.npz": "aeae4e169f803c42af7030b1395e1db2459fffaf363d4729075302df46d07f26",
|
| 122 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_37.npz": "c3787ea553f8bc3d239269f7e329e6f39d02b4ec63d06022405286dc640fbf2e",
|
| 123 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_38.npz": "1a486331012e24ff643d0dfec19f2d71bbf29e5de7df3706d7e16da33541ffca",
|
| 124 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_39.npz": "811e7c07267a7c80dbbcd51462fa01bae403341761ee9d89fce1ec67d32eec31",
|
| 125 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_40.npz": "712935236730a5e4142b5cf4da8dca2ba78416a358d79001b2787ee2a0f5b552",
|
| 126 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_41.npz": "e5fcd78bedc942ba734c28e16127c63fb02280ee279a8aee70143e2ec6d16c6e",
|
| 127 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_42.npz": "324b4909e5317f0217290669a78b1ef6f21db35665f13dbac7a78650a166a9b5",
|
| 128 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_43.npz": "b990182a89192ae58b4cfc82b3dd530bdfefbea01dd517a8d361c28c47980370",
|
| 129 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_44.npz": "cdddf30edd971309d21935c342fc353d0fe7b7a45a12c404cf8b7ed5d7a7dece",
|
| 130 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_45.npz": "525ef223869db6a843bca47df8fec0d8cf7f08f21d3626f8af03f4a33f0e896a",
|
| 131 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_46.npz": "8c4001f106b27f2646c61809d701a29245ee9997b0e291a7487f091e2651a7e3",
|
| 132 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_47.npz": "ce82107d4c1575654f41c9bc0bb070aade6d316c5a3bfd73bf84a75e9713fdeb",
|
| 133 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_48.npz": "d0a12442d24126677dd652b204f6910aff45d31ed76d368ce0efaff95943069f",
|
| 134 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_49.npz": "1566580cf72cf4ffc51c7cf5d9238bb09e92236512e92a7ae4a52fd887d5ac2e",
|
| 135 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_50.npz": "55be42b2dee8fecee914584711e405a48bd1be061aee8f8840ef43ed9a8ffe5b",
|
| 136 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_51.npz": "eca0da0b8d7522bc39c587cbd4c3f2ee88e7fc7f27370ff1d1b57c2e96d45565",
|
| 137 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_52.npz": "53491b6e515ddd86a49a3c4b5dcd5d00cd9687a0eabc595248a0cd28a85b2c3f",
|
| 138 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_53.npz": "561e99733a55b073842560ee9847b6b7dcd798d7bc769f3504ee373613708877",
|
| 139 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_54.npz": "3d0d33114524ff3a30e83b4d4f5c80aa44b031fd0e4da7fc776eb63fe89ff322",
|
| 140 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_55.npz": "3166fea4ab9810e08c4c05e5960c3fb338be355a0b9a583d9dd2a2e121c64982",
|
| 141 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_56.npz": "f4123b5dd6c2f722d029139b6f82919ee3c34de8d67326d819c8a6ca65931ff3",
|
| 142 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_57.npz": "e32c77a50be665998bd0682742535bd10b6c2478894f42779d141581b09788c0",
|
| 143 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_58.npz": "69430a65fc86a0166ccc9193746aa61fb4682ae65983944f8304907b34509bbd",
|
| 144 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_59.npz": "c35ebeaf949917077fb132e219558381c9638ef4de3819b95f289918cc4f9c3f",
|
| 145 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_60.npz": "6dc83319d29beb864c6f36d92a11ec898870e6fdf173917f0c9e73f6b18cac5c",
|
| 146 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_61.npz": "69c7b7440c70f7c4383000d1a798b5ef7eb9d3d9a7c6f28c49bd124f484f5fb5",
|
| 147 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_62.npz": "3bf69359b4ad241cb4e6dfdc00205fd637e1f7f959e35ec40eca480f6c41f80b",
|
| 148 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_63.npz": "280ad144e82b822b7ebf60af0d367dfcd261531bbccebd5791152aa75ddcf237",
|
| 149 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_64.npz": "1303b6bb6d7fad77472424cbc3696cb3368cc9e8e3f442dba0f362885b283a49",
|
| 150 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_65.npz": "f396c0ed74ed729685ff8681b2d2029acde74c28db559386b4cf15fff7cb8769",
|
| 151 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_66.npz": "b4d92dc502dc505d873f079044a09b8c2c0c310e95bdb2602a8cd572603c9953",
|
| 152 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_67.npz": "aa4765dd3124747d07ea1e06d353c12bdf2cf3f67777af71556c81d51e047a48",
|
| 153 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_68.npz": "f8df75fb71d9459a84789008ce0b0dc5ee8f0cafce7dad057e6fd6b544da9131",
|
| 154 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_69.npz": "339f59f847d5d98d52dd986eada9092a04036f82c4e570af4db66b7c51eb2000",
|
| 155 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_70.npz": "1156e221583b3435eb14c4116ef2629e89f84a8a0979253d88556d57424afec0",
|
| 156 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_71.npz": "473f3146f0ba8dbd340b44665954354d4d8fac6b5acb9ebcc952cf3f298b345c",
|
| 157 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_72.npz": "58179fe2da6eed5b63eac4710ef8bc54ceff4ff9ba2d8108942643585d21b744",
|
| 158 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_73.npz": "73fcb0af07e5a5accaa1d165446302cf22c595ffacd45015a460524d8e9cd9b1",
|
| 159 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_74.npz": "6e50287f298b566b96d44fe9961ebc7c9f0010c81b035fc4ff7456833ba86b46",
|
| 160 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_75.npz": "1a219f6efabaf773a6268b08f82716dc6177d257c728f70f81b0694d01d65cdf",
|
| 161 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_76.npz": "5aade66c63087e535fcd2c52bd6d8e075f439185a1745a8c8644bd4133a02eaf",
|
| 162 |
+
"/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_77.npz": "88e5ec9012bd07e07ef84d675db34d5fccfc45361ecf0ccbef2aca0a2a7e8bda"
|
| 163 |
+
},
|
| 164 |
+
"limitation": "per-expert additive error proxy; excludes cross-expert covariance and end-to-end model quality",
|
| 165 |
+
"candidate_codebook_scope": "expert"
|
| 166 |
+
},
|
| 167 |
+
"hub_revisions": {
|
| 168 |
+
"zai-org/GLM-5.3": "aca966e4e02791568aa6a4ced368624b3d897f42",
|
| 169 |
+
"RadixArk/GLM-5.3-NVFP4": "11af4cba759e6559eda70358a5778bd1bddddd78",
|
| 170 |
+
"jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m": "2b883d28bb9dd13a9511e2bd45a8ad1cbacbad74"
|
| 171 |
+
},
|
| 172 |
+
"cold_source": "zai-org/GLM-5.3",
|
| 173 |
+
"source_nonexpert_mtp": "nvfp4_donor",
|
| 174 |
+
"source_vision": "base_hybrid",
|
| 175 |
+
"initial_fit_partition": "train only",
|
| 176 |
+
"pv_reassign_every": 40,
|
| 177 |
+
"format": "rvq256_256x8_expert"
|
| 178 |
+
}
|
reproduce/historical_metadata/calibration_corpus.json
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"source": "/tmp/glm53-unc-traces",
|
| 3 |
+
"reasoning_rollouts": true,
|
| 4 |
+
"counts": {
|
| 5 |
+
"train": {
|
| 6 |
+
"problems": 856,
|
| 7 |
+
"tokens": 3048596
|
| 8 |
+
},
|
| 9 |
+
"validation": {
|
| 10 |
+
"problems": 89,
|
| 11 |
+
"tokens": 345875
|
| 12 |
+
},
|
| 13 |
+
"audit": {
|
| 14 |
+
"problems": 104,
|
| 15 |
+
"tokens": 322277
|
| 16 |
+
}
|
| 17 |
+
},
|
| 18 |
+
"partitions": "disjoint prompt hashes; separately captured",
|
| 19 |
+
"limitations": "Text-only captures in 2048-token windows; source datasets overlap earlier calibration. Teacher answers are not correctness-verified.",
|
| 20 |
+
"initial_fit_partition": "train only",
|
| 21 |
+
"capture_teacher": "RadixArk/GLM-5.3-NVFP4",
|
| 22 |
+
"rollout_teacher": "unc NVFP4 teacher",
|
| 23 |
+
"same_token_corpus_as": "/tmp/glm53-kernel-aware-pv/full"
|
| 24 |
+
}
|
reproduce/historical_metadata/cold_assignment.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
reproduce/historical_metadata/pv_progress.json
ADDED
|
@@ -0,0 +1,187 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"layers_pv_complete": [
|
| 3 |
+
3,
|
| 4 |
+
4,
|
| 5 |
+
5,
|
| 6 |
+
6,
|
| 7 |
+
7,
|
| 8 |
+
8,
|
| 9 |
+
9,
|
| 10 |
+
10,
|
| 11 |
+
11,
|
| 12 |
+
12,
|
| 13 |
+
13,
|
| 14 |
+
14,
|
| 15 |
+
15,
|
| 16 |
+
16,
|
| 17 |
+
17,
|
| 18 |
+
18,
|
| 19 |
+
19,
|
| 20 |
+
20,
|
| 21 |
+
21,
|
| 22 |
+
22,
|
| 23 |
+
23,
|
| 24 |
+
24,
|
| 25 |
+
25,
|
| 26 |
+
26,
|
| 27 |
+
27,
|
| 28 |
+
28,
|
| 29 |
+
29,
|
| 30 |
+
30,
|
| 31 |
+
31,
|
| 32 |
+
32,
|
| 33 |
+
33,
|
| 34 |
+
34,
|
| 35 |
+
35,
|
| 36 |
+
36,
|
| 37 |
+
37,
|
| 38 |
+
38,
|
| 39 |
+
39,
|
| 40 |
+
40,
|
| 41 |
+
41,
|
| 42 |
+
42,
|
| 43 |
+
43,
|
| 44 |
+
44,
|
| 45 |
+
45,
|
| 46 |
+
46,
|
| 47 |
+
47,
|
| 48 |
+
48,
|
| 49 |
+
49,
|
| 50 |
+
50,
|
| 51 |
+
51,
|
| 52 |
+
52,
|
| 53 |
+
53,
|
| 54 |
+
54,
|
| 55 |
+
55,
|
| 56 |
+
56,
|
| 57 |
+
57,
|
| 58 |
+
58,
|
| 59 |
+
59,
|
| 60 |
+
60,
|
| 61 |
+
61,
|
| 62 |
+
62,
|
| 63 |
+
63,
|
| 64 |
+
64,
|
| 65 |
+
65,
|
| 66 |
+
66,
|
| 67 |
+
67,
|
| 68 |
+
68,
|
| 69 |
+
69,
|
| 70 |
+
70,
|
| 71 |
+
71,
|
| 72 |
+
72,
|
| 73 |
+
73,
|
| 74 |
+
74,
|
| 75 |
+
75,
|
| 76 |
+
76,
|
| 77 |
+
77
|
| 78 |
+
],
|
| 79 |
+
"total_layers": 75,
|
| 80 |
+
"recipe": "full_corpus_sequential_lr_decay",
|
| 81 |
+
"training_tokens": 18001846,
|
| 82 |
+
"batch_tokens": 262144,
|
| 83 |
+
"microbatch_tokens": 65536,
|
| 84 |
+
"max_updates_per_layer": 69,
|
| 85 |
+
"adaptive_schedule": {
|
| 86 |
+
"from_layer": 27,
|
| 87 |
+
"relative_worsening": 0.001,
|
| 88 |
+
"checks": 3,
|
| 89 |
+
"patience_updates": 15,
|
| 90 |
+
"max_updates": 69,
|
| 91 |
+
"selection": "existing fixed reasoning validation; retain best checkpoint",
|
| 92 |
+
"description": "Drop LR4x after three consecutive validation checks >0.1% worse than best, or after45 at latest. After reduction stop following15 updates without a new best."
|
| 93 |
+
},
|
| 94 |
+
"lr_decay_after": 45,
|
| 95 |
+
"lr_decay_factor": 0.25,
|
| 96 |
+
"layer3_exception": "45 high-LR updates plus retained lower-LR refinement checkpoint",
|
| 97 |
+
"remaining_layers": "retain their previously published weights",
|
| 98 |
+
"previous_campaign_progress": {
|
| 99 |
+
"layers_pv_complete": [
|
| 100 |
+
3,
|
| 101 |
+
4,
|
| 102 |
+
5,
|
| 103 |
+
6,
|
| 104 |
+
7,
|
| 105 |
+
8
|
| 106 |
+
],
|
| 107 |
+
"total_layers": 75,
|
| 108 |
+
"recipe": "sequential_reference_gradient_pv_v1",
|
| 109 |
+
"legacy_pv_layers": [
|
| 110 |
+
9,
|
| 111 |
+
10,
|
| 112 |
+
11,
|
| 113 |
+
12,
|
| 114 |
+
13,
|
| 115 |
+
14
|
| 116 |
+
],
|
| 117 |
+
"initial_fit_layers": [
|
| 118 |
+
15,
|
| 119 |
+
16,
|
| 120 |
+
17,
|
| 121 |
+
18,
|
| 122 |
+
19,
|
| 123 |
+
20,
|
| 124 |
+
21,
|
| 125 |
+
22,
|
| 126 |
+
23,
|
| 127 |
+
24,
|
| 128 |
+
25,
|
| 129 |
+
26,
|
| 130 |
+
27,
|
| 131 |
+
28,
|
| 132 |
+
29,
|
| 133 |
+
30,
|
| 134 |
+
31,
|
| 135 |
+
32,
|
| 136 |
+
33,
|
| 137 |
+
34,
|
| 138 |
+
35,
|
| 139 |
+
36,
|
| 140 |
+
37,
|
| 141 |
+
38,
|
| 142 |
+
39,
|
| 143 |
+
40,
|
| 144 |
+
41,
|
| 145 |
+
42,
|
| 146 |
+
43,
|
| 147 |
+
44,
|
| 148 |
+
45,
|
| 149 |
+
46,
|
| 150 |
+
47,
|
| 151 |
+
48,
|
| 152 |
+
49,
|
| 153 |
+
50,
|
| 154 |
+
51,
|
| 155 |
+
52,
|
| 156 |
+
53,
|
| 157 |
+
54,
|
| 158 |
+
55,
|
| 159 |
+
56,
|
| 160 |
+
57,
|
| 161 |
+
58,
|
| 162 |
+
59,
|
| 163 |
+
60,
|
| 164 |
+
61,
|
| 165 |
+
62,
|
| 166 |
+
63,
|
| 167 |
+
64,
|
| 168 |
+
65,
|
| 169 |
+
66,
|
| 170 |
+
67,
|
| 171 |
+
68,
|
| 172 |
+
69,
|
| 173 |
+
70,
|
| 174 |
+
71,
|
| 175 |
+
72,
|
| 176 |
+
73,
|
| 177 |
+
74,
|
| 178 |
+
75,
|
| 179 |
+
76,
|
| 180 |
+
77
|
| 181 |
+
],
|
| 182 |
+
"format": "rvq256_256x8_expert",
|
| 183 |
+
"full_model_quality": "pending"
|
| 184 |
+
},
|
| 185 |
+
"format": "rvq256_256x8_expert",
|
| 186 |
+
"full_model_quality": "pending"
|
| 187 |
+
}
|
reproduce/historical_metadata/source.json
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"repo": "zai-org/GLM-5.3",
|
| 3 |
+
"revision": "aca966e4e02791568aa6a4ced368624b3d897f42",
|
| 4 |
+
"meta": {
|
| 5 |
+
"kind": "block_fp8",
|
| 6 |
+
"block": [
|
| 7 |
+
128,
|
| 8 |
+
128
|
| 9 |
+
],
|
| 10 |
+
"fmt": "e4m3"
|
| 11 |
+
},
|
| 12 |
+
"cache": "/tmp/glm53-fp8-cold",
|
| 13 |
+
"legacy_cache_revision_unverified": true
|
| 14 |
+
}
|
reproduce/layer_recipe_summary.json
ADDED
|
@@ -0,0 +1,1727 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[
|
| 2 |
+
{
|
| 3 |
+
"layer": 3,
|
| 4 |
+
"target": "same_input",
|
| 5 |
+
"block_scale_dtype": "fp16",
|
| 6 |
+
"training_sequence_count": 14657,
|
| 7 |
+
"training_tokens_seen": 15007754,
|
| 8 |
+
"training_passes": 1.0,
|
| 9 |
+
"best_step": 50,
|
| 10 |
+
"codebook_lr": 0.048,
|
| 11 |
+
"scale_lr": 0.032,
|
| 12 |
+
"lr_decay_after": 25,
|
| 13 |
+
"adaptive_schedule": {
|
| 14 |
+
"best": 0.004300016159961927,
|
| 15 |
+
"drop_after": 20,
|
| 16 |
+
"relative_worsening": 0.001,
|
| 17 |
+
"checks": 3,
|
| 18 |
+
"patience_updates": 10,
|
| 19 |
+
"bad_checks": 3,
|
| 20 |
+
"last_best": 50,
|
| 21 |
+
"early_drop": true
|
| 22 |
+
},
|
| 23 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 24 |
+
},
|
| 25 |
+
{
|
| 26 |
+
"layer": 4,
|
| 27 |
+
"target": "same_input",
|
| 28 |
+
"block_scale_dtype": "fp16",
|
| 29 |
+
"training_sequence_count": 14657,
|
| 30 |
+
"training_tokens_seen": 7864320,
|
| 31 |
+
"training_passes": 0.5240171180844249,
|
| 32 |
+
"best_step": 5,
|
| 33 |
+
"codebook_lr": 0.048,
|
| 34 |
+
"scale_lr": 0.032,
|
| 35 |
+
"lr_decay_after": 25,
|
| 36 |
+
"adaptive_schedule": {
|
| 37 |
+
"best": 0.011536361290655902,
|
| 38 |
+
"drop_after": 20,
|
| 39 |
+
"relative_worsening": 0.001,
|
| 40 |
+
"checks": 3,
|
| 41 |
+
"patience_updates": 10,
|
| 42 |
+
"bad_checks": 3,
|
| 43 |
+
"last_best": 5,
|
| 44 |
+
"early_drop": true
|
| 45 |
+
},
|
| 46 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 47 |
+
},
|
| 48 |
+
{
|
| 49 |
+
"layer": 5,
|
| 50 |
+
"target": "same_input",
|
| 51 |
+
"block_scale_dtype": "fp16",
|
| 52 |
+
"training_sequence_count": 14657,
|
| 53 |
+
"training_tokens_seen": 7864320,
|
| 54 |
+
"training_passes": 0.5240171180844249,
|
| 55 |
+
"best_step": 5,
|
| 56 |
+
"codebook_lr": 0.048,
|
| 57 |
+
"scale_lr": 0.032,
|
| 58 |
+
"lr_decay_after": 25,
|
| 59 |
+
"adaptive_schedule": {
|
| 60 |
+
"best": 0.017510809110214375,
|
| 61 |
+
"drop_after": 20,
|
| 62 |
+
"relative_worsening": 0.001,
|
| 63 |
+
"checks": 3,
|
| 64 |
+
"patience_updates": 10,
|
| 65 |
+
"bad_checks": 3,
|
| 66 |
+
"last_best": 5,
|
| 67 |
+
"early_drop": true
|
| 68 |
+
},
|
| 69 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 70 |
+
},
|
| 71 |
+
{
|
| 72 |
+
"layer": 6,
|
| 73 |
+
"target": "same_input",
|
| 74 |
+
"block_scale_dtype": "fp16",
|
| 75 |
+
"training_sequence_count": 14657,
|
| 76 |
+
"training_tokens_seen": 7864320,
|
| 77 |
+
"training_passes": 0.5240171180844249,
|
| 78 |
+
"best_step": 5,
|
| 79 |
+
"codebook_lr": 0.048,
|
| 80 |
+
"scale_lr": 0.032,
|
| 81 |
+
"lr_decay_after": 25,
|
| 82 |
+
"adaptive_schedule": {
|
| 83 |
+
"best": 0.018967027905769453,
|
| 84 |
+
"drop_after": 20,
|
| 85 |
+
"relative_worsening": 0.001,
|
| 86 |
+
"checks": 3,
|
| 87 |
+
"patience_updates": 10,
|
| 88 |
+
"bad_checks": 3,
|
| 89 |
+
"last_best": 5,
|
| 90 |
+
"early_drop": true
|
| 91 |
+
},
|
| 92 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 93 |
+
},
|
| 94 |
+
{
|
| 95 |
+
"layer": 7,
|
| 96 |
+
"target": "same_input",
|
| 97 |
+
"block_scale_dtype": "fp16",
|
| 98 |
+
"training_sequence_count": 14657,
|
| 99 |
+
"training_tokens_seen": 7864320,
|
| 100 |
+
"training_passes": 0.5240171180844249,
|
| 101 |
+
"best_step": 5,
|
| 102 |
+
"codebook_lr": 0.048,
|
| 103 |
+
"scale_lr": 0.032,
|
| 104 |
+
"lr_decay_after": 25,
|
| 105 |
+
"adaptive_schedule": {
|
| 106 |
+
"best": 0.02562102057711342,
|
| 107 |
+
"drop_after": 20,
|
| 108 |
+
"relative_worsening": 0.001,
|
| 109 |
+
"checks": 3,
|
| 110 |
+
"patience_updates": 10,
|
| 111 |
+
"bad_checks": 3,
|
| 112 |
+
"last_best": 5,
|
| 113 |
+
"early_drop": true
|
| 114 |
+
},
|
| 115 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 116 |
+
},
|
| 117 |
+
{
|
| 118 |
+
"layer": 8,
|
| 119 |
+
"target": "same_input",
|
| 120 |
+
"block_scale_dtype": "fp16",
|
| 121 |
+
"training_sequence_count": 14657,
|
| 122 |
+
"training_tokens_seen": 13107200,
|
| 123 |
+
"training_passes": 0.8733618634740414,
|
| 124 |
+
"best_step": 40,
|
| 125 |
+
"codebook_lr": 0.048,
|
| 126 |
+
"scale_lr": 0.032,
|
| 127 |
+
"lr_decay_after": 25,
|
| 128 |
+
"adaptive_schedule": {
|
| 129 |
+
"best": 0.022436852159198863,
|
| 130 |
+
"drop_after": 25,
|
| 131 |
+
"relative_worsening": 0.001,
|
| 132 |
+
"checks": 3,
|
| 133 |
+
"patience_updates": 10,
|
| 134 |
+
"bad_checks": 0,
|
| 135 |
+
"last_best": 40,
|
| 136 |
+
"early_drop": false
|
| 137 |
+
},
|
| 138 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 139 |
+
},
|
| 140 |
+
{
|
| 141 |
+
"layer": 9,
|
| 142 |
+
"target": "same_input",
|
| 143 |
+
"block_scale_dtype": "fp16",
|
| 144 |
+
"training_sequence_count": 14657,
|
| 145 |
+
"training_tokens_seen": 15007754,
|
| 146 |
+
"training_passes": 1.0,
|
| 147 |
+
"best_step": 58,
|
| 148 |
+
"codebook_lr": 0.048,
|
| 149 |
+
"scale_lr": 0.032,
|
| 150 |
+
"lr_decay_after": 25,
|
| 151 |
+
"adaptive_schedule": {
|
| 152 |
+
"best": 0.0017143418153641949,
|
| 153 |
+
"drop_after": 25,
|
| 154 |
+
"relative_worsening": 0.001,
|
| 155 |
+
"checks": 3,
|
| 156 |
+
"patience_updates": 10,
|
| 157 |
+
"bad_checks": 0,
|
| 158 |
+
"last_best": 58,
|
| 159 |
+
"early_drop": false
|
| 160 |
+
},
|
| 161 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 162 |
+
},
|
| 163 |
+
{
|
| 164 |
+
"layer": 10,
|
| 165 |
+
"target": "same_input",
|
| 166 |
+
"block_scale_dtype": "fp16",
|
| 167 |
+
"training_sequence_count": 14657,
|
| 168 |
+
"training_tokens_seen": 7864320,
|
| 169 |
+
"training_passes": 0.5240171180844249,
|
| 170 |
+
"best_step": 5,
|
| 171 |
+
"codebook_lr": 0.048,
|
| 172 |
+
"scale_lr": 0.032,
|
| 173 |
+
"lr_decay_after": 25,
|
| 174 |
+
"adaptive_schedule": {
|
| 175 |
+
"best": 0.0013590989109528494,
|
| 176 |
+
"drop_after": 20,
|
| 177 |
+
"relative_worsening": 0.001,
|
| 178 |
+
"checks": 3,
|
| 179 |
+
"patience_updates": 10,
|
| 180 |
+
"bad_checks": 3,
|
| 181 |
+
"last_best": 5,
|
| 182 |
+
"early_drop": true
|
| 183 |
+
},
|
| 184 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 185 |
+
},
|
| 186 |
+
{
|
| 187 |
+
"layer": 11,
|
| 188 |
+
"target": "same_input",
|
| 189 |
+
"block_scale_dtype": "fp16",
|
| 190 |
+
"training_sequence_count": 14657,
|
| 191 |
+
"training_tokens_seen": 7864320,
|
| 192 |
+
"training_passes": 0.5240171180844249,
|
| 193 |
+
"best_step": 5,
|
| 194 |
+
"codebook_lr": 0.048,
|
| 195 |
+
"scale_lr": 0.032,
|
| 196 |
+
"lr_decay_after": 25,
|
| 197 |
+
"adaptive_schedule": {
|
| 198 |
+
"best": 0.0013980529941903324,
|
| 199 |
+
"drop_after": 20,
|
| 200 |
+
"relative_worsening": 0.001,
|
| 201 |
+
"checks": 3,
|
| 202 |
+
"patience_updates": 10,
|
| 203 |
+
"bad_checks": 3,
|
| 204 |
+
"last_best": 5,
|
| 205 |
+
"early_drop": true
|
| 206 |
+
},
|
| 207 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 208 |
+
},
|
| 209 |
+
{
|
| 210 |
+
"layer": 12,
|
| 211 |
+
"target": "same_input",
|
| 212 |
+
"block_scale_dtype": "fp16",
|
| 213 |
+
"training_sequence_count": 14657,
|
| 214 |
+
"training_tokens_seen": 7864320,
|
| 215 |
+
"training_passes": 0.5240171180844249,
|
| 216 |
+
"best_step": 5,
|
| 217 |
+
"codebook_lr": 0.048,
|
| 218 |
+
"scale_lr": 0.032,
|
| 219 |
+
"lr_decay_after": 25,
|
| 220 |
+
"adaptive_schedule": {
|
| 221 |
+
"best": 0.001638401368651591,
|
| 222 |
+
"drop_after": 20,
|
| 223 |
+
"relative_worsening": 0.001,
|
| 224 |
+
"checks": 3,
|
| 225 |
+
"patience_updates": 10,
|
| 226 |
+
"bad_checks": 3,
|
| 227 |
+
"last_best": 5,
|
| 228 |
+
"early_drop": true
|
| 229 |
+
},
|
| 230 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 231 |
+
},
|
| 232 |
+
{
|
| 233 |
+
"layer": 13,
|
| 234 |
+
"target": "same_input",
|
| 235 |
+
"block_scale_dtype": "fp16",
|
| 236 |
+
"training_sequence_count": 14657,
|
| 237 |
+
"training_tokens_seen": 7864320,
|
| 238 |
+
"training_passes": 0.5240171180844249,
|
| 239 |
+
"best_step": 5,
|
| 240 |
+
"codebook_lr": 0.048,
|
| 241 |
+
"scale_lr": 0.032,
|
| 242 |
+
"lr_decay_after": 25,
|
| 243 |
+
"adaptive_schedule": {
|
| 244 |
+
"best": 0.00271562211069347,
|
| 245 |
+
"drop_after": 20,
|
| 246 |
+
"relative_worsening": 0.001,
|
| 247 |
+
"checks": 3,
|
| 248 |
+
"patience_updates": 10,
|
| 249 |
+
"bad_checks": 3,
|
| 250 |
+
"last_best": 5,
|
| 251 |
+
"early_drop": true
|
| 252 |
+
},
|
| 253 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 254 |
+
},
|
| 255 |
+
{
|
| 256 |
+
"layer": 14,
|
| 257 |
+
"target": "same_input",
|
| 258 |
+
"block_scale_dtype": "fp16",
|
| 259 |
+
"training_sequence_count": 14657,
|
| 260 |
+
"training_tokens_seen": 7864320,
|
| 261 |
+
"training_passes": 0.5240171180844249,
|
| 262 |
+
"best_step": 5,
|
| 263 |
+
"codebook_lr": 0.048,
|
| 264 |
+
"scale_lr": 0.032,
|
| 265 |
+
"lr_decay_after": 25,
|
| 266 |
+
"adaptive_schedule": {
|
| 267 |
+
"best": 0.0030091782715398136,
|
| 268 |
+
"drop_after": 20,
|
| 269 |
+
"relative_worsening": 0.001,
|
| 270 |
+
"checks": 3,
|
| 271 |
+
"patience_updates": 10,
|
| 272 |
+
"bad_checks": 3,
|
| 273 |
+
"last_best": 5,
|
| 274 |
+
"early_drop": true
|
| 275 |
+
},
|
| 276 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 277 |
+
},
|
| 278 |
+
{
|
| 279 |
+
"layer": 15,
|
| 280 |
+
"target": "same_input",
|
| 281 |
+
"block_scale_dtype": "fp16",
|
| 282 |
+
"training_sequence_count": 14657,
|
| 283 |
+
"training_tokens_seen": 7864320,
|
| 284 |
+
"training_passes": 0.5240171180844249,
|
| 285 |
+
"best_step": 5,
|
| 286 |
+
"codebook_lr": 0.048,
|
| 287 |
+
"scale_lr": 0.032,
|
| 288 |
+
"lr_decay_after": 25,
|
| 289 |
+
"adaptive_schedule": {
|
| 290 |
+
"best": 0.003168485440774852,
|
| 291 |
+
"drop_after": 20,
|
| 292 |
+
"relative_worsening": 0.001,
|
| 293 |
+
"checks": 3,
|
| 294 |
+
"patience_updates": 10,
|
| 295 |
+
"bad_checks": 3,
|
| 296 |
+
"last_best": 5,
|
| 297 |
+
"early_drop": true
|
| 298 |
+
},
|
| 299 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 300 |
+
},
|
| 301 |
+
{
|
| 302 |
+
"layer": 16,
|
| 303 |
+
"target": "same_input",
|
| 304 |
+
"block_scale_dtype": "fp16",
|
| 305 |
+
"training_sequence_count": 14657,
|
| 306 |
+
"training_tokens_seen": 7864320,
|
| 307 |
+
"training_passes": 0.5240171180844249,
|
| 308 |
+
"best_step": 5,
|
| 309 |
+
"codebook_lr": 0.048,
|
| 310 |
+
"scale_lr": 0.032,
|
| 311 |
+
"lr_decay_after": 25,
|
| 312 |
+
"adaptive_schedule": {
|
| 313 |
+
"best": 0.004101365709508885,
|
| 314 |
+
"drop_after": 20,
|
| 315 |
+
"relative_worsening": 0.001,
|
| 316 |
+
"checks": 3,
|
| 317 |
+
"patience_updates": 10,
|
| 318 |
+
"bad_checks": 3,
|
| 319 |
+
"last_best": 5,
|
| 320 |
+
"early_drop": true
|
| 321 |
+
},
|
| 322 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 323 |
+
},
|
| 324 |
+
{
|
| 325 |
+
"layer": 17,
|
| 326 |
+
"target": "same_input",
|
| 327 |
+
"block_scale_dtype": "fp16",
|
| 328 |
+
"training_sequence_count": 14657,
|
| 329 |
+
"training_tokens_seen": 7864320,
|
| 330 |
+
"training_passes": 0.5240171180844249,
|
| 331 |
+
"best_step": 5,
|
| 332 |
+
"codebook_lr": 0.048,
|
| 333 |
+
"scale_lr": 0.032,
|
| 334 |
+
"lr_decay_after": 25,
|
| 335 |
+
"adaptive_schedule": {
|
| 336 |
+
"best": 0.003969997484537619,
|
| 337 |
+
"drop_after": 20,
|
| 338 |
+
"relative_worsening": 0.001,
|
| 339 |
+
"checks": 3,
|
| 340 |
+
"patience_updates": 10,
|
| 341 |
+
"bad_checks": 3,
|
| 342 |
+
"last_best": 5,
|
| 343 |
+
"early_drop": true
|
| 344 |
+
},
|
| 345 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 346 |
+
},
|
| 347 |
+
{
|
| 348 |
+
"layer": 18,
|
| 349 |
+
"target": "same_input",
|
| 350 |
+
"block_scale_dtype": "fp16",
|
| 351 |
+
"training_sequence_count": 14657,
|
| 352 |
+
"training_tokens_seen": 7864320,
|
| 353 |
+
"training_passes": 0.5240171180844249,
|
| 354 |
+
"best_step": 5,
|
| 355 |
+
"codebook_lr": 0.048,
|
| 356 |
+
"scale_lr": 0.032,
|
| 357 |
+
"lr_decay_after": 25,
|
| 358 |
+
"adaptive_schedule": {
|
| 359 |
+
"best": 0.00536306292142981,
|
| 360 |
+
"drop_after": 20,
|
| 361 |
+
"relative_worsening": 0.001,
|
| 362 |
+
"checks": 3,
|
| 363 |
+
"patience_updates": 10,
|
| 364 |
+
"bad_checks": 3,
|
| 365 |
+
"last_best": 5,
|
| 366 |
+
"early_drop": true
|
| 367 |
+
},
|
| 368 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 369 |
+
},
|
| 370 |
+
{
|
| 371 |
+
"layer": 19,
|
| 372 |
+
"target": "same_input",
|
| 373 |
+
"block_scale_dtype": "fp16",
|
| 374 |
+
"training_sequence_count": 14657,
|
| 375 |
+
"training_tokens_seen": 7864320,
|
| 376 |
+
"training_passes": 0.5240171180844249,
|
| 377 |
+
"best_step": 5,
|
| 378 |
+
"codebook_lr": 0.048,
|
| 379 |
+
"scale_lr": 0.032,
|
| 380 |
+
"lr_decay_after": 25,
|
| 381 |
+
"adaptive_schedule": {
|
| 382 |
+
"best": 0.005900326847720345,
|
| 383 |
+
"drop_after": 20,
|
| 384 |
+
"relative_worsening": 0.001,
|
| 385 |
+
"checks": 3,
|
| 386 |
+
"patience_updates": 10,
|
| 387 |
+
"bad_checks": 3,
|
| 388 |
+
"last_best": 5,
|
| 389 |
+
"early_drop": true
|
| 390 |
+
},
|
| 391 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 392 |
+
},
|
| 393 |
+
{
|
| 394 |
+
"layer": 20,
|
| 395 |
+
"target": "same_input",
|
| 396 |
+
"block_scale_dtype": "fp16",
|
| 397 |
+
"training_sequence_count": 14657,
|
| 398 |
+
"training_tokens_seen": 7864320,
|
| 399 |
+
"training_passes": 0.5240171180844249,
|
| 400 |
+
"best_step": 5,
|
| 401 |
+
"codebook_lr": 0.048,
|
| 402 |
+
"scale_lr": 0.032,
|
| 403 |
+
"lr_decay_after": 25,
|
| 404 |
+
"adaptive_schedule": {
|
| 405 |
+
"best": 0.00827687275281147,
|
| 406 |
+
"drop_after": 20,
|
| 407 |
+
"relative_worsening": 0.001,
|
| 408 |
+
"checks": 3,
|
| 409 |
+
"patience_updates": 10,
|
| 410 |
+
"bad_checks": 3,
|
| 411 |
+
"last_best": 5,
|
| 412 |
+
"early_drop": true
|
| 413 |
+
},
|
| 414 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 415 |
+
},
|
| 416 |
+
{
|
| 417 |
+
"layer": 21,
|
| 418 |
+
"target": "same_input",
|
| 419 |
+
"block_scale_dtype": "fp16",
|
| 420 |
+
"training_sequence_count": 14657,
|
| 421 |
+
"training_tokens_seen": 7864320,
|
| 422 |
+
"training_passes": 0.5240171180844249,
|
| 423 |
+
"best_step": 5,
|
| 424 |
+
"codebook_lr": 0.048,
|
| 425 |
+
"scale_lr": 0.032,
|
| 426 |
+
"lr_decay_after": 25,
|
| 427 |
+
"adaptive_schedule": {
|
| 428 |
+
"best": 0.007670614413287488,
|
| 429 |
+
"drop_after": 20,
|
| 430 |
+
"relative_worsening": 0.001,
|
| 431 |
+
"checks": 3,
|
| 432 |
+
"patience_updates": 10,
|
| 433 |
+
"bad_checks": 3,
|
| 434 |
+
"last_best": 5,
|
| 435 |
+
"early_drop": true
|
| 436 |
+
},
|
| 437 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 438 |
+
},
|
| 439 |
+
{
|
| 440 |
+
"layer": 22,
|
| 441 |
+
"target": "same_input",
|
| 442 |
+
"block_scale_dtype": "fp16",
|
| 443 |
+
"training_sequence_count": 14657,
|
| 444 |
+
"training_tokens_seen": 7864320,
|
| 445 |
+
"training_passes": 0.5240171180844249,
|
| 446 |
+
"best_step": 5,
|
| 447 |
+
"codebook_lr": 0.048,
|
| 448 |
+
"scale_lr": 0.032,
|
| 449 |
+
"lr_decay_after": 25,
|
| 450 |
+
"adaptive_schedule": {
|
| 451 |
+
"best": 0.009540958133708182,
|
| 452 |
+
"drop_after": 20,
|
| 453 |
+
"relative_worsening": 0.001,
|
| 454 |
+
"checks": 3,
|
| 455 |
+
"patience_updates": 10,
|
| 456 |
+
"bad_checks": 3,
|
| 457 |
+
"last_best": 5,
|
| 458 |
+
"early_drop": true
|
| 459 |
+
},
|
| 460 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 461 |
+
},
|
| 462 |
+
{
|
| 463 |
+
"layer": 23,
|
| 464 |
+
"target": "same_input",
|
| 465 |
+
"block_scale_dtype": "fp16",
|
| 466 |
+
"training_sequence_count": 14657,
|
| 467 |
+
"training_tokens_seen": 7864320,
|
| 468 |
+
"training_passes": 0.5240171180844249,
|
| 469 |
+
"best_step": 5,
|
| 470 |
+
"codebook_lr": 0.048,
|
| 471 |
+
"scale_lr": 0.032,
|
| 472 |
+
"lr_decay_after": 25,
|
| 473 |
+
"adaptive_schedule": {
|
| 474 |
+
"best": 0.011060020056539258,
|
| 475 |
+
"drop_after": 20,
|
| 476 |
+
"relative_worsening": 0.001,
|
| 477 |
+
"checks": 3,
|
| 478 |
+
"patience_updates": 10,
|
| 479 |
+
"bad_checks": 3,
|
| 480 |
+
"last_best": 5,
|
| 481 |
+
"early_drop": true
|
| 482 |
+
},
|
| 483 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 484 |
+
},
|
| 485 |
+
{
|
| 486 |
+
"layer": 24,
|
| 487 |
+
"target": "same_input",
|
| 488 |
+
"block_scale_dtype": "fp16",
|
| 489 |
+
"training_sequence_count": 14657,
|
| 490 |
+
"training_tokens_seen": 7864320,
|
| 491 |
+
"training_passes": 0.5240171180844249,
|
| 492 |
+
"best_step": 5,
|
| 493 |
+
"codebook_lr": 0.048,
|
| 494 |
+
"scale_lr": 0.032,
|
| 495 |
+
"lr_decay_after": 25,
|
| 496 |
+
"adaptive_schedule": {
|
| 497 |
+
"best": 0.009994059305569054,
|
| 498 |
+
"drop_after": 20,
|
| 499 |
+
"relative_worsening": 0.001,
|
| 500 |
+
"checks": 3,
|
| 501 |
+
"patience_updates": 10,
|
| 502 |
+
"bad_checks": 3,
|
| 503 |
+
"last_best": 5,
|
| 504 |
+
"early_drop": true
|
| 505 |
+
},
|
| 506 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 507 |
+
},
|
| 508 |
+
{
|
| 509 |
+
"layer": 25,
|
| 510 |
+
"target": "same_input",
|
| 511 |
+
"block_scale_dtype": "fp16",
|
| 512 |
+
"training_sequence_count": 14657,
|
| 513 |
+
"training_tokens_seen": 7864320,
|
| 514 |
+
"training_passes": 0.5240171180844249,
|
| 515 |
+
"best_step": 5,
|
| 516 |
+
"codebook_lr": 0.048,
|
| 517 |
+
"scale_lr": 0.032,
|
| 518 |
+
"lr_decay_after": 25,
|
| 519 |
+
"adaptive_schedule": {
|
| 520 |
+
"best": 0.010544123217574992,
|
| 521 |
+
"drop_after": 20,
|
| 522 |
+
"relative_worsening": 0.001,
|
| 523 |
+
"checks": 3,
|
| 524 |
+
"patience_updates": 10,
|
| 525 |
+
"bad_checks": 3,
|
| 526 |
+
"last_best": 5,
|
| 527 |
+
"early_drop": true
|
| 528 |
+
},
|
| 529 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 530 |
+
},
|
| 531 |
+
{
|
| 532 |
+
"layer": 26,
|
| 533 |
+
"target": "same_input",
|
| 534 |
+
"block_scale_dtype": "fp16",
|
| 535 |
+
"training_sequence_count": 14657,
|
| 536 |
+
"training_tokens_seen": 7864320,
|
| 537 |
+
"training_passes": 0.5240171180844249,
|
| 538 |
+
"best_step": 5,
|
| 539 |
+
"codebook_lr": 0.048,
|
| 540 |
+
"scale_lr": 0.032,
|
| 541 |
+
"lr_decay_after": 25,
|
| 542 |
+
"adaptive_schedule": {
|
| 543 |
+
"best": 0.012252411564001522,
|
| 544 |
+
"drop_after": 20,
|
| 545 |
+
"relative_worsening": 0.001,
|
| 546 |
+
"checks": 3,
|
| 547 |
+
"patience_updates": 10,
|
| 548 |
+
"bad_checks": 3,
|
| 549 |
+
"last_best": 5,
|
| 550 |
+
"early_drop": true
|
| 551 |
+
},
|
| 552 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 553 |
+
},
|
| 554 |
+
{
|
| 555 |
+
"layer": 27,
|
| 556 |
+
"target": "same_input",
|
| 557 |
+
"block_scale_dtype": "fp16",
|
| 558 |
+
"training_sequence_count": 14657,
|
| 559 |
+
"training_tokens_seen": 7864320,
|
| 560 |
+
"training_passes": 0.5240171180844249,
|
| 561 |
+
"best_step": 5,
|
| 562 |
+
"codebook_lr": 0.048,
|
| 563 |
+
"scale_lr": 0.032,
|
| 564 |
+
"lr_decay_after": 25,
|
| 565 |
+
"adaptive_schedule": {
|
| 566 |
+
"best": 0.017041586186703876,
|
| 567 |
+
"drop_after": 20,
|
| 568 |
+
"relative_worsening": 0.001,
|
| 569 |
+
"checks": 3,
|
| 570 |
+
"patience_updates": 10,
|
| 571 |
+
"bad_checks": 3,
|
| 572 |
+
"last_best": 5,
|
| 573 |
+
"early_drop": true
|
| 574 |
+
},
|
| 575 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 576 |
+
},
|
| 577 |
+
{
|
| 578 |
+
"layer": 28,
|
| 579 |
+
"target": "same_input",
|
| 580 |
+
"block_scale_dtype": "fp16",
|
| 581 |
+
"training_sequence_count": 14657,
|
| 582 |
+
"training_tokens_seen": 9175040,
|
| 583 |
+
"training_passes": 0.611353304431829,
|
| 584 |
+
"best_step": 25,
|
| 585 |
+
"codebook_lr": 0.048,
|
| 586 |
+
"scale_lr": 0.032,
|
| 587 |
+
"lr_decay_after": 25,
|
| 588 |
+
"adaptive_schedule": {
|
| 589 |
+
"best": 0.018901359302169598,
|
| 590 |
+
"drop_after": 20,
|
| 591 |
+
"relative_worsening": 0.001,
|
| 592 |
+
"checks": 3,
|
| 593 |
+
"patience_updates": 10,
|
| 594 |
+
"bad_checks": 3,
|
| 595 |
+
"last_best": 25,
|
| 596 |
+
"early_drop": true
|
| 597 |
+
},
|
| 598 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 599 |
+
},
|
| 600 |
+
{
|
| 601 |
+
"layer": 29,
|
| 602 |
+
"target": "same_input",
|
| 603 |
+
"block_scale_dtype": "fp16",
|
| 604 |
+
"training_sequence_count": 14657,
|
| 605 |
+
"training_tokens_seen": 7864320,
|
| 606 |
+
"training_passes": 0.5240171180844249,
|
| 607 |
+
"best_step": 5,
|
| 608 |
+
"codebook_lr": 0.048,
|
| 609 |
+
"scale_lr": 0.032,
|
| 610 |
+
"lr_decay_after": 25,
|
| 611 |
+
"adaptive_schedule": {
|
| 612 |
+
"best": 0.01730259021141979,
|
| 613 |
+
"drop_after": 20,
|
| 614 |
+
"relative_worsening": 0.001,
|
| 615 |
+
"checks": 3,
|
| 616 |
+
"patience_updates": 10,
|
| 617 |
+
"bad_checks": 3,
|
| 618 |
+
"last_best": 5,
|
| 619 |
+
"early_drop": true
|
| 620 |
+
},
|
| 621 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 622 |
+
},
|
| 623 |
+
{
|
| 624 |
+
"layer": 30,
|
| 625 |
+
"target": "same_input",
|
| 626 |
+
"block_scale_dtype": "fp16",
|
| 627 |
+
"training_sequence_count": 14657,
|
| 628 |
+
"training_tokens_seen": 9175040,
|
| 629 |
+
"training_passes": 0.611353304431829,
|
| 630 |
+
"best_step": 25,
|
| 631 |
+
"codebook_lr": 0.048,
|
| 632 |
+
"scale_lr": 0.032,
|
| 633 |
+
"lr_decay_after": 25,
|
| 634 |
+
"adaptive_schedule": {
|
| 635 |
+
"best": 0.01953328825434068,
|
| 636 |
+
"drop_after": 20,
|
| 637 |
+
"relative_worsening": 0.001,
|
| 638 |
+
"checks": 3,
|
| 639 |
+
"patience_updates": 10,
|
| 640 |
+
"bad_checks": 3,
|
| 641 |
+
"last_best": 25,
|
| 642 |
+
"early_drop": true
|
| 643 |
+
},
|
| 644 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 645 |
+
},
|
| 646 |
+
{
|
| 647 |
+
"layer": 31,
|
| 648 |
+
"target": "same_input",
|
| 649 |
+
"block_scale_dtype": "fp16",
|
| 650 |
+
"training_sequence_count": 14657,
|
| 651 |
+
"training_tokens_seen": 7864320,
|
| 652 |
+
"training_passes": 0.5240171180844249,
|
| 653 |
+
"best_step": 5,
|
| 654 |
+
"codebook_lr": 0.048,
|
| 655 |
+
"scale_lr": 0.032,
|
| 656 |
+
"lr_decay_after": 25,
|
| 657 |
+
"adaptive_schedule": {
|
| 658 |
+
"best": 0.0185387897556427,
|
| 659 |
+
"drop_after": 20,
|
| 660 |
+
"relative_worsening": 0.001,
|
| 661 |
+
"checks": 3,
|
| 662 |
+
"patience_updates": 10,
|
| 663 |
+
"bad_checks": 3,
|
| 664 |
+
"last_best": 5,
|
| 665 |
+
"early_drop": true
|
| 666 |
+
},
|
| 667 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 668 |
+
},
|
| 669 |
+
{
|
| 670 |
+
"layer": 32,
|
| 671 |
+
"target": "same_input",
|
| 672 |
+
"block_scale_dtype": "fp16",
|
| 673 |
+
"training_sequence_count": 14657,
|
| 674 |
+
"training_tokens_seen": 7864320,
|
| 675 |
+
"training_passes": 0.5240171180844249,
|
| 676 |
+
"best_step": 5,
|
| 677 |
+
"codebook_lr": 0.048,
|
| 678 |
+
"scale_lr": 0.032,
|
| 679 |
+
"lr_decay_after": 25,
|
| 680 |
+
"adaptive_schedule": {
|
| 681 |
+
"best": 0.02538125707815895,
|
| 682 |
+
"drop_after": 20,
|
| 683 |
+
"relative_worsening": 0.001,
|
| 684 |
+
"checks": 3,
|
| 685 |
+
"patience_updates": 10,
|
| 686 |
+
"bad_checks": 3,
|
| 687 |
+
"last_best": 5,
|
| 688 |
+
"early_drop": true
|
| 689 |
+
},
|
| 690 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 691 |
+
},
|
| 692 |
+
{
|
| 693 |
+
"layer": 33,
|
| 694 |
+
"target": "same_input",
|
| 695 |
+
"block_scale_dtype": "fp16",
|
| 696 |
+
"training_sequence_count": 14657,
|
| 697 |
+
"training_tokens_seen": 7864320,
|
| 698 |
+
"training_passes": 0.5240171180844249,
|
| 699 |
+
"best_step": 5,
|
| 700 |
+
"codebook_lr": 0.048,
|
| 701 |
+
"scale_lr": 0.032,
|
| 702 |
+
"lr_decay_after": 25,
|
| 703 |
+
"adaptive_schedule": {
|
| 704 |
+
"best": 0.02889075275987553,
|
| 705 |
+
"drop_after": 20,
|
| 706 |
+
"relative_worsening": 0.001,
|
| 707 |
+
"checks": 3,
|
| 708 |
+
"patience_updates": 10,
|
| 709 |
+
"bad_checks": 3,
|
| 710 |
+
"last_best": 5,
|
| 711 |
+
"early_drop": true
|
| 712 |
+
},
|
| 713 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 714 |
+
},
|
| 715 |
+
{
|
| 716 |
+
"layer": 34,
|
| 717 |
+
"target": "same_input",
|
| 718 |
+
"block_scale_dtype": "fp16",
|
| 719 |
+
"training_sequence_count": 14657,
|
| 720 |
+
"training_tokens_seen": 7864320,
|
| 721 |
+
"training_passes": 0.5240171180844249,
|
| 722 |
+
"best_step": 5,
|
| 723 |
+
"codebook_lr": 0.048,
|
| 724 |
+
"scale_lr": 0.032,
|
| 725 |
+
"lr_decay_after": 25,
|
| 726 |
+
"adaptive_schedule": {
|
| 727 |
+
"best": 0.03261215982052183,
|
| 728 |
+
"drop_after": 20,
|
| 729 |
+
"relative_worsening": 0.001,
|
| 730 |
+
"checks": 3,
|
| 731 |
+
"patience_updates": 10,
|
| 732 |
+
"bad_checks": 3,
|
| 733 |
+
"last_best": 5,
|
| 734 |
+
"early_drop": true
|
| 735 |
+
},
|
| 736 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 737 |
+
},
|
| 738 |
+
{
|
| 739 |
+
"layer": 35,
|
| 740 |
+
"target": "same_input",
|
| 741 |
+
"block_scale_dtype": "fp16",
|
| 742 |
+
"training_sequence_count": 14657,
|
| 743 |
+
"training_tokens_seen": 7864320,
|
| 744 |
+
"training_passes": 0.5240171180844249,
|
| 745 |
+
"best_step": 5,
|
| 746 |
+
"codebook_lr": 0.048,
|
| 747 |
+
"scale_lr": 0.032,
|
| 748 |
+
"lr_decay_after": 25,
|
| 749 |
+
"adaptive_schedule": {
|
| 750 |
+
"best": 0.02849563835939984,
|
| 751 |
+
"drop_after": 20,
|
| 752 |
+
"relative_worsening": 0.001,
|
| 753 |
+
"checks": 3,
|
| 754 |
+
"patience_updates": 10,
|
| 755 |
+
"bad_checks": 3,
|
| 756 |
+
"last_best": 5,
|
| 757 |
+
"early_drop": true
|
| 758 |
+
},
|
| 759 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 760 |
+
},
|
| 761 |
+
{
|
| 762 |
+
"layer": 36,
|
| 763 |
+
"target": "same_input",
|
| 764 |
+
"block_scale_dtype": "fp16",
|
| 765 |
+
"training_sequence_count": 14657,
|
| 766 |
+
"training_tokens_seen": 7864320,
|
| 767 |
+
"training_passes": 0.5240171180844249,
|
| 768 |
+
"best_step": 5,
|
| 769 |
+
"codebook_lr": 0.048,
|
| 770 |
+
"scale_lr": 0.032,
|
| 771 |
+
"lr_decay_after": 25,
|
| 772 |
+
"adaptive_schedule": {
|
| 773 |
+
"best": 0.033030121256077974,
|
| 774 |
+
"drop_after": 20,
|
| 775 |
+
"relative_worsening": 0.001,
|
| 776 |
+
"checks": 3,
|
| 777 |
+
"patience_updates": 10,
|
| 778 |
+
"bad_checks": 3,
|
| 779 |
+
"last_best": 5,
|
| 780 |
+
"early_drop": true
|
| 781 |
+
},
|
| 782 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 783 |
+
},
|
| 784 |
+
{
|
| 785 |
+
"layer": 37,
|
| 786 |
+
"target": "same_input",
|
| 787 |
+
"block_scale_dtype": "fp16",
|
| 788 |
+
"training_sequence_count": 14657,
|
| 789 |
+
"training_tokens_seen": 7864320,
|
| 790 |
+
"training_passes": 0.5240171180844249,
|
| 791 |
+
"best_step": 5,
|
| 792 |
+
"codebook_lr": 0.048,
|
| 793 |
+
"scale_lr": 0.032,
|
| 794 |
+
"lr_decay_after": 25,
|
| 795 |
+
"adaptive_schedule": {
|
| 796 |
+
"best": 0.035430789227037796,
|
| 797 |
+
"drop_after": 20,
|
| 798 |
+
"relative_worsening": 0.001,
|
| 799 |
+
"checks": 3,
|
| 800 |
+
"patience_updates": 10,
|
| 801 |
+
"bad_checks": 3,
|
| 802 |
+
"last_best": 5,
|
| 803 |
+
"early_drop": true
|
| 804 |
+
},
|
| 805 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 806 |
+
},
|
| 807 |
+
{
|
| 808 |
+
"layer": 38,
|
| 809 |
+
"target": "same_input",
|
| 810 |
+
"block_scale_dtype": "fp16",
|
| 811 |
+
"training_sequence_count": 14657,
|
| 812 |
+
"training_tokens_seen": 7864320,
|
| 813 |
+
"training_passes": 0.5240171180844249,
|
| 814 |
+
"best_step": 5,
|
| 815 |
+
"codebook_lr": 0.048,
|
| 816 |
+
"scale_lr": 0.032,
|
| 817 |
+
"lr_decay_after": 25,
|
| 818 |
+
"adaptive_schedule": {
|
| 819 |
+
"best": 0.03938840653001622,
|
| 820 |
+
"drop_after": 20,
|
| 821 |
+
"relative_worsening": 0.001,
|
| 822 |
+
"checks": 3,
|
| 823 |
+
"patience_updates": 10,
|
| 824 |
+
"bad_checks": 3,
|
| 825 |
+
"last_best": 5,
|
| 826 |
+
"early_drop": true
|
| 827 |
+
},
|
| 828 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 829 |
+
},
|
| 830 |
+
{
|
| 831 |
+
"layer": 39,
|
| 832 |
+
"target": "same_input",
|
| 833 |
+
"block_scale_dtype": "fp16",
|
| 834 |
+
"training_sequence_count": 14657,
|
| 835 |
+
"training_tokens_seen": 7864320,
|
| 836 |
+
"training_passes": 0.5240171180844249,
|
| 837 |
+
"best_step": 5,
|
| 838 |
+
"codebook_lr": 0.048,
|
| 839 |
+
"scale_lr": 0.032,
|
| 840 |
+
"lr_decay_after": 25,
|
| 841 |
+
"adaptive_schedule": {
|
| 842 |
+
"best": 0.04198908163931699,
|
| 843 |
+
"drop_after": 20,
|
| 844 |
+
"relative_worsening": 0.001,
|
| 845 |
+
"checks": 3,
|
| 846 |
+
"patience_updates": 10,
|
| 847 |
+
"bad_checks": 3,
|
| 848 |
+
"last_best": 5,
|
| 849 |
+
"early_drop": true
|
| 850 |
+
},
|
| 851 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 852 |
+
},
|
| 853 |
+
{
|
| 854 |
+
"layer": 40,
|
| 855 |
+
"target": "same_input",
|
| 856 |
+
"block_scale_dtype": "fp16",
|
| 857 |
+
"training_sequence_count": 14657,
|
| 858 |
+
"training_tokens_seen": 7864320,
|
| 859 |
+
"training_passes": 0.5240171180844249,
|
| 860 |
+
"best_step": 5,
|
| 861 |
+
"codebook_lr": 0.048,
|
| 862 |
+
"scale_lr": 0.032,
|
| 863 |
+
"lr_decay_after": 25,
|
| 864 |
+
"adaptive_schedule": {
|
| 865 |
+
"best": 0.040921155033096596,
|
| 866 |
+
"drop_after": 20,
|
| 867 |
+
"relative_worsening": 0.001,
|
| 868 |
+
"checks": 3,
|
| 869 |
+
"patience_updates": 10,
|
| 870 |
+
"bad_checks": 3,
|
| 871 |
+
"last_best": 5,
|
| 872 |
+
"early_drop": true
|
| 873 |
+
},
|
| 874 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 875 |
+
},
|
| 876 |
+
{
|
| 877 |
+
"layer": 41,
|
| 878 |
+
"target": "same_input",
|
| 879 |
+
"block_scale_dtype": "fp16",
|
| 880 |
+
"training_sequence_count": 14657,
|
| 881 |
+
"training_tokens_seen": 7864320,
|
| 882 |
+
"training_passes": 0.5240171180844249,
|
| 883 |
+
"best_step": 5,
|
| 884 |
+
"codebook_lr": 0.048,
|
| 885 |
+
"scale_lr": 0.032,
|
| 886 |
+
"lr_decay_after": 25,
|
| 887 |
+
"adaptive_schedule": {
|
| 888 |
+
"best": 0.04285828934035643,
|
| 889 |
+
"drop_after": 20,
|
| 890 |
+
"relative_worsening": 0.001,
|
| 891 |
+
"checks": 3,
|
| 892 |
+
"patience_updates": 10,
|
| 893 |
+
"bad_checks": 3,
|
| 894 |
+
"last_best": 5,
|
| 895 |
+
"early_drop": true
|
| 896 |
+
},
|
| 897 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 898 |
+
},
|
| 899 |
+
{
|
| 900 |
+
"layer": 42,
|
| 901 |
+
"target": "same_input",
|
| 902 |
+
"block_scale_dtype": "fp16",
|
| 903 |
+
"training_sequence_count": 14657,
|
| 904 |
+
"training_tokens_seen": 7864320,
|
| 905 |
+
"training_passes": 0.5240171180844249,
|
| 906 |
+
"best_step": 5,
|
| 907 |
+
"codebook_lr": 0.048,
|
| 908 |
+
"scale_lr": 0.032,
|
| 909 |
+
"lr_decay_after": 25,
|
| 910 |
+
"adaptive_schedule": {
|
| 911 |
+
"best": 0.04590353936917018,
|
| 912 |
+
"drop_after": 20,
|
| 913 |
+
"relative_worsening": 0.001,
|
| 914 |
+
"checks": 3,
|
| 915 |
+
"patience_updates": 10,
|
| 916 |
+
"bad_checks": 3,
|
| 917 |
+
"last_best": 5,
|
| 918 |
+
"early_drop": true
|
| 919 |
+
},
|
| 920 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 921 |
+
},
|
| 922 |
+
{
|
| 923 |
+
"layer": 43,
|
| 924 |
+
"target": "same_input",
|
| 925 |
+
"block_scale_dtype": "fp16",
|
| 926 |
+
"training_sequence_count": 14657,
|
| 927 |
+
"training_tokens_seen": 7864320,
|
| 928 |
+
"training_passes": 0.5240171180844249,
|
| 929 |
+
"best_step": 5,
|
| 930 |
+
"codebook_lr": 0.048,
|
| 931 |
+
"scale_lr": 0.032,
|
| 932 |
+
"lr_decay_after": 25,
|
| 933 |
+
"adaptive_schedule": {
|
| 934 |
+
"best": 0.0445766345634676,
|
| 935 |
+
"drop_after": 20,
|
| 936 |
+
"relative_worsening": 0.001,
|
| 937 |
+
"checks": 3,
|
| 938 |
+
"patience_updates": 10,
|
| 939 |
+
"bad_checks": 3,
|
| 940 |
+
"last_best": 5,
|
| 941 |
+
"early_drop": true
|
| 942 |
+
},
|
| 943 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 944 |
+
},
|
| 945 |
+
{
|
| 946 |
+
"layer": 44,
|
| 947 |
+
"target": "same_input",
|
| 948 |
+
"block_scale_dtype": "fp16",
|
| 949 |
+
"training_sequence_count": 14657,
|
| 950 |
+
"training_tokens_seen": 7864320,
|
| 951 |
+
"training_passes": 0.5240171180844249,
|
| 952 |
+
"best_step": 5,
|
| 953 |
+
"codebook_lr": 0.048,
|
| 954 |
+
"scale_lr": 0.032,
|
| 955 |
+
"lr_decay_after": 25,
|
| 956 |
+
"adaptive_schedule": {
|
| 957 |
+
"best": 0.04553051376240955,
|
| 958 |
+
"drop_after": 20,
|
| 959 |
+
"relative_worsening": 0.001,
|
| 960 |
+
"checks": 3,
|
| 961 |
+
"patience_updates": 10,
|
| 962 |
+
"bad_checks": 3,
|
| 963 |
+
"last_best": 5,
|
| 964 |
+
"early_drop": true
|
| 965 |
+
},
|
| 966 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 967 |
+
},
|
| 968 |
+
{
|
| 969 |
+
"layer": 45,
|
| 970 |
+
"target": "same_input",
|
| 971 |
+
"block_scale_dtype": "fp16",
|
| 972 |
+
"training_sequence_count": 14657,
|
| 973 |
+
"training_tokens_seen": 7864320,
|
| 974 |
+
"training_passes": 0.5240171180844249,
|
| 975 |
+
"best_step": 5,
|
| 976 |
+
"codebook_lr": 0.048,
|
| 977 |
+
"scale_lr": 0.032,
|
| 978 |
+
"lr_decay_after": 25,
|
| 979 |
+
"adaptive_schedule": {
|
| 980 |
+
"best": 0.04488142473736064,
|
| 981 |
+
"drop_after": 20,
|
| 982 |
+
"relative_worsening": 0.001,
|
| 983 |
+
"checks": 3,
|
| 984 |
+
"patience_updates": 10,
|
| 985 |
+
"bad_checks": 3,
|
| 986 |
+
"last_best": 5,
|
| 987 |
+
"early_drop": true
|
| 988 |
+
},
|
| 989 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 990 |
+
},
|
| 991 |
+
{
|
| 992 |
+
"layer": 46,
|
| 993 |
+
"target": "same_input",
|
| 994 |
+
"block_scale_dtype": "fp16",
|
| 995 |
+
"training_sequence_count": 14657,
|
| 996 |
+
"training_tokens_seen": 7864320,
|
| 997 |
+
"training_passes": 0.5240171180844249,
|
| 998 |
+
"best_step": 5,
|
| 999 |
+
"codebook_lr": 0.048,
|
| 1000 |
+
"scale_lr": 0.032,
|
| 1001 |
+
"lr_decay_after": 25,
|
| 1002 |
+
"adaptive_schedule": {
|
| 1003 |
+
"best": 0.053625666510220674,
|
| 1004 |
+
"drop_after": 20,
|
| 1005 |
+
"relative_worsening": 0.001,
|
| 1006 |
+
"checks": 3,
|
| 1007 |
+
"patience_updates": 10,
|
| 1008 |
+
"bad_checks": 3,
|
| 1009 |
+
"last_best": 5,
|
| 1010 |
+
"early_drop": true
|
| 1011 |
+
},
|
| 1012 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1013 |
+
},
|
| 1014 |
+
{
|
| 1015 |
+
"layer": 47,
|
| 1016 |
+
"target": "same_input",
|
| 1017 |
+
"block_scale_dtype": "fp16",
|
| 1018 |
+
"training_sequence_count": 14657,
|
| 1019 |
+
"training_tokens_seen": 7864320,
|
| 1020 |
+
"training_passes": 0.5240171180844249,
|
| 1021 |
+
"best_step": 5,
|
| 1022 |
+
"codebook_lr": 0.048,
|
| 1023 |
+
"scale_lr": 0.032,
|
| 1024 |
+
"lr_decay_after": 25,
|
| 1025 |
+
"adaptive_schedule": {
|
| 1026 |
+
"best": 0.05097487104130676,
|
| 1027 |
+
"drop_after": 20,
|
| 1028 |
+
"relative_worsening": 0.001,
|
| 1029 |
+
"checks": 3,
|
| 1030 |
+
"patience_updates": 10,
|
| 1031 |
+
"bad_checks": 3,
|
| 1032 |
+
"last_best": 5,
|
| 1033 |
+
"early_drop": true
|
| 1034 |
+
},
|
| 1035 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1036 |
+
},
|
| 1037 |
+
{
|
| 1038 |
+
"layer": 48,
|
| 1039 |
+
"target": "same_input",
|
| 1040 |
+
"block_scale_dtype": "fp16",
|
| 1041 |
+
"training_sequence_count": 14657,
|
| 1042 |
+
"training_tokens_seen": 7864320,
|
| 1043 |
+
"training_passes": 0.5240171180844249,
|
| 1044 |
+
"best_step": 5,
|
| 1045 |
+
"codebook_lr": 0.048,
|
| 1046 |
+
"scale_lr": 0.032,
|
| 1047 |
+
"lr_decay_after": 25,
|
| 1048 |
+
"adaptive_schedule": {
|
| 1049 |
+
"best": 0.05253228264238492,
|
| 1050 |
+
"drop_after": 20,
|
| 1051 |
+
"relative_worsening": 0.001,
|
| 1052 |
+
"checks": 3,
|
| 1053 |
+
"patience_updates": 10,
|
| 1054 |
+
"bad_checks": 3,
|
| 1055 |
+
"last_best": 5,
|
| 1056 |
+
"early_drop": true
|
| 1057 |
+
},
|
| 1058 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1059 |
+
},
|
| 1060 |
+
{
|
| 1061 |
+
"layer": 49,
|
| 1062 |
+
"target": "same_input",
|
| 1063 |
+
"block_scale_dtype": "fp16",
|
| 1064 |
+
"training_sequence_count": 14657,
|
| 1065 |
+
"training_tokens_seen": 7864320,
|
| 1066 |
+
"training_passes": 0.5240171180844249,
|
| 1067 |
+
"best_step": 5,
|
| 1068 |
+
"codebook_lr": 0.048,
|
| 1069 |
+
"scale_lr": 0.032,
|
| 1070 |
+
"lr_decay_after": 25,
|
| 1071 |
+
"adaptive_schedule": {
|
| 1072 |
+
"best": 0.054531234896217015,
|
| 1073 |
+
"drop_after": 20,
|
| 1074 |
+
"relative_worsening": 0.001,
|
| 1075 |
+
"checks": 3,
|
| 1076 |
+
"patience_updates": 10,
|
| 1077 |
+
"bad_checks": 3,
|
| 1078 |
+
"last_best": 5,
|
| 1079 |
+
"early_drop": true
|
| 1080 |
+
},
|
| 1081 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1082 |
+
},
|
| 1083 |
+
{
|
| 1084 |
+
"layer": 50,
|
| 1085 |
+
"target": "same_input",
|
| 1086 |
+
"block_scale_dtype": "fp16",
|
| 1087 |
+
"training_sequence_count": 14657,
|
| 1088 |
+
"training_tokens_seen": 7864320,
|
| 1089 |
+
"training_passes": 0.5240171180844249,
|
| 1090 |
+
"best_step": 5,
|
| 1091 |
+
"codebook_lr": 0.048,
|
| 1092 |
+
"scale_lr": 0.032,
|
| 1093 |
+
"lr_decay_after": 25,
|
| 1094 |
+
"adaptive_schedule": {
|
| 1095 |
+
"best": 0.05778833536756532,
|
| 1096 |
+
"drop_after": 20,
|
| 1097 |
+
"relative_worsening": 0.001,
|
| 1098 |
+
"checks": 3,
|
| 1099 |
+
"patience_updates": 10,
|
| 1100 |
+
"bad_checks": 3,
|
| 1101 |
+
"last_best": 5,
|
| 1102 |
+
"early_drop": true
|
| 1103 |
+
},
|
| 1104 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1105 |
+
},
|
| 1106 |
+
{
|
| 1107 |
+
"layer": 51,
|
| 1108 |
+
"target": "same_input",
|
| 1109 |
+
"block_scale_dtype": "fp16",
|
| 1110 |
+
"training_sequence_count": 14657,
|
| 1111 |
+
"training_tokens_seen": 7864320,
|
| 1112 |
+
"training_passes": 0.5240171180844249,
|
| 1113 |
+
"best_step": 5,
|
| 1114 |
+
"codebook_lr": 0.048,
|
| 1115 |
+
"scale_lr": 0.032,
|
| 1116 |
+
"lr_decay_after": 25,
|
| 1117 |
+
"adaptive_schedule": {
|
| 1118 |
+
"best": 0.046839461506473105,
|
| 1119 |
+
"drop_after": 20,
|
| 1120 |
+
"relative_worsening": 0.001,
|
| 1121 |
+
"checks": 3,
|
| 1122 |
+
"patience_updates": 10,
|
| 1123 |
+
"bad_checks": 3,
|
| 1124 |
+
"last_best": 5,
|
| 1125 |
+
"early_drop": true
|
| 1126 |
+
},
|
| 1127 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1128 |
+
},
|
| 1129 |
+
{
|
| 1130 |
+
"layer": 52,
|
| 1131 |
+
"target": "same_input",
|
| 1132 |
+
"block_scale_dtype": "fp16",
|
| 1133 |
+
"training_sequence_count": 14657,
|
| 1134 |
+
"training_tokens_seen": 7864320,
|
| 1135 |
+
"training_passes": 0.5240171180844249,
|
| 1136 |
+
"best_step": 5,
|
| 1137 |
+
"codebook_lr": 0.048,
|
| 1138 |
+
"scale_lr": 0.032,
|
| 1139 |
+
"lr_decay_after": 25,
|
| 1140 |
+
"adaptive_schedule": {
|
| 1141 |
+
"best": 0.05742379963086141,
|
| 1142 |
+
"drop_after": 20,
|
| 1143 |
+
"relative_worsening": 0.001,
|
| 1144 |
+
"checks": 3,
|
| 1145 |
+
"patience_updates": 10,
|
| 1146 |
+
"bad_checks": 3,
|
| 1147 |
+
"last_best": 5,
|
| 1148 |
+
"early_drop": true
|
| 1149 |
+
},
|
| 1150 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1151 |
+
},
|
| 1152 |
+
{
|
| 1153 |
+
"layer": 53,
|
| 1154 |
+
"target": "same_input",
|
| 1155 |
+
"block_scale_dtype": "fp16",
|
| 1156 |
+
"training_sequence_count": 14657,
|
| 1157 |
+
"training_tokens_seen": 7864320,
|
| 1158 |
+
"training_passes": 0.5240171180844249,
|
| 1159 |
+
"best_step": 5,
|
| 1160 |
+
"codebook_lr": 0.048,
|
| 1161 |
+
"scale_lr": 0.032,
|
| 1162 |
+
"lr_decay_after": 25,
|
| 1163 |
+
"adaptive_schedule": {
|
| 1164 |
+
"best": 0.05401293698903637,
|
| 1165 |
+
"drop_after": 20,
|
| 1166 |
+
"relative_worsening": 0.001,
|
| 1167 |
+
"checks": 3,
|
| 1168 |
+
"patience_updates": 10,
|
| 1169 |
+
"bad_checks": 3,
|
| 1170 |
+
"last_best": 5,
|
| 1171 |
+
"early_drop": true
|
| 1172 |
+
},
|
| 1173 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1174 |
+
},
|
| 1175 |
+
{
|
| 1176 |
+
"layer": 54,
|
| 1177 |
+
"target": "same_input",
|
| 1178 |
+
"block_scale_dtype": "fp16",
|
| 1179 |
+
"training_sequence_count": 14657,
|
| 1180 |
+
"training_tokens_seen": 7864320,
|
| 1181 |
+
"training_passes": 0.5240171180844249,
|
| 1182 |
+
"best_step": 5,
|
| 1183 |
+
"codebook_lr": 0.048,
|
| 1184 |
+
"scale_lr": 0.032,
|
| 1185 |
+
"lr_decay_after": 25,
|
| 1186 |
+
"adaptive_schedule": {
|
| 1187 |
+
"best": 0.05698656719338924,
|
| 1188 |
+
"drop_after": 20,
|
| 1189 |
+
"relative_worsening": 0.001,
|
| 1190 |
+
"checks": 3,
|
| 1191 |
+
"patience_updates": 10,
|
| 1192 |
+
"bad_checks": 3,
|
| 1193 |
+
"last_best": 5,
|
| 1194 |
+
"early_drop": true
|
| 1195 |
+
},
|
| 1196 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1197 |
+
},
|
| 1198 |
+
{
|
| 1199 |
+
"layer": 55,
|
| 1200 |
+
"target": "same_input",
|
| 1201 |
+
"block_scale_dtype": "fp16",
|
| 1202 |
+
"training_sequence_count": 14657,
|
| 1203 |
+
"training_tokens_seen": 7864320,
|
| 1204 |
+
"training_passes": 0.5240171180844249,
|
| 1205 |
+
"best_step": 5,
|
| 1206 |
+
"codebook_lr": 0.048,
|
| 1207 |
+
"scale_lr": 0.032,
|
| 1208 |
+
"lr_decay_after": 25,
|
| 1209 |
+
"adaptive_schedule": {
|
| 1210 |
+
"best": 0.05936456247157864,
|
| 1211 |
+
"drop_after": 20,
|
| 1212 |
+
"relative_worsening": 0.001,
|
| 1213 |
+
"checks": 3,
|
| 1214 |
+
"patience_updates": 10,
|
| 1215 |
+
"bad_checks": 3,
|
| 1216 |
+
"last_best": 5,
|
| 1217 |
+
"early_drop": true
|
| 1218 |
+
},
|
| 1219 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1220 |
+
},
|
| 1221 |
+
{
|
| 1222 |
+
"layer": 56,
|
| 1223 |
+
"target": "same_input",
|
| 1224 |
+
"block_scale_dtype": "fp16",
|
| 1225 |
+
"training_sequence_count": 14657,
|
| 1226 |
+
"training_tokens_seen": 7864320,
|
| 1227 |
+
"training_passes": 0.5240171180844249,
|
| 1228 |
+
"best_step": 5,
|
| 1229 |
+
"codebook_lr": 0.048,
|
| 1230 |
+
"scale_lr": 0.032,
|
| 1231 |
+
"lr_decay_after": 25,
|
| 1232 |
+
"adaptive_schedule": {
|
| 1233 |
+
"best": 0.0527071627420072,
|
| 1234 |
+
"drop_after": 20,
|
| 1235 |
+
"relative_worsening": 0.001,
|
| 1236 |
+
"checks": 3,
|
| 1237 |
+
"patience_updates": 10,
|
| 1238 |
+
"bad_checks": 3,
|
| 1239 |
+
"last_best": 5,
|
| 1240 |
+
"early_drop": true
|
| 1241 |
+
},
|
| 1242 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1243 |
+
},
|
| 1244 |
+
{
|
| 1245 |
+
"layer": 57,
|
| 1246 |
+
"target": "same_input",
|
| 1247 |
+
"block_scale_dtype": "fp16",
|
| 1248 |
+
"training_sequence_count": 14657,
|
| 1249 |
+
"training_tokens_seen": 7864320,
|
| 1250 |
+
"training_passes": 0.5240171180844249,
|
| 1251 |
+
"best_step": 5,
|
| 1252 |
+
"codebook_lr": 0.048,
|
| 1253 |
+
"scale_lr": 0.032,
|
| 1254 |
+
"lr_decay_after": 25,
|
| 1255 |
+
"adaptive_schedule": {
|
| 1256 |
+
"best": 0.04895510529481134,
|
| 1257 |
+
"drop_after": 20,
|
| 1258 |
+
"relative_worsening": 0.001,
|
| 1259 |
+
"checks": 3,
|
| 1260 |
+
"patience_updates": 10,
|
| 1261 |
+
"bad_checks": 3,
|
| 1262 |
+
"last_best": 5,
|
| 1263 |
+
"early_drop": true
|
| 1264 |
+
},
|
| 1265 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1266 |
+
},
|
| 1267 |
+
{
|
| 1268 |
+
"layer": 58,
|
| 1269 |
+
"target": "same_input",
|
| 1270 |
+
"block_scale_dtype": "fp16",
|
| 1271 |
+
"training_sequence_count": 14657,
|
| 1272 |
+
"training_tokens_seen": 7864320,
|
| 1273 |
+
"training_passes": 0.5240171180844249,
|
| 1274 |
+
"best_step": 5,
|
| 1275 |
+
"codebook_lr": 0.048,
|
| 1276 |
+
"scale_lr": 0.032,
|
| 1277 |
+
"lr_decay_after": 25,
|
| 1278 |
+
"adaptive_schedule": {
|
| 1279 |
+
"best": 0.05838005265713665,
|
| 1280 |
+
"drop_after": 20,
|
| 1281 |
+
"relative_worsening": 0.001,
|
| 1282 |
+
"checks": 3,
|
| 1283 |
+
"patience_updates": 10,
|
| 1284 |
+
"bad_checks": 3,
|
| 1285 |
+
"last_best": 5,
|
| 1286 |
+
"early_drop": true
|
| 1287 |
+
},
|
| 1288 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1289 |
+
},
|
| 1290 |
+
{
|
| 1291 |
+
"layer": 59,
|
| 1292 |
+
"target": "same_input",
|
| 1293 |
+
"block_scale_dtype": "fp16",
|
| 1294 |
+
"training_sequence_count": 14657,
|
| 1295 |
+
"training_tokens_seen": 7864320,
|
| 1296 |
+
"training_passes": 0.5240171180844249,
|
| 1297 |
+
"best_step": 5,
|
| 1298 |
+
"codebook_lr": 0.048,
|
| 1299 |
+
"scale_lr": 0.032,
|
| 1300 |
+
"lr_decay_after": 25,
|
| 1301 |
+
"adaptive_schedule": {
|
| 1302 |
+
"best": 0.06076521669703373,
|
| 1303 |
+
"drop_after": 20,
|
| 1304 |
+
"relative_worsening": 0.001,
|
| 1305 |
+
"checks": 3,
|
| 1306 |
+
"patience_updates": 10,
|
| 1307 |
+
"bad_checks": 3,
|
| 1308 |
+
"last_best": 5,
|
| 1309 |
+
"early_drop": true
|
| 1310 |
+
},
|
| 1311 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1312 |
+
},
|
| 1313 |
+
{
|
| 1314 |
+
"layer": 60,
|
| 1315 |
+
"target": "same_input",
|
| 1316 |
+
"block_scale_dtype": "fp16",
|
| 1317 |
+
"training_sequence_count": 14657,
|
| 1318 |
+
"training_tokens_seen": 7864320,
|
| 1319 |
+
"training_passes": 0.5240171180844249,
|
| 1320 |
+
"best_step": 5,
|
| 1321 |
+
"codebook_lr": 0.048,
|
| 1322 |
+
"scale_lr": 0.032,
|
| 1323 |
+
"lr_decay_after": 25,
|
| 1324 |
+
"adaptive_schedule": {
|
| 1325 |
+
"best": 0.05275984301503296,
|
| 1326 |
+
"drop_after": 20,
|
| 1327 |
+
"relative_worsening": 0.001,
|
| 1328 |
+
"checks": 3,
|
| 1329 |
+
"patience_updates": 10,
|
| 1330 |
+
"bad_checks": 3,
|
| 1331 |
+
"last_best": 5,
|
| 1332 |
+
"early_drop": true
|
| 1333 |
+
},
|
| 1334 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1335 |
+
},
|
| 1336 |
+
{
|
| 1337 |
+
"layer": 61,
|
| 1338 |
+
"target": "same_input",
|
| 1339 |
+
"block_scale_dtype": "fp16",
|
| 1340 |
+
"training_sequence_count": 14657,
|
| 1341 |
+
"training_tokens_seen": 7864320,
|
| 1342 |
+
"training_passes": 0.5240171180844249,
|
| 1343 |
+
"best_step": 5,
|
| 1344 |
+
"codebook_lr": 0.048,
|
| 1345 |
+
"scale_lr": 0.032,
|
| 1346 |
+
"lr_decay_after": 25,
|
| 1347 |
+
"adaptive_schedule": {
|
| 1348 |
+
"best": 0.05175780703646056,
|
| 1349 |
+
"drop_after": 20,
|
| 1350 |
+
"relative_worsening": 0.001,
|
| 1351 |
+
"checks": 3,
|
| 1352 |
+
"patience_updates": 10,
|
| 1353 |
+
"bad_checks": 3,
|
| 1354 |
+
"last_best": 5,
|
| 1355 |
+
"early_drop": true
|
| 1356 |
+
},
|
| 1357 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1358 |
+
},
|
| 1359 |
+
{
|
| 1360 |
+
"layer": 62,
|
| 1361 |
+
"target": "same_input",
|
| 1362 |
+
"block_scale_dtype": "fp16",
|
| 1363 |
+
"training_sequence_count": 14657,
|
| 1364 |
+
"training_tokens_seen": 7864320,
|
| 1365 |
+
"training_passes": 0.5240171180844249,
|
| 1366 |
+
"best_step": 5,
|
| 1367 |
+
"codebook_lr": 0.048,
|
| 1368 |
+
"scale_lr": 0.032,
|
| 1369 |
+
"lr_decay_after": 25,
|
| 1370 |
+
"adaptive_schedule": {
|
| 1371 |
+
"best": 0.05643135330733986,
|
| 1372 |
+
"drop_after": 20,
|
| 1373 |
+
"relative_worsening": 0.001,
|
| 1374 |
+
"checks": 3,
|
| 1375 |
+
"patience_updates": 10,
|
| 1376 |
+
"bad_checks": 3,
|
| 1377 |
+
"last_best": 5,
|
| 1378 |
+
"early_drop": true
|
| 1379 |
+
},
|
| 1380 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1381 |
+
},
|
| 1382 |
+
{
|
| 1383 |
+
"layer": 63,
|
| 1384 |
+
"target": "same_input",
|
| 1385 |
+
"block_scale_dtype": "fp16",
|
| 1386 |
+
"training_sequence_count": 14657,
|
| 1387 |
+
"training_tokens_seen": 7864320,
|
| 1388 |
+
"training_passes": 0.5240171180844249,
|
| 1389 |
+
"best_step": 5,
|
| 1390 |
+
"codebook_lr": 0.048,
|
| 1391 |
+
"scale_lr": 0.032,
|
| 1392 |
+
"lr_decay_after": 25,
|
| 1393 |
+
"adaptive_schedule": {
|
| 1394 |
+
"best": 0.0563437113825566,
|
| 1395 |
+
"drop_after": 20,
|
| 1396 |
+
"relative_worsening": 0.001,
|
| 1397 |
+
"checks": 3,
|
| 1398 |
+
"patience_updates": 10,
|
| 1399 |
+
"bad_checks": 3,
|
| 1400 |
+
"last_best": 5,
|
| 1401 |
+
"early_drop": true
|
| 1402 |
+
},
|
| 1403 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1404 |
+
},
|
| 1405 |
+
{
|
| 1406 |
+
"layer": 64,
|
| 1407 |
+
"target": "same_input",
|
| 1408 |
+
"block_scale_dtype": "fp16",
|
| 1409 |
+
"training_sequence_count": 14657,
|
| 1410 |
+
"training_tokens_seen": 7864320,
|
| 1411 |
+
"training_passes": 0.5240171180844249,
|
| 1412 |
+
"best_step": 5,
|
| 1413 |
+
"codebook_lr": 0.048,
|
| 1414 |
+
"scale_lr": 0.032,
|
| 1415 |
+
"lr_decay_after": 25,
|
| 1416 |
+
"adaptive_schedule": {
|
| 1417 |
+
"best": 0.048833171244225114,
|
| 1418 |
+
"drop_after": 20,
|
| 1419 |
+
"relative_worsening": 0.001,
|
| 1420 |
+
"checks": 3,
|
| 1421 |
+
"patience_updates": 10,
|
| 1422 |
+
"bad_checks": 3,
|
| 1423 |
+
"last_best": 5,
|
| 1424 |
+
"early_drop": true
|
| 1425 |
+
},
|
| 1426 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1427 |
+
},
|
| 1428 |
+
{
|
| 1429 |
+
"layer": 65,
|
| 1430 |
+
"target": "same_input",
|
| 1431 |
+
"block_scale_dtype": "fp16",
|
| 1432 |
+
"training_sequence_count": 14657,
|
| 1433 |
+
"training_tokens_seen": 7864320,
|
| 1434 |
+
"training_passes": 0.5240171180844249,
|
| 1435 |
+
"best_step": 5,
|
| 1436 |
+
"codebook_lr": 0.048,
|
| 1437 |
+
"scale_lr": 0.032,
|
| 1438 |
+
"lr_decay_after": 25,
|
| 1439 |
+
"adaptive_schedule": {
|
| 1440 |
+
"best": 0.060005061741192876,
|
| 1441 |
+
"drop_after": 20,
|
| 1442 |
+
"relative_worsening": 0.001,
|
| 1443 |
+
"checks": 3,
|
| 1444 |
+
"patience_updates": 10,
|
| 1445 |
+
"bad_checks": 3,
|
| 1446 |
+
"last_best": 5,
|
| 1447 |
+
"early_drop": true
|
| 1448 |
+
},
|
| 1449 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1450 |
+
},
|
| 1451 |
+
{
|
| 1452 |
+
"layer": 66,
|
| 1453 |
+
"target": "same_input",
|
| 1454 |
+
"block_scale_dtype": "fp16",
|
| 1455 |
+
"training_sequence_count": 14657,
|
| 1456 |
+
"training_tokens_seen": 1310720,
|
| 1457 |
+
"training_passes": 0.08733618634740414,
|
| 1458 |
+
"best_step": 5,
|
| 1459 |
+
"codebook_lr": 0.048,
|
| 1460 |
+
"scale_lr": 0.032,
|
| 1461 |
+
"lr_decay_after": 25,
|
| 1462 |
+
"adaptive_schedule": {
|
| 1463 |
+
"best": 0.05761269836156414,
|
| 1464 |
+
"drop_after": 25,
|
| 1465 |
+
"relative_worsening": 0.001,
|
| 1466 |
+
"checks": 3,
|
| 1467 |
+
"patience_updates": 10,
|
| 1468 |
+
"bad_checks": 0,
|
| 1469 |
+
"last_best": 5,
|
| 1470 |
+
"early_drop": false
|
| 1471 |
+
},
|
| 1472 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1473 |
+
},
|
| 1474 |
+
{
|
| 1475 |
+
"layer": 67,
|
| 1476 |
+
"target": "same_input",
|
| 1477 |
+
"block_scale_dtype": "fp16",
|
| 1478 |
+
"training_sequence_count": 14657,
|
| 1479 |
+
"training_tokens_seen": 7864320,
|
| 1480 |
+
"training_passes": 0.5240171180844249,
|
| 1481 |
+
"best_step": 5,
|
| 1482 |
+
"codebook_lr": 0.048,
|
| 1483 |
+
"scale_lr": 0.032,
|
| 1484 |
+
"lr_decay_after": 25,
|
| 1485 |
+
"adaptive_schedule": {
|
| 1486 |
+
"best": 0.05480543365221943,
|
| 1487 |
+
"drop_after": 20,
|
| 1488 |
+
"relative_worsening": 0.001,
|
| 1489 |
+
"checks": 3,
|
| 1490 |
+
"patience_updates": 10,
|
| 1491 |
+
"bad_checks": 3,
|
| 1492 |
+
"last_best": 5,
|
| 1493 |
+
"early_drop": true
|
| 1494 |
+
},
|
| 1495 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1496 |
+
},
|
| 1497 |
+
{
|
| 1498 |
+
"layer": 68,
|
| 1499 |
+
"target": "same_input",
|
| 1500 |
+
"block_scale_dtype": "fp16",
|
| 1501 |
+
"training_sequence_count": 14657,
|
| 1502 |
+
"training_tokens_seen": 7864320,
|
| 1503 |
+
"training_passes": 0.5240171180844249,
|
| 1504 |
+
"best_step": 5,
|
| 1505 |
+
"codebook_lr": 0.048,
|
| 1506 |
+
"scale_lr": 0.032,
|
| 1507 |
+
"lr_decay_after": 25,
|
| 1508 |
+
"adaptive_schedule": {
|
| 1509 |
+
"best": 0.062059367848489574,
|
| 1510 |
+
"drop_after": 20,
|
| 1511 |
+
"relative_worsening": 0.001,
|
| 1512 |
+
"checks": 3,
|
| 1513 |
+
"patience_updates": 10,
|
| 1514 |
+
"bad_checks": 3,
|
| 1515 |
+
"last_best": 5,
|
| 1516 |
+
"early_drop": true
|
| 1517 |
+
},
|
| 1518 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1519 |
+
},
|
| 1520 |
+
{
|
| 1521 |
+
"layer": 69,
|
| 1522 |
+
"target": "same_input",
|
| 1523 |
+
"block_scale_dtype": "fp16",
|
| 1524 |
+
"training_sequence_count": 14657,
|
| 1525 |
+
"training_tokens_seen": 10485760,
|
| 1526 |
+
"training_passes": 0.6986894907792331,
|
| 1527 |
+
"best_step": 30,
|
| 1528 |
+
"codebook_lr": 0.048,
|
| 1529 |
+
"scale_lr": 0.032,
|
| 1530 |
+
"lr_decay_after": 25,
|
| 1531 |
+
"adaptive_schedule": {
|
| 1532 |
+
"best": 0.056006381621796365,
|
| 1533 |
+
"drop_after": 20,
|
| 1534 |
+
"relative_worsening": 0.001,
|
| 1535 |
+
"checks": 3,
|
| 1536 |
+
"patience_updates": 10,
|
| 1537 |
+
"bad_checks": 3,
|
| 1538 |
+
"last_best": 30,
|
| 1539 |
+
"early_drop": true
|
| 1540 |
+
},
|
| 1541 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1542 |
+
},
|
| 1543 |
+
{
|
| 1544 |
+
"layer": 70,
|
| 1545 |
+
"target": "same_input",
|
| 1546 |
+
"block_scale_dtype": "fp16",
|
| 1547 |
+
"training_sequence_count": 14657,
|
| 1548 |
+
"training_tokens_seen": 7864320,
|
| 1549 |
+
"training_passes": 0.5240171180844249,
|
| 1550 |
+
"best_step": 5,
|
| 1551 |
+
"codebook_lr": 0.048,
|
| 1552 |
+
"scale_lr": 0.032,
|
| 1553 |
+
"lr_decay_after": 25,
|
| 1554 |
+
"adaptive_schedule": {
|
| 1555 |
+
"best": 0.0733148360257097,
|
| 1556 |
+
"drop_after": 20,
|
| 1557 |
+
"relative_worsening": 0.001,
|
| 1558 |
+
"checks": 3,
|
| 1559 |
+
"patience_updates": 10,
|
| 1560 |
+
"bad_checks": 3,
|
| 1561 |
+
"last_best": 5,
|
| 1562 |
+
"early_drop": true
|
| 1563 |
+
},
|
| 1564 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1565 |
+
},
|
| 1566 |
+
{
|
| 1567 |
+
"layer": 71,
|
| 1568 |
+
"target": "same_input",
|
| 1569 |
+
"block_scale_dtype": "fp16",
|
| 1570 |
+
"training_sequence_count": 14657,
|
| 1571 |
+
"training_tokens_seen": 7864320,
|
| 1572 |
+
"training_passes": 0.5240171180844249,
|
| 1573 |
+
"best_step": 5,
|
| 1574 |
+
"codebook_lr": 0.048,
|
| 1575 |
+
"scale_lr": 0.032,
|
| 1576 |
+
"lr_decay_after": 25,
|
| 1577 |
+
"adaptive_schedule": {
|
| 1578 |
+
"best": 0.06478566802319799,
|
| 1579 |
+
"drop_after": 20,
|
| 1580 |
+
"relative_worsening": 0.001,
|
| 1581 |
+
"checks": 3,
|
| 1582 |
+
"patience_updates": 10,
|
| 1583 |
+
"bad_checks": 3,
|
| 1584 |
+
"last_best": 5,
|
| 1585 |
+
"early_drop": true
|
| 1586 |
+
},
|
| 1587 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1588 |
+
},
|
| 1589 |
+
{
|
| 1590 |
+
"layer": 72,
|
| 1591 |
+
"target": "same_input",
|
| 1592 |
+
"block_scale_dtype": "fp16",
|
| 1593 |
+
"training_sequence_count": 14657,
|
| 1594 |
+
"training_tokens_seen": 7864320,
|
| 1595 |
+
"training_passes": 0.5240171180844249,
|
| 1596 |
+
"best_step": 5,
|
| 1597 |
+
"codebook_lr": 0.048,
|
| 1598 |
+
"scale_lr": 0.032,
|
| 1599 |
+
"lr_decay_after": 25,
|
| 1600 |
+
"adaptive_schedule": {
|
| 1601 |
+
"best": 0.05970967496138404,
|
| 1602 |
+
"drop_after": 20,
|
| 1603 |
+
"relative_worsening": 0.001,
|
| 1604 |
+
"checks": 3,
|
| 1605 |
+
"patience_updates": 10,
|
| 1606 |
+
"bad_checks": 3,
|
| 1607 |
+
"last_best": 5,
|
| 1608 |
+
"early_drop": true
|
| 1609 |
+
},
|
| 1610 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1611 |
+
},
|
| 1612 |
+
{
|
| 1613 |
+
"layer": 73,
|
| 1614 |
+
"target": "same_input",
|
| 1615 |
+
"block_scale_dtype": "fp16",
|
| 1616 |
+
"training_sequence_count": 14657,
|
| 1617 |
+
"training_tokens_seen": 7864320,
|
| 1618 |
+
"training_passes": 0.5240171180844249,
|
| 1619 |
+
"best_step": 5,
|
| 1620 |
+
"codebook_lr": 0.048,
|
| 1621 |
+
"scale_lr": 0.032,
|
| 1622 |
+
"lr_decay_after": 25,
|
| 1623 |
+
"adaptive_schedule": {
|
| 1624 |
+
"best": 0.0664773336538204,
|
| 1625 |
+
"drop_after": 20,
|
| 1626 |
+
"relative_worsening": 0.001,
|
| 1627 |
+
"checks": 3,
|
| 1628 |
+
"patience_updates": 10,
|
| 1629 |
+
"bad_checks": 3,
|
| 1630 |
+
"last_best": 5,
|
| 1631 |
+
"early_drop": true
|
| 1632 |
+
},
|
| 1633 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1634 |
+
},
|
| 1635 |
+
{
|
| 1636 |
+
"layer": 74,
|
| 1637 |
+
"target": "same_input",
|
| 1638 |
+
"block_scale_dtype": "fp16",
|
| 1639 |
+
"training_sequence_count": 14657,
|
| 1640 |
+
"training_tokens_seen": 7864320,
|
| 1641 |
+
"training_passes": 0.5240171180844249,
|
| 1642 |
+
"best_step": 5,
|
| 1643 |
+
"codebook_lr": 0.048,
|
| 1644 |
+
"scale_lr": 0.032,
|
| 1645 |
+
"lr_decay_after": 25,
|
| 1646 |
+
"adaptive_schedule": {
|
| 1647 |
+
"best": 0.05450847628729126,
|
| 1648 |
+
"drop_after": 20,
|
| 1649 |
+
"relative_worsening": 0.001,
|
| 1650 |
+
"checks": 3,
|
| 1651 |
+
"patience_updates": 10,
|
| 1652 |
+
"bad_checks": 3,
|
| 1653 |
+
"last_best": 5,
|
| 1654 |
+
"early_drop": true
|
| 1655 |
+
},
|
| 1656 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1657 |
+
},
|
| 1658 |
+
{
|
| 1659 |
+
"layer": 75,
|
| 1660 |
+
"target": "same_input",
|
| 1661 |
+
"block_scale_dtype": "fp16",
|
| 1662 |
+
"training_sequence_count": 14657,
|
| 1663 |
+
"training_tokens_seen": 9175040,
|
| 1664 |
+
"training_passes": 0.611353304431829,
|
| 1665 |
+
"best_step": 25,
|
| 1666 |
+
"codebook_lr": 0.048,
|
| 1667 |
+
"scale_lr": 0.032,
|
| 1668 |
+
"lr_decay_after": 25,
|
| 1669 |
+
"adaptive_schedule": {
|
| 1670 |
+
"best": 0.04963358343974084,
|
| 1671 |
+
"drop_after": 20,
|
| 1672 |
+
"relative_worsening": 0.001,
|
| 1673 |
+
"checks": 3,
|
| 1674 |
+
"patience_updates": 10,
|
| 1675 |
+
"bad_checks": 3,
|
| 1676 |
+
"last_best": 25,
|
| 1677 |
+
"early_drop": true
|
| 1678 |
+
},
|
| 1679 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1680 |
+
},
|
| 1681 |
+
{
|
| 1682 |
+
"layer": 76,
|
| 1683 |
+
"target": "same_input",
|
| 1684 |
+
"block_scale_dtype": "fp16",
|
| 1685 |
+
"training_sequence_count": 14657,
|
| 1686 |
+
"training_tokens_seen": 7864320,
|
| 1687 |
+
"training_passes": 0.5240171180844249,
|
| 1688 |
+
"best_step": 5,
|
| 1689 |
+
"codebook_lr": 0.048,
|
| 1690 |
+
"scale_lr": 0.032,
|
| 1691 |
+
"lr_decay_after": 25,
|
| 1692 |
+
"adaptive_schedule": {
|
| 1693 |
+
"best": 0.05268687462948389,
|
| 1694 |
+
"drop_after": 20,
|
| 1695 |
+
"relative_worsening": 0.001,
|
| 1696 |
+
"checks": 3,
|
| 1697 |
+
"patience_updates": 10,
|
| 1698 |
+
"bad_checks": 3,
|
| 1699 |
+
"last_best": 5,
|
| 1700 |
+
"early_drop": true
|
| 1701 |
+
},
|
| 1702 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1703 |
+
},
|
| 1704 |
+
{
|
| 1705 |
+
"layer": 77,
|
| 1706 |
+
"target": "same_input",
|
| 1707 |
+
"block_scale_dtype": "fp16",
|
| 1708 |
+
"training_sequence_count": 14657,
|
| 1709 |
+
"training_tokens_seen": 14417920,
|
| 1710 |
+
"training_passes": 0.9606980498214457,
|
| 1711 |
+
"best_step": 45,
|
| 1712 |
+
"codebook_lr": 0.048,
|
| 1713 |
+
"scale_lr": 0.032,
|
| 1714 |
+
"lr_decay_after": 25,
|
| 1715 |
+
"adaptive_schedule": {
|
| 1716 |
+
"best": 0.03900097511967364,
|
| 1717 |
+
"drop_after": 25,
|
| 1718 |
+
"relative_worsening": 0.001,
|
| 1719 |
+
"checks": 3,
|
| 1720 |
+
"patience_updates": 10,
|
| 1721 |
+
"bad_checks": 0,
|
| 1722 |
+
"last_best": 45,
|
| 1723 |
+
"early_drop": false
|
| 1724 |
+
},
|
| 1725 |
+
"checkpoint_selection": "held-out original reference trajectory for both arms"
|
| 1726 |
+
}
|
| 1727 |
+
]
|
reproduce/package_manifest.json
ADDED
|
@@ -0,0 +1,247 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"files": {
|
| 3 |
+
"arvq88_reference.py": "b0606bf8ae91bab655ed5b26bb9b7437f0e28718301fdb1b86f475b1b31924e1",
|
| 4 |
+
"METHODOLOGY.md": "45ad59ea38585133a8005e169898a4dcd7ae56b9a4ce8ff3ba7e9f1ca00e9253",
|
| 5 |
+
"pv_progress.json": "583c9ebe0f26dfd4ffee434f1a7abc66af61a149b4f818ab5ccc4d5d78e7e3a9",
|
| 6 |
+
"README.md": "6d219086224a5c4fdc8085d04bae3292747673408d52275fa7ce17561c0a9c4b",
|
| 7 |
+
"calibration_corpus.json": "43517a1e4a3580b4da7c628019d0ad5d0a610aa469cfddcafc05f86bff47a770",
|
| 8 |
+
"reproduce/source_manifest.json": "eadfaac3a4552c341d64293b25f702b66358ba29450bbdba0e9b80f473ac9819",
|
| 9 |
+
"reproduce/environment_observed.json": "e9ae2da2456329999caf5ff468067e3c47caabd754bd80ae27b74f3c6c6f16e3",
|
| 10 |
+
"reproduce/requirements.in": "3f58e110a3d1816aef49af4fe3e44e2ff36e4bffd8c3187a099d556ef690d671",
|
| 11 |
+
"reproduce/data_artifacts.json": "bf1e037960df3435a829526d7145149036553f6e2ebe514d35c308a37e9a4126",
|
| 12 |
+
"reproduce/checkpoint_before_audit.json": "a8822199bb0af460349ba8c930e7abdbdf23096ba68276f429068c94b325fcce",
|
| 13 |
+
"reproduce/layer_recipe_summary.json": "1bdb40cd40eb3a80e9b0fc4e7b162be0143cd42cfbdcc592170e24e9f918a68e",
|
| 14 |
+
"reproduce/recipe.json": "4a6e876f0320909178001af53b086063a1951191e2ddfc26e7c696c76cdbc524",
|
| 15 |
+
"reproduce/README.md": "007935e4deb8d402f7d43bcd2dc7e2d2d27cf78dfde39fad35a58b4eae96126e",
|
| 16 |
+
"reproduce/RUNBOOK.md": "d13c303ef1c80eba3bc0aea0df43a263b042a5cd3fd76500e88432fe561b6407",
|
| 17 |
+
"reproduce/prepare_workspace.py": "cc679deae776de7beacece214c4868e51e35bc4a61a7a2725a4d05d1bb8324e5",
|
| 18 |
+
"reproduce/verify.py": "418df7c6dc54de146f0ff465618d42e69f8489eecc7743abc9464f8e15fa96e2",
|
| 19 |
+
"reproduce/test_roundtrip.py": "fdce4358d869c1fb5fe01a7a9a878d6a29710af65779bbf0637be52eefd84735",
|
| 20 |
+
"reproduce/external_sources.json": "42d948405368409e11c646a313f055e417271f4a828a907743d86a3094bba82a",
|
| 21 |
+
"reproduce/THIRD_PARTY.md": "b26dc2e38b00f1b7e17881535567ac8e811725abbdc709342f67f2eb08b93ac5",
|
| 22 |
+
"reproduce/tensor_header_sample.json": "7f804b22fa108706614d8fb7d12302eae7e4c537c746299c7ca00dc5670f4916",
|
| 23 |
+
"reproduce/AUDIT.md": "ae0e0e3447269352176307811ed4d07c4b9b91525f4d06724e9840cce3af1c8b",
|
| 24 |
+
"reproduce/historical_metadata/source.json": "c797ff4755b4ef9bf4b1fe8143ea7cfd55e6b52383fe11debaf4a23bf4cb1ffc",
|
| 25 |
+
"reproduce/historical_metadata/backbone_sources.json": "b9fc2ca4ed0da88b4061aff2d6cf53d59afae80a1d291a1ac8952f9a2063cd51",
|
| 26 |
+
"reproduce/historical_metadata/cold_assignment.json": "2bc853d8c282e2e7b7a3cb6fe390bda59d912d046063c6ee235cfa551014f695",
|
| 27 |
+
"reproduce/historical_metadata/calibration_corpus.json": "5c3e0f16c03632cf987d4e6cc4af731dd51abd7f12e18366285dbe0766150743",
|
| 28 |
+
"reproduce/historical_metadata/build_provenance.json": "d137642397e2c8cab6da61f1a7975453b5bf9d010ba1217a2d2b683d9cf75dbd",
|
| 29 |
+
"reproduce/historical_metadata/pv_progress.json": "ebec214f80ea0b16365f1aa819f6723a9461f5bf14a80f977f705025140ac6bc",
|
| 30 |
+
"reproduce/historical_metadata/METHODOLOGY.md": "43591534ca9a1e8db9015fd05368177a37efb1400b5d2ed699e2844e17485805",
|
| 31 |
+
"reproduce/source/tools/gguf_remote.py": "8befc8a77f2e7964c9bca0c5732a919b58dc5598fe33b00d4708bf0d8547232a",
|
| 32 |
+
"reproduce/source/tools/budget.py": "562dbf3a01822eb500f92d46f9710c987ba48bd2faf4fe1f0a12474ed1cf08cb",
|
| 33 |
+
"reproduce/source/tools/make_manifests.py": "b1c92a95a3e49bc5d27207e915a402da6a8713facbae1e09c00b5d1a932a4740",
|
| 34 |
+
"reproduce/source/tools/range_download.py": "6b8462c4af6372b5899bdae608537a76e5a470fc4938bc3d67ea5de7f16e9581",
|
| 35 |
+
"reproduce/source/tools/validate_serve.py": "e328e06a20246d58b0225ddb0d6cc00fa7450623c7ebb464da41cb58f29c8ea6",
|
| 36 |
+
"reproduce/source/tools/build_checkpoint.py": "af00a9e5c1a1b96c302bf7ea7eb7f5f9b5f32060de2c186085067b2bd5868ee5",
|
| 37 |
+
"reproduce/source/tools/patch_checkpoint.py": "cb5bdf47be506c0293b96055b83cf38f0ca807414aa0b9c640c16c2186246a3b",
|
| 38 |
+
"reproduce/source/tools/strip_stale.py": "3684269f45eeb2ea1b6698c80774a4a74cf616dea4832004be884ef1d5483b8e",
|
| 39 |
+
"reproduce/source/tools/collect_expert_stats.py": "979d22afbaf18e920a576857e6aedadbb74a120fdea593f2609a2e813342d16f",
|
| 40 |
+
"reproduce/source/tools/build_checkpoint_v3.py": "e0dc5a53b242a4a7d54ba7412798792bfc75f95ccd31632f7e589949a27500c2",
|
| 41 |
+
"reproduce/source/tools/solve_assignment.py": "f95d133a3dd96a797bec1a589e49e405056ffe9a525174994b95ec0a9059bba2",
|
| 42 |
+
"reproduce/source/tools/collect_expert_stats_v2.py": "7837bca536acd8513b2a7f6f127460b2e490bb615eabcc81fa06675d4915e1c4",
|
| 43 |
+
"reproduce/source/tools/patch_checkpoint_v3.py": "8e0479ec80071339dcd1ca06a9de4cbc12e661750142c037c24ee3f51df5a60e",
|
| 44 |
+
"reproduce/source/tools/solve_assignment_2tier.py": "c0de546bfe8ec4698eb8c0ff5c73a7f270513f217670ed44481cf7913cec3430",
|
| 45 |
+
"reproduce/source/tools/make_hot_manifest.py": "310b088822bbcf423bc4fc67ccd1033b0c8bfbb761f778f9b13f95ec303822d5",
|
| 46 |
+
"reproduce/source/tools/build_checkpoint_v4.py": "46624935e36542442b15b65d4f27d7cb4fb4ee40759ec59a9823b69785c7d4d6",
|
| 47 |
+
"reproduce/source/tools/build_checkpoint_v5.py": "b72c276e1977b67ab31f7e269dc1c82529d0431e45df1c1676cca378db19fd29",
|
| 48 |
+
"reproduce/source/tools/make_hot_manifest2.py": "a7b8ddc26edf59956413e2ba51297b694a9edc33797fd726596798d249637dd6",
|
| 49 |
+
"reproduce/source/tools/build_checkpoint_v6.py": "bd707467d8d8ff0738eddc4ecb324125b7b6c094fab6532addacca9e7cd26725",
|
| 50 |
+
"reproduce/source/tools/run_rtx6000.sh": "6ab7fdcffa0fd1f4768134d0f98d670259935fafcef9e478204579b6376ac478",
|
| 51 |
+
"reproduce/source/tools/capture_golden.py": "353997aca54a26b622d831dc6224199bf9c73bfcdb31be61a9abdb4bfc6bb100",
|
| 52 |
+
"reproduce/source/tools/make_kernel_vectors.py": "633dcde4ca640a19754baddaa98eea065ee8c45af80c68c075f4cb9ed542f395",
|
| 53 |
+
"reproduce/source/tools/verify_sm120.py": "51f4f52f28e6424bc094273055c6ad62d138d70780c7632edfba0415524d5433",
|
| 54 |
+
"reproduce/source/tools/capture_acts.py": "907206bce4a843da2c5d3d3c153dc09c0787a70bae62fd388664fd2d114da632",
|
| 55 |
+
"reproduce/source/tools/aqlm_converge.py": "9cb4c093a6c8dfd7639b46bd68f98c7b113d39666642a4c117748776cd625cd5",
|
| 56 |
+
"reproduce/source/tools/CONVERGER.md": "3394b8437e2eb6038dfd8d7159768634dd4c87137847293a38dfb2b252472894",
|
| 57 |
+
"reproduce/source/tools/score_experts_reap.py": "21e675cc5ad0b2727ccedd5695341541a98a8575fb4ce9797ae8110a9d175b23",
|
| 58 |
+
"reproduce/source/tools/build_checkpoint_v7.py": "54e4838a08d6b6c8b6e4e9e8cc83c7636b46a1f4511e85140c72dd4280c2d911",
|
| 59 |
+
"reproduce/source/tools/build_calib_v3.py": "38f51b1d4ada52631d35c6afab1983b3118e4dadcf1a67c5f7ff3ede30ac499c",
|
| 60 |
+
"reproduce/source/tools/phase15_fit.py": "b2c4fe789e9fc355abcbdecb9a89fe53922f58564cb53af381ff27fb1abb6346",
|
| 61 |
+
"reproduce/source/tools/build_checkpoint_v8.py": "c30f8265ea2c6a982dcb6e2b432b28f34c4d765e1e55ad06ad42a275992b11b7",
|
| 62 |
+
"reproduce/source/tools/capture_acts_v2.py": "15ee0237b58c0e1cc68dbfcf832dd7e741fdfb0270678ea30bd1819749d01b8c",
|
| 63 |
+
"reproduce/source/tools/aqlm_full.py": "37b3757632e3be9e9e88ebb7145769c96ca956835593e1c039a5b36d79b84c0d",
|
| 64 |
+
"reproduce/source/tools/make_bf16_manifest.py": "7716ffabdd110ad606fe69992d1bc9380d7f4713c1f2dc52c1163e28f1b7be61",
|
| 65 |
+
"reproduce/source/tools/phase21_pipeline.py": "45fa1b0a34863d5a0112f4d8ced733b4a284dc6a2fe3b4d74308cbb64411b8ad",
|
| 66 |
+
"reproduce/source/tools/build_heldout_xl.py": "a331c916257890166729bb6b4de470c93a735f06194fd740a58e23f25233bd3a",
|
| 67 |
+
"reproduce/source/tools/bf16_stream.py": "67b7cdcb13c621f7c53d60d1c4d2d3aef6f64044a7c5e29e19ae6b42955360c4",
|
| 68 |
+
"reproduce/source/tools/pv_tune.py": "878ee5749eca7cc88ca3591a0b5d1f5d9c2f33e1baf584f7a1ae9631517f5038",
|
| 69 |
+
"reproduce/source/tools/capture_acts_fresh.py": "e46b3148902f8cee8a82a145ccfeac9a3c21332773b9f8b0e37ffe96b328d9b7",
|
| 70 |
+
"reproduce/source/tools/build_checkpoint_v9.py": "3f4e177466d57661d3f4f5fa75af22a11d89b46401c4e53a0cea970a82d46131",
|
| 71 |
+
"reproduce/source/tools/convert_fp8_attn.py": "20ed2e8f8398b1098e54ce953b8234f1484a75e4bd5508bbe2bc4ddd98d4af77",
|
| 72 |
+
"reproduce/source/tools/fetch_nvfp4_experts.py": "79c8f59d25880eaf288b2403a95c67b95df48144b8ae5a283cb55ff0a65bebaa",
|
| 73 |
+
"reproduce/source/tools/build_checkpoint_v10.py": "d84b3a38b74a7f6bef2913a03e710eac13a6d434a1a2adc96863b06de45863f8",
|
| 74 |
+
"reproduce/source/tools/encode_onto_codebook.py": "f9e3343a6cd5d95da2a48f79431cedb70b7c5c311dfd1392e398f8c425425626",
|
| 75 |
+
"reproduce/source/tools/run_b200_tp8.sh": "a71e0d82b6a3e9cc485f87098c8cb8f9fb86e0755d0bf3845521452f781dffda",
|
| 76 |
+
"reproduce/source/tools/aqlm_quantize.py": "da0cea94862d7703c028a139bc15fbc41d719390c09633303a1b678e6fcd31cd",
|
| 77 |
+
"reproduce/source/tools/build_calib_v31.py": "b87f2d9415380e28cfba2a0ce75f876ee48e37bb6e865c4a02d43e2300698fa7",
|
| 78 |
+
"reproduce/source/tools/reasoning_slice.py": "6878c1f6162cb97a98ee11679f97e39f616c0e61818312670bd8a4ffb5639f23",
|
| 79 |
+
"reproduce/source/tools/build_calib_v32.py": "6b369e36e210351fbdead6754170a83c6410e9692292e83a29926e50d23d1b08",
|
| 80 |
+
"reproduce/source/tools/build_trace_refit_v32.py": "0d4bdfa5345c0051bf03a7bcf9798922cb4cccdf3e3a379526e541e2af4f3ab9",
|
| 81 |
+
"reproduce/source/btx53/pilot_expert.py": "fb09fd073d8168d3b563ed5affb184be22bf6954b15d721fb039a5dd1fb40399",
|
| 82 |
+
"reproduce/source/btx53/retry_prefetch.sh": "9c27f3c724b803b284bc79dcd8a5d2fba6d145c3e4f7be43aebbf9f09b981da9",
|
| 83 |
+
"reproduce/source/btx53/pipeline.py": "091bcbe282ac930930704865bb02b2738103c5f42abbb1d2d1e8aed6d2849f24",
|
| 84 |
+
"reproduce/source/btx53/build_baseline.py": "0fd3b364f5ed470354723342ae766e8695b1804d7a50495adfe68941cfaf9344",
|
| 85 |
+
"reproduce/source/btx53/build_v2_seed_checkpoint.sh": "d0b4445045ee7900f4f423a54e2c1adcf4634756ba056f73b227b7bd228f7a56",
|
| 86 |
+
"reproduce/source/btx53/arvq_gate_loop.sh": "c7c8c9527efbafd16d9e11a604e51e505fa8fbbff6079edb9bfbdcc2476837fd",
|
| 87 |
+
"reproduce/source/btx53/arvq_gate.sh": "530ce8c4f6755a164862cfd1fce5c839e57d523288401bed0ab2ee25b26e1f65",
|
| 88 |
+
"reproduce/source/btx53/arvq_watchdog.sh": "415cb80615b9d6d0e2918cbfad978c4f312d49b271f37db633eb7e51b978ad60",
|
| 89 |
+
"reproduce/source/tools/sanity/sc1_schema.py": "78bff17dea376fb982ce8e58e77ce34772b3891b35e0acde5c70ba9a9e3f5cc1",
|
| 90 |
+
"reproduce/source/tools/sanity/sc2_dequant_stats.py": "f8440c5a4982482a5d5efb10d0b9e390c30bcace65d809ae8da0e4be8c7abc78",
|
| 91 |
+
"reproduce/source/tools/sanity/sc6_ppl.py": "7cafb6d5c6d9222e7582241c1533bfc6389e55a35079e1c6fee1235d3c7882a1",
|
| 92 |
+
"reproduce/source/tools/sanity/bench_small.py": "75bb8bec5854d8415aaa7b806ada089af29be7cc36283c6fbbe14a7be65f7647",
|
| 93 |
+
"reproduce/source/tools/evalsets/build_heldout_53.py": "e913a8b1223044c86665769497bf551be8da10a88e1ae94f64f97c4970619883",
|
| 94 |
+
"reproduce/source/tools/evalsets/build_neutral_sets.py": "21ec003f850b4397345cbd30ef320feb8c7db7f85ecc73c9fc17f22c824c087e",
|
| 95 |
+
"reproduce/source/tools/vision/fetch_tensor.py": "40b10e1b5e37096b23663884f0f3d7c9dc7326c2b6a09918e9f53beaf879934f",
|
| 96 |
+
"reproduce/source/tools/vision/compare_embeds.py": "ba890a855820b3f9cbc1c4e042da183ac1c521b3705921f917439d06b7a708dd",
|
| 97 |
+
"reproduce/source/tools/vision/compare_tower.py": "2cf6db6c278107837ec16239a4c2dcbe3a93a7cab99d9c592bf599e9d4d685d5",
|
| 98 |
+
"reproduce/source/tools/capture53/launch_stats.sh": "d4f4453b0ffb48cb7fcbd441bd1cac1e6af132c982cdd4eeb276bd91c10b0d01",
|
| 99 |
+
"reproduce/source/tools/capture53/solve_assignment53.py": "5cbdb5a27c8b34b194b04083a120d6e0f9bde64171a19d2629ace3349c4e77e8",
|
| 100 |
+
"reproduce/source/tools/capture53/mtp_requant53.py": "ce982688427c49e3f04ddd1e18e89b6dfa5ee9dcbf454fb9c2e184ded72e6ee6",
|
| 101 |
+
"reproduce/source/tools/capture53/vision_probe53.py": "b6a772b02b808bf2f72920b5f2678cf09210e91d6210b9e7c501c2691c88c7cc",
|
| 102 |
+
"reproduce/source/tools/capture53/score_reap53.py": "c711b388826f3fac6a0684d16772e98c54bbc29b35876513ff98feef79a45e58",
|
| 103 |
+
"reproduce/source/tools/capture53/pv_tune53.py": "9f2c246411878d6af6902c4b7074fc997a70e3f9064e6a235796da2630d03fc7",
|
| 104 |
+
"reproduce/source/tools/capture53/eval_hybrid53.py": "dd291b3ff47935773ae02b6c22d574065db5af54bc779d3605e9757eb8f98047",
|
| 105 |
+
"reproduce/source/tools/capture53/solve_blend53.py": "bda421790c9e42292cc4d6c1a023d0d8b60a2777262f9165e54a378eeed86420",
|
| 106 |
+
"reproduce/source/tools/capture53/build_checkpoint53.py": "f4ef159add9890d34580e62f153b9e8db72e9f8411396477da249c45e4510d40",
|
| 107 |
+
"reproduce/source/tools/capture53/solve_reap53.py": "29d35e7abbe23985e1adbeeb439015884f8523ba36bd602cecb817a5c9fd0308",
|
| 108 |
+
"reproduce/source/tools/capture53/build_calib_mm.py": "cbc0364753cef10607327fa3126caa825207882a05d0315d23f52dcc9a03862c",
|
| 109 |
+
"reproduce/source/tools/capture53/graft_vision53.py": "04fbca9c6bec0c935dad247ac2f06ee3df583ba5053c46235b7ad7b94504b85d",
|
| 110 |
+
"reproduce/source/tools/capture53/aqlm_converge53.py": "b3ccd5521e0ad3a82d38e56a6ce070625c2a99528d38cd932ac57aadc64a1af0",
|
| 111 |
+
"reproduce/source/tools/capture53/solve_r3.py": "5979122eab06cbcc1568f0942896f8d056773378064bf24551dc3c1ccb9b6c09",
|
| 112 |
+
"reproduce/source/tools/capture53/capture_mm53.py": "9201ddd9cbfa89092016f4f51b7859a6cb20d25ab13c575f0bb270d7b786bc4c",
|
| 113 |
+
"reproduce/source/tools/capture53/stream_capture53.py": "9ca7120673abb2ac0b4fbeaf25811dc36ca25da9d7791cceb65ae53d935d8a08",
|
| 114 |
+
"reproduce/source/tools/capture53/prepare_unc_trace_model.py": "bb08ea3058856ec31b595f2cf4aa24eb541cbf1ee2fb7f77f2008cc1e2d954af",
|
| 115 |
+
"reproduce/source/tools/capture53/eval3_kld.py": "0c270ef4822f47cfbc92721456297e74a4ef115608a6d25d7eaba1d91e55cfe7",
|
| 116 |
+
"reproduce/source/btx53/arvq88/__init__.py": "0bf1139931af3c4357fa592ee2f2abc669bc2ce17826b4ab9cb4f593f773d018",
|
| 117 |
+
"reproduce/source/btx53/arvq88/inputs.py": "4a8ebeef4ad6a07e47b7d8e28474dd2ea48a192f755cd14abada993c052305a0",
|
| 118 |
+
"reproduce/source/btx53/arvq88/fit.py": "7288bbe580b307dbc399b161ee5899ea0cab9d57e57d40224ebce74152575497",
|
| 119 |
+
"reproduce/source/btx53/arvq88/checkpoint.py": "4d1194767fb46102085732eb114e524bf052653ad82f80367a655ba1155b7b90",
|
| 120 |
+
"reproduce/source/btx53/arvq88/cli.py": "452f32eb2eaa6fde3a8270c79599e185e36ab11e6993901cc7f94f56d5903840",
|
| 121 |
+
"reproduce/source/btx53/arvq88/gate_worker.py": "05199a15d687bf943ba0e8b9f8d16958a195a5874aff9d1b42ae00bbebda697f",
|
| 122 |
+
"reproduce/source/btx53/arvq88/reap_worker.py": "ac39da41fb132a026d666c892e300948b7606492610816068b0223848b14049d",
|
| 123 |
+
"reproduce/source/btx53/arvq88/reap.py": "9e69dbfcf48987b53c8694aa8c32f29e04bea270481a7c98a98b0a2e0a277704",
|
| 124 |
+
"reproduce/source/btx53/arvq88/pv_campaign.py": "95a483dedc27f61fb767cb8020471a04fd5cfa05078f816efbfc6fdb9cbd15cb",
|
| 125 |
+
"reproduce/source/btx53/arvq88/validation.py": "829c6a4ece05c189653a18dfe9429f3a48cb29c54e373fcd398706e3439a47ce",
|
| 126 |
+
"reproduce/source/btx53/arvq88/promote_initial.py": "27b7600139e9f5bf30ce0b3db0cb587284c16a79e77baa1cd3049a03b0e04dc0",
|
| 127 |
+
"reproduce/source/btx53/arvq88/local_source.py": "0c57365308afcf75bf4b485a1e86fd1ca169ce7c4ddb999f80f2c3f938bb1bd9",
|
| 128 |
+
"reproduce/source/btx53/arvq88/backbone.py": "88e88930ea9ecec52259f5ab648915dc9b3a4d8144b302dadfbeb736c9dcc45b",
|
| 129 |
+
"reproduce/source/btx53/arvq88/activation.py": "cf13e2170085459b97907a270abb32d694af3ce69f723b332119a62be00e6641",
|
| 130 |
+
"reproduce/source/btx53/arvq88/alternating.py": "19e38a8230dba40bc022180ed4234b5a60f53765d36a417677e35711e9f0c80c",
|
| 131 |
+
"reproduce/source/btx53/arvq88/arvq_reap.py": "c39faf937462d85b2ede04f19f14c2cc8ef875c6788a1dab2445a03834882cfd",
|
| 132 |
+
"reproduce/source/btx53/arvq88/jobs.py": "4b9021e5763ba15b45cf4c0cf76d6f62f705ac175b076cdae380032a2d9ac59d",
|
| 133 |
+
"reproduce/source/btx53/arvq88/experiment.py": "65272954df9ef44f88467c05f5ce5f12c8045b38765f67b773353890d7877de4",
|
| 134 |
+
"reproduce/source/btx53/arvq88/incremental_publish.py": "70526567046728bd2905d059bf475f6a1d4d74d4242f7288cd64091e9137f910",
|
| 135 |
+
"reproduce/source/btx53/arvq88/gradient_indices.py": "09bd66d77575deec66d1732677355c406cb9b228abfc5002e79cbae5adb72e83",
|
| 136 |
+
"reproduce/source/btx53/arvq88/sequential_pv.py": "807609be6d53c94ed578ae91f81e74d3be60451d436d6d5f1b5053004ea821f0",
|
| 137 |
+
"reproduce/source/btx53/arvq88/sequential_pilot.py": "030f3bd1d83ea2f9bcd97ee8514d01dc69d269e7d7b7f87159da667b46b76710",
|
| 138 |
+
"reproduce/source/btx53/arvq88/encoder.py": "6c20c56fc7a308724ab4faf4174c4e8e32cba5a43c3fccd3f3e6d4da1f0b583e",
|
| 139 |
+
"reproduce/source/btx53/arvq88/sequential_capture.py": "bcbb6cb450243411e064624c70b1ec460e052680eb7f9958f418b25577d9e671",
|
| 140 |
+
"reproduce/source/btx53/arvq88/final_publish.py": "e7dcb2859de2e08a5c2bea7b21b9aae61e96bb2361b28ba4270cc145fb0da188",
|
| 141 |
+
"reproduce/source/btx53/arvq88/pv.py": "82f3f559e44b5db2480e041fa45d983bf1e6afc2d74a68c0511af0d174f9ed68",
|
| 142 |
+
"reproduce/source/btx53/arvq88/pack.py": "b0606bf8ae91bab655ed5b26bb9b7437f0e28718301fdb1b86f475b1b31924e1",
|
| 143 |
+
"reproduce/source/btx53/arvqprep/__init__.py": "921c16de3c268ed12dcac8329852f47fb6875d926a2ea665a03e69c81059b23d",
|
| 144 |
+
"reproduce/source/btx53/arvqprep/pack.py": "ce459847f07a558377872e0f1f83620d9f26d9eec7ec94d08247052876bab656",
|
| 145 |
+
"reproduce/source/btx53/arvqprep/encoder.py": "ace5d22c886994a0327e8320c6674f71c1caab779d215e3d05295d1f490bfec1",
|
| 146 |
+
"reproduce/source/btx53/arvqprep/smoke.py": "fb445c69e081b1c23acaf03feef234e215c61cceaf1b6d0b4c47aa63f9032a34",
|
| 147 |
+
"reproduce/source/btx53/arvqprep/container.py": "3f5ab2a065a1a44826cd0183ac39ba1d016ac87be655ed3b9502b1c3137ae7ad",
|
| 148 |
+
"reproduce/source/btx53/arvqprep/workflow.py": "778d6782e95cf81ea7d7c44f8f508aab65706e64a18cc3d15adc57eb61e5e095",
|
| 149 |
+
"reproduce/source/btx53/arvqprep/publish.py": "3249526e87beba17e69612d034cc400c57cf88a09eea3b3622eb8e4f43b868ce",
|
| 150 |
+
"reproduce/source/btx53/arvqprep/pv_workflow.py": "5a57e8ae4e76f938f6cda3b245321c4812e5e7c16e52b3976adcfa96b3cb9187",
|
| 151 |
+
"reproduce/source/btx53/btxprep/core.py": "b4fad5c192efa2e9e0088427ad8e90dacac59b9220479809b81ca826556c8348",
|
| 152 |
+
"reproduce/source/btx53/btxprep/transforms.py": "2f0bda3c93e39836ae6bca82deceb876e87b54797dd696c194a99041f6bd6d1c",
|
| 153 |
+
"reproduce/source/btx53/btxprep/container.py": "f260e70e4e6fda4c20b8cb85469620ebe67d343983776c370b0bfe7bdb2f967e",
|
| 154 |
+
"reproduce/source/btx53/btxprep/__init__.py": "f8b773e399d3d5b859d05a85347d8f6a91eed2808447c934d917c1de39d43a40",
|
| 155 |
+
"reproduce/source/btx53/btxprep/encoder.py": "ac8745d1ba1590817a16bc4ed48acc9af980a7103ad22e32aeb179b48b5813f2",
|
| 156 |
+
"reproduce/source/btx53/tools/remote_st.py": "4a27ac143fff1cc8cdf6fb349ac58bba1d22933c77c562779bb111b727f0031e",
|
| 157 |
+
"reproduce/source/btx53/tools/ingest.py": "e05630716d805903f1bfe658512da2d3f63abbd280a3af87ca9baadb5becd52c",
|
| 158 |
+
"reproduce/source/btx53/tools/upload_btx.py": "72bc32e8f2c2755e9ec5c713d6226bd7021265d1cddaf1e927b3682b508e5cba",
|
| 159 |
+
"reproduce/source/btx53/tools/filter_shards.py": "749d7b6f8a5b6ba24442508ce66e4108086b837e89c20234c3bf4939b27ada7d",
|
| 160 |
+
"reproduce/source/btx53/tools/fetch_hyb_kinds.py": "aaa37fd535d26b3772c4e467c55502c810ba39b73ca998078c487aeb1a7be846",
|
| 161 |
+
"reproduce/source/btx53/tools/gates_btx.py": "bcd2d8b5e4ef9e3779bfbec625d7af5b0727d954502db63a46147de882de170c",
|
| 162 |
+
"reproduce/source/btx53/tools/upload_btx_full.py": "ee8b46dd631629e15d1f1bb86c0b06d4ad22e573c674e129e58c04b5e7e1dc8d",
|
| 163 |
+
"reproduce/source/btx53/tools/launch_btx_workers.sh": "4e8484172fe4dea0b48390c72959d81bc4d25ac05265187b9bb90df9c5bc592a",
|
| 164 |
+
"reproduce/source/btx53/tools/upload_hybrid.py": "5b3e4be4ccf076d569ea5a48f0e79d6168236d4d6eccf84053a4bf183a4cc2f6",
|
| 165 |
+
"reproduce/source/btx53/tools/gates_arvq.py": "ce6b09045cc8f1c9cea7647083e133a52885b284c09879ccf4b9090d58d27453",
|
| 166 |
+
"reproduce/source/btx53/tools/launch_arvq_workers.sh": "78b013b3c8f5e2750fc6a3b7e88d58ddc5a61aafdea0d51eb666d0d3dba2000c",
|
| 167 |
+
"reproduce/source/btx53/tools/arvq_driver.sh": "13d72054ed28f5e112b98ddd7db0cdb7de00e0a4ead5cbd54b03bc0386860cc8",
|
| 168 |
+
"reproduce/source/btx53/tools/pv_arvq.py": "82dd57c204c057eaacd17aefcd37faa1d6b90ab7f6931aa2872b3acf65822858",
|
| 169 |
+
"reproduce/source/btx53/tools/prefetch_cold_fp8.py": "5825b1069ee10681516d69a703699abae4ac46011f6e71d6537acfece4345921",
|
| 170 |
+
"reproduce/source/btx53/arvq88/tests/test_workflow.py": "d9bc78d2a898a9c8dab96ea2a79dcf183cf182fb535992ff106ff452c8961979",
|
| 171 |
+
"reproduce/source/btx53/arvq88/tests/test_final_publish.py": "6c2b104ec40284415763fea5005510df0b63055bf99f081345e8248de0f7eb24",
|
| 172 |
+
"reproduce/source/btx53/arvq88/tests/test_validation.py": "f563a433c66412085a191deb8c693865632aedf1de19948567c710d11c3606b7",
|
| 173 |
+
"reproduce/source/btx53/arvq88/tests/test_local_source.py": "544513082fe0c666ade7ed00a837aeeeb915c52fb1325362a488e1c8c3a050ed",
|
| 174 |
+
"reproduce/source/btx53/arvq88/tests/test_donor.py": "1948eca1225198bed8964ff9d56e5f2b98c86847537a108670cf59260a20d76e",
|
| 175 |
+
"reproduce/source/btx53/arvq88/tests/test_backbone.py": "58f3875440e45b9bfd5be915452a7f1a18cd36688909f737d5830e1a5062fa85",
|
| 176 |
+
"reproduce/source/btx53/arvq88/tests/test_activation.py": "eff65a3e37b8b4b77b8fbb8bd979f6ca5012680a9643716f257f4e8985662944",
|
| 177 |
+
"reproduce/source/btx53/arvq88/tests/test_expert_books.py": "5942045e2e09a0a50ea1bae852a2f9a411b8dfd949e78e8e67151b8e119ec33f",
|
| 178 |
+
"reproduce/source/btx53/arvq88/tests/test_incremental_publish.py": "d0296210c666b4c556dcc6139500a1a0a36646186ee7a019d0ca2d41b9a80368",
|
| 179 |
+
"reproduce/source/btx53/arvq88/tests/test_gradient_indices.py": "6f1527d21b73c882500bb8ff229c54a023abe89d43ad23c6c3176122c188347e",
|
| 180 |
+
"reproduce/source/btx53/arvq88/tests/test_sequential_pv.py": "7d3c65cf41648b1b564a1939bcbee3d495ec4bd9d0af7028a5d788b052fadd18",
|
| 181 |
+
"reproduce/source/btx53/arvq88/docs/per_expert_handoff.md": "554425132e253e53a7890ba6f270c2cbfcec4e7115921db2983036c5b6d20664",
|
| 182 |
+
"reproduce/source/btx53/arvq88/docs/expert_experiments.md": "6d78a3ee627255f13e2be7995bf21df4c36f9703b192bc5020a5c1c7a7a7336e",
|
| 183 |
+
"reproduce/source/btx53/arvq88/docs/gradient_pv_recipe.md": "77686dc9d94c3a4c5559024598dbbcafe7ca55f17858e7305fb2ad069a4f274f",
|
| 184 |
+
"reproduce/source/btx53/arvq88/docs/sequential_pilot.md": "079342a060d5fc384e00b724ca3e1189c169ef4b7c590df12b1739c1dceba6d5",
|
| 185 |
+
"reproduce/source/btx53/arvq88/perf/graph_expert.py": "9d663b49ef7bb23db843f78068f146e2c178f23137360acfaa1aa544c6942d73",
|
| 186 |
+
"reproduce/source/btx53/arvq88/perf/bench_graph_expert.py": "9b6b5759376fb771caed63e908e256d9b187fec93e0d6bcb6a16a16c97cdb42e",
|
| 187 |
+
"reproduce/source/btx53/arvq88/perf/sparse_delta.py": "52565e3bdaa477d1676fffaf10bdffd10e1ad2d1b4275ded2cdd1d210bd0f81e",
|
| 188 |
+
"reproduce/source/btx53/arvq88/perf/test_perf.py": "3f9221722a3d0aa3360ba02ca5823612162aefb27d084f0d4a3bacf4bac45ef9",
|
| 189 |
+
"reproduce/source/btx53/arvq88/perf/README.md": "ebeeb89781cd162e1711837bab91a9077df496a876bc99ab6d6a0fe53779c21b",
|
| 190 |
+
"reproduce/source/btx53/arvq88/perf/rerun_pilot.py": "691665c2f6d482d8518f7153124814f8333a810d44c8c03269756ef2bbf8199f",
|
| 191 |
+
"reproduce/source/btx53/arvq88/perf/full_reference.py": "cabbd41c1be1f2aa77e3b298b5958a79aed88cc73263ec358ef2c64a9d4846b5",
|
| 192 |
+
"reproduce/source/btx53/arvq88/perf/sequential_pv_repair.py": "74ece4df20d6573ae8bf752736adb12217968098c49e6f5d3bfbb7ebc1d3b3f3",
|
| 193 |
+
"reproduce/source/btx53/arvq88/perf/test_repair.py": "78112f17972eadf45c7f48f578b8c75170c70e7fe0aa75f63bf861c6161f952c",
|
| 194 |
+
"reproduce/source/btx53/arvq88/perf/sequential_pv_cached.py": "9c629c8083fe5cc2e760e1d6afc3dc1f60c1320ac30b4dcd95db35dc93d8b163",
|
| 195 |
+
"reproduce/source/btx53/arvq88/perf/batch_benchmark.py": "c78c5608221c051b93817b373c94c7ca844416851f72d3905f4e231272776829",
|
| 196 |
+
"reproduce/source/btx53/arvq88/perf/capture_full_layer3.py": "3359edb2956d41de364c32e209275c1d87d047bbfe48ee84f29d76d9ee535a1b",
|
| 197 |
+
"reproduce/source/btx53/arvq88/perf/test_full_corpus.py": "98491ab3943d70b4ed2536ecbd72bd163561002d6fba61124f829adcbcf5bf1c",
|
| 198 |
+
"reproduce/source/btx53/arvq88/perf/async_shards.py": "b536e99bf5ea3badd4291f3fbc9b22a0594d334e9b551f3146a71288991d7529",
|
| 199 |
+
"reproduce/source/btx53/arvq88/perf/test_async_shards.py": "ecee760fa718fa0542a8650bf8a5b7d1b7ccade05b6682f0660401cac4208f34",
|
| 200 |
+
"reproduce/source/btx53/arvq88/perf/FULL_CORPUS_PIPELINE.md": "2c842e5833b64c1de55b9b3d33756ddb5cc6dfaef808ccd652cbc37f07280040",
|
| 201 |
+
"reproduce/source/btx53/arvq88/perf/distributed_batch.py": "ff4fcfdc9c74bc46f06bb51a126d2764aa4d983da5b9eebc06885081330fa39e",
|
| 202 |
+
"reproduce/source/btx53/arvq88/perf/test_distributed_batch.py": "36bc25e59d2ec9045751409ccb74a2f0464a2de502fa5adb4affff5fec3a56cc",
|
| 203 |
+
"reproduce/source/btx53/arvq88/perf/check_recapture.py": "9b49570cf9dca3d42ee9262822b7992b2d1ddf6025838f4fbe3d9ff5ab090593",
|
| 204 |
+
"reproduce/source/btx53/arvq88/perf/publish_full_corpus.py": "a361a0fe959a214dc8f7550c90248697b46f1d1fa3aeba1f149503723617b0b7",
|
| 205 |
+
"reproduce/source/btx53/arvq88/perf/capture_transfer.py": "aa6f8702cbaaf3ebcc9dad65053c422f3ec7f1e1d9c4ac3fc60834769c38b386",
|
| 206 |
+
"reproduce/source/btx53/arvq88/perf/test_capture_transfer.py": "c33c0ffda4ca906930760cd100d7e0d283d32eeba1fbebd47af34bffa680abfa",
|
| 207 |
+
"reproduce/source/btx53/arvq88/perf/monitor_output_error.py": "cce2dd878e8e164e8b3a296cc70699cc8eb6e3225bac3f9b8c5171e063af40cb",
|
| 208 |
+
"reproduce/source/btx53/arvq88/perf/capture_matched_validation.py": "5d0b9d1892923acbd066f0f88030221db39aecee6ca62e8a1f094a29833a473d",
|
| 209 |
+
"reproduce/source/btx53/arvq88/perf/adaptive_schedule.py": "504f54400931422d706544029ca1c3673224f33f77d7801c89dda1d4ba53dd30",
|
| 210 |
+
"reproduce/source/btx53/arvq88/perf/test_adaptive_schedule.py": "456bc945a66c6ff1893f86a972a63b2efb44ae3a386e09e02df23673780cc537",
|
| 211 |
+
"reproduce/source/btx53/arvq88/perf/campaign_watchdog.py": "548d78758937c79d17601772f2c5424a8bbf5d9952144b6fbbde396a8de4e499",
|
| 212 |
+
"reproduce/source/btx53/arvq88/perf/row_scale_pilot.py": "a84aec6f9aa2824aaa9f00e8d53c19b5f5deca76ddd4c0da69230bebe95c0181",
|
| 213 |
+
"reproduce/source/btx53/arvq88/perf/four_layer_row_pilot.py": "de817289405b9f2a4eb34e28b315e779dcadadfafc401c1e28d54c75b5a0acf7",
|
| 214 |
+
"reproduce/source/btx53/arvq88/perf/row_scale_combined_eval.py": "6e2084b496058983fc4b143142ec3ebcb9703bed5801bb7bb0a599017976e17f",
|
| 215 |
+
"reproduce/source/btx53/arvq88/perf/FP16_BLOCK_SCALE_HANDOFF.md": "9304e796d7fbff92157737b9aa6cf2a6e29e66d16656409af4ab06c5f18e8cfa",
|
| 216 |
+
"reproduce/source/btx53/arvq88/perf/sequential_pv_same_input.py": "c6fb2c7a2f565c45f326837da2546ebd3fd01d2496353539d3a3303c9c370f13",
|
| 217 |
+
"reproduce/source/btx53/arvq88/perf/full_corpus.py": "db9e8cb6b3d7ea22fc8490fc355fd9421dfc7429c600431cc37f52017f9811c2",
|
| 218 |
+
"reproduce/source/btx53/arvq88/perf/full_pipeline.py": "c35476f87e9c2231cc1acc0059e17f91d0c51f04a99d464423054d54b4d793a4",
|
| 219 |
+
"reproduce/source/btx53/arvq88/perf/fused_handoff.py": "6c62bf5606ed48fa8a1a9d27a5f694f353f44c7a02f651475522f2e431a5c37a",
|
| 220 |
+
"reproduce/source/btx53/arvq88/perf/fused_evaluation_capture.py": "10ef6f00129bcea2915b41febdb08ba2f942c4f132e61ca4a774d8f913bf8daa",
|
| 221 |
+
"reproduce/source/btx53/arvq88/perf/capture_full_sequential.py": "ee221434a21fb4818f9ab20925ea66c31a396c8b3557dda1a45174fabd284927",
|
| 222 |
+
"reproduce/source/btx53/arvq88/perf/publish_v2_streaming.py": "07ea5e0fb78b462590f422895e05c823208da79d929980e3a6b0dde071f69395",
|
| 223 |
+
"reproduce/source/btx53/arvq88/perf/sequential_pv_fast.py": "a9b04ba27ebfd00f10bfbf96791c816858cef27d560086610e3bf6906fde9356",
|
| 224 |
+
"reproduce/source/btx53/arvq88/perf/propagate_full.py": "efb52d2784d1fca48beb3404c462255756c3fe499f68a418baa8a790d2018bf9",
|
| 225 |
+
"reproduce/source/btx53/arvq88/perf/baseline_propagation.py": "a0169695c1993f9c57da6c078e6e22aaefd404e1108a39a1989ceb1142d1cf35",
|
| 226 |
+
"reproduce/source/btx53/arvq88/perf/sequential_pv_full_corpus.py": "297df5e8e276e2caaae2af049b4dc0bf2e6185e0a9844bfdf705272aa4f96c3c",
|
| 227 |
+
"reproduce/source/btx53/arvq88/perf/publish_same_input.py": "bc75d242557ad1cac3b2cdc316b5760a501da5b221f7ed1f439a39d09b3de978",
|
| 228 |
+
"reproduce/source/btx53/arvq88/perf/publish_reference.py": "204d86bd3393bd91b5b43a5f7eb8f5c40d2cb3811962b772ab5cfffca1f63bab",
|
| 229 |
+
"reproduce/source/btx53/arvq88/perf/sequential_pv_mcbook.py": "f60b89d8325f4afd07349980c9190f78caf7ca0c05494805b04a1032de8fc749",
|
| 230 |
+
"reproduce/source/btx53/arvqprep/tests/test_workflow.py": "9250f47f793f13098ea48ce9ec51baec71ff65c410b3beac4609f32f8decefeb",
|
| 231 |
+
"reproduce/source/btx53/btxprep/tests/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 232 |
+
"reproduce/source/btx53/btxprep/tests/test_roundtrip.py": "fdce4358d869c1fb5fe01a7a9a878d6a29710af65779bbf0637be52eefd84735"
|
| 233 |
+
},
|
| 234 |
+
"weights_changed": false,
|
| 235 |
+
"validated": "2026-09-26",
|
| 236 |
+
"checks": [
|
| 237 |
+
"source SHA256",
|
| 238 |
+
"all Python parses",
|
| 239 |
+
"CPU FP8/FP16 decode fixtures",
|
| 240 |
+
"50x next-token boundary semantics",
|
| 241 |
+
"relocated workspace smoke"
|
| 242 |
+
],
|
| 243 |
+
"not_validated": [
|
| 244 |
+
"full historical retraining",
|
| 245 |
+
"native SM120 execution"
|
| 246 |
+
]
|
| 247 |
+
}
|
reproduce/prepare_workspace.py
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Materialize a fresh portable source copy; no fitting or publication is run."""
|
| 2 |
+
import argparse,json,shutil,sys,re,hashlib
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
p=argparse.ArgumentParser();p.add_argument('--dest',type=Path,required=True);a=p.parse_args();dest=a.dest.resolve();root=Path(__file__).resolve().parent
|
| 5 |
+
if dest.exists():raise SystemExit('Destination must not exist; refusing to modify an existing workspace')
|
| 6 |
+
shutil.copytree(root/'source',dest)
|
| 7 |
+
mapping={'/home/coder/git/glm52':str(dest),'/data':str(dest/'data'),'/tmp':str(dest/'tmp')}
|
| 8 |
+
pattern="(?:"+"|".join(re.escape(k) for k in mapping)+")(?=/|$|[\\s\"' ])"
|
| 9 |
+
changed={}
|
| 10 |
+
for f in dest.rglob('*'):
|
| 11 |
+
if not f.is_file() or f.suffix not in ('.py','.sh','.md'):continue
|
| 12 |
+
s=f.read_text();t=s
|
| 13 |
+
# Interpreter locations must resolve to the activated environment.
|
| 14 |
+
t=re.sub(r'/tmp/(?:venv[^/]*|[^/]+/venv)/bin/python(?:3)?','__GLM53_PYTHON__',t)
|
| 15 |
+
t=re.sub(pattern,lambda m:mapping[m.group()],t).replace('__GLM53_PYTHON__',sys.executable)
|
| 16 |
+
if t!=s:f.write_text(t);changed[str(f.relative_to(dest))]=hashlib.sha256(f.read_bytes()).hexdigest()
|
| 17 |
+
(dest/'data').mkdir(exist_ok=True);(dest/'tmp').mkdir(exist_ok=True);(dest/'.venv/bin').mkdir(parents=True,exist_ok=True);(dest/'.venv/bin/python').symlink_to(sys.executable)
|
| 18 |
+
cfg=json.loads((root/'recipe.json').read_text())
|
| 19 |
+
def relocate(v):
|
| 20 |
+
if isinstance(v,str):
|
| 21 |
+
return re.sub(pattern,lambda m:mapping[m.group()],v)
|
| 22 |
+
if isinstance(v,list):return [relocate(x) for x in v]
|
| 23 |
+
if isinstance(v,dict):return {k:relocate(x) for k,x in v.items()}
|
| 24 |
+
return v
|
| 25 |
+
(dest/'recipe.json').write_text(json.dumps(relocate(cfg),indent=2));(dest/'relocation.json').write_text(json.dumps({'mapping':mapping,'python':sys.executable,'changed_sha256':changed},indent=2))
|
| 26 |
+
print(f'Prepared {dest}. Populate donor/data/cache paths, inspect recipe.json, then follow RUNBOOK.md. No job launched.')
|
reproduce/recipe.json
ADDED
|
@@ -0,0 +1,116 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"baseline": "/tmp/glm53-vision-expert-full",
|
| 3 |
+
"corpus": "/tmp/glm53-vision-trace-refit",
|
| 4 |
+
"source_cache": "/tmp/glm53-fp8-cold",
|
| 5 |
+
"sequence_length": 1024,
|
| 6 |
+
"sequences": [
|
| 7 |
+
64,
|
| 8 |
+
16,
|
| 9 |
+
16
|
| 10 |
+
],
|
| 11 |
+
"world_size": 8,
|
| 12 |
+
"record_initial_audit": true,
|
| 13 |
+
"train_token_file": "/tmp/glm53-layer3-full-corpus/train.npy",
|
| 14 |
+
"batch_tokens": 262144,
|
| 15 |
+
"microbatch_tokens": 65536,
|
| 16 |
+
"codebook_lr": 0.048,
|
| 17 |
+
"scale_lr": 0.032,
|
| 18 |
+
"reference_only": false,
|
| 19 |
+
"target": "same_input",
|
| 20 |
+
"block_scale_dtype": "fp16",
|
| 21 |
+
"codebook_dtype": "fp4_grid",
|
| 22 |
+
"boundary_boost": 50.0,
|
| 23 |
+
"boundary_token_ids": [
|
| 24 |
+
154842,
|
| 25 |
+
154820
|
| 26 |
+
],
|
| 27 |
+
"reassign_every": 10,
|
| 28 |
+
"reassign_max_fraction": 0.005,
|
| 29 |
+
"reassign_trust_ratio": 0.02,
|
| 30 |
+
"reassign_target_ratio": 0.1,
|
| 31 |
+
"lr_decay_after": 25,
|
| 32 |
+
"lr_decay_factor": 0.25,
|
| 33 |
+
"adaptive_schedule": {
|
| 34 |
+
"relative_worsening": 0.001,
|
| 35 |
+
"checks": 3,
|
| 36 |
+
"patience_updates": 10
|
| 37 |
+
},
|
| 38 |
+
"fuse_next_capture": true,
|
| 39 |
+
"layers": [
|
| 40 |
+
3,
|
| 41 |
+
4,
|
| 42 |
+
5,
|
| 43 |
+
6,
|
| 44 |
+
7,
|
| 45 |
+
8,
|
| 46 |
+
9,
|
| 47 |
+
10,
|
| 48 |
+
11,
|
| 49 |
+
12,
|
| 50 |
+
13,
|
| 51 |
+
14,
|
| 52 |
+
15,
|
| 53 |
+
16,
|
| 54 |
+
17,
|
| 55 |
+
18,
|
| 56 |
+
19,
|
| 57 |
+
20,
|
| 58 |
+
21,
|
| 59 |
+
22,
|
| 60 |
+
23,
|
| 61 |
+
24,
|
| 62 |
+
25,
|
| 63 |
+
26,
|
| 64 |
+
27,
|
| 65 |
+
28,
|
| 66 |
+
29,
|
| 67 |
+
30,
|
| 68 |
+
31,
|
| 69 |
+
32,
|
| 70 |
+
33,
|
| 71 |
+
34,
|
| 72 |
+
35,
|
| 73 |
+
36,
|
| 74 |
+
37,
|
| 75 |
+
38,
|
| 76 |
+
39,
|
| 77 |
+
40,
|
| 78 |
+
41,
|
| 79 |
+
42,
|
| 80 |
+
43,
|
| 81 |
+
44,
|
| 82 |
+
45,
|
| 83 |
+
46,
|
| 84 |
+
47,
|
| 85 |
+
48,
|
| 86 |
+
49,
|
| 87 |
+
50,
|
| 88 |
+
51,
|
| 89 |
+
52,
|
| 90 |
+
53,
|
| 91 |
+
54,
|
| 92 |
+
55,
|
| 93 |
+
56,
|
| 94 |
+
57,
|
| 95 |
+
58,
|
| 96 |
+
59,
|
| 97 |
+
60,
|
| 98 |
+
61,
|
| 99 |
+
62,
|
| 100 |
+
63,
|
| 101 |
+
64,
|
| 102 |
+
65,
|
| 103 |
+
66,
|
| 104 |
+
67,
|
| 105 |
+
68,
|
| 106 |
+
69,
|
| 107 |
+
70,
|
| 108 |
+
71,
|
| 109 |
+
72,
|
| 110 |
+
73,
|
| 111 |
+
74,
|
| 112 |
+
75,
|
| 113 |
+
76,
|
| 114 |
+
77
|
| 115 |
+
]
|
| 116 |
+
}
|
reproduce/requirements.in
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
torch
|
| 2 |
+
numpy
|
| 3 |
+
transformers
|
| 4 |
+
datasets
|
| 5 |
+
safetensors
|
| 6 |
+
huggingface_hub
|
| 7 |
+
accelerate
|
| 8 |
+
Pillow
|
| 9 |
+
scipy
|
| 10 |
+
triton
|
reproduce/source/btx53/arvq88/__init__.py
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
"""GLM-5.3 256+256 FP4 additive-vector quantization workflow."""
|
reproduce/source/btx53/arvq88/activation.py
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Serving activation-plane reference, pinned to b1380cf74d170b69a8708f1b1287cc09a6eb7df4.
|
| 2 |
+
|
| 3 |
+
Matches hybrid.cu pack_planes and the default FP16 SwiGLU path. This emulates
|
| 4 |
+
quantization boundaries, not SM120 MMA instruction accumulation bit-for-bit.
|
| 5 |
+
"""
|
| 6 |
+
import torch
|
| 7 |
+
REFERENCE_REVISION='b1380cf74d170b69a8708f1b1287cc09a6eb7df4'
|
| 8 |
+
ARITHMETIC='fp4_planes4_fp16_boundaries_v1'
|
| 9 |
+
|
| 10 |
+
def fp16_ste(x):
|
| 11 |
+
rounded=x.to(torch.float16).float()
|
| 12 |
+
return x+(rounded-x).detach() if x.requires_grad else rounded
|
| 13 |
+
|
| 14 |
+
def planes(x):
|
| 15 |
+
"""Return FP4 codes, FP8 scale bytes, and decoded *unweighted* planes."""
|
| 16 |
+
if x.shape[-1]%16:raise ValueError('Activation input dimension must be divisible by 16')
|
| 17 |
+
v=x.to(torch.float16).float().reshape(*x.shape[:-1],-1,16)
|
| 18 |
+
if not torch.isfinite(v).all():raise ValueError('Nonfinite FP16 serving activation')
|
| 19 |
+
cuts=torch.tensor([.25,.75,1.25,1.75,2.5,3.5,5.],device=x.device)
|
| 20 |
+
table=torch.tensor([0.,.5,1.,1.5,2.,3.,4.,6.],device=x.device)
|
| 21 |
+
codes=[];scales=[];decoded=[]
|
| 22 |
+
for _ in range(4):
|
| 23 |
+
peak=v.abs().amax(-1,keepdim=True)
|
| 24 |
+
exponent=torch.ceil(torch.log2((peak/6).clamp_min(2**-20))).clamp(-6,8)
|
| 25 |
+
scale=torch.exp2(exponent)
|
| 26 |
+
# CUDA compares strictly > each midpoint: ties go toward zero, not ties-to-even.
|
| 27 |
+
q=(v.abs().div(scale).unsqueeze(-1)>cuts).sum(-1)
|
| 28 |
+
dec=table[q]*scale*torch.where(v<0,-1.,1.)
|
| 29 |
+
codes.append((q|((v<0).long()<<3)).to(torch.uint8).reshape(x.shape))
|
| 30 |
+
scales.append(((exponent.squeeze(-1).long()+7)<<3).to(torch.uint8))
|
| 31 |
+
decoded.append(dec.reshape(x.shape))
|
| 32 |
+
v=(v-dec)*16
|
| 33 |
+
return torch.stack(codes),torch.stack(scales),torch.stack(decoded)
|
| 34 |
+
|
| 35 |
+
def activation_ste(x):
|
| 36 |
+
with torch.no_grad():
|
| 37 |
+
_,_,p=planes(x)
|
| 38 |
+
result=p[0]+p[1]/16+p[2]/256+p[3]/4096
|
| 39 |
+
return x+(result-x).detach() if x.requires_grad else result
|
| 40 |
+
|
| 41 |
+
def swiglu_ste(gu):
|
| 42 |
+
# Serving: FP32 projection -> FP16; SiLU FP16 result; FP16 product.
|
| 43 |
+
gu=fp16_ste(gu);gate,up=gu.chunk(2,-1)
|
| 44 |
+
return fp16_ste(fp16_ste(torch.nn.functional.silu(gate))*up)
|
reproduce/source/btx53/arvq88/alternating.py
ADDED
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Discrete LDLQ proposals interleaved with continuous output matching.
|
| 2 |
+
|
| 3 |
+
This is a bounded alternating optimizer, not a reproduction of the PV-Tuning
|
| 4 |
+
paper. Proposals use training activation Hessians and are retained only if the
|
| 5 |
+
training routed-output objective decreases. Validation still selects checkpoints.
|
| 6 |
+
"""
|
| 7 |
+
import math
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
import torch
|
| 10 |
+
from .encoder import cholesky_inv_upper, hessian_fc1, sweep_expert
|
| 11 |
+
from .activation import ARITHMETIC, activation_ste, swiglu_ste
|
| 12 |
+
|
| 13 |
+
|
| 14 |
+
def accept_candidate(before, after):
|
| 15 |
+
return math.isfinite(after) and after < before - 1e-8
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
@torch.no_grad()
|
| 19 |
+
def reassign_indices(p13, p2, x, ids, train, cold, args, meta, evaluate_train):
|
| 20 |
+
from ingest import load_expert_bf16
|
| 21 |
+
before = evaluate_train()
|
| 22 |
+
# Codes must participate in both validation-best snapshots and rollback.
|
| 23 |
+
old = [(p.codes_a.clone(), p.codes_b.clone()) for p in (p13, p2)]
|
| 24 |
+
changed = total = 0
|
| 25 |
+
try:
|
| 26 |
+
for e, eid in enumerate(cold):
|
| 27 |
+
rows = train[(ids[train] == eid).any(1)]
|
| 28 |
+
if len(rows) < 32:
|
| 29 |
+
continue # No held-out-row fallback for rare experts.
|
| 30 |
+
z = x[rows[:args.reassign_hessian_tokens]]
|
| 31 |
+
wg, wu, wd = [w.to(x.device).float() for w in load_expert_bf16(
|
| 32 |
+
str(Path(args.src_cache)/f'layer_{args.layer}'/f'expert_{eid}.safetensors'), meta)]
|
| 33 |
+
z = activation_ste(z) if args.arithmetic == ARITHMETIC else z
|
| 34 |
+
# Down sees the current quantized gate/up, not teacher intermediates.
|
| 35 |
+
gu = z @ p13.weight(e).T
|
| 36 |
+
if args.arithmetic == ARITHMETIC:
|
| 37 |
+
down_x = activation_ste(swiglu_ste(gu))
|
| 38 |
+
else:
|
| 39 |
+
g, u = gu.chunk(2, -1)
|
| 40 |
+
down_x = torch.nn.functional.silu(g) * u
|
| 41 |
+
for p, W, inp in ((p13, torch.cat([wg, wu]), z), (p2, wd, down_x)):
|
| 42 |
+
H = hessian_fc1(inp)
|
| 43 |
+
hv = cholesky_inv_upper(H)
|
| 44 |
+
cb, scale, glob = p.quantized(e)
|
| 45 |
+
a, b, _ = sweep_expert(W, hv, cb[:256], cb[256:], scale,
|
| 46 |
+
float(glob), col_block=128, refine=2)
|
| 47 |
+
changed += int(((a != p.codes_a[e]) | (b != p.codes_b[e])).sum())
|
| 48 |
+
total += a.numel()
|
| 49 |
+
p.codes_a[e].copy_(a); p.codes_b[e].copy_(b)
|
| 50 |
+
del H, hv, a, b
|
| 51 |
+
del wg, wu, wd, gu, down_x, z
|
| 52 |
+
after = evaluate_train()
|
| 53 |
+
accepted = accept_candidate(before, after)
|
| 54 |
+
if not accepted:
|
| 55 |
+
for p, (a, b) in zip((p13, p2), old):
|
| 56 |
+
p.codes_a.copy_(a); p.codes_b.copy_(b)
|
| 57 |
+
return {'method': 'LDLQ training-Hessian proposal with routed-training-output acceptance',
|
| 58 |
+
'training_before': before, 'training_candidate': after,
|
| 59 |
+
'accepted': accepted, 'proposed_changed_groups': changed,
|
| 60 |
+
'groups_considered': total}
|
| 61 |
+
except BaseException:
|
| 62 |
+
for p, (a, b) in zip((p13, p2), old):
|
| 63 |
+
p.codes_a.copy_(a); p.codes_b.copy_(b)
|
| 64 |
+
raise
|
reproduce/source/btx53/arvq88/arvq_reap.py
ADDED
|
@@ -0,0 +1,174 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Score actual ARVQ vs NVFP4 expert outputs on training captures.
|
| 2 |
+
|
| 3 |
+
Uses fitted ARVQ candidates for ALL 256 experts, including currently hot experts.
|
| 4 |
+
The score is the positive reduction in routing-weighted squared output error
|
| 5 |
+
against original FP8 when an expert is kept NVFP4 instead of ARVQ. Cross-expert
|
| 6 |
+
error covariance and downstream model quality remain outside this proxy.
|
| 7 |
+
"""
|
| 8 |
+
import argparse
|
| 9 |
+
import json
|
| 10 |
+
import sys
|
| 11 |
+
from pathlib import Path
|
| 12 |
+
import numpy as np
|
| 13 |
+
import torch
|
| 14 |
+
|
| 15 |
+
ROOT = Path(__file__).resolve().parents[1]
|
| 16 |
+
sys.path[:0] = [str(ROOT), str(ROOT/'tools'), str(ROOT.parent/'tools')]
|
| 17 |
+
from arvq88.activation import activation_ste, swiglu_ste, ARITHMETIC
|
| 18 |
+
from arvq88.encoder import EncodedLayerProj, reconstruct
|
| 19 |
+
from arvq88.inputs import write
|
| 20 |
+
from arvq88.pack import sha
|
| 21 |
+
from arvq88.reap import blend_allocate
|
| 22 |
+
|
| 23 |
+
|
| 24 |
+
def hot_benefit(reference, cold_output, hot_output, gates):
|
| 25 |
+
if reference.shape != cold_output.shape or reference.shape != hot_output.shape:
|
| 26 |
+
raise ValueError('Expert outputs must have matching shapes')
|
| 27 |
+
weight = gates.double().square()
|
| 28 |
+
cold = (cold_output.double()-reference.double()).square().sum(-1)
|
| 29 |
+
hot = (hot_output.double()-reference.double()).square().sum(-1)
|
| 30 |
+
return float((weight*(cold-hot)).sum())
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def score_output(x, w13, w2, quantized):
|
| 34 |
+
if quantized:
|
| 35 |
+
return activation_ste(swiglu_ste(activation_ste(x) @ w13.T)) @ w2.T
|
| 36 |
+
g,u=(x @ w13.T).chunk(2,-1)
|
| 37 |
+
return (torch.nn.functional.silu(g)*u) @ w2.T
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
@torch.no_grad()
|
| 41 |
+
def score_layer(args):
|
| 42 |
+
from ingest import SourceMeta, load_expert_bf16
|
| 43 |
+
from aqlm_quantize import ShardReader, dequant_nvfp4, FP4_LUT
|
| 44 |
+
torch.set_num_threads(2)
|
| 45 |
+
dev=args.device; L=args.layer; source=Path(args.fit_dir)/f'layer_{L:05d}'
|
| 46 |
+
manifest=json.loads((source/'arvq-manifest.json').read_text())
|
| 47 |
+
if manifest['cold_expert_ids']!=list(range(256)):
|
| 48 |
+
raise ValueError('ARVQ allocation scoring requires candidates for all 256 experts')
|
| 49 |
+
store={}
|
| 50 |
+
for name in ('w13','w2'):
|
| 51 |
+
store.update(torch.load(source/f'{name}.pt',map_location='cpu',weights_only=True))
|
| 52 |
+
enc={k:EncodedLayerProj(**store[k]) for k in ('w13','w2')}
|
| 53 |
+
captures=torch.load(args.capture,map_location='cpu',weights_only=True)
|
| 54 |
+
if 'pv_split' in captures:
|
| 55 |
+
mask=captures['pv_split']==0
|
| 56 |
+
captures={k:v[mask] for k,v in captures.items() if k in ('x','topk_ids','topk_weights')}
|
| 57 |
+
x=captures['x'].to(dev).float();ids=captures['topk_ids'].to(dev);g=captures['topk_weights'].to(dev).float()
|
| 58 |
+
meta=SourceMeta(**json.loads(Path(args.source_meta).read_text())['meta'])
|
| 59 |
+
reader=ShardReader(args.donor);lut=torch.tensor(FP4_LUT,device=dev,dtype=torch.float32)
|
| 60 |
+
benefits=np.zeros(256); counts=np.zeros(256,dtype=np.int64)
|
| 61 |
+
for e in range(256):
|
| 62 |
+
rows=(ids==e).any(1).nonzero().flatten()
|
| 63 |
+
counts[e]=len(rows)
|
| 64 |
+
if not len(rows):continue
|
| 65 |
+
cold13=reconstruct(enc['w13'],e,dev);cold2=reconstruct(enc['w2'],e,dev)
|
| 66 |
+
wg,wu,wd=[w.to(dev).float() for w in load_expert_bf16(str(Path(args.src_cache)/f'layer_{L}'/f'expert_{e}.safetensors'),meta)]
|
| 67 |
+
pre=f'model.layers.{L}.mlp.experts.{e}'
|
| 68 |
+
hot=[dequant_nvfp4(reader,pre+'.'+proj,dev,lut).float() for proj in ('gate_proj','up_proj','down_proj')]
|
| 69 |
+
hot13=torch.cat(hot[:2]);hot2=hot[2];ref13=torch.cat([wg,wu])
|
| 70 |
+
for batch in rows.split(args.batch):
|
| 71 |
+
z=x[batch];weight=(g[batch]*(ids[batch]==e)).sum(1)
|
| 72 |
+
ref=score_output(z,ref13,wd,False)
|
| 73 |
+
cold=score_output(z,cold13,cold2,True)
|
| 74 |
+
hot_out=score_output(z,hot13,hot2,True)
|
| 75 |
+
benefits[e]+=hot_benefit(ref,cold,hot_out,weight)
|
| 76 |
+
del cold13,cold2,wg,wu,wd,hot,hot13,hot2,ref13
|
| 77 |
+
out=Path(args.out);out.parent.mkdir(parents=True,exist_ok=True)
|
| 78 |
+
np.savez(out,scores=benefits.clip(min=0),signed_hot_benefit=benefits,routed_rows=counts,layer=L)
|
| 79 |
+
write(out.with_suffix('.json'),{'layer':L,'metric':'sum routing_weight^2 * (ARVQ squared error - NVFP4 squared error), clipped at zero',
|
| 80 |
+
'reference':'original block-FP8 weights','arithmetic':ARITHMETIC,'capture_sha256':sha(args.capture),
|
| 81 |
+
'candidate_manifest_sha256':sha(source/'arvq-manifest.json'),'scope':'all 256 experts; training tokens only',
|
| 82 |
+
'limitation':'FP32 GEMMs with serving rounding emulation; no native MMA or cross-expert covariance'})
|
| 83 |
+
print('SCORED',L,args.capture,flush=True)
|
| 84 |
+
|
| 85 |
+
|
| 86 |
+
def allocate_scores(text_files,mm_files,hot_count=5750):
|
| 87 |
+
if len(text_files)!=75 or len(mm_files)!=75:raise ValueError('Need both modalities for all 75 layers')
|
| 88 |
+
def read(files):
|
| 89 |
+
scores=[]
|
| 90 |
+
for L,p in zip(range(3,78),files):
|
| 91 |
+
d=np.load(p)
|
| 92 |
+
if int(d['layer'])!=L:raise ValueError('Wrong score layer order')
|
| 93 |
+
scores.append(d['scores'])
|
| 94 |
+
return np.stack(scores)
|
| 95 |
+
text,mm=read(text_files),read(mm_files)
|
| 96 |
+
result=blend_allocate(text,mm,hot_count)
|
| 97 |
+
result['provenance']={'mode':'ARVQ output-benefit allocation','text_weight':.75,'mm_weight':.25,
|
| 98 |
+
'hot_count':hot_count,'floor':8,'cap':176,'text_partition':'train only',
|
| 99 |
+
'metric':'positive routing-weighted squared-error reduction from retaining NVFP4 over fitted ARVQ',
|
| 100 |
+
'scores_sha256':{str(p):sha(p) for p in [*text_files,*mm_files]},
|
| 101 |
+
'limitation':'per-expert additive error proxy; excludes cross-expert covariance and end-to-end model quality'}
|
| 102 |
+
return result
|
| 103 |
+
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def recompute(args, hub, work):
|
| 108 |
+
"""Fit all expert candidates, measure both modalities, then spend hot budget."""
|
| 109 |
+
from types import SimpleNamespace
|
| 110 |
+
from arvq88.inputs import ensure_source, LAYERS
|
| 111 |
+
from arvq88.jobs import run_jobs
|
| 112 |
+
if not getattr(args,'mm_acts',None):
|
| 113 |
+
raise ValueError('ARVQ REAP requires --mm-acts with matching donor multimodal captures')
|
| 114 |
+
mm=Path(args.mm_acts)
|
| 115 |
+
if not all((mm/f'acts_layer{L}.pt').is_file() for L in LAYERS):
|
| 116 |
+
raise ValueError('Incomplete multimodal captures; refusing text-only ARVQ REAP')
|
| 117 |
+
prefit=work/'arvq_reap';prefit.mkdir(exist_ok=True)
|
| 118 |
+
a=SimpleNamespace(**vars(args));a.workdir=str(prefit)
|
| 119 |
+
# These are candidates, not a valid final 5750-hot allocation.
|
| 120 |
+
all_experts={'layers':{str(L):{'cold_5750':list(range(256)),'hot':[]} for L in LAYERS}}
|
| 121 |
+
write(prefit/'status.json',{'stage':'preparing_original_fp8_sources','experts':19200})
|
| 122 |
+
ensure_source(a,hub,prefit,all_experts)
|
| 123 |
+
write(prefit/'assignment.json',all_experts);write(prefit/'run_config.json',vars(a))
|
| 124 |
+
write(prefit/'status.json',{'stage':'fitting_all_expert_candidates','layers':75,'experts':19200})
|
| 125 |
+
run_jobs([(f'fit-{L}',lambda gpu,L=L:[sys.executable,'-u',str(ROOT/'arvq88/fit.py'),str(prefit),str(L),f'cuda:{gpu}']) for L in LAYERS],args.gpus,prefit/'logs')
|
| 126 |
+
jobs=[]
|
| 127 |
+
for modality,folder in [('text',Path(args.calib)),('mm',mm)]:
|
| 128 |
+
for L in LAYERS:
|
| 129 |
+
output=prefit/'scores'/modality/f'layer_{L}.npz'
|
| 130 |
+
jobs.append((f'score-{modality}-{L}',lambda gpu,L=L,folder=folder,output=output:[
|
| 131 |
+
sys.executable,'-u',str(ROOT/'arvq88/arvq_reap.py'),
|
| 132 |
+
'--fit-dir',str(prefit/'initial'),'--capture',str(folder/f'acts_layer{L}.pt'),
|
| 133 |
+
'--src-cache',args.src_cache,'--source-meta',str(prefit/'source.json'),
|
| 134 |
+
'--donor',args.donor_cache,'--layer',str(L),'--device',f'cuda:{gpu}','--out',str(output)]))
|
| 135 |
+
write(prefit/'status.json',{'stage':'scoring_text_and_multimodal','jobs':150})
|
| 136 |
+
run_jobs(jobs,args.gpus,prefit/'logs')
|
| 137 |
+
result=allocate_scores([prefit/'scores/text'/f'layer_{L}.npz' for L in LAYERS],
|
| 138 |
+
[prefit/'scores/mm'/f'layer_{L}.npz' for L in LAYERS],args.hot_count)
|
| 139 |
+
result['provenance']['candidate_codebook_scope']=args.codebook_scope
|
| 140 |
+
write(prefit/'status.json',{'stage':'allocation_complete','hot':args.hot_count,'cold':19200-args.hot_count})
|
| 141 |
+
return result
|
| 142 |
+
|
| 143 |
+
|
| 144 |
+
def select_candidates(prefit,work,allocation):
|
| 145 |
+
"""Keep exactly the fitted candidates used to calculate hot-slot benefit."""
|
| 146 |
+
from arvq88.pack import export_layer
|
| 147 |
+
prefit,work=Path(prefit),Path(work)
|
| 148 |
+
for L in range(3,78):
|
| 149 |
+
source=prefit/'initial'/f'layer_{L:05d}';dest=work/'initial'/f'layer_{L:05d}'
|
| 150 |
+
dest.mkdir(parents=True,exist_ok=True)
|
| 151 |
+
manifest=json.loads((source/'arvq-manifest.json').read_text())
|
| 152 |
+
if manifest['cold_expert_ids']!=list(range(256)):raise ValueError('Candidates must cover all experts')
|
| 153 |
+
cold=allocation['layers'][str(L)]['cold_5750'];files={}
|
| 154 |
+
for proj in ['w13','w2']:
|
| 155 |
+
d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj]
|
| 156 |
+
for key in ['a','b','s']:d[key]=d[key][cold]
|
| 157 |
+
if d['c0'].ndim==3:
|
| 158 |
+
for key in ['c0','c1']:d[key]=d[key][cold]
|
| 159 |
+
path=dest/f'{proj}.pt';tmp=path.with_suffix('.tmp');torch.save({proj:d},tmp);tmp.replace(path)
|
| 160 |
+
files[proj]={'file':path.name,'sha256':sha(path)}
|
| 161 |
+
write(dest/'arvq-manifest.json',{**manifest,'cold_expert_ids':cold,'files':files,
|
| 162 |
+
'selection':'Exact subset of all-expert ARVQ candidates scored for allocation',
|
| 163 |
+
'candidate_manifest_sha256':sha(source/'arvq-manifest.json')})
|
| 164 |
+
export_layer(dest,work/'cold',L)
|
| 165 |
+
|
| 166 |
+
|
| 167 |
+
if __name__=='__main__':
|
| 168 |
+
ap=argparse.ArgumentParser(description=__doc__)
|
| 169 |
+
ap.add_argument('--fit-dir',required=True);ap.add_argument('--capture',required=True)
|
| 170 |
+
ap.add_argument('--src-cache',required=True);ap.add_argument('--source-meta',required=True)
|
| 171 |
+
ap.add_argument('--donor',required=True);ap.add_argument('--layer',type=int,required=True)
|
| 172 |
+
ap.add_argument('--out',required=True);ap.add_argument('--device',default='cuda:0')
|
| 173 |
+
ap.add_argument('--batch',type=int,default=256)
|
| 174 |
+
score_layer(ap.parse_args())
|
reproduce/source/btx53/arvq88/backbone.py
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Explicit tensor provenance: language/MTP from donor, vision from base."""
|
| 2 |
+
import copy,json,re,os
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
|
| 5 |
+
def vision_tensor(name):
|
| 6 |
+
return name.startswith(('vision_tower.','mm_projector.'))
|
| 7 |
+
def routed_tensor(name):
|
| 8 |
+
m=re.match(r'model.layers\.(\d+)\.mlp\.experts\.',name)
|
| 9 |
+
return bool(m and 3<=int(m[1])<=77)
|
| 10 |
+
|
| 11 |
+
def copy_backbone(base,donor,dest,save):
|
| 12 |
+
from safetensors import safe_open
|
| 13 |
+
from .pack import sha
|
| 14 |
+
provenance={}
|
| 15 |
+
for root,label,choose in [(base,'default-vision',vision_tensor),(donor,'nvfp4-donor',lambda n:not vision_tensor(n) and not routed_tensor(n))]:
|
| 16 |
+
index=json.loads((root/'model.safetensors.index.json').read_text())['weight_map']
|
| 17 |
+
for i,filename in enumerate(sorted(set(index.values()))):
|
| 18 |
+
names=sorted(n for n,f in index.items() if f==filename and choose(n))
|
| 19 |
+
if not names:continue
|
| 20 |
+
target=f'{label}-{i:05d}.safetensors'
|
| 21 |
+
with safe_open(str(root/filename),framework='pt') as stream:
|
| 22 |
+
save(target,{n:stream.get_tensor(n) for n in names})
|
| 23 |
+
# save_file preserves dtype and values; gate the resulting immutable shard identity.
|
| 24 |
+
provenance[target]={'source':str(root/filename),'tensor_names':names,'sha256':sha(dest/target),'role':label}
|
| 25 |
+
expected={n for n in json.loads((base/'model.safetensors.index.json').read_text())['weight_map'] if not routed_tensor(n) and not vision_tensor(n) and not n.endswith(('.weight_scale','.weight_scale_2','.input_scale','.input_scale_2'))}
|
| 26 |
+
actual={n for f in provenance.values() if f['role']=='nvfp4-donor' for n in f['tensor_names']}
|
| 27 |
+
if expected-actual:raise ValueError(f'Donor missing language/MTP tensors: {sorted(expected-actual)[:5]}')
|
| 28 |
+
return provenance
|
| 29 |
+
|
| 30 |
+
def hybrid_config(base,donor):
|
| 31 |
+
config=copy.deepcopy(base);text=copy.deepcopy(donor.get('text_config',donor))
|
| 32 |
+
qc=copy.deepcopy(base.get('quantization_config',{}))
|
| 33 |
+
donor_qc=text.pop('quantization_config',{})
|
| 34 |
+
if donor_qc.get('quant_method')=='modelopt':qc['nvfp4']=donor_qc
|
| 35 |
+
elif donor_qc.get('nvfp4'):qc['nvfp4']=donor_qc['nvfp4']
|
| 36 |
+
else:raise ValueError('Expected NVFP4 donor quantization config')
|
| 37 |
+
if 'text_config' in config:
|
| 38 |
+
config['text_config']=text
|
| 39 |
+
for name in ['pad_token_id','eos_token_id','tie_word_embeddings','dtype']:
|
| 40 |
+
if name in text:config[name]=text[name]
|
| 41 |
+
else:config=text
|
| 42 |
+
config['quantization_config']=copy.deepcopy(qc)
|
| 43 |
+
config.get('text_config',config)['quantization_config']=copy.deepcopy(qc)
|
| 44 |
+
return config
|
reproduce/source/btx53/arvq88/checkpoint.py
ADDED
|
@@ -0,0 +1,168 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Assemble/gate/publish an explicitly versioned 8+8 hybrid checkpoint."""
|
| 2 |
+
import json,re,shutil,os
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
from .inputs import LAYERS,complete_snapshot,write,token
|
| 5 |
+
from .pack import sha
|
| 6 |
+
|
| 7 |
+
def build(args,hub,work,allocation):
|
| 8 |
+
# Refresh the card renderer after long PV runs so reviewed provenance text is current.
|
| 9 |
+
import importlib,sys
|
| 10 |
+
if 'arvq88.pv_campaign' in sys.modules:importlib.reload(sys.modules['arvq88.pv_campaign'])
|
| 11 |
+
import torch
|
| 12 |
+
from safetensors import safe_open
|
| 13 |
+
from safetensors.torch import save_file
|
| 14 |
+
base=Path(args.local_hot)
|
| 15 |
+
if not complete_snapshot(base):base=hub.snapshot(args.base_hybrid,work/'base-hybrid')
|
| 16 |
+
dest=work/'checkpoint'
|
| 17 |
+
if dest.resolve()==base.resolve():raise ValueError('Output checkpoint cannot overwrite the local hot donor')
|
| 18 |
+
dest.mkdir(exist_ok=True)
|
| 19 |
+
manifests=[json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text()) for L in LAYERS]
|
| 20 |
+
formats={(m['format'],m['version']) for m in manifests}
|
| 21 |
+
if len(formats)!=1 or not formats<= {('rvq256_256x8',2),('rvq256_256x8_expert',3)}:
|
| 22 |
+
raise ValueError('Mixed or unsupported cold formats')
|
| 23 |
+
cold_format,cold_version=next(iter(formats))
|
| 24 |
+
index=json.loads((base/'model.safetensors.index.json').read_text())['weight_map'];wm={};total=0
|
| 25 |
+
def save(name,tensors):
|
| 26 |
+
nonlocal total
|
| 27 |
+
if not tensors:return
|
| 28 |
+
save_file(tensors,str(dest/name));wm.update({k:name for k in tensors});total+=sum(t.numel()*t.element_size() for t in tensors.values())
|
| 29 |
+
from .backbone import copy_backbone,hybrid_config
|
| 30 |
+
donor=Path(args.donor_cache)
|
| 31 |
+
if not complete_snapshot(donor):donor=hub.snapshot(args.nvfp4_donor,donor)
|
| 32 |
+
if dest.resolve()==donor.resolve():raise ValueError('Output cannot overwrite donor')
|
| 33 |
+
# A stale assembled checkpoint must not retain weights from the previous provenance rule.
|
| 34 |
+
marker=dest/'backbone_sources.json'
|
| 35 |
+
if any(dest.glob('*.safetensors')) and not marker.exists():raise ValueError('Use a fresh checkpoint directory for donor-backbone assembly')
|
| 36 |
+
provenance=copy_backbone(base,donor,dest,save)
|
| 37 |
+
write(marker,provenance)
|
| 38 |
+
dwm=json.loads((donor/'model.safetensors.index.json').read_text())['weight_map']
|
| 39 |
+
def donor_get(name):
|
| 40 |
+
nonlocal dwm,donor
|
| 41 |
+
if dwm is None:
|
| 42 |
+
if not complete_snapshot(donor):donor=hub.snapshot(args.nvfp4_donor,donor)
|
| 43 |
+
dwm=json.loads((donor/'model.safetensors.index.json').read_text())['weight_map']
|
| 44 |
+
with safe_open(str(donor/dwm[name]),framework='pt') as f:return f.get_tensor(name)
|
| 45 |
+
for L in LAYERS:
|
| 46 |
+
pre=f'model.layers.{L}.mlp.experts.';row=allocation['layers'][str(L)];cold=row['cold_5750'];hot=row['hot']
|
| 47 |
+
with safe_open(str(base/index[pre+'hyb_kind']),framework='pt') as f:old=f.get_tensor(pre+'hyb_kind')
|
| 48 |
+
tensors={};expected=torch.zeros(256,dtype=torch.int8);expected[cold]=2;tensors[pre+'hyb_kind']=expected
|
| 49 |
+
suffixes=['nvfp4_w13_packed','nvfp4_w13_bscale','nvfp4_w13_scale2','nvfp4_w2_packed','nvfp4_w2_bscale','nvfp4_w2_scale2']
|
| 50 |
+
if torch.equal(old,expected) and args.nvfp4_donor=='RadixArk/GLM-5.3-NVFP4':
|
| 51 |
+
for suffix in suffixes:
|
| 52 |
+
name=pre+suffix
|
| 53 |
+
with safe_open(str(base/index[name]),framework='pt') as f:tensors[name]=f.get_tensor(name)
|
| 54 |
+
else:
|
| 55 |
+
parts={k:[] for k in suffixes}
|
| 56 |
+
for e in hot:
|
| 57 |
+
ep=pre+str(e)+'.'
|
| 58 |
+
for suffix,out_suffix in [('weight','packed'),('weight_scale','bscale')]:
|
| 59 |
+
gu=torch.cat([donor_get(ep+p+'.'+suffix) for p in ['gate_proj','up_proj']]).view(torch.uint8)
|
| 60 |
+
dn=donor_get(ep+'down_proj.'+suffix).view(torch.uint8)
|
| 61 |
+
parts['nvfp4_w13_'+out_suffix].append(gu);parts['nvfp4_w2_'+out_suffix].append(dn)
|
| 62 |
+
parts['nvfp4_w13_scale2'].append(torch.stack([donor_get(ep+p+'.weight_scale_2').float().reshape(()) for p in ['gate_proj','up_proj']]))
|
| 63 |
+
parts['nvfp4_w2_scale2'].append(donor_get(ep+'down_proj.weight_scale_2').float().reshape(1))
|
| 64 |
+
tensors.update({pre+k:torch.stack(v) for k,v in parts.items()})
|
| 65 |
+
save(f'hot-layer-{L:03d}.safetensors',tensors)
|
| 66 |
+
manifest=json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text())
|
| 67 |
+
if manifest['cold_expert_ids']!=cold:raise ValueError('Cold fit/assignment mismatch')
|
| 68 |
+
for f in manifest['files'].values():
|
| 69 |
+
src=work/'cold'/f['file']
|
| 70 |
+
if sha(src)!=f['sha256']:raise ValueError('Cold shard identity mismatch')
|
| 71 |
+
target=dest/src.name
|
| 72 |
+
if not target.exists():os.link(src,target)
|
| 73 |
+
elif sha(target)!=f['sha256']:raise ValueError('Stale assembled cold shard')
|
| 74 |
+
with safe_open(str(target),framework='pt') as stream:
|
| 75 |
+
for name in stream.keys():
|
| 76 |
+
t=stream.get_tensor(name);wm[name]=src.name;total+=t.numel()*t.element_size()
|
| 77 |
+
for p in base.iterdir():
|
| 78 |
+
if p.is_file() and p.suffix in ['.json','.py','.jinja'] and p.name not in ['config.json','model.safetensors.index.json','backbone_sources.json'] and 'report' not in p.name and 'provenance' not in p.name:shutil.copy2(p,dest/p.name)
|
| 79 |
+
# Keep the vision wrapper/processor; language config and tokenizer follow donor.
|
| 80 |
+
for p in donor.iterdir():
|
| 81 |
+
if p.is_file() and (p.name.startswith(('tokenizer','special_tokens','added_tokens','chat_template','generation_config')) or p.name.endswith('.model')):shutil.copy2(p,dest/p.name)
|
| 82 |
+
config=hybrid_config(json.loads((base/'config.json').read_text()),json.loads((donor/'config.json').read_text()))
|
| 83 |
+
books={str(L):{'n_nvfp4':len(allocation['layers'][str(L)]['hot']),'n_base':0,'n_cold':len(allocation['layers'][str(L)]['cold_5750'])} for L in LAYERS}
|
| 84 |
+
for cfg in [config,config.get('text_config',config)]:
|
| 85 |
+
qc=cfg.setdefault('quantization_config',{});qc['quant_method']='nvfp4_arvq_hybrid'
|
| 86 |
+
qc['arvq']={'format':cold_format,'version':cold_version,'activation_planes':4,'weight_scale_group':128,
|
| 87 |
+
'codebook_scope':'expert' if cold_version==3 else 'layer','codebook_sizes':[256,256]}
|
| 88 |
+
qc['aqlm_layer_books']=books
|
| 89 |
+
write(dest/'config.json',config);write(dest/'model.safetensors.index.json',{'metadata':{'total_size':total},'weight_map':wm})
|
| 90 |
+
shutil.copy2(work/'assignment.json',dest/'cold_assignment.json');shutil.copy2(work/'source.json',dest/'source.json')
|
| 91 |
+
shutil.copytree(work/'cold',dest/'cold_manifests',dirs_exist_ok=True,ignore=shutil.ignore_patterns('*.safetensors'))
|
| 92 |
+
return dest
|
| 93 |
+
|
| 94 |
+
def gates(work,layers=None,device="cpu",report_path=None):
|
| 95 |
+
import torch
|
| 96 |
+
from safetensors import safe_open
|
| 97 |
+
from .pack import decode
|
| 98 |
+
torch.set_num_threads(2);layers=LAYERS if layers is None else layers;ck=work/'checkpoint';wm=json.loads((ck/'model.safetensors.index.json').read_text())['weight_map'];errors=[];hashes={}
|
| 99 |
+
if any(k.endswith(('.w13_codes','.w2c_codes','.w2m_codes')) for k in wm):errors.append('Old AQLM indices leaked')
|
| 100 |
+
for L in layers:
|
| 101 |
+
m=json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text());E=len(m['cold_expert_ids'])
|
| 102 |
+
pre=f'model.layers.{L}.mlp.experts.'
|
| 103 |
+
with safe_open(str(ck/wm[pre+'hyb_kind']),framework='pt') as stream:
|
| 104 |
+
kind=stream.get_tensor(pre+'hyb_kind')
|
| 105 |
+
if (kind==2).nonzero().flatten().tolist()!=m['cold_expert_ids']:errors.append('Hot/cold assignment mismatch')
|
| 106 |
+
for suffix,shape in [('nvfp4_w13_packed',(256-E,4096,3072)),('nvfp4_w2_packed',(256-E,6144,1024)),('nvfp4_w13_bscale',(256-E,4096,384)),('nvfp4_w2_bscale',(256-E,6144,128)),('nvfp4_w13_scale2',(256-E,2)),('nvfp4_w2_scale2',(256-E,1))]:
|
| 107 |
+
if tuple(stream.get_slice(pre+suffix).get_shape())!=shape:errors.append(f'Hot shape mismatch {L} {suffix}')
|
| 108 |
+
for proj in ['w13','w2']:
|
| 109 |
+
f=m['files'][proj];path=ck/f['file'];digest=sha(path);hashes[f['file']]=digest
|
| 110 |
+
if digest!=f['sha256']:errors.append(f'SHA mismatch {path.name}')
|
| 111 |
+
with safe_open(str(path),framework='pt') as stream:
|
| 112 |
+
t={k:stream.get_tensor(k) for k in stream.keys()}
|
| 113 |
+
N,K=(4096,6144) if proj=='w13' else (6144,2048);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
|
| 114 |
+
cbshape=(E,512) if m.get('version')==3 else (512,)
|
| 115 |
+
config=json.loads((ck/'config.json').read_text())
|
| 116 |
+
qconfig=config.get('text_config',config)['quantization_config']['arvq']
|
| 117 |
+
if qconfig['format']!=m['format'] or qconfig['version']!=m['version']:errors.append(f'Format metadata mismatch {L}')
|
| 118 |
+
for suffix,shape,dtype in [('packed',(E,N//16,K//64,64),torch.uint32),('codebooks',cbshape,torch.uint32),('scales',(E,N//16,K//128,16),torch.uint8),('global',(1,),torch.float32)]:
|
| 119 |
+
if t[pre+suffix].shape!=shape or t[pre+suffix].dtype!=dtype:errors.append(f'Shape/dtype {L} {suffix}')
|
| 120 |
+
if wm.get(pre+suffix)!=f['file']:errors.append('Index mismatch')
|
| 121 |
+
scales=t[pre+'scales'].view(torch.float8_e4m3fn).float()
|
| 122 |
+
if not torch.isfinite(scales).all() or (scales<0).any() or not torch.isfinite(t[pre+'global']).all() or not (t[pre+'global']>0).all():errors.append(f'Invalid scales {L} {proj}')
|
| 123 |
+
if not torch.isfinite(decode({k:v.to(device) for k,v in t.items()},L,proj,0)).all():errors.append('Nonfinite decoded sample')
|
| 124 |
+
del t,scales
|
| 125 |
+
backbone=json.loads((ck/'backbone_sources.json').read_text())
|
| 126 |
+
for name,entry in backbone.items():
|
| 127 |
+
if sha(ck/name)!=entry['sha256']:errors.append(f'Backbone SHA mismatch {name}')
|
| 128 |
+
if any(wm.get(n)!=name for n in entry['tensor_names']):errors.append(f'Backbone index mismatch {name}')
|
| 129 |
+
report={'passed':not errors,'layers':len(layers),'layer_ids':layers,'errors':errors,'cold_file_sha256':hashes,'scope':'Initial-fit export has exhaustive index roundtrips and sampled weight parity; checkpoint hashes, shapes, scales and decoded finiteness checked. No SM120 serving or model quality gate.'}
|
| 130 |
+
report['metadata_sha256']={n:sha(ck/n) for n in ['config.json','model.safetensors.index.json','backbone_sources.json']}
|
| 131 |
+
write(report_path or ck/'gates_report.json',report)
|
| 132 |
+
if errors:raise RuntimeError(f'8+8 gates failed: {errors[:3]}')
|
| 133 |
+
return report
|
| 134 |
+
|
| 135 |
+
def parallel_gates(work,gpus):
|
| 136 |
+
import subprocess,sys
|
| 137 |
+
from .inputs import ROOT
|
| 138 |
+
folder=work/'gate_parts';folder.mkdir(exist_ok=True);jobs=[]
|
| 139 |
+
for i,gpu in enumerate(gpus):
|
| 140 |
+
layers=LAYERS[i::len(gpus)]
|
| 141 |
+
if not layers:continue
|
| 142 |
+
result=folder/f'gpu{gpu}.json'
|
| 143 |
+
with open(folder/f'gpu{gpu}.log','a') as log:
|
| 144 |
+
p=subprocess.Popen([sys.executable,str(ROOT/'arvq88/gate_worker.py'),str(work),','.join(map(str,layers)),f'cuda:{gpu}',str(result)],stdout=log,stderr=subprocess.STDOUT,start_new_session=True)
|
| 145 |
+
jobs.append((p,result))
|
| 146 |
+
codes=[p.wait() for p,_ in jobs]
|
| 147 |
+
if any(codes):raise RuntimeError(f'Gate workers failed: {codes}; inspect {folder}')
|
| 148 |
+
parts=[json.loads(r.read_text()) for _,r in jobs];ids=[l for p in parts for l in p['layer_ids']]
|
| 149 |
+
if sorted(ids)!=LAYERS:raise RuntimeError('Gate layer coverage mismatch')
|
| 150 |
+
if any(p['metadata_sha256']!=parts[0]['metadata_sha256'] for p in parts):raise RuntimeError('Checkpoint metadata changed during gates')
|
| 151 |
+
report={**parts[0],'passed':all(p['passed'] for p in parts),'layers':75,'layer_ids':LAYERS,
|
| 152 |
+
'errors':[e for p in parts for e in p['errors']],
|
| 153 |
+
'cold_file_sha256':{k:v for p in parts for k,v in p['cold_file_sha256'].items()},'workers':len(parts)}
|
| 154 |
+
write(work/'checkpoint/gates_report.json',report)
|
| 155 |
+
if not report['passed']:raise RuntimeError('Gates did not pass')
|
| 156 |
+
return report
|
| 157 |
+
|
| 158 |
+
def publish(args,hub,work):
|
| 159 |
+
ck=work/'checkpoint';report=json.loads((ck/'gates_report.json').read_text())
|
| 160 |
+
if not report['passed'] or len(report['cold_file_sha256'])!=150:raise RuntimeError('Incomplete initial checkpoint gates')
|
| 161 |
+
for name,digest in {**report['cold_file_sha256'],**report['metadata_sha256']}.items():
|
| 162 |
+
if sha(ck/name)!=digest:raise RuntimeError('Weights changed after gates')
|
| 163 |
+
parent=hub.api.model_info(args.dst).sha
|
| 164 |
+
commit=hub.api.upload_folder(repo_id=args.dst,folder_path=str(ck),revision='main',parent_commit=parent,commit_message='GLM-5.3 NVFP4 + ARVQ 8+8 activation-calibrated initial fit')
|
| 165 |
+
remote={s.rfilename:s for s in hub.api.model_info(args.dst,revision=commit.oid,files_metadata=True).siblings}
|
| 166 |
+
for name,digest in report['cold_file_sha256'].items():
|
| 167 |
+
if remote[name].lfs.sha256!=digest:raise RuntimeError('Remote cold-weight hash mismatch')
|
| 168 |
+
write(work/'upload.json',{'repo':args.dst,'revision':commit.oid,'cold_files_verified':150});return commit.commit_url
|
reproduce/source/btx53/arvq88/cli.py
ADDED
|
@@ -0,0 +1,154 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Local-or-Hub GLM-5.3 to NVFP4/ARVQ 8+8 with audited PV tuning."""
|
| 2 |
+
import argparse,fcntl,json,os,subprocess,sys,time,traceback,shutil
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
from .inputs import ROOT,LAYERS,Hub,assignment,ensure_captures,ensure_source,write
|
| 5 |
+
|
| 6 |
+
def parser():
|
| 7 |
+
p=argparse.ArgumentParser(description=__doc__)
|
| 8 |
+
p.add_argument('--arvq88',action='store_true',help=argparse.SUPPRESS)
|
| 9 |
+
p.add_argument('--src',default='zai-org/GLM-5.3',help='Hugging Face repo ID or local model directory');p.add_argument('--dst',required=True)
|
| 10 |
+
p.add_argument('--workdir',default='/tmp/glm53-arvq88');p.add_argument('--gpus',default='0,1,2,3,4,5,6,7')
|
| 11 |
+
p.add_argument('--assign',default=str(ROOT/'cold_assignment.json'));p.add_argument('--reap-scores')
|
| 12 |
+
p.add_argument('--mm-stats',help='Cached multimodal salience NPZ for r3 allocation')
|
| 13 |
+
p.add_argument('--mm-calib',default='/tmp/glm53-calib-mm')
|
| 14 |
+
p.add_argument('--vision-dir',default='/tmp/glm53-vision')
|
| 15 |
+
p.add_argument('--reap-parts',help='Historical all-expert 1x16 AQLM warm starts, allocation proxy only')
|
| 16 |
+
p.add_argument('--reap-metric',choices=['aqlm','arvq'],default='aqlm',help='ARVQ recomputes hot allocation from fitted ARVQ vs NVFP4 output error on both modalities')
|
| 17 |
+
p.add_argument('--mm-acts',help='Matching donor multimodal activation directory for ARVQ allocation scoring')
|
| 18 |
+
p.add_argument('--hot-count',type=int,default=5750);p.add_argument('--calib',default='/tmp/glm53-capture/acts')
|
| 19 |
+
p.add_argument('--calib-tokens',default='/tmp/glm52-calib-v3');p.add_argument('--calib-token-count',type=int,default=15_000_000)
|
| 20 |
+
p.add_argument('--src-cache',default='/tmp/glm53-fp8-cold')
|
| 21 |
+
p.add_argument('--nvfp4-donor',default='RadixArk/GLM-5.3-NVFP4',help='NVFP4 Hugging Face repo ID or local sharded checkpoint directory; download cache is automatic')
|
| 22 |
+
p.add_argument('--base-hybrid',default='jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m')
|
| 23 |
+
p.add_argument('--local-hot',default='/tmp/glm53-hybrid-pipeline/checkpoint')
|
| 24 |
+
p.add_argument('--download-workers',type=int,default=8);p.add_argument('--threads',type=int,default=2)
|
| 25 |
+
p.add_argument('--col-block',type=int,default=128);p.add_argument('--cb-iters',type=int,default=8)
|
| 26 |
+
p.add_argument('--sweep-passes',type=int,default=2);p.add_argument('--refine',type=int,default=1)
|
| 27 |
+
p.add_argument('--codebook-scope',choices=['layer','expert'],default='layer')
|
| 28 |
+
p.add_argument('--pv-reassign-every',type=int,default=0,help='Alternate index proposals every N PV steps; 0 freezes indices')
|
| 29 |
+
from .activation import ARITHMETIC
|
| 30 |
+
p.add_argument('--pv-arithmetic',choices=['fp32',ARITHMETIC],default=ARITHMETIC,help='PV forward arithmetic: serving activation emulation or legacy FP32')
|
| 31 |
+
p.add_argument('--pv-steps',type=int,default=200,help='PV steps per layer; 0 produces initial fits only')
|
| 32 |
+
p.add_argument('--clean-repo',action='store_true',help='After verified final PV upload, remove obsolete repository files')
|
| 33 |
+
p.add_argument('--detach',action='store_true');p.add_argument('--do-upload',action='store_true')
|
| 34 |
+
p.add_argument('--plan',action='store_true',help='Show resolved workflow without downloading, fitting or uploading')
|
| 35 |
+
return p
|
| 36 |
+
|
| 37 |
+
def resolve_donor(args,work):
|
| 38 |
+
from .local_source import local_directory,fingerprint
|
| 39 |
+
from .inputs import complete_snapshot
|
| 40 |
+
local=local_directory(args.nvfp4_donor)
|
| 41 |
+
args.donor_identity=None
|
| 42 |
+
if local:
|
| 43 |
+
if not complete_snapshot(local) or not (local/'config.json').is_file():
|
| 44 |
+
raise ValueError('Local NVFP4 donor requires config.json, model.safetensors.index.json and all indexed shards')
|
| 45 |
+
args.nvfp4_donor=str(local);args.donor_cache=str(local)
|
| 46 |
+
args.donor_identity=fingerprint(local)
|
| 47 |
+
else:
|
| 48 |
+
import hashlib
|
| 49 |
+
legacy=Path('/tmp/glm53-dl/nvfp4-full')
|
| 50 |
+
args.donor_cache=str(legacy if args.nvfp4_donor=='RadixArk/GLM-5.3-NVFP4' and complete_snapshot(legacy)
|
| 51 |
+
else work/'donors'/hashlib.sha256(args.nvfp4_donor.encode()).hexdigest()[:16])
|
| 52 |
+
|
| 53 |
+
def main():
|
| 54 |
+
args=parser().parse_args()
|
| 55 |
+
from .local_source import local_directory,fingerprint
|
| 56 |
+
local=local_directory(args.src)
|
| 57 |
+
if local:
|
| 58 |
+
args.src=str(local);fingerprint(local) # Validate config/index/shards before capture or GPU work.
|
| 59 |
+
args.gpus=[int(x) for x in args.gpus.split(',')]
|
| 60 |
+
if not args.gpus or len(args.gpus)!=len(set(args.gpus)) or min(args.gpus)<0:raise ValueError('GPUs must be a unique nonnegative list')
|
| 61 |
+
if args.pv_steps!=0 and args.pv_steps<40:raise ValueError('PV steps must be 0 or at least 40')
|
| 62 |
+
if args.pv_reassign_every<0:raise ValueError('Invalid index reassignment interval')
|
| 63 |
+
if args.reap_metric=='arvq' and not args.mm_acts:raise ValueError('ARVQ scoring requires --mm-acts')
|
| 64 |
+
if args.clean_repo and args.pv_steps==0:raise ValueError('Repository cleanup requires final PV weights')
|
| 65 |
+
if args.threads<1 or args.download_workers<1 or args.calib_token_count<2048:raise ValueError('Invalid CPU/download/calibration count')
|
| 66 |
+
if args.col_block<=0 or args.col_block%8 or 2048%args.col_block or args.cb_iters<1 or args.sweep_passes<1 or args.refine<0:raise ValueError('Invalid encoder settings')
|
| 67 |
+
work=Path(args.workdir);work.mkdir(parents=True,exist_ok=True)
|
| 68 |
+
resolve_donor(args,work)
|
| 69 |
+
# Do not silently reuse this machine's original-source artifacts for another input.
|
| 70 |
+
if args.src!='zai-org/GLM-5.3':
|
| 71 |
+
for name,default,replacement in [('src_cache','/tmp/glm53-fp8-cold',work/'source'),('calib','/tmp/glm53-capture/acts',work/'capture/acts'),('calib_tokens','/tmp/glm52-calib-v3',work/'calib-tokens'),('assign',str(ROOT/'cold_assignment.json'),work/'recomputed-assignment.json')]:
|
| 72 |
+
if getattr(args,name)==default:setattr(args,name,str(replacement))
|
| 73 |
+
if args.nvfp4_donor!='RadixArk/GLM-5.3-NVFP4':
|
| 74 |
+
if args.calib=='/tmp/glm53-capture/acts':args.calib=str(work/'capture/acts')
|
| 75 |
+
if args.assign==str(ROOT/'cold_assignment.json'):args.assign=str(work/'recomputed-assignment.json')
|
| 76 |
+
if args.base_hybrid!='jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m' and args.local_hot=='/tmp/glm53-hybrid-pipeline/checkpoint':args.local_hot=str(work/'base-hybrid')
|
| 77 |
+
if args.plan:
|
| 78 |
+
print(json.dumps({'format':'rvq256_256x8_expert' if args.codebook_scope=='expert' else 'rvq256_256x8','stage':'PV-tuned' if args.pv_steps else 'initial fit (no PV)','args':vars(args),'steps':['reuse/download text calibration','reuse/capture NVFP4-teacher activations','reuse allocation or scores, otherwise recalculate 75% text REAP + 25% multimodal salience allocation','reuse/fetch original cold source tensors','fit 8+8 with per-layer cache identities',*(['PV with train/validation/audit splits, early stopping and audit fallback'] if args.pv_steps else []),'assemble donor hot/nonexpert/MTP + ARVQ cold + default vision','gate packed format and hashes','upload main only with --do-upload'],'fork_changes':False},indent=2));return
|
| 79 |
+
compat='/tmp/compat13/usr/local/cuda-13.2/compat'
|
| 80 |
+
if Path(compat).exists() and compat not in os.environ.get('LD_LIBRARY_PATH','').split(':'):
|
| 81 |
+
os.execvpe(sys.executable,[sys.executable]+sys.argv,{**os.environ,'LD_LIBRARY_PATH':compat+':'+os.environ.get('LD_LIBRARY_PATH','')})
|
| 82 |
+
if args.detach:
|
| 83 |
+
with open(work/'driver.log','a') as log:
|
| 84 |
+
child=subprocess.Popen([sys.executable,'-u']+[a for a in sys.argv if a!='--detach'],stdin=subprocess.DEVNULL,stdout=log,stderr=subprocess.STDOUT,start_new_session=True)
|
| 85 |
+
print(f'Detached 8+8 pipeline PID {child.pid}; {work}/driver.log');return
|
| 86 |
+
lock=open(work/'.workflow.lock','a');fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB)
|
| 87 |
+
import torch
|
| 88 |
+
torch.set_num_threads(args.threads)
|
| 89 |
+
if not torch.cuda.is_available():raise RuntimeError('CUDA required for 8+8 pipeline')
|
| 90 |
+
# Never overbook another user's or an existing detached worker's GPU.
|
| 91 |
+
result=subprocess.run(['nvidia-smi','--query-compute-apps=gpu_uuid,pid','--format=csv,noheader'],capture_output=True,text=True,check=True)
|
| 92 |
+
gpuinfo=subprocess.run(['nvidia-smi','--query-gpu=index,uuid','--format=csv,noheader'],capture_output=True,text=True,check=True)
|
| 93 |
+
selected={line.split(',')[1].strip() for line in gpuinfo.stdout.splitlines() if int(line.split(',')[0]) in args.gpus}
|
| 94 |
+
if any(line.split(',')[0].strip() in selected for line in result.stdout.splitlines() if line.strip()):raise RuntimeError('Requested GPUs have active compute jobs; choose free --gpus or wait. No processes were stopped.')
|
| 95 |
+
try:
|
| 96 |
+
binding={'src':args.src,'nvfp4_donor':args.nvfp4_donor,'base_hybrid':args.base_hybrid,'format':'rvq256_256x8_expert' if args.codebook_scope=='expert' else 'rvq256_256x8',
|
| 97 |
+
'reap_metric':args.reap_metric,'pv_reassign_every':args.pv_reassign_every}
|
| 98 |
+
if args.donor_identity:binding['local_donor_identity']=args.donor_identity
|
| 99 |
+
if (work/'binding.json').exists() and json.loads((work/'binding.json').read_text())!=binding:raise ValueError('Workdir belongs to another input/donor/format')
|
| 100 |
+
write(work/'binding.json',binding);hub=Hub(work/'hfcache')
|
| 101 |
+
ensure_captures(args,hub,work)
|
| 102 |
+
allocation=assignment(args,hub,work)
|
| 103 |
+
ensure_source(args,hub,work,allocation)
|
| 104 |
+
write(work/'run_config.json',vars(args));pending=list(LAYERS);running={}
|
| 105 |
+
if args.reap_metric=='arvq':
|
| 106 |
+
from .arvq_reap import select_candidates
|
| 107 |
+
select_candidates(work/'arvq_reap',work,allocation);pending=[]
|
| 108 |
+
while pending or running:
|
| 109 |
+
for gpu in args.gpus:
|
| 110 |
+
if gpu in running or not pending:continue
|
| 111 |
+
L=pending.pop(0)
|
| 112 |
+
with open(work/f'fit-layer{L}.log','a') as log:
|
| 113 |
+
child=subprocess.Popen([sys.executable,'-u',str(ROOT/'arvq88/fit.py'),str(work),str(L),f'cuda:{gpu}'],stdout=log,stderr=subprocess.STDOUT,stdin=subprocess.DEVNULL,start_new_session=True)
|
| 114 |
+
running[gpu]=(L,child);print('FIT',L,'GPU',gpu,'PID',child.pid,flush=True)
|
| 115 |
+
time.sleep(15)
|
| 116 |
+
for gpu,(L,child) in list(running.items()):
|
| 117 |
+
rc=child.poll()
|
| 118 |
+
if rc is None:continue
|
| 119 |
+
if rc:raise RuntimeError(f'Fit layer {L} failed; see fit-layer{L}.log')
|
| 120 |
+
del running[gpu];print('FIT DONE',L,flush=True)
|
| 121 |
+
if args.pv_steps:
|
| 122 |
+
from .pv_campaign import run as run_pv
|
| 123 |
+
run_pv(args,work,work/'initial')
|
| 124 |
+
from .checkpoint import build,parallel_gates,publish
|
| 125 |
+
ck=build(args,hub,work,allocation);parallel_gates(work,args.gpus)
|
| 126 |
+
write(ck/'build_provenance.json',{'inputs':binding,'hub_revisions':hub.revisions,'allocation':allocation['provenance'],'capture_dir':args.calib,'calib_tokens':args.calib_tokens,'stage':'PV-tuned' if args.pv_steps else 'initial fit, no PV','format':binding['format'],'source_nonexpert_mtp':'nvfp4_donor','source_vision':'base_hybrid','capture_teacher':'nvfp4_donor','allocation_recipe':allocation['provenance']})
|
| 127 |
+
(ck/'README.md').write_text(f'''---
|
| 128 |
+
base_model: {args.src}
|
| 129 |
+
tags: [arvq, nvfp4, experimental]
|
| 130 |
+
---
|
| 131 |
+
# GLM-5.3 NVFP4 / ARVQ 8+8 hybrid
|
| 132 |
+
|
| 133 |
+
Cold experts use two shared FP4 dictionaries of 256 entries, 8-dimensional groups,
|
| 134 |
+
16 index bits/group and FP8 scales/128 weights (2.0625 bpw plus codebooks).
|
| 135 |
+
This is the activation-calibrated initial Hessian fit, **before PV**.
|
| 136 |
+
Hot experts use the selected NVFP4 donor. Nonexpert and MTP tensors come from the NVFP4 donor; vision comes from the
|
| 137 |
+
base hybrid. See build_provenance.json for every donor and allocation method.
|
| 138 |
+
|
| 139 |
+
Codebook scope: {args.codebook_scope}. Requires a loader/kernel supporting
|
| 140 |
+
{binding['format']} version {3 if args.codebook_scope=='expert' else 2}: 512 codewords per pair,
|
| 141 |
+
64 uint32 index words per tile and 8+8 index extraction. The prior 8+7 loader is
|
| 142 |
+
incompatible. This pipeline does not modify the serving fork.
|
| 143 |
+
Offline packing/structure checks passed; serving and model-level quality are untested.
|
| 144 |
+
''')
|
| 145 |
+
shutil.copy2(ROOT/'arvq88/pack.py',ck/'arvq88_reference.py')
|
| 146 |
+
if args.pv_steps:
|
| 147 |
+
from .pv_campaign import write_card
|
| 148 |
+
write_card(args,work)
|
| 149 |
+
from .final_publish import publish
|
| 150 |
+
write(work/'complete.json',{'checkpoint':str(ck),'uploaded':False})
|
| 151 |
+
if args.do_upload:print('UPLOADED',publish(args,hub,work),flush=True)
|
| 152 |
+
else:print('Checkpoint ready:',ck,'; add --do-upload to publish',flush=True)
|
| 153 |
+
except BaseException:
|
| 154 |
+
(work/'run.failed').write_text(traceback.format_exc());raise
|
reproduce/source/btx53/arvq88/docs/expert_experiments.md
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Per-expert books, alternating updates, and ARVQ allocation
|
| 2 |
+
|
| 3 |
+
These changes are experimental. Production v2 defaults remain shared books and
|
| 4 |
+
fixed indices. No new weights are uploaded by the ablation runner.
|
| 5 |
+
|
| 6 |
+
The serving-agent prompt is in `per_expert_handoff.md`. Version 3 changes only
|
| 7 |
+
the packed codebook tensor from `[512]` to `[E,512]` per layer/projection.
|
| 8 |
+
The paired index tensor, block scales and single projection global scale retain
|
| 9 |
+
the v2 layout. Expert slots follow the ordered cold-expert manifest.
|
| 10 |
+
|
| 11 |
+
`pv.py --reassign-every 40` interleaves continuous updates with LDLQ discrete
|
| 12 |
+
proposals computed from training activation Hessians. Proposals are accepted
|
| 13 |
+
only if a fixed training subset improves in routed cold-output error. This is
|
| 14 |
+
an alternating optimizer, not a reproduction of the PV-Tuning paper. Best-state
|
| 15 |
+
snapshots include code indices, so validation selection and audit fallback
|
| 16 |
+
restore complete models, not just codebooks. Rare experts with fewer than 32
|
| 17 |
+
training rows are skipped during discrete proposals rather than using held-out
|
| 18 |
+
rows. This can limit improvement for rare experts.
|
| 19 |
+
|
| 20 |
+
`fit.py` accepts `codebook_scope: expert` in run_config.json and learns separate
|
| 21 |
+
books for each expert with a common projection global scale. Fitting filters
|
| 22 |
+
merged captures to `pv_split == 0` when labels are present.
|
| 23 |
+
|
| 24 |
+
Run the controlled pilot:
|
| 25 |
+
|
| 26 |
+
```bash
|
| 27 |
+
LD_LIBRARY_PATH=/tmp/compat13/usr/local/cuda-13.2/compat \
|
| 28 |
+
.venv/bin/python -u btx53/arvq88/experiment.py \
|
| 29 |
+
--base /tmp/glm53-vision-trace-refit \
|
| 30 |
+
--workdir /tmp/glm53-arvq-expert-ablation --layers 3,40,70
|
| 31 |
+
```
|
| 32 |
+
|
| 33 |
+
This waits for all selected GPUs to be free, fits independent books, then runs
|
| 34 |
+
shared-fixed/shared-alternating/expert-fixed/expert-alternating variants with
|
| 35 |
+
identical source, split captures, expert selection and continuous step budgets.
|
| 36 |
+
Alternating variants have additional discrete-search work; this is not an
|
| 37 |
+
iso-compute benchmark. Outputs are comparison.json and packed/<variant>/ shards.
|
| 38 |
+
The original shared fit is reused as a baseline; the per-expert fit is fresh.
|
| 39 |
+
No production promotion is automatic. Native serving parity and task evaluation
|
| 40 |
+
are owned by the serving agent. The reused audit corpus is not a new independent
|
| 41 |
+
benchmark; do not tune repeatedly against audit results.
|
| 42 |
+
|
| 43 |
+
ARVQ allocation is available through these new pipeline arguments:
|
| 44 |
+
|
| 45 |
+
```bash
|
| 46 |
+
.venv/bin/python btx53/pipeline.py \
|
| 47 |
+
--src zai-org/GLM-5.3 --dst jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-hybrid \
|
| 48 |
+
--workdir /tmp/glm53-arvq-expert-full \
|
| 49 |
+
--calib /tmp/glm53-vision-trace-refit/trace_capture/acts \
|
| 50 |
+
--calib-tokens /tmp/glm53-vision-trace-refit/train/tokens \
|
| 51 |
+
--codebook-scope expert --pv-reassign-every 40 --reap-metric arvq \
|
| 52 |
+
--mm-acts /tmp/glm53-vision-trace-refit/capture-mm-rlianv4b/acts_mm
|
| 53 |
+
```
|
| 54 |
+
|
| 55 |
+
Run a full campaign only after the pilot and serving tests warrant it; the above
|
| 56 |
+
command is documented, not launched. It fits candidate ARVQ weights for all 256
|
| 57 |
+
experts in all 75 layers, fetches missing original FP8 experts, scores candidates
|
| 58 |
+
on both text and multimodal captures, then selects exactly 5750 hot experts using
|
| 59 |
+
the 75/25 layer-normalized blend, floor 8, cap 176. The final initial weights are
|
| 60 |
+
an exact subset of those candidates; allocation is followed by alternating PV.
|
| 61 |
+
|
| 62 |
+
For each expert the new score is:
|
| 63 |
+
|
| 64 |
+
max(0, sum_t g[t,e]^2 * (||ARVQ_e(x_t)-FP8_e(x_t)||^2
|
| 65 |
+
- ||NVFP4_e(x_t)-FP8_e(x_t)||^2))
|
| 66 |
+
|
| 67 |
+
Both quantized candidates use four-plane / FP16-boundary emulation, not native
|
| 68 |
+
SM120 MMA. The score measures the local benefit of retaining the NVFP4 expert,
|
| 69 |
+
with no AQLM error proxy. It omits cross-expert error covariance, routing changes,
|
| 70 |
+
and downstream error accumulation. Both modalities are required; an empty or
|
| 71 |
+
nonpositive modality layer fails instead of silently using text-only allocation.
|
| 72 |
+
|
| 73 |
+
The initial implementation does not use full-model cross-entropy/distillation
|
| 74 |
+
and does not establish end-to-end quality. The old corpus captures use fixed
|
| 75 |
+
teacher routing and 2048-token windows. Independent reasoning/task evaluations
|
| 76 |
+
remain necessary even if all layer audits pass.
|
reproduce/source/btx53/arvq88/docs/gradient_pv_recipe.md
ADDED
|
@@ -0,0 +1,180 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Proposed Vision ARVQ v3 tuning recipe
|
| 2 |
+
|
| 3 |
+
Prepared 18 September 2026 (Melbourne). Status: plan, not a production restart.
|
| 4 |
+
The prior full-run driver and publisher were stopped at the user's request.
|
| 5 |
+
Initial fits, REAP allocation, local results, and published weights are preserved.
|
| 6 |
+
|
| 7 |
+
## Findings verified against code and logs
|
| 8 |
+
|
| 9 |
+
- Claude transcript `d7bce691-1871-46a6-af7a-61eae60194fc.jsonl`,
|
| 10 |
+
2026-09-01T22:31:07Z: AQLM converge optimized Hessian-weighted weight error;
|
| 11 |
+
subsequent PV optimized layer output MSE with fixed indices.
|
| 12 |
+
- `tools/capture53/pv_tune53.py` inherits `tools/pv_tune.py`: per-layer,
|
| 13 |
+
expert-coordinate gradient updates; cached routing; NVFP4-dequant teacher;
|
| 14 |
+
validation best restore. Its `reseed()` body is a no-op. This is not the
|
| 15 |
+
full-model algorithm in the PV-Tuning paper.
|
| 16 |
+
- `PLAN53.md` and `tools/capture53/eval_hybrid53.py`: the old pipeline also
|
| 17 |
+
performed full-model, layer-streamed perplexity evaluation. The old evaluator
|
| 18 |
+
decodes AQLM and needs an ARVQ v3 loader and arithmetic changes.
|
| 19 |
+
- Current ARVQ trains all codebooks and per-128-weight scales. Per expert:
|
| 20 |
+
8,192 dictionary values + 294,912 scales; 13,450 cold experts expose
|
| 21 |
+
3,966,566,400 scale values. Low storage precision is not immunity to overfit.
|
| 22 |
+
- Current merged capture: 98,304 rows/layer, 32,768 each for train/val/audit.
|
| 23 |
+
It contains x, topk_ids, topk_weights, pv_split, but no per-row prompt IDs.
|
| 24 |
+
- Current objective matches cold FP8 contributions only. The old AQLM code
|
| 25 |
+
explicitly subtracted frozen student-hot contributions from the full target.
|
| 26 |
+
With FP8 reference versus NVFP4 hot weights these are not interchangeable.
|
| 27 |
+
- New sparse gradient prototype passes the 42-test package suite, including
|
| 28 |
+
gradient direction, trust limit, training-row isolation and improvement tests.
|
| 29 |
+
A real layer-3/two-expert/two-step smoke completed packing/reload audit, but
|
| 30 |
+
rejected all 8 expert/projection proposals. This is not quality qualification.
|
| 31 |
+
|
| 32 |
+
## Fixed scope and baselines
|
| 33 |
+
|
| 34 |
+
Preserve expert-specific books, v3 serialization, original FP8 weight source,
|
| 35 |
+
current 5,750/13,450 REAP allocation, and existing corpus. No refitting/re-download
|
| 36 |
+
is necessary for this optimizer change. Preserve all old PV artifacts, including
|
| 37 |
+
12 previously published layers. Compare proposed replacements against both the
|
| 38 |
+
initial fit and the currently published layer. If both are candidates, select on
|
| 39 |
+
validation and reserve the new final audit for the frozen recipe.
|
| 40 |
+
|
| 41 |
+
This remains layer-local alternating optimization, not formal EM and not a
|
| 42 |
+
claim to reproduce full-model published PV-Tuning.
|
| 43 |
+
|
| 44 |
+
## Phase 1: data and trustworthy objective
|
| 45 |
+
|
| 46 |
+
1. Recapture or expand from the same training problems to an initial target of
|
| 47 |
+
131,072 representative training rows/layer. Keep separate prompt-defined
|
| 48 |
+
validation; retain prompt ID, token position, domain, modality and capture
|
| 49 |
+
provenance per row. Use reservoir/stratified sampling across conversations,
|
| 50 |
+
not the first rows or unrestricted token-level resplitting.
|
| 51 |
+
2. Log routed token counts AND distinct prompt counts for every expert. Tentative
|
| 52 |
+
minimum for expert-specific discrete tuning: 256 training routes from at least
|
| 53 |
+
16 prompts; below this freeze indices and scale corrections, and initially
|
| 54 |
+
freeze its books too. Pilot coverage decides whether more capture is needed.
|
| 55 |
+
3. Include multimodal TRAINING captures, not just multimodal REAP scores. Start
|
| 56 |
+
with the existing 75/25 text/MM mixture as a hypothesis, report both domains
|
| 57 |
+
separately, and do not treat previously allocation-used MM rows as held-out.
|
| 58 |
+
4. Build full routed FP8 reference outputs, subtract frozen student-hot output,
|
| 59 |
+
then fit the cold sum to that residual. Report both full routed error and
|
| 60 |
+
cold-only error. Match actual hot arithmetic or clearly label the emulator.
|
| 61 |
+
5. Normalize training MSE with a fixed training-set reference energy rather than
|
| 62 |
+
changing denominators per minibatch. Preserve natural routing weights;
|
| 63 |
+
expert-enriched samples require sampling corrections to avoid silently
|
| 64 |
+
changing the deployment objective.
|
| 65 |
+
6. Audit forward arithmetic against the serving commit reported as b8db885d9,
|
| 66 |
+
including routing weights, intermediate rounding, scale layout, cold slots,
|
| 67 |
+
accumulation and TP reductions. A100 surrogate gradients are acceptable;
|
| 68 |
+
native forward parity needs the SM120 environment, not a claim of bit parity.
|
| 69 |
+
|
| 70 |
+
## Phase 2: constrained continuous updates
|
| 71 |
+
|
| 72 |
+
- Start with books + regularized per-output-row multiplicative scale corrections.
|
| 73 |
+
Fold corrections into the existing E4M3 per-128 scales at export: no format or
|
| 74 |
+
kernel change. Freeze projection-global scales to remove avoidable scale
|
| 75 |
+
ambiguity. This exposes 10,240 row corrections per expert instead of 294,912
|
| 76 |
+
independent block-scale adjustments.
|
| 77 |
+
- Use FP32 latent parameters, FP4/E4M3 quantized forward and STE backward.
|
| 78 |
+
- Penalize drift from initial quantized weights/books and log-scale corrections;
|
| 79 |
+
use stronger shrinkage for poorly supported experts. Do not describe this as
|
| 80 |
+
a guarantee. The tiny prototype's uniform latent anchor is only scaffolding.
|
| 81 |
+
- Pilot conservative book learning rates around 3e-4 versus the prior 3e-3,
|
| 82 |
+
and row-log-scale LR around 1e-4, with warmup then decay. These are starting
|
| 83 |
+
hypotheses, not transferred AQLM hyperparameters: parameter units differ.
|
| 84 |
+
- Effective batch 1,024 through four 256-row microbatches; measure wall time.
|
| 85 |
+
Unfreeze individual block scales only if a later bounded ablation establishes
|
| 86 |
+
an independent-validation benefit. Capacity should be earned by evidence.
|
| 87 |
+
|
| 88 |
+
## Phase 3: output-gradient discrete updates
|
| 89 |
+
|
| 90 |
+
- After a short continuous warmup, propose every 20 continuous updates initially.
|
| 91 |
+
- Compute gradients of the SAME routed-output objective with respect to one
|
| 92 |
+
expert's decoded gate/up or down matrix at a time, with other experts fixed.
|
| 93 |
+
The full routed residual retains interactions between experts.
|
| 94 |
+
- Precondition/damp the gradient and construct a temporary nearby target weight.
|
| 95 |
+
Search code pairs against this gradient-stepped target, NOT original FP8
|
| 96 |
+
matrices. Weight-space projection itself is legitimate: the target matters.
|
| 97 |
+
- Use beam search over both books (small beam e.g. 4, retain current pair),
|
| 98 |
+
shortlist by predicted improvement and consider at most 0.1% of groups per
|
| 99 |
+
expert/projection initially. Hard ceiling 1%; also bound decoded weight change.
|
| 100 |
+
- Backtrack the target step and/or number of changed groups when actual loss
|
| 101 |
+
worsens. Raw gradient direction plus a global norm constraint was insufficient
|
| 102 |
+
in the smoke test. No forced replacements and no all-layer all-or-nothing sweep.
|
| 103 |
+
- Accept each expert/projection separately using actual quantized forward loss
|
| 104 |
+
on proposal-training rows and a second rotating training check batch. Never
|
| 105 |
+
query validation to accept individual proposals. Both batches remain training.
|
| 106 |
+
- Exploit the down projection's linearity while gate/up is fixed: for a candidate
|
| 107 |
+
compute its routed output delta and exact squared-loss change
|
| 108 |
+
`2 * <residual, delta_y> + ||delta_y||^2`. This is a more faithful ranking
|
| 109 |
+
than weight distance. Batched changes still require a combined check because
|
| 110 |
+
their output deltas interact. Gate/up proposals require nonlinear forward checks.
|
| 111 |
+
- For gate/up changes recompute downstream activations. Following acceptance,
|
| 112 |
+
refresh gradients/residuals before the next projection. Restore codes and any
|
| 113 |
+
affected optimizer state on rollback; codebooks/scales are fixed in this step.
|
| 114 |
+
- Stream one expert/projection's dense gradient/temporary target. Do not allocate
|
| 115 |
+
persistent dense Adam moments or shadow weights for all cold experts: those
|
| 116 |
+
are impractical at this model size. Sparse/factorized state needs memory tests.
|
| 117 |
+
|
| 118 |
+
## Phase 4: stopping, validation and final audit
|
| 119 |
+
|
| 120 |
+
- Report fixed training-probe and validation loss every 25-32 continuous steps,
|
| 121 |
+
by domain and at prompt level. Include accepted proposals, actual/predicted
|
| 122 |
+
improvement, quantized value changes, and coverage. Batch loss alone is not
|
| 123 |
+
evidence of a train/validation gap.
|
| 124 |
+
- First pilot uses 200 steps for a matched comparison. For the chosen full recipe
|
| 125 |
+
use a bounded 2-4 passes over the expanded training cache, rather than assuming
|
| 126 |
+
200 steps is convergence. Always retain the best validation state.
|
| 127 |
+
- Do not stop for a short plateau. After >=2 passes, require sustained training
|
| 128 |
+
improvement AND validation deterioration over >=3 checks, exceeding uncertainty
|
| 129 |
+
from paired prompt-level resampling. Separate compute-budget exhaustion and
|
| 130 |
+
plateau from demonstrated overfitting. Small fixed percentage thresholds in the
|
| 131 |
+
prototype are provisional, not statistical proof.
|
| 132 |
+
- Existing audit data has already influenced several accepted runs. Build a new
|
| 133 |
+
disjoint final audit and do not inspect it during optimizer/recipe selection.
|
| 134 |
+
One final pass/fail; do not tune until that audit passes.
|
| 135 |
+
|
| 136 |
+
## Phase 5: pilot, rollout and full-model verification
|
| 137 |
+
|
| 138 |
+
1. Three full representative layers (3,40,70), baseline initial/current weights.
|
| 139 |
+
2. Compare continuous-only against continuous + gradient-guided indices with
|
| 140 |
+
identical data, scale parameterization, regularization and step budget. This
|
| 141 |
+
isolates whether discrete optimization actually helps; no shared-book trials.
|
| 142 |
+
3. Qualify on held-out error including text/MM slices, actual serialized forward,
|
| 143 |
+
proposal improvement, no malformed tensors, deterministic rollback and peak
|
| 144 |
+
memory/runtime. Rejection rate alone is neither success nor failure.
|
| 145 |
+
4. Launch all 75 only after the pilot demonstrates a held-out benefit or useful
|
| 146 |
+
nonregression at realistic cost. Incremental uploads retain truthful recipe
|
| 147 |
+
provenance and current validation status; no stale old-run result reuse.
|
| 148 |
+
5. Port the existing streamed evaluator to ARVQ v3. Compare NVFP4 donor, initial
|
| 149 |
+
per-expert checkpoint, and final candidate on identical held-out/neutral text
|
| 150 |
+
and multimodal checks. Weight-only decoded evaluation is a diagnostic; the
|
| 151 |
+
activation-quantized forward and native SM120 tests remain required.
|
| 152 |
+
6. Measure teacher-input versus student-input activation/routing drift. If local
|
| 153 |
+
improvements fail to transfer, use one sequential refinement sweep: feed each
|
| 154 |
+
layer with already-quantized upstream student activations, pair with reference
|
| 155 |
+
trajectories and recompute appropriate routing/targets. Preserve held-out
|
| 156 |
+
prompt separation. Do not recapture the teacher and label it student capture.
|
| 157 |
+
7. Consider short multi-layer reconstruction or full-model KL distillation only
|
| 158 |
+
if measured cross-layer drift justifies it and a memory/runtime prototype fits.
|
| 159 |
+
Independent layer tuning cannot be claimed to minimize end-to-end loss.
|
| 160 |
+
|
| 161 |
+
## Code scope
|
| 162 |
+
|
| 163 |
+
- `gradient_indices.py`: damped/beam proposals, backtracking, two-batch acceptance,
|
| 164 |
+
memory-bounded state, exact per-expert rollback and diagnostics. Prototype exists.
|
| 165 |
+
- `pv.py`: target residual, restricted scale parameterization, fixed normalization,
|
| 166 |
+
gradient accumulation, regularization, uncertainty-aware stopping. Partial
|
| 167 |
+
prototype exists; not approved as the final recipe by its smoke result.
|
| 168 |
+
- New `pv_data.py` / capture extension: prompt provenance, coverage and multimodal
|
| 169 |
+
partitions, larger representative reservoirs, student-trajectory capture.
|
| 170 |
+
- `validation.py`: per-prompt paired metrics, separate overfit versus plateau,
|
| 171 |
+
fresh final audit and baseline comparisons.
|
| 172 |
+
- `pv_campaign.py`: recipe hash, unique workdir, pilot selection and true optimizer
|
| 173 |
+
resume. Restart from initial fits by default; keep old tuned artifacts intact.
|
| 174 |
+
- `incremental_publish.py` / `write_card`: truthful optimizer/coverage/stopping
|
| 175 |
+
provenance and mixed-revision status; no old Hessian method described as new PV.
|
| 176 |
+
- New `eval_arvq53.py`: adapt `tools/capture53/eval_hybrid53.py`, with explicit
|
| 177 |
+
weight-only versus activation-emulated modes and serving parity fixtures.
|
| 178 |
+
|
| 179 |
+
No new completion ETA until the selected full-layer pilot measures iteration,
|
| 180 |
+
proposal, validation and upload costs. All user-facing times: Melbourne.
|
reproduce/source/btx53/arvq88/docs/per_expert_handoff.md
ADDED
|
@@ -0,0 +1,78 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Implement per-expert ARVQ 8+8 codebook support in the active GLM serving fork.
|
| 2 |
+
|
| 3 |
+
Coordinate with the fitting agent in /home/coder/git/glm52. That agent owns fitting,
|
| 4 |
+
alternating index updates, ARVQ-specific REAP scoring, and checkpoint export. You own
|
| 5 |
+
the serving loader/kernel and actual SM120 tests. Do not duplicate fitter changes.
|
| 6 |
+
|
| 7 |
+
Context
|
| 8 |
+
- Existing 8+8 checkpoints share two FP4 books per (layer, projection).
|
| 9 |
+
- New experiments give each cold expert its own pair of 256 x 8 FP4 books.
|
| 10 |
+
- gate/up (w13) share a pair; down (w2) has a separate pair.
|
| 11 |
+
- New tensor layout is an explicit version 3 format. Existing version 2 must keep
|
| 12 |
+
working unchanged. The stale /tmp/arvq-inspect2 checkout only supported 8+7 at
|
| 13 |
+
inspection time: work from the actual current 8+8 serving branch, not that clone.
|
| 14 |
+
- No new production checkpoint is ready yet. Start with synthetic parity fixtures.
|
| 15 |
+
|
| 16 |
+
Exact proposed format contract (fitter will export this)
|
| 17 |
+
config.quantization_config.arvq, also under text_config when present:
|
| 18 |
+
format: rvq256_256x8_expert
|
| 19 |
+
version: 3
|
| 20 |
+
codebook_scope: expert
|
| 21 |
+
codebook_sizes: [256, 256]
|
| 22 |
+
activation_planes: 4
|
| 23 |
+
weight_scale_group: 128
|
| 24 |
+
|
| 25 |
+
For each layer L, prefix model.layers.L.mlp.experts.arvq_{w13|w2}_:
|
| 26 |
+
packed: uint32 [E, N/16, K/64, 64] (identical to v2)
|
| 27 |
+
codebooks: uint32 [E, 512] (NEW expert axis)
|
| 28 |
+
scales: uint8 [E, N/16, K/128, 16] (identical to v2)
|
| 29 |
+
global: float32 [1] (identical to v2; not per expert)
|
| 30 |
+
E is the number of cold experts in this layer.
|
| 31 |
+
w13: N=4096, K=6144. w2: N=6144, K=2048.
|
| 32 |
+
Each book row packs 8 FP4 E2M1 values into one uint32: component j occupies
|
| 33 |
+
bits [4*j,4*j+3]. Codes 0..7 represent [0,.5,1,1.5,2,3,4,6], bit 3 is sign.
|
| 34 |
+
Rows 0..255 are book 0; rows 256..511 are book 1.
|
| 35 |
+
The representation is W[e,r,k] = global * scale[e,r,k//128] *
|
| 36 |
+
(C[e,a[e,r,k//8],k%8] + C[e,256+b[e,r,k//8],k%8]).
|
| 37 |
+
Index pairs retain a | (b << 8), packed in the existing v2 MMA-fragment order.
|
| 38 |
+
Use the existing v2 row/column packing, not a new natural-order interpretation.
|
| 39 |
+
See btx53/arvq88/pack.py for the exporter/reference decoder.
|
| 40 |
+
|
| 41 |
+
Critical expert mapping
|
| 42 |
+
- Book row e is the COLD SLOT, matching packed[e], scales[e], and the ordered
|
| 43 |
+
cold_expert_ids manifest. It is not the original global expert ID.
|
| 44 |
+
- TP shards retain the full local codebook pair for every locally represented
|
| 45 |
+
cold expert; books are not sliced along input/output axes. Existing N/K index
|
| 46 |
+
and scale sharding remains unchanged.
|
| 47 |
+
- EP/expert reordering must reorder codebooks with the other expert tensors.
|
| 48 |
+
- Preserve the shared [512] v2 codebook path. Reject incompatible format/shape
|
| 49 |
+
combinations rather than silently interpreting expert books as shared.
|
| 50 |
+
|
| 51 |
+
Kernel changes
|
| 52 |
+
- Select cb + cold_slot * 512 for v3; shared cb base for v2.
|
| 53 |
+
- Keep FP4 two-book MMA algebra and four activation planes unchanged.
|
| 54 |
+
- Audit all direct, LUT, fused, prefill/decode, and CUDA-graph paths. A lookup table
|
| 55 |
+
computed once from layer-shared books cannot be reused across different experts.
|
| 56 |
+
Make caches/LUT keys include the relevant expert or dispatch to a correct path.
|
| 57 |
+
- Account for any shared-memory sizing that assumed exactly one shared book pair.
|
| 58 |
+
- No additional per-weight index bits are required; actual throughput/cache effects
|
| 59 |
+
must be measured, not inferred from unchanged MMA count.
|
| 60 |
+
|
| 61 |
+
Required tests before claiming correctness
|
| 62 |
+
1. Two experts with identical indices/scales and deliberately DIFFERENT books must
|
| 63 |
+
decode to different known outputs. Exercise permuted original expert IDs.
|
| 64 |
+
2. Compare packed decode to an independent CPU oracle, including a,b=0 and 255,
|
| 65 |
+
negative FP4 values, extreme valid scales, both projections, and TP slicing.
|
| 66 |
+
3. Duplicated per-expert books must agree with the existing shared-book path within
|
| 67 |
+
the same kernel's numerical tolerance; test full MoE outputs and routing weights.
|
| 68 |
+
4. On SM120 compare gate/up outputs, FP16 SwiGLU, down outputs, and summed MoE
|
| 69 |
+
outputs to the FP4-plane reference on identical inputs. Report max abs and
|
| 70 |
+
relative L2 discrepancies for TP1 and actual deployment TP.
|
| 71 |
+
5. Test both decode and prefill, graphs on/off, expert permutations, and LUT/direct
|
| 72 |
+
dispatch. Benchmark latency/throughput and memory against version 2.
|
| 73 |
+
6. Run the actual failing model tests once the fitting agent supplies candidate
|
| 74 |
+
weights. Offline layer audit success is not evidence of end-to-end quality.
|
| 75 |
+
|
| 76 |
+
Scope: implement and test serving support; return commit/patch, commands, parity
|
| 77 |
+
results, and benchmark results. Do not upload, overwrite, or deploy a production
|
| 78 |
+
checkpoint. Report hardware-dependent checks you cannot run as untested.
|
reproduce/source/btx53/arvq88/docs/sequential_pilot.md
ADDED
|
@@ -0,0 +1,65 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Two-block reference-trajectory pilot
|
| 2 |
+
|
| 3 |
+
Work directory: `/tmp/glm53-sequential-pilot`.
|
| 4 |
+
Driver: `python btx53/arvq88/sequential_pilot.py --work /tmp/glm53-sequential-pilot`.
|
| 5 |
+
No Hugging Face publication is performed by this pilot.
|
| 6 |
+
|
| 7 |
+
The first routed layers are 3 and 4. Frozen donor layers 0–2 generate the prefix.
|
| 8 |
+
Both reference and student use this prefix and unchanged donor attention, router,
|
| 9 |
+
shared experts and norms. Reference routed experts in layers 3/4 come from the
|
| 10 |
+
original FP8 source. This is a controlled hybrid reference, not a claim to run
|
| 11 |
+
all-original FP8 nonexpert weights or native SM120 kernels.
|
| 12 |
+
|
| 13 |
+
Capture uses 64 training, 16 validation, 16 development-audit windows of 1,024
|
| 14 |
+
tokens, sampled deterministically from the existing prompt-separated corpus.
|
| 15 |
+
Within-split windows are not guaranteed to represent distinct prompts. The
|
| 16 |
+
historical audit split is development data, not a fresh untouched final test.
|
| 17 |
+
Attention is dense-equivalent at this context length (<= index_topk).
|
| 18 |
+
|
| 19 |
+
Layer 3 has identical same-input/reference-trajectory targets and is trained once.
|
| 20 |
+
Its retained serialized outputs become the student's layer-4 inputs. The
|
| 21 |
+
reference layer-3 outputs independently propagate into reference layer 4.
|
| 22 |
+
Layer 4 then receives two matched 200-step runs from identical initial weights:
|
| 23 |
+
|
| 24 |
+
- `same_input`: FP8 reference experts on the student's own block input.
|
| 25 |
+
- `reference`: the original reference trajectory's complete block output.
|
| 26 |
+
|
| 27 |
+
Both subtract the frozen student's residual/attention/shared/hot contributions
|
| 28 |
+
before fitting the cold sum. Loss includes the complete block residual mismatch.
|
| 29 |
+
The same fixed loss normalization, training minibatches, seed, optimizer settings,
|
| 30 |
+
validation split, allocation and initial codebooks are used in both arms.
|
| 31 |
+
|
| 32 |
+
Eight GPUs own disjoint cold-expert subsets within each layer. Their partial
|
| 33 |
+
routed outputs are summed, and each rank differentiates the same coupled loss
|
| 34 |
+
with respect to its local experts. No averaging of local expert targets replaces
|
| 35 |
+
this objective. Continuous updates use effective batch 1,024 in four microbatches.
|
| 36 |
+
Books use LR 3e-4, row-log-scale corrections 1e-4, and an initialization anchor.
|
| 37 |
+
Block scales and projection globals are frozen; row corrections fold into the
|
| 38 |
+
existing E4M3 scale layout. Experts with <256 routed training rows are frozen.
|
| 39 |
+
|
| 40 |
+
Every 20 steps, expert-local output gradients propose sparse code changes.
|
| 41 |
+
Per expert/projection proposals are generated concurrently; their output deltas
|
| 42 |
+
are then accepted in deterministic rank order against the up-to-date coupled
|
| 43 |
+
residual. Prefix backtracking tries smaller changes. Acceptance requires lower
|
| 44 |
+
error on the proposal-training batch and nonregression on a separate training
|
| 45 |
+
check batch. Validation never accepts individual proposals.
|
| 46 |
+
|
| 47 |
+
Quantized forward emulates FP4 cold books, E4M3 weight scales, four activation
|
| 48 |
+
planes and FP16 SwiGLU boundaries. Hot experts use NVFP4-dequantized weights with
|
| 49 |
+
float32 GEMMs. Cross-GPU sums and donor BF16 backbone are surrogates for native
|
| 50 |
+
serving numerics; record this limitation with results.
|
| 51 |
+
|
| 52 |
+
Validation every 25 steps retains the best state. Budget is 200 steps. Early
|
| 53 |
+
stopping is allowed only after >=150 steps and 3 consecutive checks where
|
| 54 |
+
training improves >=0.5% while validation worsens >=0.5% versus the validation
|
| 55 |
+
best. This is a conservative heuristic, not statistical proof of overfitting.
|
| 56 |
+
|
| 57 |
+
`comparison.json` compares both layer-4 candidates against the SAME reference
|
| 58 |
+
trajectory on held-out windows; includes a paired-window bootstrap diagnostic.
|
| 59 |
+
Output errors are full-block relative L2, so they are not numerically comparable
|
| 60 |
+
to earlier cold-contribution relative errors (~0.2–0.3).
|
| 61 |
+
|
| 62 |
+
Preflight: 8-GPU one-step layer-3 smoke completed; 304/388 proposals accepted,
|
| 63 |
+
118,355 groups changed. Held-out full-block relative L2 0.0054971166 ->
|
| 64 |
+
0.0054560048. Saved/reloaded forward matched. These are smoke measurements,
|
| 65 |
+
not results of the 200-step target comparison.
|
reproduce/source/btx53/arvq88/encoder.py
ADDED
|
@@ -0,0 +1,321 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Activation-Hessian-tuned ARVQ fit for one (layer, projection).
|
| 2 |
+
|
| 3 |
+
Codebooks (c0[256,8], c1[256,8], FP4-constrained) are SHARED across all cold
|
| 4 |
+
experts of a (layer, projection), so this is a joint per-layer fit:
|
| 5 |
+
|
| 6 |
+
1. per expert: raw-basis Hessian H_e, escalating-damp Cholesky hinv_e,
|
| 7 |
+
per-(row,128block) init scale s_e, global scalar (fixed).
|
| 8 |
+
2. codebook EM on a Hessian-importance-weighted subsample of normalized
|
| 9 |
+
8-dim weight groups (k-means warm start -> alternating refine, FP4 project
|
| 10 |
+
after every M-step). <- "fit codebooks to data" (a)
|
| 11 |
+
3. per expert final assignment via a GPTQ/LDLQ error-feedback column sweep
|
| 12 |
+
with the fixed codebooks (plain-L2 inner assignment; cross-group feedback
|
| 13 |
+
carries the off-diagonal Hessian), then LS refit of the E4M3 scales.
|
| 14 |
+
<- "Hessian-aware code selection + error feedback" (b)
|
| 15 |
+
|
| 16 |
+
No incoherence rotation: Hessians and targets are in the raw weight basis
|
| 17 |
+
(the SM120 kernel feeds raw activations). Weighting is in the EM importance +
|
| 18 |
+
error feedback + scale LS, never a sqrt(diag(H)) whitening (the FP4 grid on the
|
| 19 |
+
codebook forbids rescaling the target space).
|
| 20 |
+
"""
|
| 21 |
+
from __future__ import annotations
|
| 22 |
+
|
| 23 |
+
from dataclasses import dataclass
|
| 24 |
+
|
| 25 |
+
import torch
|
| 26 |
+
import torch.nn.functional as F
|
| 27 |
+
|
| 28 |
+
from arvqprep.pack import project_to_fp4
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
# --------------------------------------------------------------- Hessians
|
| 32 |
+
def hessian_fc1(x: torch.Tensor) -> torch.Tensor:
|
| 33 |
+
"""gate/up share E[x x^T] over the H axis. x: [T, H] (routed tokens)."""
|
| 34 |
+
x = x.float()
|
| 35 |
+
return (x.t() @ x) / max(x.shape[0], 1)
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def hessian_fc2(x: torch.Tensor, wg: torch.Tensor, wu: torch.Tensor) -> torch.Tensor:
|
| 39 |
+
"""down E[m m^T], m = silu(x Wg^T) * (x Wu^T) over the I axis. x:[T,H]."""
|
| 40 |
+
x = x.float()
|
| 41 |
+
m = F.silu(x @ wg.float().t()) * (x @ wu.float().t()) # [T, I]
|
| 42 |
+
return (m.t() @ m) / max(x.shape[0], 1)
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
def cholesky_inv_upper(hess: torch.Tensor, damp_start: float = 1e-3) -> torch.Tensor:
|
| 46 |
+
"""Escalating-damp upper Cholesky of H^-1 (same recipe as btxprep/encoder)."""
|
| 47 |
+
k = hess.shape[0]
|
| 48 |
+
mean_diag = hess.diagonal().mean().clamp(min=1e-12)
|
| 49 |
+
damp = damp_start
|
| 50 |
+
eye = torch.eye(k, device=hess.device, dtype=hess.dtype)
|
| 51 |
+
for _ in range(12):
|
| 52 |
+
try:
|
| 53 |
+
return torch.linalg.cholesky(
|
| 54 |
+
torch.linalg.inv(hess + damp * mean_diag * eye), upper=True)
|
| 55 |
+
except Exception: # noqa: BLE001
|
| 56 |
+
damp *= 4.0
|
| 57 |
+
raise RuntimeError("ARVQ Cholesky failed at maximum damp")
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
# ---------------------------------------------------------- codebook EM
|
| 61 |
+
def _nearest(x, c):
|
| 62 |
+
"""argmin_k ||x - c_k||^2 (plain L2). x:[M,8], c:[K,8] -> [M]."""
|
| 63 |
+
return (c.square().sum(1)[None, :] - 2 * x @ c.t()).argmin(1)
|
| 64 |
+
|
| 65 |
+
|
| 66 |
+
def _wmeans(x, w, ids, count, prev):
|
| 67 |
+
"""Weighted per-cluster mean of x (scalar sample weights w); keep prev if empty."""
|
| 68 |
+
sums = torch.zeros_like(prev)
|
| 69 |
+
sums.index_add_(0, ids, x * w[:, None])
|
| 70 |
+
wsum = torch.zeros(count, device=x.device, dtype=x.dtype)
|
| 71 |
+
wsum.index_add_(0, ids, w)
|
| 72 |
+
return torch.where(wsum[:, None] > 0, sums / wsum.clamp_min(1e-20)[:, None], prev)
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
def fit_codebooks(x, w, *, iters=8, seed=0):
|
| 76 |
+
"""Weighted FP4 additive-RVQ codebook fit on a subsample.
|
| 77 |
+
x:[M,8] normalized group vectors, w:[M] importance weights.
|
| 78 |
+
Returns c0[256,8], c1[256,8] (FP4-on-grid), best weighted relative-L2."""
|
| 79 |
+
g = torch.Generator(device=x.device).manual_seed(seed)
|
| 80 |
+
perm = torch.randperm(x.shape[0], generator=g, device=x.device)
|
| 81 |
+
c0 = project_to_fp4(x[perm[:256]].clone())
|
| 82 |
+
i = _nearest(x, c0)
|
| 83 |
+
c1 = project_to_fp4(_kmeans(x - c0[i], 256, g))
|
| 84 |
+
j = _nearest(x - c0[i], c1)
|
| 85 |
+
wsum = (w * x.square().sum(1)).sum().clamp_min(1e-20)
|
| 86 |
+
best = None
|
| 87 |
+
for _ in range(iters):
|
| 88 |
+
i = _nearest(x - c1[j], c0)
|
| 89 |
+
c0 = project_to_fp4(_wmeans(x - c1[j], w, i, 256, c0))
|
| 90 |
+
j = _nearest(x - c0[i], c1)
|
| 91 |
+
c1 = project_to_fp4(_wmeans(x - c0[i], w, j, 256, c1))
|
| 92 |
+
err = ((w * (x - c0[i] - c1[j]).square().sum(1)).sum() / wsum).item()
|
| 93 |
+
if best is None or err < best[0]:
|
| 94 |
+
best = (err, c0.clone(), c1.clone())
|
| 95 |
+
err, c0, c1 = best
|
| 96 |
+
return c0, c1, err ** 0.5
|
| 97 |
+
|
| 98 |
+
|
| 99 |
+
def _kmeans(x, count, gen, iters=12):
|
| 100 |
+
c = x[torch.randperm(x.shape[0], generator=gen, device=x.device)[:count]].clone()
|
| 101 |
+
for _ in range(iters):
|
| 102 |
+
ids = _nearest(x, c)
|
| 103 |
+
sums = torch.zeros_like(c)
|
| 104 |
+
sums.index_add_(0, ids, x)
|
| 105 |
+
cnt = torch.bincount(ids, minlength=count).float().clamp_min(1)[:, None]
|
| 106 |
+
c = torch.where(torch.bincount(ids, minlength=count)[:, None] > 0, sums / cnt, c)
|
| 107 |
+
return c
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
# ----------------------------------------------------- per-expert sweep
|
| 111 |
+
def _assign(tn, c0, c0n, c1, c1n, refine):
|
| 112 |
+
"""Plain-L2 additive assignment. tn:[N,8]. Returns a,b (long)."""
|
| 113 |
+
a = (c0n[None, :] - 2 * tn @ c0.t()).argmin(1)
|
| 114 |
+
r = tn - c0[a]
|
| 115 |
+
b = (c1n[None, :] - 2 * r @ c1.t()).argmin(1)
|
| 116 |
+
for _ in range(refine):
|
| 117 |
+
r = tn - c1[b]
|
| 118 |
+
a = (c0n[None, :] - 2 * r @ c0.t()).argmin(1)
|
| 119 |
+
r = tn - c0[a]
|
| 120 |
+
b = (c1n[None, :] - 2 * r @ c1.t()).argmin(1)
|
| 121 |
+
return a, b
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
def sweep_expert(W, hinv, c0, c1, s, glob, *, col_block, refine):
|
| 125 |
+
"""LDLQ error-feedback assignment for one expert.
|
| 126 |
+
|
| 127 |
+
W:[N,K] fp32 target; hinv:[K,K] upper fp32; s:[N,K/128] fp32 (E4M3-valued)
|
| 128 |
+
scale; glob: scalar. Returns a,b uint8 [N,K/8] and reconstruction q [N,K]."""
|
| 129 |
+
N, K = W.shape
|
| 130 |
+
w = W.clone()
|
| 131 |
+
q = torch.empty_like(w)
|
| 132 |
+
a_full = torch.empty(N, K // 8, dtype=torch.uint8, device=W.device)
|
| 133 |
+
b_full = torch.empty(N, K // 8, dtype=torch.uint8, device=W.device)
|
| 134 |
+
c0n = c0.square().sum(1)
|
| 135 |
+
c1n = c1.square().sum(1)
|
| 136 |
+
for j0 in range(0, K, col_block):
|
| 137 |
+
j1 = min(j0 + col_block, K)
|
| 138 |
+
blk = w[:, j0:j1]
|
| 139 |
+
qhat = torch.empty_like(blk)
|
| 140 |
+
for gg in range((j1 - j0) // 8):
|
| 141 |
+
c = j0 + gg * 8
|
| 142 |
+
sden = (glob * s[:, c // 128]).clamp_min(1e-20)[:, None] # [N,1]
|
| 143 |
+
tn = blk[:, gg * 8:gg * 8 + 8] / sden
|
| 144 |
+
a, b = _assign(tn, c0, c0n, c1, c1n, refine)
|
| 145 |
+
qhat[:, gg * 8:gg * 8 + 8] = sden * (c0[a] + c1[b])
|
| 146 |
+
a_full[:, c // 8] = a.to(torch.uint8)
|
| 147 |
+
b_full[:, c // 8] = b.to(torch.uint8)
|
| 148 |
+
q[:, j0:j1] = qhat
|
| 149 |
+
err = (blk - qhat) @ torch.linalg.inv(hinv[j0:j1, j0:j1])
|
| 150 |
+
if j1 < K:
|
| 151 |
+
w[:, j1:] -= err @ hinv[j0:j1, j1:]
|
| 152 |
+
return a_full, b_full, q
|
| 153 |
+
|
| 154 |
+
|
| 155 |
+
def refit_scales(W, a, b, c0, c1, glob, hdiag, K):
|
| 156 |
+
"""LS per-(row,128block) scale given codes; E4M3-quantized. Returns [N,K/128]."""
|
| 157 |
+
N = W.shape[0]
|
| 158 |
+
u = (c0[a.long()] + c1[b.long()]).reshape(N, K // 8, 8) * glob # unit recon
|
| 159 |
+
Wg = W.reshape(N, K // 8, 8)
|
| 160 |
+
hd = hdiag.reshape(K // 8, 8)[None] # [1,K/8,8]
|
| 161 |
+
num = (hd * Wg * u).reshape(N, K // 128, 128).sum(-1)
|
| 162 |
+
den = (hd * u * u).reshape(N, K // 128, 128).sum(-1).clamp_min(1e-20)
|
| 163 |
+
s = (num / den)
|
| 164 |
+
s = s.clamp_min(0).to(torch.float8_e4m3fn).float()
|
| 165 |
+
return s
|
| 166 |
+
|
| 167 |
+
|
| 168 |
+
@dataclass
|
| 169 |
+
class EncodedLayerProj:
|
| 170 |
+
c0: torch.Tensor # [256,8] fp32 (FP4 grid)
|
| 171 |
+
c1: torch.Tensor # [256,8] fp32; [E,256,8] for expert scope
|
| 172 |
+
glob: float
|
| 173 |
+
a: torch.Tensor # [E,N,K/8] uint8
|
| 174 |
+
b: torch.Tensor # [E,N,K/8] uint8
|
| 175 |
+
s: torch.Tensor # [E,N,K/128] fp32 (E4M3-valued)
|
| 176 |
+
N: int
|
| 177 |
+
K: int
|
| 178 |
+
cb_rel_l2: float
|
| 179 |
+
recon_rel_fro: float # weight-space, Hessian-weighted, over experts
|
| 180 |
+
scale_dtype: str = "fp8_e4m3"
|
| 181 |
+
codebook_dtype: str = "fp4_grid" # 'fp4_grid' (v2/v3) or 'fp16' (free atoms)
|
| 182 |
+
|
| 183 |
+
|
| 184 |
+
def _e4m3(x):
|
| 185 |
+
return x.to(torch.float8_e4m3fn).float()
|
| 186 |
+
|
| 187 |
+
|
| 188 |
+
def fit_layer_projection(W, H, *, col_block=128, cb_iters=8, sweep_passes=2,
|
| 189 |
+
refine=1, subsample_per_expert=20000, seed=0,
|
| 190 |
+
device="cuda", verbose=False, global_scale=None,
|
| 191 |
+
codebook_scope="layer"):
|
| 192 |
+
"""W: list of E weight targets [N,K] (fp16/fp32, CPU or GPU).
|
| 193 |
+
H: list of E Hessians [K,K] (fp32). Returns EncodedLayerProj."""
|
| 194 |
+
dev = torch.device(device)
|
| 195 |
+
E = len(W)
|
| 196 |
+
N, K = W[0].shape
|
| 197 |
+
gen = torch.Generator(device=dev).manual_seed(seed)
|
| 198 |
+
if codebook_scope not in ('layer','expert'):
|
| 199 |
+
raise ValueError('Unknown codebook scope')
|
| 200 |
+
if codebook_scope == 'expert':
|
| 201 |
+
# Retain one global scale per projection so the v2 index/scale layout
|
| 202 |
+
# and arithmetic stay unchanged. Only the dictionaries gain an axis.
|
| 203 |
+
if global_scale is None:
|
| 204 |
+
pool=[]
|
| 205 |
+
for weight in W:
|
| 206 |
+
rms=weight.to(dev).float().reshape(N,K//128,128).square().mean(-1).sqrt().flatten()
|
| 207 |
+
pick=torch.randperm(len(rms),generator=gen,device=dev)[:4096]
|
| 208 |
+
pool.append(rms[pick])
|
| 209 |
+
global_scale=float(torch.cat(pool).median().clamp_min(1e-8))
|
| 210 |
+
del pool,rms
|
| 211 |
+
results=[]
|
| 212 |
+
for e in range(E):
|
| 213 |
+
result=fit_layer_projection([W[e]],[H[e]],col_block=col_block,
|
| 214 |
+
cb_iters=cb_iters,sweep_passes=sweep_passes,refine=refine,
|
| 215 |
+
subsample_per_expert=subsample_per_expert,seed=seed+e,
|
| 216 |
+
device=device,verbose=verbose,global_scale=global_scale)
|
| 217 |
+
results.append(result)
|
| 218 |
+
return EncodedLayerProj(torch.stack([r.c0 for r in results]),
|
| 219 |
+
torch.stack([r.c1 for r in results]),float(global_scale),
|
| 220 |
+
torch.cat([r.a for r in results]),torch.cat([r.b for r in results]),
|
| 221 |
+
torch.cat([r.s for r in results]),N,K,
|
| 222 |
+
sum(r.cb_rel_l2 for r in results)/E,
|
| 223 |
+
sum(r.recon_rel_fro for r in results)/E)
|
| 224 |
+
|
| 225 |
+
# ---- Phase A: hinv, hdiag, init scale, global
|
| 226 |
+
hinv, hdiag, s = [], [], []
|
| 227 |
+
rms_pool = []
|
| 228 |
+
for e in range(E):
|
| 229 |
+
We = W[e].to(dev).float()
|
| 230 |
+
hd = H[e].to(dev).float().diagonal().clamp_min(1e-12)
|
| 231 |
+
hinv.append(cholesky_inv_upper(H[e].to(dev).float()).half())
|
| 232 |
+
hdiag.append(hd.half())
|
| 233 |
+
blk_rms = We.reshape(N, K // 128, 128).square().mean(-1).sqrt() # [N,K/128]
|
| 234 |
+
s.append(blk_rms) # tmp (pre-global)
|
| 235 |
+
rms_pool.append(blk_rms.flatten()[torch.randperm(N * (K // 128),
|
| 236 |
+
generator=gen, device=dev)[:4096]])
|
| 237 |
+
del We
|
| 238 |
+
glob = (float(torch.cat(rms_pool).median().clamp_min(1e-8).item())
|
| 239 |
+
if global_scale is None else float(global_scale))
|
| 240 |
+
if not __import__('math').isfinite(glob) or glob <= 0:
|
| 241 |
+
raise ValueError('Global scale must be finite and positive')
|
| 242 |
+
for e in range(E):
|
| 243 |
+
s[e] = _e4m3(s[e] / glob) # E4M3 scale
|
| 244 |
+
|
| 245 |
+
# ---- Phase B: importance-weighted subsample of normalized group vectors
|
| 246 |
+
xs, ws = [], []
|
| 247 |
+
for e in range(E):
|
| 248 |
+
We = W[e].to(dev).float().reshape(N, K // 8, 8)
|
| 249 |
+
sden = (glob * s[e]).clamp_min(1e-20) # [N,K/128]
|
| 250 |
+
sden = sden.repeat_interleave(16, dim=1) # [N,K/8]
|
| 251 |
+
tn = (We / sden[:, :, None]).reshape(-1, 8) # [N*K/8,8]
|
| 252 |
+
imp = (hdiag[e].float().reshape(K // 8, 8).mean(1)[None, :]
|
| 253 |
+
* (s[e].repeat_interleave(16, dim=1) ** 2)).reshape(-1) # [N*K/8]
|
| 254 |
+
m = tn.shape[0]
|
| 255 |
+
idx = torch.randperm(m, generator=gen, device=dev)[:subsample_per_expert]
|
| 256 |
+
xs.append(tn[idx])
|
| 257 |
+
ws.append(imp[idx])
|
| 258 |
+
del We, tn, imp
|
| 259 |
+
xs = torch.cat(xs)
|
| 260 |
+
ws = torch.cat(ws).clamp_min(1e-12)
|
| 261 |
+
c0, c1, cb_rel = fit_codebooks(xs, ws, iters=cb_iters, seed=seed)
|
| 262 |
+
if verbose:
|
| 263 |
+
print(f" codebook subsample rel-L2 {cb_rel:.4f} (M={xs.shape[0]})",
|
| 264 |
+
flush=True)
|
| 265 |
+
del xs, ws
|
| 266 |
+
|
| 267 |
+
# ---- Phase D: per-expert error-feedback sweep + scale refit
|
| 268 |
+
a_all = torch.empty(E, N, K // 8, dtype=torch.uint8)
|
| 269 |
+
b_all = torch.empty(E, N, K // 8, dtype=torch.uint8)
|
| 270 |
+
s_all = torch.empty(E, N, K // 128, dtype=torch.float32)
|
| 271 |
+
num_w = den_w = 0.0
|
| 272 |
+
for e in range(E):
|
| 273 |
+
We = W[e].to(dev).float()
|
| 274 |
+
hv = hinv[e].float()
|
| 275 |
+
hd = hdiag[e].float()
|
| 276 |
+
se = s[e]
|
| 277 |
+
a = b = q = None
|
| 278 |
+
for p in range(sweep_passes):
|
| 279 |
+
a, b, q = sweep_expert(We, hv, c0, c1, se, glob,
|
| 280 |
+
col_block=col_block, refine=refine)
|
| 281 |
+
if p + 1 < sweep_passes:
|
| 282 |
+
se = refit_scales(We, a, b, c0, c1, glob, hd, K)
|
| 283 |
+
se = refit_scales(We, a, b, c0, c1, glob, hd, K)
|
| 284 |
+
a, b, q = sweep_expert(We, hv, c0, c1, se, glob,
|
| 285 |
+
col_block=col_block, refine=refine)
|
| 286 |
+
a_all[e], b_all[e], s_all[e] = a.cpu(), b.cpu(), se.cpu()
|
| 287 |
+
d = (We - q)
|
| 288 |
+
# Half-stored Hessians can lose positive semidefiniteness. Evaluate
|
| 289 |
+
# against the stabilized positive metric used by the sweep instead:
|
| 290 |
+
# H_eff^-1 = U.T @ U, hence d H_eff d.T = ||d U^-1||^2.
|
| 291 |
+
num_w += positive_quadratic(d, hv)
|
| 292 |
+
den_w += positive_quadratic(We, hv)
|
| 293 |
+
del We, hv, q, d
|
| 294 |
+
hinv[e] = None
|
| 295 |
+
recon = (num_w / max(den_w, 1e-20)) ** 0.5
|
| 296 |
+
return EncodedLayerProj(c0.cpu(), c1.cpu(), glob, a_all, b_all, s_all,
|
| 297 |
+
N, K, cb_rel, recon)
|
| 298 |
+
|
| 299 |
+
|
| 300 |
+
def positive_quadratic(weight, inverse_cholesky):
|
| 301 |
+
"""Nonnegative squared norm under the regularized fitting Hessian."""
|
| 302 |
+
z = torch.linalg.solve_triangular(inverse_cholesky.T, weight.T,
|
| 303 |
+
upper=False)
|
| 304 |
+
value = float(z.double().square().sum())
|
| 305 |
+
if not __import__('math').isfinite(value):
|
| 306 |
+
raise ValueError('Nonfinite regularized Hessian diagnostic')
|
| 307 |
+
return value
|
| 308 |
+
|
| 309 |
+
|
| 310 |
+
def reconstruct(enc: EncodedLayerProj, e: int, device="cpu") -> torch.Tensor:
|
| 311 |
+
"""Decode one expert's weight [N,K] from stored codes/scale/codebooks."""
|
| 312 |
+
dev = torch.device(device)
|
| 313 |
+
c0, c1 = enc.c0.to(dev), enc.c1.to(dev)
|
| 314 |
+
if c0.ndim == 3:
|
| 315 |
+
c0, c1 = c0[e], c1[e]
|
| 316 |
+
a = enc.a[e].to(dev).long()
|
| 317 |
+
b = enc.b[e].to(dev).long()
|
| 318 |
+
s = enc.s[e].to(dev).repeat_interleave(16, dim=1) # [N,K/8]
|
| 319 |
+
cb = (c0[a] + c1[b]) # [N,K/8,8]
|
| 320 |
+
W = (enc.glob * s[:, :, None] * cb).reshape(enc.N, enc.K)
|
| 321 |
+
return W
|
reproduce/source/btx53/arvq88/experiment.py
ADDED
|
@@ -0,0 +1,129 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Matched layer ablation for shared/expert books and fixed/alternating indices.
|
| 2 |
+
|
| 3 |
+
Waits for unoccupied GPUs; never stops another job. Exports local experimental
|
| 4 |
+
weights only. Full-model serving evaluation must precede a production campaign.
|
| 5 |
+
"""
|
| 6 |
+
import argparse
|
| 7 |
+
import fcntl
|
| 8 |
+
import json
|
| 9 |
+
import os
|
| 10 |
+
import shutil
|
| 11 |
+
import subprocess
|
| 12 |
+
import sys
|
| 13 |
+
import time
|
| 14 |
+
import traceback
|
| 15 |
+
from pathlib import Path
|
| 16 |
+
|
| 17 |
+
ROOT=Path(__file__).resolve().parents[1]
|
| 18 |
+
sys.path[:0]=[str(ROOT),str(ROOT/'tools')]
|
| 19 |
+
from arvq88.inputs import write
|
| 20 |
+
from arvq88.jobs import run_jobs
|
| 21 |
+
from arvq88.pack import export_layer,sha
|
| 22 |
+
from arvq88.activation import ARITHMETIC
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
def summary(work,layers):
|
| 26 |
+
rows={}
|
| 27 |
+
for variant in ['shared_fixed','shared_alternating','expert_fixed','expert_alternating']:
|
| 28 |
+
rows[variant]={}
|
| 29 |
+
for L in layers:
|
| 30 |
+
p=work/variant/f'layer_{L:05d}'/'report.json'
|
| 31 |
+
if not p.exists():continue
|
| 32 |
+
r=json.loads(p.read_text())
|
| 33 |
+
rows[variant][str(L)]={k:r[k] for k in ['initial_output_rel','serialized_output_rel','initial_audit_rel','serialized_audit_rel','audit_passed','audit_candidate_accepted','steps_run','codebook_scope','indices_frozen','index_reassignment']}
|
| 34 |
+
write(work/'comparison.json',{'variants':rows,'scope':'Same cold allocation, corpus splits, source weights and PV budgets; isolated layer ablation. Not full-model task quality.',
|
| 35 |
+
'allocation_next_step':'Use --reap-metric arvq --mm-acts with all-expert fitted candidates, after serving parity and task evaluation.'})
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def main():
|
| 39 |
+
ap=argparse.ArgumentParser(description=__doc__)
|
| 40 |
+
ap.add_argument('--base',default='/tmp/glm53-vision-trace-refit')
|
| 41 |
+
ap.add_argument('--workdir',default='/tmp/glm53-arvq-expert-ablation')
|
| 42 |
+
ap.add_argument('--layers',default='3,40,70');ap.add_argument('--gpus',default='0,1,2,3,4,5,6,7')
|
| 43 |
+
ap.add_argument('--steps',type=int,default=200);ap.add_argument('--reassign-every',type=int,default=40)
|
| 44 |
+
ap.add_argument('--check',action='store_true')
|
| 45 |
+
a=ap.parse_args();base=Path(a.base);work=Path(a.workdir);work.mkdir(parents=True,exist_ok=True)
|
| 46 |
+
layers=[int(v) for v in a.layers.split(',')];gpus=[int(v) for v in a.gpus.split(',')]
|
| 47 |
+
if not layers or len(set(layers))!=len(layers) or any(L not in range(3,78) for L in layers):raise ValueError('Invalid layers')
|
| 48 |
+
if not gpus or len(set(gpus))!=len(gpus) or min(gpus)<0 or a.reassign_every<=0 or a.steps<40:raise ValueError('Invalid run budget')
|
| 49 |
+
config=json.loads((base/'config.json').read_text())
|
| 50 |
+
train=base/'train/capture/acts';captures=base/'trace_capture/acts'
|
| 51 |
+
for L in layers:
|
| 52 |
+
for p in [train/f'acts_layer{L}.pt',captures/f'acts_layer{L}.pt',base/'initial'/f'layer_{L:05d}/arvq-manifest.json']:
|
| 53 |
+
if not p.is_file():raise FileNotFoundError(p)
|
| 54 |
+
identity={'base':str(base),'assignment_sha256':sha(base/'assignment.json'),'source_sha256':sha(base/'source.json'),
|
| 55 |
+
'layers':layers,'steps':a.steps,'reassign_every':a.reassign_every,'format':'rvq256_256x8_expert',
|
| 56 |
+
'inputs_sha256':{str(p):sha(p) for L in layers for p in [train/f'acts_layer{L}.pt',captures/f'acts_layer{L}.pt',base/'initial'/f'layer_{L:05d}/arvq-manifest.json']},
|
| 57 |
+
'code_sha256':{name:sha(ROOT/'arvq88'/name) for name in ['pv.py','alternating.py','encoder.py','pack.py','fit.py']}}
|
| 58 |
+
if (work/'binding.json').exists() and json.loads((work/'binding.json').read_text())!=identity:
|
| 59 |
+
raise ValueError('Experiment inputs or code changed; choose a fresh workdir')
|
| 60 |
+
if a.check:
|
| 61 |
+
print('PREFLIGHT OK',json.dumps(identity),flush=True);return
|
| 62 |
+
lock=open(work/'.experiment.lock','a');fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB)
|
| 63 |
+
write(work/'binding.json',identity)
|
| 64 |
+
if (work/'complete.json').exists():return
|
| 65 |
+
def status(stage,**kw):write(work/'status.json',{'stage':stage,'updated_utc':time.strftime('%Y-%m-%dT%H:%M:%SZ',time.gmtime()),**kw});print(stage,json.dumps(kw),flush=True)
|
| 66 |
+
try:
|
| 67 |
+
status('waiting_for_free_gpus',gpus=gpus,layers=layers)
|
| 68 |
+
while True:
|
| 69 |
+
rows=subprocess.check_output(['nvidia-smi','--query-gpu=index,memory.used','--format=csv,noheader,nounits'],text=True)
|
| 70 |
+
mem={int(i):int(m) for i,m in (r.split(',') for r in rows.splitlines())}
|
| 71 |
+
if any(g not in mem for g in gpus):raise ValueError('Selected GPU does not exist')
|
| 72 |
+
if all(mem[g]<1024 for g in gpus):break
|
| 73 |
+
time.sleep(30)
|
| 74 |
+
expert=work/'expert_initial';expert.mkdir(exist_ok=True)
|
| 75 |
+
for name in ['assignment.json','source.json']:shutil.copy2(base/name,expert/name)
|
| 76 |
+
config.update(workdir=str(expert),calib=str(train),codebook_scope='expert',do_upload=False)
|
| 77 |
+
write(expert/'run_config.json',config)
|
| 78 |
+
status('fitting_per_expert_books',layers=layers)
|
| 79 |
+
pending=[]
|
| 80 |
+
for L in layers:
|
| 81 |
+
folder=expert/'initial'/f'layer_{L:05d}'
|
| 82 |
+
path=folder/'arvq-manifest.json'
|
| 83 |
+
if path.exists():
|
| 84 |
+
m=json.loads(path.read_text());fit_id=m['source_identity']
|
| 85 |
+
expected=json.loads((expert/'assignment.json').read_text())['layers'][str(L)]['cold_5750']
|
| 86 |
+
if (m['cold_expert_ids']!=expected or fit_id['capture_sha256']!=sha(train/f'acts_layer{L}.pt')
|
| 87 |
+
or fit_id['source']!=json.loads((expert/'source.json').read_text())
|
| 88 |
+
or fit_id['settings'].get('codebook_scope')!='expert'):
|
| 89 |
+
raise ValueError('Completed fit inputs do not match experiment')
|
| 90 |
+
for entry in m['files'].values():
|
| 91 |
+
if sha(folder/entry['file'])!=entry['sha256']:raise ValueError('Completed fit weights changed')
|
| 92 |
+
print('REUSE VALIDATED FIT',L,flush=True)
|
| 93 |
+
else:pending.append(L)
|
| 94 |
+
run_jobs([(f'fit-{L}',lambda gpu,L=L:[sys.executable,'-u',str(ROOT/'arvq88/fit.py'),str(expert),str(L),f'cuda:{gpu}']) for L in pending],gpus,work/'logs')
|
| 95 |
+
jobs=[]
|
| 96 |
+
for variant in ['shared_fixed','shared_alternating','expert_fixed','expert_alternating']:
|
| 97 |
+
initial=expert/'initial' if variant.startswith('expert') else base/'initial'
|
| 98 |
+
reassign=a.reassign_every if variant.endswith('alternating') else 0
|
| 99 |
+
for L in layers:
|
| 100 |
+
out=work/variant/f'layer_{L:05d}'
|
| 101 |
+
if (out/'report.json').exists():continue
|
| 102 |
+
jobs.append((f'{variant}-{L}',lambda gpu,L=L,out=out,initial=initial,reassign=reassign:[
|
| 103 |
+
sys.executable,'-u',str(ROOT/'arvq88/pv.py'),'--arvq-dir',str(initial),
|
| 104 |
+
'--layer',str(L),'--experts','0','--steps',str(a.steps),'--device',f'cuda:{gpu}',
|
| 105 |
+
'--src-cache',config['src_cache'],'--source-meta',str(base/'source.json'),
|
| 106 |
+
'--calib',str(captures),'--codebook-lr','0.003','--scale-lr','0.002',
|
| 107 |
+
'--arithmetic',ARITHMETIC,'--reassign-every',str(reassign),'--out',str(out)]))
|
| 108 |
+
status('matched_pv_ablation',variants=4,layers=layers)
|
| 109 |
+
run_jobs(jobs,gpus,work/'logs')
|
| 110 |
+
status('exporting_experimental_weights')
|
| 111 |
+
for variant in ['shared_fixed','shared_alternating','expert_fixed','expert_alternating']:
|
| 112 |
+
initial=expert/'initial' if variant.startswith('expert') else base/'initial'
|
| 113 |
+
for L in layers:
|
| 114 |
+
out=work/variant/f'layer_{L:05d}';r=json.loads((out/'report.json').read_text())
|
| 115 |
+
manifest=json.loads((initial/f'layer_{L:05d}/arvq-manifest.json').read_text())
|
| 116 |
+
stage=work/'exports'/variant/f'layer_{L:05d}';stage.mkdir(parents=True,exist_ok=True)
|
| 117 |
+
for proj in ['w13','w2']:
|
| 118 |
+
target=stage/f'{proj}.pt'
|
| 119 |
+
if not target.exists():os.link(out/'offline_weights.pt',target)
|
| 120 |
+
write(stage/'arvq-manifest.json',{**manifest,'fit':variant,'pv_report':r})
|
| 121 |
+
export_layer(stage,work/'packed'/variant,L)
|
| 122 |
+
summary(work,layers)
|
| 123 |
+
write(work/'complete.json',{'uploaded':False,'comparison':str(work/'comparison.json'),'serving_validation':'pending'})
|
| 124 |
+
status('complete',uploaded=False)
|
| 125 |
+
except BaseException:
|
| 126 |
+
(work/'run.failed').write_text(traceback.format_exc());status('failed',report=str(work/'run.failed'));raise
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
if __name__=='__main__':main()
|
reproduce/source/btx53/arvq88/final_publish.py
ADDED
|
@@ -0,0 +1,69 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Publish final PV artifacts, verify them, then remove obsolete repo files."""
|
| 2 |
+
import hashlib,json
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
from .pack import sha
|
| 5 |
+
from .inputs import write
|
| 6 |
+
from .validation import require_audit
|
| 7 |
+
|
| 8 |
+
def inventory(root):
|
| 9 |
+
result={}
|
| 10 |
+
for p in sorted(root.rglob('*')):
|
| 11 |
+
if not p.is_file():continue
|
| 12 |
+
name=str(p.relative_to(root))
|
| 13 |
+
# upload_folder ignores .cache/huggingface/ by default, so it never reaches the remote;
|
| 14 |
+
# including it here would make verify() fail on a file that was intentionally not uploaded.
|
| 15 |
+
if name.startswith('.cache/'):continue
|
| 16 |
+
size=p.stat().st_size
|
| 17 |
+
entry={'size':size,'sha256':sha(p)}
|
| 18 |
+
if size<10*1024*1024:
|
| 19 |
+
data=p.read_bytes();entry['blob_id']=hashlib.sha1(f'blob {len(data)}\0'.encode()+data).hexdigest()
|
| 20 |
+
result[name]=entry
|
| 21 |
+
return result
|
| 22 |
+
|
| 23 |
+
def verify(api,repo,revision,expected):
|
| 24 |
+
files={s.rfilename:s for s in api.model_info(repo,revision=revision,files_metadata=True).siblings}
|
| 25 |
+
for name,e in expected.items():
|
| 26 |
+
if name not in files:raise RuntimeError(f'Remote file missing: {name}')
|
| 27 |
+
s=files[name]
|
| 28 |
+
if s.size!=e['size']:raise RuntimeError(f'Remote size mismatch: {name}')
|
| 29 |
+
if s.lfs:
|
| 30 |
+
if s.lfs.sha256!=e['sha256']:raise RuntimeError(f'Remote SHA mismatch: {name}')
|
| 31 |
+
elif s.blob_id!=e.get('blob_id'):raise RuntimeError(f'Remote Git blob mismatch: {name}')
|
| 32 |
+
return files
|
| 33 |
+
|
| 34 |
+
def publish(args,hub,work):
|
| 35 |
+
ck=work/'checkpoint';pv=json.loads((work/'pv_report.json').read_text());gates=json.loads((ck/'gates_report.json').read_text())
|
| 36 |
+
if not pv['passed'] or set(pv['layers_completed'])!=set(range(3,78)) or not gates['passed']:raise RuntimeError('Final PV/gates incomplete')
|
| 37 |
+
if set(map(int,pv.get('layers',{})))!=set(range(3,78)):raise RuntimeError('Missing per-layer audits')
|
| 38 |
+
for report in pv['layers'].values():require_audit(report)
|
| 39 |
+
for name,digest in {**gates['cold_file_sha256'],**gates['metadata_sha256']}.items():
|
| 40 |
+
if sha(ck/name)!=digest:raise RuntimeError('Checkpoint changed after gates')
|
| 41 |
+
return publish_verified(args,hub,work,'complete GLM-5.3 NVFP4 / ARVQ 8+8 PV-tuned hybrid')
|
| 42 |
+
|
| 43 |
+
def publish_verified(args,hub,work,description):
|
| 44 |
+
from huggingface_hub import CommitOperationDelete
|
| 45 |
+
ck=work/'checkpoint'
|
| 46 |
+
expected=inventory(ck);write(work/'final_upload_inventory.json',expected)
|
| 47 |
+
from huggingface_hub.errors import RepositoryNotFoundError
|
| 48 |
+
try:before=hub.api.model_info(args.dst,files_metadata=True)
|
| 49 |
+
except RepositoryNotFoundError:
|
| 50 |
+
hub.api.create_repo(repo_id=args.dst,repo_type='model',exist_ok=True)
|
| 51 |
+
before=hub.api.model_info(args.dst,files_metadata=True)
|
| 52 |
+
write(work/'repo_before_final.json',{'revision':before.sha,'files':[s.rfilename for s in before.siblings]})
|
| 53 |
+
commit=hub.api.upload_folder(repo_id=args.dst,folder_path=str(ck),revision='main',parent_commit=before.sha,commit_message='Publish '+description)
|
| 54 |
+
files=verify(hub.api,args.dst,commit.oid,expected)
|
| 55 |
+
write(work/'final_upload_verified.json',{'revision':commit.oid,'files':len(expected)})
|
| 56 |
+
revision=commit.oid
|
| 57 |
+
if args.clean_repo:
|
| 58 |
+
obsolete=sorted(set(files)-set(expected)-{'.gitattributes'})
|
| 59 |
+
write(work/'repo_cleanup_plan.json',{'verified_revision':revision,'delete':obsolete,'keep':sorted(expected)})
|
| 60 |
+
if obsolete:
|
| 61 |
+
cleanup=hub.api.create_commit(repo_id=args.dst,revision='main',parent_commit=revision,
|
| 62 |
+
operations=[CommitOperationDelete(path_in_repo=name) for name in obsolete],
|
| 63 |
+
commit_message='Keep only current 8+8 checkpoint, metadata and validation reports')
|
| 64 |
+
revision=cleanup.oid
|
| 65 |
+
remaining=verify(hub.api,args.dst,revision,expected)
|
| 66 |
+
if set(remaining)-set(expected)-{'.gitattributes'}:raise RuntimeError('Obsolete files remain after cleanup')
|
| 67 |
+
write(work/'all.done',{'repo':args.dst,'revision':revision,'files_verified':len(expected),'clean_repo':args.clean_repo})
|
| 68 |
+
print('FINAL PUBLISHED AND VERIFIED',args.dst,revision,flush=True)
|
| 69 |
+
return revision
|
reproduce/source/btx53/arvq88/fit.py
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""One 8+8 initial-fit worker. Source and cache identity are pinned by the driver."""
|
| 2 |
+
import sys,json,time,fcntl
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
ROOT=Path(__file__).resolve().parents[1];sys.path[:0]=[str(ROOT),str(ROOT/'tools')]
|
| 5 |
+
import torch
|
| 6 |
+
from arvq88 import encoder as enc
|
| 7 |
+
from arvq88.pack import export_layer
|
| 8 |
+
from arvq88.inputs import write
|
| 9 |
+
from arvqprep.workflow import sha256
|
| 10 |
+
from ingest import SourceMeta,load_expert_bf16
|
| 11 |
+
|
| 12 |
+
def fit(work,L,dev):
|
| 13 |
+
args=json.loads((work/'run_config.json').read_text());allocation=json.loads((work/'assignment.json').read_text())
|
| 14 |
+
cold=allocation['layers'][str(L)]['cold_5750'];base=work/'initial'/f'layer_{L:05d}';base.mkdir(parents=True,exist_ok=True)
|
| 15 |
+
lock=open(base/'.lock','a');fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB)
|
| 16 |
+
torch.set_num_threads(args['threads']);source=json.loads((work/'source.json').read_text());meta=SourceMeta(**source['meta'])
|
| 17 |
+
act_path=Path(args['calib'])/f'acts_layer{L}.pt'
|
| 18 |
+
settings={k:args[k] for k in ['cb_iters','sweep_passes','refine','col_block']}
|
| 19 |
+
settings['codebook_scope']=args.get('codebook_scope','layer')
|
| 20 |
+
fingerprint={'source':source,'capture_sha256':sha256(act_path),'cold_ids':cold,'encoder_sha256':sha256(Path(enc.__file__)),'settings':settings}
|
| 21 |
+
identity=base/'fit_identity.json'
|
| 22 |
+
if identity.exists() and json.loads(identity.read_text())!=fingerprint:raise ValueError('Cached initial fit inputs/settings changed')
|
| 23 |
+
if not identity.exists() and any(base.glob('*.pt')):raise ValueError('Unidentified fit cache; use import-cache command or fresh workdir')
|
| 24 |
+
write(identity,fingerprint)
|
| 25 |
+
if (base/'arvq-manifest.json').exists():
|
| 26 |
+
export_layer(base,work/'cold',L);return
|
| 27 |
+
acts=torch.load(act_path,weights_only=True,map_location='cpu');x=acts['x'].float();ids=acts['topk_ids'].long()
|
| 28 |
+
if 'pv_split' in acts:
|
| 29 |
+
train=acts['pv_split']==0;x=x[train];ids=ids[train]
|
| 30 |
+
if not len(x):raise ValueError('No training activations for initial fit')
|
| 31 |
+
files={};metrics={}
|
| 32 |
+
for proj in ['w13','w2']:
|
| 33 |
+
path=base/f'{proj}.pt';mp=base/f'{proj}.json'
|
| 34 |
+
if not path.exists() or not mp.exists():
|
| 35 |
+
W=[];H=[];start=time.time()
|
| 36 |
+
for eid in cold:
|
| 37 |
+
wg,wu,wd=[w.to(dev) for w in load_expert_bf16(str(Path(args['src_cache'])/f'layer_{L}'/f'expert_{eid}.safetensors'),meta)]
|
| 38 |
+
xr=x[(ids==eid).any(1)];xr=(x if len(xr)<32 else xr).to(dev)
|
| 39 |
+
H.append((enc.hessian_fc1(xr) if proj=='w13' else enc.hessian_fc2(xr,wg,wu)).half())
|
| 40 |
+
W.append((torch.cat([wg,wu]) if proj=='w13' else wd).half());del wg,wu,wd,xr
|
| 41 |
+
result=enc.fit_layer_projection(W,H,device=dev,verbose=True,seed=0,**fingerprint['settings'])
|
| 42 |
+
tmp=path.with_suffix('.tmp');torch.save({proj:vars(result)},tmp);tmp.replace(path)
|
| 43 |
+
write(mp,{'cb_rel_l2':result.cb_rel_l2,'hess_recon_rel_fro':result.recon_rel_fro,'diagnostic_metric':'regularized fitting Hessian; mean expert relative error for expert scope','seconds':time.time()-start})
|
| 44 |
+
del result,W,H;torch.cuda.empty_cache()
|
| 45 |
+
files[proj]={'file':path.name,'sha256':sha256(path)};metrics[proj]=json.loads(mp.read_text())
|
| 46 |
+
write(base/'arvq-manifest.json',{'layer':L,'cold_expert_ids':cold,'files':files,'per_proj':metrics,'no_rotation':True,'source_identity':fingerprint,'format':'rvq256_256x8'})
|
| 47 |
+
export_layer(base,work/'cold',L)
|
| 48 |
+
if __name__=='__main__':fit(Path(sys.argv[1]),int(sys.argv[2]),sys.argv[3])
|
reproduce/source/btx53/arvq88/gate_worker.py
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import sys
|
| 2 |
+
from pathlib import Path
|
| 3 |
+
root=Path(__file__).resolve().parents[1];sys.path[:0]=[str(root),str(root/'tools')]
|
| 4 |
+
from arvq88.checkpoint import gates
|
| 5 |
+
gates(Path(sys.argv[1]),[int(x) for x in sys.argv[2].split(',')],sys.argv[3],Path(sys.argv[4]))
|
reproduce/source/btx53/arvq88/gradient_indices.py
ADDED
|
@@ -0,0 +1,120 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Sparse proximal code proposals driven by routed-output gradients.
|
| 2 |
+
|
| 3 |
+
This is a layer-local optimizer, not full-model published PV-Tuning. Each
|
| 4 |
+
expert/projection proposal is evaluated against actual training output loss.
|
| 5 |
+
"""
|
| 6 |
+
import math
|
| 7 |
+
import torch
|
| 8 |
+
from .activation import ARITHMETIC, activation_ste, swiglu_ste
|
| 9 |
+
|
| 10 |
+
METHOD = 'routed_output_gradient_sparse_v1'
|
| 11 |
+
|
| 12 |
+
@torch.no_grad()
|
| 13 |
+
def propose_codes(p, e, grad, max_fraction=.01, trust_ratio=.01, target_ratio=.1):
|
| 14 |
+
"""Search single-book alternatives near a gradient step; cap count and norm.
|
| 15 |
+
|
| 16 |
+
The candidate pool is the highest-gradient 4*budget groups. Search all 256
|
| 17 |
+
choices in either book, with the other book fixed. Score using the proximal
|
| 18 |
+
linearization g*delta + ||delta||^2/(2*eta). Original FP8 weights are never used.
|
| 19 |
+
"""
|
| 20 |
+
w=p.weight(e).detach().reshape(-1,8);g=grad.detach().reshape(-1,8)
|
| 21 |
+
if not torch.isfinite(g).all():raise ValueError('Nonfinite index gradient')
|
| 22 |
+
gn=g.norm();wn=w.norm();budget=int(len(w)*max_fraction)
|
| 23 |
+
if budget<1 or float(gn)==0 or float(wn)==0:return None
|
| 24 |
+
eta=target_ratio*wn/gn
|
| 25 |
+
pool=min(len(w),4*budget)
|
| 26 |
+
selected=g.square().sum(1).topk(pool,sorted=False).indices
|
| 27 |
+
cb,s,glob=p.quantized(e);cb=cb.detach();factor=(s.detach().repeat_interleave(16,1)*glob.detach()).flatten()
|
| 28 |
+
old_a=p.codes_a[e].flatten();old_b=p.codes_b[e].flatten()
|
| 29 |
+
scores=[];aa=[];bb=[];norms=[]
|
| 30 |
+
for ix in selected.split(2048):
|
| 31 |
+
a=old_a[ix].long();b=old_b[ix].long();f=factor[ix,None]
|
| 32 |
+
current=w[ix];target=current-eta*g[ix]
|
| 33 |
+
# Distances need only [chunk,256], not [chunk,256,8].
|
| 34 |
+
best=[]
|
| 35 |
+
for book,other in ((cb[:256],cb[256+b]),(cb[256:],cb[a])):
|
| 36 |
+
residual=target-f*other
|
| 37 |
+
distances=f.square()*book.square().sum(1)[None,:]-2*f*(residual@book.T)
|
| 38 |
+
choice=distances.argmin(1)
|
| 39 |
+
candidate=f*(book[choice]+other);delta=candidate-current
|
| 40 |
+
cost=(g[ix]*delta).sum(1)+delta.square().sum(1)/(2*eta)
|
| 41 |
+
best.append((choice,cost,delta.square().sum(1)))
|
| 42 |
+
choose_a=best[0][1]<=best[1][1]
|
| 43 |
+
scores.append(torch.minimum(best[0][1],best[1][1]))
|
| 44 |
+
aa.append(torch.where(choose_a,best[0][0],a));bb.append(torch.where(choose_a,b,best[1][0]))
|
| 45 |
+
norms.append(torch.where(choose_a,best[0][2],best[1][2]))
|
| 46 |
+
score=torch.cat(scores);a=torch.cat(aa);b=torch.cat(bb);delta2=torch.cat(norms)
|
| 47 |
+
valid=((score<0)&(delta2>0)).nonzero().flatten()
|
| 48 |
+
if not len(valid):return None
|
| 49 |
+
order=valid[score[valid].argsort()[:budget]]
|
| 50 |
+
# Strict trust radius for the reconstructed projection, not latent values.
|
| 51 |
+
order=order[delta2[order].cumsum(0)<=(trust_ratio*wn).square()]
|
| 52 |
+
if not len(order):return None
|
| 53 |
+
ix=selected[order]
|
| 54 |
+
return {'positions':ix,'a':a[order].to(old_a.dtype),'b':b[order].to(old_b.dtype),
|
| 55 |
+
'old_a':old_a[ix].clone(),'old_b':old_b[ix].clone(),
|
| 56 |
+
'predicted_change':float(score[order].sum()),
|
| 57 |
+
'relative_weight_change':float(delta2[order].sum().sqrt()/wn)}
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
def expert_output(z,w13,w2,arithmetic):
|
| 61 |
+
if arithmetic==ARITHMETIC:
|
| 62 |
+
gu=activation_ste(z)@w13.T
|
| 63 |
+
return activation_ste(swiglu_ste(gu))@w2.T
|
| 64 |
+
gu=z@w13.T;gate,up=gu.chunk(2,-1)
|
| 65 |
+
return (torch.nn.functional.silu(gate)*up)@w2.T
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
@torch.no_grad()
|
| 69 |
+
def reassign_indices(p13,p2,x,rows,cold,args,target,prediction,routes):
|
| 70 |
+
"""Sequential expert/projection updates with exact routed residual checks.
|
| 71 |
+
|
| 72 |
+
Only the provided training rows enter gradients, candidate scoring or acceptance.
|
| 73 |
+
The routed residual includes every cold expert, so gates and cross-expert
|
| 74 |
+
error interactions are included. Validation/audit are inaccessible here.
|
| 75 |
+
"""
|
| 76 |
+
ref=target[rows];residual=prediction(rows)-ref
|
| 77 |
+
denominator=ref.square().sum().clamp_min(1e-20)
|
| 78 |
+
before=float((residual.square().sum()/denominator).sqrt())
|
| 79 |
+
accepted=proposed=changed=skipped=0;details=[]
|
| 80 |
+
for e,eid in enumerate(cold):
|
| 81 |
+
weights=routes[e][rows];use=weights!=0
|
| 82 |
+
if int(use.sum())<args.reassign_min_rows:skipped+=1;continue
|
| 83 |
+
z=x[rows[use]];gate=weights[use,None]
|
| 84 |
+
for p,name in ((p13,'w13'),(p2,'w2')):
|
| 85 |
+
w13=p13.weight(e).detach();w2=p2.weight(e).detach()
|
| 86 |
+
with torch.enable_grad():
|
| 87 |
+
weight=w13 if name=='w13' else w2;weight.requires_grad_(True)
|
| 88 |
+
old_output=expert_output(z,w13,w2,args.arithmetic)
|
| 89 |
+
# At current weights this equals the actual coupled output loss.
|
| 90 |
+
routed=residual[use].detach()+gate*(old_output-old_output.detach())
|
| 91 |
+
loss=routed.square().sum()/denominator
|
| 92 |
+
grad,=torch.autograd.grad(loss,weight)
|
| 93 |
+
old_output=old_output.detach();w13=w13.detach();w2=w2.detach()
|
| 94 |
+
proposal=propose_codes(p,e,grad,args.reassign_max_fraction,
|
| 95 |
+
args.reassign_trust_ratio,args.reassign_target_ratio)
|
| 96 |
+
del grad,loss,routed,weight
|
| 97 |
+
if proposal is None:continue
|
| 98 |
+
proposed+=1;ix=proposal['positions']
|
| 99 |
+
a=p.codes_a[e].view(-1);b=p.codes_b[e].view(-1)
|
| 100 |
+
a[ix]=proposal['a'];b[ix]=proposal['b']
|
| 101 |
+
try:
|
| 102 |
+
new_output=expert_output(z,p13.weight(e).detach(),p2.weight(e).detach(),args.arithmetic)
|
| 103 |
+
candidate=residual[use]+gate*(new_output-old_output)
|
| 104 |
+
old_loss=float(residual[use].square().sum());new_loss=float(candidate.square().sum())
|
| 105 |
+
keep=math.isfinite(new_loss) and new_loss<old_loss-max(1e-12,old_loss*1e-7)
|
| 106 |
+
if keep:residual[use]=candidate;accepted+=1;changed+=len(ix)
|
| 107 |
+
else:a[ix]=proposal['old_a'];b[ix]=proposal['old_b']
|
| 108 |
+
except BaseException:
|
| 109 |
+
a[ix]=proposal['old_a'];b[ix]=proposal['old_b'];raise
|
| 110 |
+
details.append({'expert':eid,'projection':name,'accepted':keep,'changed_groups':len(ix),
|
| 111 |
+
'relative_weight_change':proposal['relative_weight_change'],
|
| 112 |
+
'training_before_local':old_loss,'training_candidate_local':new_loss})
|
| 113 |
+
# Recompute rather than trust accumulated residual arithmetic. Rollback is
|
| 114 |
+
# handled by the caller's full index snapshot if the net check regresses.
|
| 115 |
+
after=float(((prediction(rows)-ref).square().sum()/denominator).sqrt())
|
| 116 |
+
return {'method':METHOD,'training_before':before,'training_candidate':after,
|
| 117 |
+
'accepted':math.isfinite(after) and after<before,
|
| 118 |
+
'accepted_proposals':accepted,'proposals':proposed,'changed_groups':changed,
|
| 119 |
+
'skipped_experts':skipped,'max_changed_fraction':args.reassign_max_fraction,
|
| 120 |
+
'trust_ratio':args.reassign_trust_ratio,'details':details}
|
reproduce/source/btx53/arvq88/incremental_publish.py
ADDED
|
@@ -0,0 +1,122 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Atomic initial v3 checkpoint publication, then audited layer replacements."""
|
| 2 |
+
import json
|
| 3 |
+
import os
|
| 4 |
+
import shutil
|
| 5 |
+
import time
|
| 6 |
+
from pathlib import Path
|
| 7 |
+
from .inputs import write
|
| 8 |
+
from .pack import sha
|
| 9 |
+
from .validation import require_audit
|
| 10 |
+
from .final_publish import inventory,verify
|
| 11 |
+
|
| 12 |
+
|
| 13 |
+
def progress_card(repo,completed):
|
| 14 |
+
return f'''---
|
| 15 |
+
library_name: vllm
|
| 16 |
+
tags: [arvq, nvfp4, experimental]
|
| 17 |
+
---
|
| 18 |
+
# {repo.rsplit('/',1)[-1]}
|
| 19 |
+
|
| 20 |
+
Complete Vision NVFP4 / per-expert ARVQ 8+8 checkpoint, format
|
| 21 |
+
`rvq256_256x8_expert`, version 3. Requires the per-expert-book serving loader.
|
| 22 |
+
The new ARVQ-based 75% text / 25% multimodal allocation and all initial fits
|
| 23 |
+
are complete. **{len(completed)}/75 layers have been replaced by audited PV
|
| 24 |
+
weights.** Remaining layers retain their initial fits. See pv_progress.json.
|
| 25 |
+
|
| 26 |
+
Each cold expert has separate two-book FP4 dictionaries for gate/up and down.
|
| 27 |
+
PV emulates four FP4 activation planes and FP16 boundaries, with discrete
|
| 28 |
+
Hessian-aware index proposals every 40 steps. Validation selects the best
|
| 29 |
+
state; audit regression restores the initial fit. Reference source is original
|
| 30 |
+
GLM-5.3 FP8; activations use the saved NVFP4-teacher conversation corpus.
|
| 31 |
+
Hot experts, language backbone and BF16 MTP come from RadixArk NVFP4; vision
|
| 32 |
+
comes from the existing vision hybrid. Indices, cold slots and hot allocation
|
| 33 |
+
are coherent at every published revision. Each layer replacement is atomic.
|
| 34 |
+
|
| 35 |
+
Offline audits/packing checks are not full-model quality or native-SM120 parity
|
| 36 |
+
benchmarks. Independent-book serving quality validation remains pending.
|
| 37 |
+
'''
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
def commit_snapshot(hub,repo,root,parent,message,delete=()):
|
| 41 |
+
from huggingface_hub import CommitOperationAdd,CommitOperationDelete
|
| 42 |
+
expected=inventory(root)
|
| 43 |
+
operations=[CommitOperationAdd(path_in_repo=name,path_or_fileobj=str(Path(root)/name)) for name in expected]
|
| 44 |
+
# Upload LFS objects first, but expose all tensor/config changes in ONE commit.
|
| 45 |
+
hub.api.preupload_lfs_files(repo,operations,revision='main',num_threads=8)
|
| 46 |
+
result=hub.api.create_commit(repo_id=repo,revision='main',parent_commit=parent,
|
| 47 |
+
operations=operations+[CommitOperationDelete(path_in_repo=name) for name in delete],
|
| 48 |
+
commit_message=message,num_threads=8)
|
| 49 |
+
verify(hub.api,repo,result.oid,expected)
|
| 50 |
+
return result.oid
|
| 51 |
+
|
| 52 |
+
|
| 53 |
+
def publish_initial(args,hub,work):
|
| 54 |
+
ck=work/'checkpoint';gates=json.loads((ck/'gates_report.json').read_text())
|
| 55 |
+
if not gates['passed']:raise ValueError('Initial checkpoint gates failed')
|
| 56 |
+
for name,digest in {**gates['cold_file_sha256'],**gates['metadata_sha256']}.items():
|
| 57 |
+
if sha(ck/name)!=digest:raise ValueError('Checkpoint changed after initial gates')
|
| 58 |
+
remote=hub.api.model_info(args.dst,files_metadata=True)
|
| 59 |
+
local={str(p.relative_to(ck)) for p in ck.rglob('*') if p.is_file()}
|
| 60 |
+
obsolete=sorted({s.rfilename for s in remote.siblings}-local-{'.gitattributes'})
|
| 61 |
+
revision=commit_snapshot(hub,args.dst,ck,remote.sha,'Replace with complete per-expert ARVQ v3 initial fit; PV follows',obsolete)
|
| 62 |
+
state={'repo':args.dst,'revision':revision,'layers':[]}
|
| 63 |
+
write(work/'incremental_upload.json',state)
|
| 64 |
+
print('INITIAL CHECKPOINT UPLOADED AND VERIFIED',revision,flush=True)
|
| 65 |
+
return state
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
def layer_snapshot(work,L,completed,repo):
|
| 69 |
+
ck=work/'checkpoint'
|
| 70 |
+
try:m=json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text())
|
| 71 |
+
except (FileNotFoundError,json.JSONDecodeError):return None
|
| 72 |
+
r=m.get('pv_report')
|
| 73 |
+
if not r:return None
|
| 74 |
+
require_audit(r)
|
| 75 |
+
if not r.get('passed') or r.get('indices_frozen',True) or r.get('codebook_scope')!='expert':
|
| 76 |
+
raise ValueError('Unexpected PV recipe for incremental upload')
|
| 77 |
+
assignment=json.loads((work/'assignment.json').read_text())['layers'][str(L)]['cold_5750']
|
| 78 |
+
if m['cold_expert_ids']!=assignment or m.get('version')!=3:raise ValueError('Layer allocation/format mismatch')
|
| 79 |
+
stage=work/'upload_layers'/f'layer_{L:05d}';stage.mkdir(parents=True,exist_ok=True)
|
| 80 |
+
gates=json.loads((ck/'gates_report.json').read_text())
|
| 81 |
+
for entry in m['files'].values():
|
| 82 |
+
source=work/'cold'/entry['file']
|
| 83 |
+
if sha(source)!=entry['sha256']:raise ValueError('Packed layer hash mismatch')
|
| 84 |
+
dst=stage/entry['file']
|
| 85 |
+
if dst.exists():dst.unlink()
|
| 86 |
+
os.link(source,dst)
|
| 87 |
+
gates['cold_file_sha256'][entry['file']]=entry['sha256']
|
| 88 |
+
gates['scope']='Initial structural gates plus incremental packing/parity and PV audits; final full gates pending'
|
| 89 |
+
write(stage/'gates_report.json',gates)
|
| 90 |
+
write(stage/'cold_manifests'/f'layer-{L:03d}-manifest.json',m)
|
| 91 |
+
write(stage/'pv_layers'/f'layer_{L:03d}.json',r)
|
| 92 |
+
write(stage/'pv_progress.json',{'layers_pv_complete':sorted(completed),'total_layers':75,
|
| 93 |
+
'remaining_layers':'initial fit','format':'rvq256_256x8_expert'})
|
| 94 |
+
(stage/'README.md').write_text(progress_card(repo,completed))
|
| 95 |
+
return stage
|
| 96 |
+
|
| 97 |
+
|
| 98 |
+
def publish_layers(args,hub,work):
|
| 99 |
+
statefile=work/'incremental_upload.json'
|
| 100 |
+
state=json.loads(statefile.read_text()) if statefile.exists() else publish_initial(args,hub,work)
|
| 101 |
+
done=set(state['layers'])
|
| 102 |
+
while len(done)<75:
|
| 103 |
+
if (work/'run.failed').exists():raise RuntimeError('Fit/PV campaign failed; completed uploads retained')
|
| 104 |
+
changed=False
|
| 105 |
+
for L in range(3,78):
|
| 106 |
+
if L in done:continue
|
| 107 |
+
stage=layer_snapshot(work,L,done|{L},args.dst)
|
| 108 |
+
if stage is None:continue
|
| 109 |
+
revision=commit_snapshot(hub,args.dst,stage,state['revision'],f'Replace layer {L} with audited per-expert alternating PV ({len(done)+1}/75)')
|
| 110 |
+
# Keep local assembly aligned; atomic replaces avoid mutating initial hardlinks.
|
| 111 |
+
for source in stage.rglob('*'):
|
| 112 |
+
if not source.is_file():continue
|
| 113 |
+
dst=work/'checkpoint'/source.relative_to(stage);dst.parent.mkdir(parents=True,exist_ok=True)
|
| 114 |
+
tmp=dst.with_suffix(dst.suffix+'.replace')
|
| 115 |
+
if tmp.exists():tmp.unlink()
|
| 116 |
+
if source.suffix=='.safetensors':os.link(source,tmp)
|
| 117 |
+
else:shutil.copy2(source,tmp)
|
| 118 |
+
tmp.replace(dst)
|
| 119 |
+
done.add(L);state.update(revision=revision,layers=sorted(done));write(statefile,state)
|
| 120 |
+
print('PV LAYER UPLOADED AND VERIFIED',L,len(done),'/75',revision,flush=True);changed=True
|
| 121 |
+
if not changed:time.sleep(20)
|
| 122 |
+
write(work/'incremental_pv.done.json',state)
|
reproduce/source/btx53/arvq88/inputs.py
ADDED
|
@@ -0,0 +1,198 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Validated cache reuse, pinned Hub input fetch and REAP allocation."""
|
| 2 |
+
import json,os,shutil,sys,subprocess
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
from concurrent.futures import ThreadPoolExecutor
|
| 5 |
+
import numpy as np
|
| 6 |
+
from .local_source import local_directory,model_index,fingerprint,LocalShard
|
| 7 |
+
from arvqprep.workflow import sha256
|
| 8 |
+
ROOT=Path(__file__).resolve().parents[1]
|
| 9 |
+
LAYERS=list(range(3,78))
|
| 10 |
+
def write(path,obj):
|
| 11 |
+
path=Path(path);path.parent.mkdir(parents=True,exist_ok=True)
|
| 12 |
+
tmp=path.with_suffix(path.suffix+'.tmp');tmp.write_text(json.dumps(obj,indent=2));tmp.replace(path)
|
| 13 |
+
def validate_assignment(obj,hot_count=5750):
|
| 14 |
+
layers=obj.get('layers',obj)
|
| 15 |
+
if set(map(int,layers))!=set(LAYERS):raise ValueError('Allocation must cover layers 3–77')
|
| 16 |
+
result={}
|
| 17 |
+
for L in LAYERS:
|
| 18 |
+
d=layers[str(L)];cold=d.get('cold_5750',d.get('cold'))
|
| 19 |
+
if cold is None or len(set(cold))!=len(cold) or any(type(e)!=int or not 0<=e<256 for e in cold):raise ValueError(f'Invalid cold IDs layer {L}')
|
| 20 |
+
hot=sorted(set(range(256))-set(cold))
|
| 21 |
+
if 'hot' in d and sorted(d['hot'])!=hot:raise ValueError('Overlapping/missing expert assignment')
|
| 22 |
+
result[str(L)]={'cold_5750':sorted(cold),'hot':hot}
|
| 23 |
+
if sum(len(v['hot']) for v in result.values())!=hot_count:raise ValueError('Allocation hot budget mismatch')
|
| 24 |
+
return {'layers':result}
|
| 25 |
+
def allocate(scores,hot_count=5750,floor=8,cap=176):
|
| 26 |
+
scores=np.asarray(scores,dtype=np.float64)
|
| 27 |
+
if scores.shape!=(75,256) or not np.isfinite(scores).all() or (scores<0).any():raise ValueError('Invalid REAP score array')
|
| 28 |
+
if not 75*floor<=hot_count<=75*cap:raise ValueError('Infeasible hot budget')
|
| 29 |
+
norm=scores/np.maximum(scores.sum(1,keepdims=True),1e-30)
|
| 30 |
+
hot=[set(np.argsort(-row,kind='stable')[:floor].tolist()) for row in norm];placed=75*floor
|
| 31 |
+
for pos in np.argsort(-norm.ravel(),kind='stable'):
|
| 32 |
+
if placed==hot_count:break
|
| 33 |
+
l,e=divmod(int(pos),256)
|
| 34 |
+
if len(hot[l])<cap and e not in hot[l]:hot[l].add(e);placed+=1
|
| 35 |
+
return validate_assignment({'layers':{str(L):{'cold':sorted(set(range(256))-hot[i])} for i,L in enumerate(LAYERS)}},hot_count)
|
| 36 |
+
def token():
|
| 37 |
+
p=Path.home()/'.cache/huggingface/token'
|
| 38 |
+
return p.read_text().strip() if p.exists() else None
|
| 39 |
+
class Hub:
|
| 40 |
+
def __init__(self,cache):
|
| 41 |
+
from huggingface_hub import HfApi
|
| 42 |
+
self.api=HfApi(token=token());self.cache=Path(cache);self.revisions={}
|
| 43 |
+
def revision(self,repo):
|
| 44 |
+
if repo not in self.revisions:
|
| 45 |
+
local=local_directory(repo)
|
| 46 |
+
self.revisions[repo]=fingerprint(local) if local else self.api.model_info(repo).sha
|
| 47 |
+
return self.revisions[repo]
|
| 48 |
+
def file(self,repo,name):
|
| 49 |
+
local=local_directory(repo)
|
| 50 |
+
if local:
|
| 51 |
+
p=local/name
|
| 52 |
+
if not p.is_file():raise FileNotFoundError(p)
|
| 53 |
+
return p
|
| 54 |
+
from huggingface_hub import hf_hub_download
|
| 55 |
+
return Path(hf_hub_download(repo,name,revision=self.revision(repo),cache_dir=str(self.cache),token=token()))
|
| 56 |
+
def snapshot(self,repo,dest):
|
| 57 |
+
local=local_directory(repo)
|
| 58 |
+
if local:return local
|
| 59 |
+
from huggingface_hub import snapshot_download
|
| 60 |
+
snapshot_download(repo,revision=self.revision(repo),local_dir=str(dest),token=token(),max_workers=8)
|
| 61 |
+
return Path(dest)
|
| 62 |
+
def complete_snapshot(path):
|
| 63 |
+
try:
|
| 64 |
+
wm=json.loads((Path(path)/'model.safetensors.index.json').read_text())['weight_map']
|
| 65 |
+
return bool(wm) and all((Path(path)/f).is_file() for f in set(wm.values()))
|
| 66 |
+
except (FileNotFoundError,KeyError,ValueError):return False
|
| 67 |
+
def valid_capture(path):
|
| 68 |
+
import torch
|
| 69 |
+
try:
|
| 70 |
+
d=torch.load(path,weights_only=True,map_location='cpu');n=len(d['x'])
|
| 71 |
+
return n>=32 and d['x'].shape==(n,6144) and d['topk_ids'].shape==(n,8) and d['topk_weights'].shape==(n,8) and bool(torch.isfinite(d['x']).all()) and bool(torch.isfinite(d['topk_weights']).all()) and bool(((d['topk_ids']>=0)&(d['topk_ids']<256)).all())
|
| 72 |
+
except (OSError,KeyError,ValueError,RuntimeError,EOFError):return False
|
| 73 |
+
def build_corpus(args,hub):
|
| 74 |
+
"""Download a bounded, deterministic text corpus when local token shards are absent."""
|
| 75 |
+
from datasets import load_dataset
|
| 76 |
+
from transformers import AutoTokenizer
|
| 77 |
+
target=Path(args.calib_tokens);target.mkdir(parents=True,exist_ok=True)
|
| 78 |
+
marker=target/'corpus.building.json'
|
| 79 |
+
if marker.exists():
|
| 80 |
+
for shard in target.glob('shard_*.npy'):shard.unlink()
|
| 81 |
+
write(marker,{'builder':'arvq88 fallback corpus','target_tokens':args.calib_token_count})
|
| 82 |
+
tok=AutoTokenizer.from_pretrained(args.src,revision=None if local_directory(args.src) else hub.revision(args.src),token=token(),trust_remote_code=True)
|
| 83 |
+
specs=[('m-a-p/CodeFeedback-Filtered-Instruction',.4),('tatsu-lab/alpaca',.4),('databricks/databricks-dolly-15k',.2)]
|
| 84 |
+
pending=[];number=0;actual={};revisions={}
|
| 85 |
+
for repo,fraction in specs:
|
| 86 |
+
budget=int(args.calib_token_count*fraction);count=0
|
| 87 |
+
revisions[repo]=hub.api.dataset_info(repo).sha
|
| 88 |
+
data=load_dataset(repo,revision=revisions[repo],split='train',streaming=True,token=token()).shuffle(seed=42,buffer_size=10000)
|
| 89 |
+
while count<budget:
|
| 90 |
+
before=count
|
| 91 |
+
for row in data:
|
| 92 |
+
text='\n'.join(str(v) for v in row.values() if isinstance(v,(str,list)))
|
| 93 |
+
ids=tok.encode(text,add_special_tokens=False)+[tok.eos_token_id]
|
| 94 |
+
ids=ids[:budget-count];pending.extend(ids);count+=len(ids)
|
| 95 |
+
while len(pending)>=1_000_000:
|
| 96 |
+
np.save(target/f'shard_{number:05d}.npy',np.asarray(pending[:1_000_000],dtype=np.uint32));pending=pending[1_000_000:];number+=1
|
| 97 |
+
if count>=budget:break
|
| 98 |
+
if count==before:raise RuntimeError(f'Empty calibration dataset {repo}')
|
| 99 |
+
actual[repo]=count
|
| 100 |
+
if pending:np.save(target/f'shard_{number:05d}.npy',np.asarray(pending,dtype=np.uint32))
|
| 101 |
+
write(target/'corpus.json',{'builder':'fallback text calibration, not legacy calib-v3','sources_token_counts':actual,'dataset_revisions':revisions,'seed':42,'reasoning_rollouts':False,'tokenizer_repo':args.src,'tokenizer_revision':hub.revision(args.src)})
|
| 102 |
+
marker.unlink()
|
| 103 |
+
def ensure_captures(args,hub,work,need_stats=False):
|
| 104 |
+
import torch
|
| 105 |
+
shards=list(Path(args.calib_tokens).glob('shard_*.npy'))
|
| 106 |
+
if not shards or (Path(args.calib_tokens)/'corpus.building.json').exists():build_corpus(args,hub)
|
| 107 |
+
else:
|
| 108 |
+
for p in shards:
|
| 109 |
+
a=np.load(p,mmap_mode='r')
|
| 110 |
+
if a.ndim!=1 or not np.issubdtype(a.dtype,np.integer) or not a.size:raise ValueError(f'Invalid token shard {p}')
|
| 111 |
+
ready=all(valid_capture(Path(args.calib)/f'acts_layer{L}.pt') for L in LAYERS)
|
| 112 |
+
stats=Path(args.calib).parent/'expert_stats53.npz'
|
| 113 |
+
if ready and (not need_stats or stats.exists()):return stats
|
| 114 |
+
if not complete_snapshot(args.donor_cache):hub.snapshot(args.nvfp4_donor,args.donor_cache)
|
| 115 |
+
# Capture to a fresh directory so partial old ranks cannot enter the merge.
|
| 116 |
+
dest=work/'capture';dest.mkdir(exist_ok=True)
|
| 117 |
+
env={**os.environ,'GLM53_DONOR':args.donor_cache,'GLM53_CALIB':args.calib_tokens,'GLM53_CAPTURE_OUT':str(dest)}
|
| 118 |
+
script=ROOT.parent/'tools/capture53/stream_capture53.py';workers=[]
|
| 119 |
+
for rank,gpu in enumerate(args.gpus):
|
| 120 |
+
log=open(work/f'capture-rank{rank}.log','a')
|
| 121 |
+
workers.append(subprocess.Popen([sys.executable,'-u',str(script),'stats'],env={**env,'CUDA_VISIBLE_DEVICES':str(gpu),'RANK':str(rank),'WORLD':str(len(args.gpus))},stdout=log,stderr=subprocess.STDOUT,start_new_session=True));log.close()
|
| 122 |
+
codes=[p.wait() for p in workers]
|
| 123 |
+
if any(codes):raise RuntimeError('Capture workers failed; see capture-rank logs')
|
| 124 |
+
subprocess.run([sys.executable,str(script),'merge'],env={**env,'WORLD':str(len(args.gpus))},check=True)
|
| 125 |
+
args.calib=str(dest/'acts')
|
| 126 |
+
if not all(valid_capture(Path(args.calib)/f'acts_layer{L}.pt') for L in LAYERS):raise RuntimeError('Capture incomplete')
|
| 127 |
+
return dest/'expert_stats53.npz'
|
| 128 |
+
def assignment(args,hub,work):
|
| 129 |
+
if getattr(args,'reap_metric','aqlm')=='arvq':
|
| 130 |
+
from .arvq_reap import recompute
|
| 131 |
+
obj=recompute(args,hub,work)
|
| 132 |
+
write(work/'assignment.json',obj);args.assign=str(work/'assignment.json');return obj
|
| 133 |
+
cached=Path(args.assign)
|
| 134 |
+
if cached.exists():
|
| 135 |
+
obj=validate_assignment(json.loads(cached.read_text()),args.hot_count)
|
| 136 |
+
obj['provenance']={'mode':'reused allocation','path':str(cached),'sha256':sha256(cached)}
|
| 137 |
+
else:
|
| 138 |
+
from .reap import recompute
|
| 139 |
+
obj=recompute(args,hub,work)
|
| 140 |
+
write(work/'assignment.json',obj);args.assign=str(work/'assignment.json');return obj
|
| 141 |
+
def ensure_source(args,hub,work,allocation):
|
| 142 |
+
import torch
|
| 143 |
+
from safetensors import safe_open
|
| 144 |
+
from safetensors.torch import save_file
|
| 145 |
+
from ingest import probe_source,validate_meta_against_header
|
| 146 |
+
from remote_st import RemoteShard
|
| 147 |
+
import remote_st
|
| 148 |
+
cfg=json.loads(hub.file(args.src,'config.json').read_text());text=cfg.get('text_config',cfg)
|
| 149 |
+
if text.get('hidden_size')!=6144 or text.get('moe_intermediate_size')!=2048 or text.get('n_routed_experts',text.get('num_local_experts'))!=256:
|
| 150 |
+
raise ValueError('Input is not the supported GLM-5.3 expert geometry')
|
| 151 |
+
meta=probe_source(text);native={**vars(meta),'block':list(meta.block) if meta.block else None}
|
| 152 |
+
revision=hub.revision(args.src);cache=Path(args.src_cache);cache.mkdir(parents=True,exist_ok=True)
|
| 153 |
+
had_existing=any(cache.glob('layer_*'))
|
| 154 |
+
marker=cache/'source_identity.json'
|
| 155 |
+
if marker.exists():
|
| 156 |
+
identity=json.loads(marker.read_text())
|
| 157 |
+
if identity!={'repo':args.src,'revision':revision}:raise ValueError('Source cache identity mismatch; use another --src-cache')
|
| 158 |
+
elif args.src!='zai-org/GLM-5.3' and any(cache.glob('layer_*')):raise ValueError('Unidentified cache cannot be reused for a different source repo')
|
| 159 |
+
local=local_directory(args.src)
|
| 160 |
+
index=model_index(local) if local else json.loads(hub.file(args.src,'model.safetensors.index.json').read_text())['weight_map']
|
| 161 |
+
remote_st.resolve_url=lambda repo,name:f'https://huggingface.co/{repo}/resolve/{revision}/{name}'
|
| 162 |
+
headers={}
|
| 163 |
+
def shard(name):
|
| 164 |
+
fn=index[name]
|
| 165 |
+
if fn not in headers:headers[fn]=LocalShard(local/fn) if local else RemoteShard(args.src,fn)
|
| 166 |
+
return headers[fn]
|
| 167 |
+
dtype={'F8_E4M3':torch.float8_e4m3fn,'BF16':torch.bfloat16,'F16':torch.float16,'F32':torch.float32,'U8':torch.uint8}
|
| 168 |
+
def fetch(pair):
|
| 169 |
+
L,e=pair;path=cache/f'layer_{L}'/f'expert_{e}.safetensors'
|
| 170 |
+
if path.exists():
|
| 171 |
+
try:
|
| 172 |
+
with safe_open(str(path),framework='pt') as f:
|
| 173 |
+
metadata=f.metadata() or {}
|
| 174 |
+
if metadata.get('source',args.src)!=args.src:raise ValueError('Expert source mismatch')
|
| 175 |
+
for proj,shape in [('gate_proj',(2048,6144)),('up_proj',(2048,6144)),('down_proj',(6144,2048))]:
|
| 176 |
+
t=f.get_slice(proj+'.weight')
|
| 177 |
+
if tuple(t.get_shape())!=shape:raise ValueError('Invalid cached expert shape')
|
| 178 |
+
validate_meta_against_header(meta,t.get_dtype())
|
| 179 |
+
if meta.kind=='block_fp8' and tuple(f.get_slice(proj+'.weight_scale_inv').get_shape())!=(shape[0]//meta.block[0],shape[1]//meta.block[1]):raise ValueError('Invalid cached scales')
|
| 180 |
+
return
|
| 181 |
+
except Exception as exc:raise ValueError(f'Invalid cached expert {path}; remove or repair it before retry') from exc
|
| 182 |
+
tensors={}
|
| 183 |
+
for proj in ['gate_proj','up_proj','down_proj']:
|
| 184 |
+
for suffix in ['weight']+(['weight_scale_inv'] if meta.kind=='block_fp8' else []):
|
| 185 |
+
name=f'model.layers.{L}.mlp.experts.{e}.{proj}.{suffix}';sh=shard(name);m=sh.tensor_meta(name)
|
| 186 |
+
raw=bytearray(sh.read_tensor_bytes(name));expected=m['data_offsets'][1]-m['data_offsets'][0]
|
| 187 |
+
if len(raw)!=expected:raise IOError('Short source tensor read')
|
| 188 |
+
t=torch.frombuffer(raw,dtype=dtype[m['dtype']]).reshape(m['shape']).clone()
|
| 189 |
+
if suffix=='weight' and meta.kind=='block_fp8' and t.dtype==torch.uint8:t=t.view(torch.float8_e4m3fn)
|
| 190 |
+
tensors[proj+'.'+suffix]=t
|
| 191 |
+
path.parent.mkdir(parents=True,exist_ok=True);tmp=path.with_suffix('.part');save_file(tensors,str(tmp),metadata={'source':args.src,'revision':revision});tmp.replace(path)
|
| 192 |
+
jobs=[(L,e) for L in LAYERS for e in allocation['layers'][str(L)]['cold_5750']]
|
| 193 |
+
with ThreadPoolExecutor(max_workers=args.download_workers) as pool:list(pool.map(fetch,jobs))
|
| 194 |
+
# Existing canonical donor files predate revision markers; record their status honestly.
|
| 195 |
+
if local and fingerprint(local)!=revision:raise ValueError('Local source changed during extraction')
|
| 196 |
+
write(work/'source.json',{'repo':args.src,'revision':revision,'meta':native,'cache':str(cache),'legacy_cache_revision_unverified':not marker.exists() and had_existing})
|
| 197 |
+
if not had_existing:write(marker,{'repo':args.src,'revision':revision})
|
| 198 |
+
return meta
|
reproduce/source/btx53/arvq88/jobs.py
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Bounded subprocess jobs on explicitly reserved GPUs."""
|
| 2 |
+
import subprocess
|
| 3 |
+
import time
|
| 4 |
+
from pathlib import Path
|
| 5 |
+
|
| 6 |
+
|
| 7 |
+
def run_jobs(jobs, gpus, logs):
|
| 8 |
+
pending=list(jobs);running={};logs=Path(logs);logs.mkdir(parents=True,exist_ok=True)
|
| 9 |
+
try:
|
| 10 |
+
while pending or running:
|
| 11 |
+
for gpu in gpus:
|
| 12 |
+
if gpu in running or not pending:continue
|
| 13 |
+
name,command=pending.pop(0)
|
| 14 |
+
with (logs/(name+'.log')).open('a') as log:
|
| 15 |
+
p=subprocess.Popen(command(gpu),stdin=subprocess.DEVNULL,stdout=log,
|
| 16 |
+
stderr=subprocess.STDOUT,start_new_session=True)
|
| 17 |
+
running[gpu]=(name,p);print('START',name,'GPU',gpu,'PID',p.pid,flush=True)
|
| 18 |
+
time.sleep(5)
|
| 19 |
+
for gpu,(name,p) in list(running.items()):
|
| 20 |
+
rc=p.poll()
|
| 21 |
+
if rc is None:continue
|
| 22 |
+
if rc:raise RuntimeError(f'{name} failed ({rc}); see {logs/name}.log')
|
| 23 |
+
del running[gpu];print('DONE',name,flush=True)
|
| 24 |
+
except BaseException:
|
| 25 |
+
for name,p in running.values():
|
| 26 |
+
if p.poll() is None:p.terminate()
|
| 27 |
+
for name,p in running.values():p.wait()
|
| 28 |
+
raise
|
reproduce/source/btx53/arvq88/local_source.py
ADDED
|
@@ -0,0 +1,43 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Read standard local safetensors snapshots without Hub requests."""
|
| 2 |
+
import hashlib,json,struct
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
|
| 5 |
+
def local_directory(source):
|
| 6 |
+
p=Path(source).expanduser()
|
| 7 |
+
if p.exists():
|
| 8 |
+
if not p.is_dir():raise ValueError(f'Model input must be a directory: {p}')
|
| 9 |
+
return p.resolve()
|
| 10 |
+
if str(source).startswith(('/', './', '../', '~')):
|
| 11 |
+
raise FileNotFoundError(f'Local model directory does not exist: {p}')
|
| 12 |
+
return None
|
| 13 |
+
|
| 14 |
+
def model_index(root):
|
| 15 |
+
p=root/'model.safetensors.index.json'
|
| 16 |
+
if p.is_file():return json.loads(p.read_text())['weight_map']
|
| 17 |
+
p=root/'model.safetensors'
|
| 18 |
+
if not p.is_file():raise FileNotFoundError(f'Missing safetensors weights/index in {root}')
|
| 19 |
+
return {k:p.name for k in LocalShard(p).header if k!='__metadata__'}
|
| 20 |
+
|
| 21 |
+
def fingerprint(root):
|
| 22 |
+
# Metadata identity, not a full content hash of a multi-hundred-GB donor.
|
| 23 |
+
h=hashlib.sha256((root/'config.json').read_bytes())
|
| 24 |
+
index=model_index(root);h.update(json.dumps(index,sort_keys=True).encode())
|
| 25 |
+
for name in sorted(set(index.values())):
|
| 26 |
+
p=root/name;st=p.stat()
|
| 27 |
+
h.update(json.dumps([name,st.st_size,st.st_mtime_ns,st.st_ctime_ns]).encode())
|
| 28 |
+
return 'local-stat-sha256:'+h.hexdigest()
|
| 29 |
+
|
| 30 |
+
class LocalShard:
|
| 31 |
+
def __init__(self,path):
|
| 32 |
+
self.path=Path(path)
|
| 33 |
+
with self.path.open('rb') as f:
|
| 34 |
+
raw=f.read(8)
|
| 35 |
+
if len(raw)!=8:raise ValueError('Truncated safetensors header')
|
| 36 |
+
n=struct.unpack('<Q',raw)[0]
|
| 37 |
+
if n>100_000_000:raise ValueError('Invalid safetensors header length')
|
| 38 |
+
self.header=json.loads(f.read(n));self.offset=8+n
|
| 39 |
+
def tensor_meta(self,name):return self.header[name]
|
| 40 |
+
def read_tensor_bytes(self,name):
|
| 41 |
+
a,b=self.tensor_meta(name)['data_offsets']
|
| 42 |
+
with self.path.open('rb') as f:
|
| 43 |
+
f.seek(self.offset+a);return f.read(b-a)
|
reproduce/source/btx53/arvq88/pack.py
ADDED
|
@@ -0,0 +1,142 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""ARVQ 8+8 test format: natural indices -> MMA fragments, 64 uint32 words/tile."""
|
| 2 |
+
import torch,json,hashlib
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
LEVELS=[0.,.5,1.,1.5,2.,3.,4.,6.,-0.,-.5,-1.,-1.5,-2.,-3.,-4.,-6.]
|
| 5 |
+
def maps(device):
|
| 6 |
+
p=torch.arange(128,device=device);j=p//32;lane=p%32
|
| 7 |
+
return lane//4+8*(j%2),(j//2)*4+lane%4
|
| 8 |
+
def pack(a,b,N,K):
|
| 9 |
+
r,k=maps(a.device);rows=torch.arange(N//16,device=a.device)[:,None,None]*16+r[None,None,:];cols=torch.arange(K//64,device=a.device)[None,:,None]*8+k[None,None,:]
|
| 10 |
+
v=(a.long()|(b.long()<<8))[rows,cols]
|
| 11 |
+
return (v[...,::2]|(v[...,1::2]<<16)).to(torch.uint32)
|
| 12 |
+
def unpack(w,N,K):
|
| 13 |
+
w=w.long();v=torch.stack([w&65535,w>>16],-1).reshape(N//16,K//64,128)
|
| 14 |
+
r,k=maps(w.device);rows=(torch.arange(N//16,device=w.device)[:,None,None]*16+r[None,None,:]).expand_as(v);cols=(torch.arange(K//64,device=w.device)[None,:,None]*8+k[None,None,:]).expand_as(v)
|
| 15 |
+
a=torch.empty(N,K//8,dtype=torch.uint8,device=w.device);b=torch.empty_like(a)
|
| 16 |
+
a[rows,cols]=(v&255).byte();b[rows,cols]=(v>>8).byte();return a,b
|
| 17 |
+
def pack_codebooks(c0,c1):
|
| 18 |
+
cb=torch.cat([c0,c1],dim=-2)
|
| 19 |
+
if cb.shape[-2:]!=(512,8) or cb.ndim not in (2,3):raise ValueError('Expected shared or per-expert 8+8 books')
|
| 20 |
+
levels=torch.tensor(LEVELS,device=cb.device);n=(cb[...,None]-levels).abs().argmin(-1)
|
| 21 |
+
if not torch.equal(levels[n],cb):raise ValueError('Codebooks must be exactly on the FP4 grid')
|
| 22 |
+
return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
|
| 23 |
+
def pack_codebooks_mb(cb,factors):
|
| 24 |
+
"""mcbook16: cb [E,256+B*256,8] in EFFECTIVE units. Rows are stored as plain-grid
|
| 25 |
+
nibbles; decode multiplies residual book m by factors[m] (exact powers of two)."""
|
| 26 |
+
B=len(factors)
|
| 27 |
+
if cb.ndim!=3 or cb.shape[1]!=256+B*256 or cb.shape[2]!=8:raise ValueError('Invalid mcbook codebook shape')
|
| 28 |
+
f=torch.cat([torch.ones(256),torch.tensor(factors).float().repeat_interleave(256)])[None,:,None]
|
| 29 |
+
grid=cb/f
|
| 30 |
+
levels=torch.tensor(LEVELS,device=cb.device);n=(grid[...,None]-levels).abs().argmin(-1)
|
| 31 |
+
if not torch.equal(levels[n]*f,cb):raise ValueError('Codebooks must be exactly on their per-book grids')
|
| 32 |
+
return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
|
| 33 |
+
def decode_mb(t,layer,proj,e,N,K):
|
| 34 |
+
pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
|
| 35 |
+
a,b=unpack(t[pre+'packed'][e],N,K)
|
| 36 |
+
cb=t[pre+'codebooks'][e].long();factors=t[pre+'book_factors'].float();B=len(factors)
|
| 37 |
+
if cb.shape!=(256+B*256,):raise ValueError('Invalid mcbook packed codebook shape')
|
| 38 |
+
values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
|
| 39 |
+
values=torch.cat([values[:256],values[256:]*factors.repeat_interleave(256)[:,None]])
|
| 40 |
+
sel=t[pre+'selectors'][e].long()
|
| 41 |
+
m=sel.repeat_interleave(16,0).repeat_interleave(8,1)
|
| 42 |
+
stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
|
| 43 |
+
if stored.dtype==torch.float16:scales=stored.float()
|
| 44 |
+
elif stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
|
| 45 |
+
else:raise ValueError('Unsupported packed block scale dtype')
|
| 46 |
+
return ((values[a.long()]+values[256+m*256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
|
| 47 |
+
def export_layer_mb(source,dest,L):
|
| 48 |
+
"""mcbook16 export (format version 5): 16 residual books, per-tile 4-bit selectors,
|
| 49 |
+
fp16 block scales; packed index stream and scales unchanged from v4."""
|
| 50 |
+
from safetensors.torch import save_file,load_file
|
| 51 |
+
source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
|
| 52 |
+
original=json.loads((source/'arvq-manifest.json').read_text());files={};first_factors=None
|
| 53 |
+
for proj,tag in [('w13','gateup'),('w2','down')]:
|
| 54 |
+
d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
|
| 55 |
+
if d.get('scale_dtype')!='fp16':raise ValueError('mcbook export requires fp16 block scales')
|
| 56 |
+
if not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
|
| 57 |
+
factors=[float(f) for f in d['book_factor']]
|
| 58 |
+
if first_factors is None:first_factors=factors
|
| 59 |
+
elif factors!=first_factors:raise ValueError('Mixed projection book factors')
|
| 60 |
+
sel=d['selector']
|
| 61 |
+
if sel.shape!=(E,N//16,K//64) or int(sel.max())>=len(factors):raise ValueError('Invalid selector tensor')
|
| 62 |
+
t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
|
| 63 |
+
pre+'codebooks':pack_codebooks_mb(d['cb'],factors),
|
| 64 |
+
pre+'selectors':sel.to(torch.uint8).contiguous(),
|
| 65 |
+
pre+'book_factors':torch.tensor(factors,dtype=torch.float32),
|
| 66 |
+
pre+'scales':d['s'].half().reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
|
| 67 |
+
pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
|
| 68 |
+
for e in range(E):
|
| 69 |
+
a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
|
| 70 |
+
for e in sorted({0,E//2,E-1}):
|
| 71 |
+
m=d['selector'][e].long().repeat_interleave(16,0).repeat_interleave(8,1)
|
| 72 |
+
cb=d['cb'][e]
|
| 73 |
+
ref=((cb[d['a'][e].long()]+cb[256+m*256+d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
|
| 74 |
+
assert torch.equal(ref,decode_mb(t,L,proj,e,N,K))
|
| 75 |
+
name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
|
| 76 |
+
save_file(t,str(tmp),metadata={'format':'pt','arvq_format':'rvq256_mb16_256x8_expert_fp16block','codebook_scope':'expert','residual_books':str(len(factors)),'book_factors':json.dumps(factors),'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires mcbook16 v5 loader/kernel (vllm-glm52-sm120 branch experiment/arvq-mcbook16)'})
|
| 77 |
+
tmp.replace(dest/name)
|
| 78 |
+
loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64) and loaded[pre+'selectors'].shape==(E,N//16,K//64);del loaded
|
| 79 |
+
files[proj]={'file':name,'sha256':sha(dest/name)};del t,d
|
| 80 |
+
m={**original,'format':'rvq256_mb16_256x8_expert_fp16block','version':5,'residual_scale_shift':0,'block_scale_dtype':'float16','codebook_scope':'expert','bits':2.125+4/1024,'codebook_sizes':[256,16*256],'index_bits':[8,8],'selector_bits_per_tile':4,'words_per_tile':64,'book_factors':first_factors,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True,'requires_mcbook16_loader':True}
|
| 81 |
+
(dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
|
| 82 |
+
def decode(t,layer,proj,e,N=None,K=None,residual_scale_shift=0):
|
| 83 |
+
pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
|
| 84 |
+
default=(4096,6144) if proj=='w13' else (6144,2048)
|
| 85 |
+
N,K=default if N is None else (N,K)
|
| 86 |
+
a,b=unpack(t[pre+'packed'][e],N,K);cb=t[pre+'codebooks'].long()
|
| 87 |
+
if cb.ndim==2:cb=cb[e]
|
| 88 |
+
if cb.shape!=(512,):raise ValueError('Invalid packed codebook shape')
|
| 89 |
+
values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
|
| 90 |
+
stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
|
| 91 |
+
if stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
|
| 92 |
+
elif stored.dtype==torch.float16:scales=stored.float()
|
| 93 |
+
else:raise ValueError('Unsupported packed block scale dtype')
|
| 94 |
+
f=2.0**(-int(residual_scale_shift))
|
| 95 |
+
return ((values[a.long()]+f*values[256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
|
| 96 |
+
def sha(p):
|
| 97 |
+
h=hashlib.sha256()
|
| 98 |
+
with open(p,'rb') as f:
|
| 99 |
+
for chunk in iter(lambda:f.read(8<<20),b''):h.update(chunk)
|
| 100 |
+
return h.hexdigest()
|
| 101 |
+
def export_layer(source,dest,L,residual_scale_shift=0):
|
| 102 |
+
from safetensors.torch import save_file,load_file
|
| 103 |
+
source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
|
| 104 |
+
probe=torch.load(source/'w13.pt',map_location='cpu',weights_only=True,mmap=True)['w13']
|
| 105 |
+
if 'selector' in probe:
|
| 106 |
+
del probe
|
| 107 |
+
if int(residual_scale_shift):raise ValueError('mcbook exports are shift-0 (per-book factors)')
|
| 108 |
+
return export_layer_mb(source,dest,L)
|
| 109 |
+
del probe
|
| 110 |
+
rss=int(residual_scale_shift);f=2.0**(-rss)
|
| 111 |
+
original=json.loads((source/'arvq-manifest.json').read_text());files={}
|
| 112 |
+
for proj,tag in [('w13','gateup'),('w2','down')]:
|
| 113 |
+
d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
|
| 114 |
+
expert=d['c0'].ndim==3
|
| 115 |
+
if expert and d['c0'].shape!=(E,256,8):raise ValueError('Codebook expert count mismatch')
|
| 116 |
+
if proj=='w13':scope='expert' if expert else 'layer'
|
| 117 |
+
elif scope!=('expert' if expert else 'layer'):raise ValueError('Mixed projection codebook scopes')
|
| 118 |
+
fp16=d.get('scale_dtype','fp8_e4m3')=='fp16'
|
| 119 |
+
if proj=='w13':fp16_blocks=fp16
|
| 120 |
+
elif fp16_blocks!=fp16:raise ValueError('Mixed projection scale precision')
|
| 121 |
+
if fp16 and not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
|
| 122 |
+
stored_scales=d['s'].half() if fp16 else d['s'].to(torch.float8_e4m3fn).view(torch.uint8)
|
| 123 |
+
t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
|
| 124 |
+
pre+'codebooks':pack_codebooks(d['c0'],d['c1']),
|
| 125 |
+
pre+'scales':stored_scales.reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
|
| 126 |
+
pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
|
| 127 |
+
for e in range(E):
|
| 128 |
+
a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
|
| 129 |
+
for e in sorted({0,E//2,E-1}):
|
| 130 |
+
c0,c1=(d['c0'][e],d['c1'][e]) if expert else (d['c0'],d['c1'])
|
| 131 |
+
ref=((c0[d['a'][e].long()]+f*c1[d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
|
| 132 |
+
assert torch.equal(ref,decode(t,L,proj,e,N,K,rss))
|
| 133 |
+
base_fmt='rvq256_256x8_expert_fp16block' if fp16 else ('rvq256_256x8_expert' if expert else 'rvq256_256x8')
|
| 134 |
+
arvq_fmt=base_fmt+(f'_rs{int(2**rss)}' if rss else '')
|
| 135 |
+
name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
|
| 136 |
+
save_file(t,str(tmp),metadata={'format':'pt','arvq_format':arvq_fmt,'residual_scale_shift':str(rss),'codebook_scope':scope,'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires FP16 block-scale v4 loader/kernel' if fp16 else ('requires per-expert v3 loader/kernel' if expert else 'requires 8+8 loader/kernel')})
|
| 137 |
+
tmp.replace(dest/name);del t,d
|
| 138 |
+
loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64);del loaded
|
| 139 |
+
files[proj]={'file':name,'sha256':sha(dest/name)}
|
| 140 |
+
base_mfmt='rvq256_256x8_expert_fp16block' if fp16_blocks else ('rvq256_256x8_expert' if scope=='expert' else 'rvq256_256x8')
|
| 141 |
+
m={**original,'format':base_mfmt+(f'_rs{int(2**rss)}' if rss else ''),'version':4 if fp16_blocks else (3 if scope=='expert' else 2),'residual_scale_shift':rss,'block_scale_dtype':'float16' if fp16_blocks else 'float8_e4m3fn','codebook_scope':scope,'bits':2.125 if fp16_blocks else 2.0625,'codebook_sizes':[256,256],'index_bits':[8,8],'words_per_tile':64,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True}
|
| 142 |
+
(dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
|
reproduce/source/btx53/arvq88/perf/FP16_BLOCK_SCALE_HANDOFF.md
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# ARVQ FP16 block-scale format (v4)
|
| 2 |
+
|
| 3 |
+
This replaces FP8 E4M3 block scales with FP16 block scales. There are no added row scales. Books remain per-expert packed FP4, indices remain two 8-bit indices per group of eight weights, and global scale remains FP32 per projection. Initial FP8 values are exactly representable in FP16.
|
| 4 |
+
|
| 5 |
+
Checkpoint tensors use the existing `arvq_{w13,w2}_scales` names and logical tiled shape `[cold_experts, N/16, K/128, 16]`, now dtype F16 instead of U8 (FP8 bytes). Manifest version4, format `rvq256_256x8_expert_fp16block`, `block_scale_dtype: float16`. Packed codebooks and indices retain v3 layout. Index+block-scale storage is2.125 bits/weight, excluding books/global metadata. Each scale belongs to one expert, one output row, and128 input weights. Scales cannot be folded into one whole-row epilogue multiplier.
|
| 6 |
+
|
| 7 |
+
Decode for a group of8 weights:
|
| 8 |
+
`W[e,r,8*g:8*g+8] = global * scale16[e,r,g//16] * (book0[e,index0] + book1[e,index1])`.
|
| 9 |
+
|
| 10 |
+
The existing native MMA instruction accepts UE4M3 weight scales, not FP16. A compatible native path must preserve FP16 scale precision, e.g. use unit weight scales in FP4 MMA, accumulate the128-weight block (two K64 tiles, both codebooks and activation planes), multiply its FP32 contribution by the FP16 block scale converted to FP32, and accumulate blocks. Do not silently cast the new scales back to E4M3. Activation quantization and hot NVFP4 paths remain unchanged. Benchmark native latency and validate arithmetic before publication.
|
| 11 |
+
|
| 12 |
+
Training: `/tmp/glm53-fp16-sequential`, sequential layers3–77 from original pre-PV initialization, full18,001,846-token corpus, batch262144/micro65536, Adam and periodic output-gradient index reassignment. Trainable latent log-scales are FP32; every forward/export projects scales to FP16. Global scale is frozen. Max69 updates; three validation checks >0.1% worse trigger a4x LR drop, otherwise drop after45; stop after15 lower-LR updates without a new best. Validation, development-audit and export-replay gates remain enforced.
|
| 13 |
+
|
| 14 |
+
Current fitting/export modules:
|
| 15 |
+
- `sequential_pv_full_corpus.py`: config `block_scale_dtype: fp16` selects precision.
|
| 16 |
+
- `arvq88/pv.py`: encoded `scale_dtype` travels with the representation and is honored during replay/propagation.
|
| 17 |
+
- `arvq88/pack.py`: F16 tiled scales, v4 metadata and exact decoded-weight roundtrip.
|
| 18 |
+
|
| 19 |
+
Local fitting and propagation can proceed. Automatic HF publication is disabled until the loader/kernel supports this format. Do not publish v4 tensors under v3 metadata.
|
reproduce/source/btx53/arvq88/perf/FULL_CORPUS_PIPELINE.md
ADDED
|
@@ -0,0 +1,80 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Full-corpus sequential PV
|
| 2 |
+
|
| 3 |
+
Run `full_pipeline.py --work WORK --first 3 --last N` with the repository Python.
|
| 4 |
+
The work directory needs config.json, the initial fixed capture3, and token file.
|
| 5 |
+
The driver runs eight GPU ranks per stage and never uploads. Set explicit last
|
| 6 |
+
layer: a bounded pilot is the default operational choice until it qualifies.
|
| 7 |
+
|
| 8 |
+
The existing layer-3 fit uses 18,001,846 training tokens, one shuffled sequence
|
| 9 |
+
pass, 262144 effective batch, 65536 microbatch, Adam codebook LR .048 and scale
|
| 10 |
+
LR .032. Gradient accumulation normalizes every microbatch by the full batch
|
| 11 |
+
size; the regularizer, clipping and Adam step execute once per effective batch.
|
| 12 |
+
Validation every five updates retains the best state. Audit is evaluated on the
|
| 13 |
+
retained state; the driver refuses propagation if validation or audit regresses.
|
| 14 |
+
|
| 15 |
+
Capture stores student normalized inputs/routing, frozen branch output, the
|
| 16 |
+
reference-minus-frozen target, and the BF16 reference output. Both full reference
|
| 17 |
+
and full student trajectories advance one block at a time. Layer 3 needs a
|
| 18 |
+
one-time recapture because the earlier training capture omitted frozen/reference
|
| 19 |
+
outputs, which cannot be reconstructed from their difference alone. Dense prefix
|
| 20 |
+
weights now remain resident during that recapture. Layers 4 onward read cached
|
| 21 |
+
boundaries and never execute the earlier prefix.
|
| 22 |
+
|
| 23 |
+
Propagation uses the retained serialized FP4 books, FP8 scales and indices with
|
| 24 |
+
the same activation emulation as training. It caches decoded expert weights per
|
| 25 |
+
GPU across all chunks. A fixed-evaluation replay gate allows relative difference
|
| 26 |
+
at most 1e-4 after BF16 output rounding, accounting for different FP32 reduction
|
| 27 |
+
order; it does not claim bitwise parity or native SM120 execution. Partial final
|
| 28 |
+
sequences are zero-padded at boundaries and masked from capture/training losses.
|
| 29 |
+
|
| 30 |
+
Rank zero alone assembles and prefetches each CPU training batch, then broadcasts
|
| 31 |
+
identical tensors to the expert ranks over NCCL; capture and propagation shard writes
|
| 32 |
+
run on one background thread with bounded outstanding data. Metadata is committed
|
| 33 |
+
only after the data file is atomically renamed. A partial smoke run cannot publish
|
| 34 |
+
a rank-complete marker. `timings.json` records rank-0 data wait, transfer, optimizer,
|
| 35 |
+
reassignment and validation/checkpoint time; these are not all-rank max timings.
|
| 36 |
+
|
| 37 |
+
The earlier monolithic 256K run exhausted GPU memory on fresh data. Its first five
|
| 38 |
+
updates took 158.9 seconds. The 64K accumulated version with prefetch took 65.0
|
| 39 |
+
seconds with validation 0.0047233582 versus 0.0047233591 (relative difference
|
| 40 |
+
~2e-7). This is an observed first-five-update comparison, not a full-run speedup.
|
| 41 |
+
|
| 42 |
+
Pending qualification: the scheduled 262144-token integration smoke exercises
|
| 43 |
+
capture3 -> serialized propagation3 -> capture4. Full-layer transition throughput
|
| 44 |
+
must be measured before extrapolating all 75 layers. CUDA graphs are not enabled:
|
| 45 |
+
variable routing and fresh sequence batches require a separately validated static
|
| 46 |
+
buffer/shape design. Native serving-kernel numerical qualification remains separate.
|
| 47 |
+
|
| 48 |
+
The driver keeps two layers of regenerable full-corpus intermediates and all fit
|
| 49 |
+
outputs/checkpoints. It prunes only directories with its own completed stage
|
| 50 |
+
receipt after a later layer passes held-out checks; it never prunes the original
|
| 51 |
+
legacy training_capture directory or fixed validation/audit captures.
|
| 52 |
+
|
| 53 |
+
New captures store normalized inputs as BF16 only after an exact FP32 roundtrip
|
| 54 |
+
assertion for every chunk; arithmetic restores FP32 before use. This halves input
|
| 55 |
+
storage and CPU assembly traffic without rounding any captured input values.
|
| 56 |
+
|
| 57 |
+
## Capture optimizations added during the live campaign
|
| 58 |
+
|
| 59 |
+
Layer-14 measurements: capture setup ~90 s, capture computation ~154 s,
|
| 60 |
+
CPU copying/statistics ~99 s, input transfers ~31 s. Separate fixed-evaluation
|
| 61 |
+
capture costs another ~100 s. Propagation CPU boundary packing costs ~27 s.
|
| 62 |
+
|
| 63 |
+
New capture uses two reusable pinned host banks and asynchronous CUDA copies,
|
| 64 |
+
with an event that the shard writer waits on before serialization. The writer's
|
| 65 |
+
single-outstanding-job contract prevents overwriting a bank still being saved.
|
| 66 |
+
Each rank compares its first staged production chunk exactly with the original
|
| 67 |
+
blocking CPU transfer. A real CUDA test also exercises six writes across bank
|
| 68 |
+
reuse and verifies every saved value. Propagation now copies the full output in
|
| 69 |
+
one transfer and reshapes complete sequences without CPU copies; partial tails
|
| 70 |
+
retain explicit zero padding and have exact parity tests.
|
| 71 |
+
|
| 72 |
+
Full training capture also produces a candidate fixed-evaluation capture using
|
| 73 |
+
resident weights. The existing standalone capture still runs for the first
|
| 74 |
+
qualification layer, compares every tensor on all eight ranks, and publishes
|
| 75 |
+
fused_eval_approved.json only on exact equality. Subsequent fixed capture stages
|
| 76 |
+
reuse that candidate instead of reloading the reference/donor weights. If parity
|
| 77 |
+
fails, standalone capture continues to supply the production tensors.
|
| 78 |
+
|
| 79 |
+
Production speedups are not yet measured. Inspect training_capture*/rank*_timings.json,
|
| 80 |
+
fused_eval_qualification, fused_eval_approved.json and pipeline stage receipts.
|
reproduce/source/btx53/arvq88/perf/README.md
ADDED
|
@@ -0,0 +1,145 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# PV performance investigation (18 September 2026, Melbourne)
|
| 2 |
+
|
| 3 |
+
The active sequential quality pilot was not modified. These files are isolated
|
| 4 |
+
prototypes and a separately invocable trainer fork; no changes were published.
|
| 5 |
+
|
| 6 |
+
## Measured single-expert forward + backward
|
| 7 |
+
|
| 8 |
+
Actual layer-3 expert 3, 64 routed training rows, original fitted per-expert
|
| 9 |
+
books, float32 GEMMs and the same FP4/FP16 activation emulation. GPU 0 was shared
|
| 10 |
+
with the running pilot. Four timing repeats per path after warmup, no optimizer
|
| 11 |
+
step/communication/validation included. Results repeated in two invocations.
|
| 12 |
+
|
| 13 |
+
| Path | Median, second invocation |
|
| 14 |
+
| --- | ---: |
|
| 15 |
+
| Original eager arithmetic, four 16-row microbatches | 386.47 ms |
|
| 16 |
+
| Original eager arithmetic, one 64-row batch | 132.02 ms |
|
| 17 |
+
| Graph-compatible arithmetic, graphed one-batch forward/backward | 119.08 ms |
|
| 18 |
+
|
| 19 |
+
Larger batch: 2.93x throughput for this expert calculation. Graph replay adds
|
| 20 |
+
about 10% lower latency relative to eager one-batch; combined factor 3.25x.
|
| 21 |
+
These are NOT full-layer or whole-campaign speedup measurements.
|
| 22 |
+
|
| 23 |
+
Graph versus eager one-batch output and parameter-gradient relative differences
|
| 24 |
+
were zero in this test. After in-place index and row-scale changes, replay still
|
| 25 |
+
matched eager output exactly. Four-microbatch versus one-batch maximum parameter
|
| 26 |
+
-gradient relative difference was 5.99e-5 (FP32 accumulation/GEMM differences),
|
| 27 |
+
with losses 0.1189905256 versus 0.1189905405. End-to-end optimization trajectories
|
| 28 |
+
therefore need comparison, not an assertion of bit-identical training.
|
| 29 |
+
|
| 30 |
+
Artifacts: `/tmp/glm53-pv-graph-benchmark.json` and
|
| 31 |
+
`/tmp/glm53-pv-graph-benchmark-check.json`. Benchmark peak allocated memory:
|
| 32 |
+
0.514 GiB, isolated expert only; not a full-layer graph memory estimate.
|
| 33 |
+
|
| 34 |
+
## Why larger batches help
|
| 35 |
+
|
| 36 |
+
The production pilot calls local_prediction separately for each of four
|
| 37 |
+
microbatches. Every call reconstructs full expert matrices and backpropagates
|
| 38 |
+
through their codebook gathers. We repeat this large decode/scatter workload
|
| 39 |
+
four times despite unchanged parameters within the optimizer step. The GPUs
|
| 40 |
+
can report high utilization while doing this redundant work.
|
| 41 |
+
|
| 42 |
+
`sequential_pv_fast.py` is a separate fork with configurable `--microbatch`
|
| 43 |
+
(default 1024 versus the original 256). Global batch, objective, selected rows,
|
| 44 |
+
regularization and update count remain unchanged. Full-layer peak-memory and
|
| 45 |
+
held-out parity/timing tests are needed before deployment.
|
| 46 |
+
|
| 47 |
+
## Communication optimization
|
| 48 |
+
|
| 49 |
+
The pilot broadcasts an 8192 x 6144 float32 delta (192 MiB) for each proposal
|
| 50 |
+
attempt, mostly zeros outside an expert's routed rows. `sparse_delta.py` sends
|
| 51 |
+
only row IDs and their nonzero payload, then reconstructs the same dense buffer
|
| 52 |
+
before the unchanged acceptance arithmetic. For eight-of-256 routing the average
|
| 53 |
+
payload is roughly 1/32 as large, though expert popularity varies.
|
| 54 |
+
|
| 55 |
+
Two-rank CPU distributed tests confirm bit-exact reconstruction versus dense
|
| 56 |
+
broadcast, for either owner and empty row sets. This optimization is included
|
| 57 |
+
in the separate fast trainer fork. Eight-GPU NCCL qualification remains pending.
|
| 58 |
+
|
| 59 |
+
## CUDA Graph integration boundary
|
| 60 |
+
|
| 61 |
+
`graph_expert.py` and `bench_graph_expert.py` demonstrate captured forward and
|
| 62 |
+
backward with live index/parameter buffers. They do not graph the production
|
| 63 |
+
trainer. Original per-expert use.any(), nonzero(), variable routed batch lengths,
|
| 64 |
+
and finite-value host checks prevent simply wrapping the loop in CUDAGraph.
|
| 65 |
+
|
| 66 |
+
Next integration: precompute routing outside capture, use fixed/padded row-count
|
| 67 |
+
buckets with correct zero masks, capture expert forward/backward, and keep the
|
| 68 |
+
coupled-output all-reduce and discrete acceptance outside the initial graph.
|
| 69 |
+
Stable parameter/code buffer addresses and graph-memory budgets are required.
|
| 70 |
+
Validate finiteness outside captured execution; do not silently remove checks.
|
| 71 |
+
|
| 72 |
+
## Verification and next qualification
|
| 73 |
+
|
| 74 |
+
- Graph-compatible activation forward and STE derivative match the existing path.
|
| 75 |
+
- Sparse communication matches dense broadcast exactly in two-rank tests.
|
| 76 |
+
- Real GPU graph benchmark checks loss, gradients, and live code/scale updates.
|
| 77 |
+
- All tests in `perf/test_perf.py` pass; new modules compile.
|
| 78 |
+
|
| 79 |
+
After the quality pilot, compare baseline and fast fork on the same eight-GPU
|
| 80 |
+
layer and sample schedule, including an actual reassignment cycle. Measure
|
| 81 |
+
continuous steps, proposals, validation, peak memory and held-out error separately.
|
| 82 |
+
Then integrate graph buckets if their additional gain justifies the complexity.
|
| 83 |
+
No full-model runtime estimate should be reduced by the single-expert 3.25x factor.
|
| 84 |
+
|
| 85 |
+
## Follow-up: specialized codebook backward
|
| 86 |
+
|
| 87 |
+
An actual CPU/CUDA operator profile (`/tmp/glm53-pv-operator-profile.txt`)
|
| 88 |
+
identified general advanced-index backward as 83% of self CUDA time in the
|
| 89 |
+
measured expert step; matrix multiplication accounted for <1%. This motivated
|
| 90 |
+
using `torch.nn.functional.embedding` for each book lookup, instead of `cb[a]`.
|
| 91 |
+
It changes the lookup/reduction implementation, not the representation or loss.
|
| 92 |
+
|
| 93 |
+
Two separate real-expert benchmarks while the pilot was active:
|
| 94 |
+
|
| 95 |
+
| Layer/expert | Four microbatches, original | One batch, original | One batch, embedding | Embedding + graphs |
|
| 96 |
+
| --- | ---: | ---: | ---: | ---: |
|
| 97 |
+
| 3 / 3 | 442.67 ms | 115.18 ms | 35.39 ms | 25.37 ms |
|
| 98 |
+
| 4 / 0 | 373.97 ms | 133.65 ms | 35.19 ms | 34.75 ms |
|
| 99 |
+
|
| 100 |
+
These include contention and are single-expert forward/backward measurements.
|
| 101 |
+
The combined ratio is about 11–17x for this path, NOT the whole PV campaign.
|
| 102 |
+
Graph benefit is variable here; specialized embedding backward and batching are
|
| 103 |
+
more consistently material than launch replay alone.
|
| 104 |
+
|
| 105 |
+
Outputs and losses match exactly between original one-batch, embedding and
|
| 106 |
+
graphed embedding. GPU parameter-gradient relative differences are ~1.5–2.2e-6,
|
| 107 |
+
consistent with a changed FP32 reduction order. CPU float64 tests independently
|
| 108 |
+
verify matching parameter gradients, including after changing index buffers.
|
| 109 |
+
Graph replay after changing indices and row corrections still matches eager.
|
| 110 |
+
|
| 111 |
+
Artifacts: `/tmp/glm53-pv-embedding-benchmark.json` and
|
| 112 |
+
`/tmp/glm53-pv-embedding-layer4.json`. `EmbeddingRowProjection` implements the
|
| 113 |
+
isolated prototype; the specialized embedding lookup is now also present in
|
| 114 |
+
`sequential_pv_fast.py`. All three performance tests pass. The active pilot still
|
| 115 |
+
uses its original code. Full-layer NCCL, memory, quality and timing qualification
|
| 116 |
+
is required before switching a production run to the faster fork.
|
| 117 |
+
|
| 118 |
+
## Qualified full sequential campaign (18 September 2026)
|
| 119 |
+
|
| 120 |
+
The completed faster rerun matched both validation and development-audit error
|
| 121 |
+
exactly for reference layers 3 and 4. All propagated BF16 values matched exactly.
|
| 122 |
+
Measured training loops (including reassignment/validation) improved from
|
| 123 |
+
1525.16 to 336.60 seconds for layer 3 and 1226.31 to 303.00 seconds for layer 4.
|
| 124 |
+
This qualifies the larger microbatch, embedding backward and sparse broadcasts;
|
| 125 |
+
CUDA graphs remain a separate prototype and are not in the full trainer.
|
| 126 |
+
|
| 127 |
+
`full_reference.py --work /tmp/glm53-sequential-reference-full --pilot
|
| 128 |
+
/tmp/glm53-sequential-pilot-fast` reuses those two fitted layers and chains capture
|
| 129 |
+
and 200-step reference tuning through layer 77. `sequential_capture.py` now reads
|
| 130 |
+
reference and retained student outputs from the immediately preceding layer.
|
| 131 |
+
It retains the original corpus, allocation, seeds and tuning hyperparameters.
|
| 132 |
+
An additional initial development-audit evaluation permits an audit regression
|
| 133 |
+
gate before propagation/publication. Validation and audit gates allow only
|
| 134 |
+
1e-6 relative numerical slack; failure stops for investigation. The audit is
|
| 135 |
+
therefore development data, not an untouched final generalization estimate.
|
| 136 |
+
|
| 137 |
+
`publish_reference.py --work /tmp/glm53-sequential-reference-full` runs separately
|
| 138 |
+
on CPU/network. It reconstructs cold-slot order across eight expert shards,
|
| 139 |
+
checks slot coverage/allocation/global metadata, invokes the existing exact
|
| 140 |
+
index roundtrip and sampled decoded-weight parity checks, then atomically
|
| 141 |
+
replaces both projection files and reports on HF with parent-commit protection.
|
| 142 |
+
Remote hashes are verified before recording completion. Progress explicitly
|
| 143 |
+
identifies remaining legacy-PV and initial-fit layers. The old campaign and
|
| 144 |
+
publisher remain stopped. Full-model quality/native-SM120 qualification remains
|
| 145 |
+
pending; neither pilot proves that outcome.
|