jarrelscy commited on
Commit
cb95231
·
verified ·
1 Parent(s): cbde0ae

Publish quantization/data reproduction source and correct recipe documentation

Browse files

Recovered source, runbooks, recipe/data/environment provenance, CPU verification and path relocation. Correct ARVQ-v2 corpus, objective, schedule, boundary weighting and FP16 reference decode; preserve model weights and original layer reports. Explicitly document historical reproduction gaps.

This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. METHODOLOGY.md +4 -339
  2. README.md +16 -22
  3. arvq88_reference.py +87 -9
  4. calibration_corpus.json +16 -23
  5. pv_progress.json +18 -101
  6. reproduce/AUDIT.md +13 -0
  7. reproduce/README.md +88 -0
  8. reproduce/RUNBOOK.md +29 -0
  9. reproduce/THIRD_PARTY.md +1 -0
  10. reproduce/checkpoint_before_audit.json +100 -0
  11. reproduce/data_artifacts.json +333 -0
  12. reproduce/environment_observed.json +16 -0
  13. reproduce/external_sources.json +18 -0
  14. reproduce/historical_metadata/METHODOLOGY.md +342 -0
  15. reproduce/historical_metadata/backbone_sources.json +0 -0
  16. reproduce/historical_metadata/build_provenance.json +178 -0
  17. reproduce/historical_metadata/calibration_corpus.json +24 -0
  18. reproduce/historical_metadata/cold_assignment.json +0 -0
  19. reproduce/historical_metadata/pv_progress.json +187 -0
  20. reproduce/historical_metadata/source.json +14 -0
  21. reproduce/layer_recipe_summary.json +1727 -0
  22. reproduce/package_manifest.json +247 -0
  23. reproduce/prepare_workspace.py +26 -0
  24. reproduce/recipe.json +116 -0
  25. reproduce/requirements.in +10 -0
  26. reproduce/source/btx53/arvq88/__init__.py +1 -0
  27. reproduce/source/btx53/arvq88/activation.py +44 -0
  28. reproduce/source/btx53/arvq88/alternating.py +64 -0
  29. reproduce/source/btx53/arvq88/arvq_reap.py +174 -0
  30. reproduce/source/btx53/arvq88/backbone.py +44 -0
  31. reproduce/source/btx53/arvq88/checkpoint.py +168 -0
  32. reproduce/source/btx53/arvq88/cli.py +154 -0
  33. reproduce/source/btx53/arvq88/docs/expert_experiments.md +76 -0
  34. reproduce/source/btx53/arvq88/docs/gradient_pv_recipe.md +180 -0
  35. reproduce/source/btx53/arvq88/docs/per_expert_handoff.md +78 -0
  36. reproduce/source/btx53/arvq88/docs/sequential_pilot.md +65 -0
  37. reproduce/source/btx53/arvq88/encoder.py +321 -0
  38. reproduce/source/btx53/arvq88/experiment.py +129 -0
  39. reproduce/source/btx53/arvq88/final_publish.py +69 -0
  40. reproduce/source/btx53/arvq88/fit.py +48 -0
  41. reproduce/source/btx53/arvq88/gate_worker.py +5 -0
  42. reproduce/source/btx53/arvq88/gradient_indices.py +120 -0
  43. reproduce/source/btx53/arvq88/incremental_publish.py +122 -0
  44. reproduce/source/btx53/arvq88/inputs.py +198 -0
  45. reproduce/source/btx53/arvq88/jobs.py +28 -0
  46. reproduce/source/btx53/arvq88/local_source.py +43 -0
  47. reproduce/source/btx53/arvq88/pack.py +142 -0
  48. reproduce/source/btx53/arvq88/perf/FP16_BLOCK_SCALE_HANDOFF.md +19 -0
  49. reproduce/source/btx53/arvq88/perf/FULL_CORPUS_PIPELINE.md +80 -0
  50. reproduce/source/btx53/arvq88/perf/README.md +145 -0
METHODOLOGY.md CHANGED
@@ -1,342 +1,7 @@
1
- # ARVQ cold-expert quantization: methodology
2
 
3
- How the cold Mixture-of-Experts weights in this repository
4
- (`GLM-5.3-Vision-NVFP4-ARVQ-hybrid`) were quantized and tuned. This document
5
- reconstructs the actual pipeline and tuning decisions from the build transcripts
6
- and the fitting/serving code; numbers are cited from those sources. It is a
7
- methodology record, not a quality claim: full-model task quality and native
8
- SM120 execution were **not** evaluated for this checkpoint.
9
 
10
- ---
11
 
12
- ## 1. What this model is (the hybrid)
13
-
14
- GLM-5.3 is a large MoE. Each MoE block routes each token to its top-8 of 256
15
- routed experts (plus a shared expert). This checkpoint is a **hybrid** in which
16
- each expert is stored in one of two ways:
17
-
18
- - **Hot experts** — kept in the NVFP4 donor format (unchanged).
19
- - **Cold experts** — re-quantized to ~2 bits with **ARVQ** (additive residual
20
- vector quantization), the subject of this document.
21
-
22
- The hot/cold split is fixed by a REAP allocation that keeps the **5,750** most
23
- important experts hot; the remaining **13,450** experts across the model are
24
- cold (per the 09-14 build transcript, which verified "13,450 cold-expert source
25
- files"). The allocation is 75% text-REAP / 25% multimodal-salience weighted and
26
- is **not** modified by this campaign.
27
-
28
- Only the **75 MoE blocks, layers 3–77**, are touched. Layers 0–2 are dense and
29
- stay frozen. Everything except the cold experts is inherited unchanged:
30
- attention, backbone, shared experts, hot NVFP4 experts, BF16 MTP, and the
31
- vision components.
32
-
33
- Cold experts are quantized per (layer, projection). The two projections are the
34
- fused gate/up `w13` (N=4096, K=6144) and the down projection `w2` (N=6144,
35
- K=2048).
36
-
37
- ---
38
-
39
- ## 2. ARVQ representation and bit budget
40
-
41
- Each cold expert weight group of **8 contiguous columns** is represented as the
42
- sum of two codebook atoms times a scale:
43
-
44
- ```
45
- w_group = global * block_scale * (c0[a] + c1[b])
46
- ```
47
-
48
- - **Two codebooks** `c0`, `c1`, each **256 entries × 8 dims** (512 codewords per
49
- pair). This is the "rvq256_256x8" format. Codebooks are **per-expert** in this
50
- v3 checkpoint (`codebook_scope = expert`, packed format
51
- `rvq256_256x8_expert`, version 3).
52
- - **Indices** `a`, `b` are one uint8 each per 8-weight group → **16 index bits
53
- per group = 2.0 bits/weight**.
54
- - **Codebook atoms are constrained to the FP4 grid** `{0, ±0.5, ±1, ±1.5, ±2,
55
- ±3, ±4, ±6}` (project-to-FP4 after every update). The FP4 grid on the atoms is
56
- why no incoherence/Hadamard rotation is used: the target space cannot be
57
- rescaled without leaving the grid (see `encoder.py` docstring).
58
- - **Block scales** are stored as **`float8_e4m3fn`**, one scale per **[row,
59
- 128-column block]** (48 blocks/row for `w13`, 16 blocks/row for `w2`). At use
60
- time each block scale is `repeat_interleave(16)` to cover the 16 eight-wide
61
- groups in a 128-column block. Adding the FP8 scale (8 bits / 128 weights =
62
- 0.0625 bit/weight) gives the reported **2.0625 bits/weight** plus the
63
- amortized per-expert codebooks.
64
- - A single fp32 `global` scalar per (layer, projection) sits above the block
65
- scales so the index/scale layout is unchanged from earlier versions.
66
-
67
- Packing (`pack.py`) writes MMA-fragment index tiles (64 uint32 words/tile) and
68
- verifies, per layer, that every expert's indices round-trip exactly and that a
69
- sample of decoded weights matches bit-for-bit before publishing.
70
-
71
- ---
72
-
73
- ## 3. Initial fit (Hessian-aware, no PV yet)
74
-
75
- Before any gradient tuning, each (layer, projection) gets an
76
- activation-Hessian-tuned initialization (`encoder.py`, `fit.py`):
77
-
78
- 1. **Per-expert raw-basis Hessian.** `w13` uses `H = E[x xᵀ]` over routed
79
- tokens; `w2` uses `H = E[m mᵀ]` where `m = silu(x·Wgᵀ)·(x·Wuᵀ)` is the BF16
80
- SwiGLU intermediate. Experts with <32 routed tokens fall back to all captured
81
- tokens. Activations are stored/computed in **float32**; Hessians are half.
82
- 2. **Escalating-damp Cholesky** of `H⁻¹`: damping starts at `1e-3 · mean(diag H)`
83
- and multiplies by 4× (up to 12 attempts) until the factorization succeeds.
84
- 3. **Codebook EM** on a Hessian-importance-weighted subsample (20,000 groups/
85
- expert) of normalized 8-dim groups: k-means warm start → 8 alternating
86
- refine iterations, FP4-projecting the atoms after every M-step.
87
- 4. **Per-expert LDLQ/GPTQ error-feedback column sweep** with the fixed
88
- codebooks (col_block=128, 2 sweep passes, 1 inner refine): plain-L2 inner
89
- assignment, off-diagonal Hessian carries cross-group error feedback.
90
- 5. **Least-squares refit of the E4M3 block scales** given the chosen codes.
91
-
92
- An earlier decision, recorded 09-14: a **tuned Hadamard rotation was tested and
93
- rejected** — a matched 48-expert down-projection refit gave 0.372 (H128 rotated)
94
- versus **0.191 unrotated**, so the unrotated raw-basis fit was adopted.
95
-
96
- Nonfinite guards are raised as errors throughout (the fit/tuner refuse to
97
- proceed on NaN/Inf rather than silently scrubbing).
98
-
99
- ---
100
-
101
- ## 4. The PV objective: `--target reference`
102
-
103
- The cold experts are then output-tuned per layer. The objective for this
104
- checkpoint is **`--target reference`** (`sequential_pv_full_corpus.py`).
105
-
106
- For each captured token position the tuner forms:
107
-
108
- - `frozen` — the block output from everything that is **not** a cold expert on
109
- the **student** input (retained tuned upstream layers → attention → shared
110
- expert → hot NVFP4 experts). This is fixed.
111
- - `reference` — the **original FP8 reference block output**: the unchanged donor
112
- backbone with the **original FP8 routed experts** (all 256), evaluated on the
113
- **original reference trajectory** for that token position.
114
- - `required = reference − frozen` — the residual the cold experts must supply.
115
-
116
- The loss is the relative squared error of the cold-expert sum against `required`:
117
-
118
- ```
119
- loss = || student_cold(x_student) − required ||² / (denominator · Σw)
120
- ```
121
-
122
- where `denominator` is the per-row mean energy of `reference − frozen` over the
123
- corpus (so the scale is comparable across layers). Selection metric is
124
- `reference_rel = ||(frozen + cold) − reference|| / ||reference||`, measured over
125
- the **whole transformer block including the residual**, not the cold-experts sum
126
- alone.
127
-
128
- **Why reference (and its known weakness).** The transcripts weighed
129
- `reference` against a `same_input` target (match the block's output on the
130
- *student's own* drifted input). The rationale for reference (09-16): "Matching
131
- the original reference trajectory is more directly aligned with preserving the
132
- model's behavior." The acknowledged risk: because each layer receives a
133
- **different (student) input** than the original, the cold experts may lack the
134
- capacity to reproduce the reference output and could "learn corrections that
135
- work on training examples but fail on unseen ones" — i.e. reference-target is
136
- the more out-of-distribution objective. It was chosen only after a matched
137
- layer-4 A/B:
138
-
139
- | Layer-4 metric | reference vs same_input |
140
- |---|---|
141
- | validation reference error | **0.21% better** with reference |
142
- | development-audit error | 0.08% **worse** with reference |
143
-
144
- A small mixed difference; reference was retained for behavioral alignment. Two
145
- important design consequences: the target is the **complete block output**
146
- (including the residual), so matching only the MoE contribution cannot leave the
147
- incoming residual error untouched; and the loss compares the **routing-weighted
148
- sum of all cold experts** jointly (gate/up and down together).
149
-
150
- ---
151
-
152
- ## 5. Sequential per-layer pipeline with trajectory coupling
153
-
154
- Layers are processed **sequentially, 3 → 77** (`full_reference.py` /
155
- `sequential_capture.py`):
156
-
157
- 1. **Capture** (8-way sequence-parallel). For layer L, the student input is the
158
- **retained tuned output of block L-1** (`student_outputs.pt`), while the
159
- reference target re-runs the **donor backbone + original FP8 experts** on the
160
- original reference hidden state. Both trajectories are carried forward in a
161
- rolling cache; layer 3 is the control (no preceding ARVQ layer, so student =
162
- reference input). This cross-layer coupling means each layer is tuned against
163
- the exact drifted inputs it will see at serving time.
164
- 2. **Fit / PV tune** the cold experts (Section 6).
165
- 3. **Gate** (Section 7).
166
- 4. **Export + serialized replay** — decode to the packed format and require the
167
- reloaded forward to match the tuned forward to <2e-6 before use.
168
- 5. **Per-layer HF publish** — `publish_reference.py` re-assembles cold slots
169
- across the 8 expert shards, checks slot coverage/allocation/global metadata,
170
- runs the exact index round-trip and sampled decoded-weight parity, then
171
- **atomically replaces that layer's two tensor files + reports in one commit**
172
- with parent-commit protection; remote hashes are re-verified.
173
-
174
- Parallelism: cold experts are **sharded across 8 GPUs** (each rank owns distinct
175
- experts); routed outputs are summed with an all-reduce so a single coupled-output
176
- objective is optimized without a dense-model replica.
177
-
178
- ---
179
-
180
- ## 6. PV tuning numerics (the published full-corpus campaign)
181
-
182
- Per the model card, all 75 layers were replaced by a **full-corpus sequential
183
- PV campaign** drawing sequentially from **18,001,846 training tokens** at
184
- context **1024** (53,248 legacy tokens were removed around 17 exact matches to
185
- held-out prompts). Fixed **validation** and **development-audit** sets each hold
186
- **16,384 tokens**.
187
-
188
- - **Optimizer:** Adam. Two trainable parameter groups — FP4-constrained
189
- per-expert **codebooks** and FP8-constrained per-block **scales**. (The
190
- faster fork trains per-row scale deltas instead; the published full-corpus
191
- fork trains the per-128-block log-scales.)
192
- - **Effective batch:** 262,144 tokens, accumulated as **four 65,536-token
193
- microbatch passes** (256 whole 1024-token sequences), one shuffled
194
- no-replacement pass over the corpus.
195
- - **Learning-rate schedule:** book/scale LR **0.048 / 0.032** through update 45,
196
- then dropped by ×0.25 to **0.012 / 0.008** (`lr_decay_after=45`,
197
- `lr_decay_factor=0.25`). (Note: these campaign LRs are higher than the
198
- in-repo code defaults of 0.003/0.002, consistent with the much larger
199
- 262k-token batch.)
200
- - **Update budget:** layers 4–26 use a fixed **69 updates**. From layer 27, 69
201
- is the maximum with an **adaptive early stop**: three validation checks more
202
- than 0.1% worse than best trigger an earlier LR reduction; 15 updates without
203
- improvement after that reduction permit stopping.
204
- - **Validation every 5 updates**, retaining the best checkpoint.
205
- - **Index reassignment every 20 updates** (Section 6.1).
206
- - **Regularization:** codebook drift-from-anchor penalty + scale drift penalty,
207
- weighted 0.01; grad-norm clip 1.0; codebook atoms clamped to [-6, 6];
208
- log-scales clamped to within ±0.35 of their initial value.
209
- - **Coverage gate:** experts with fewer than **256 routed training rows** are
210
- frozen (gradient masked) and keep their initial books/scales.
211
-
212
- ### 6.1 Output-gradient index reassignment
213
-
214
- Continuous Adam cannot move the discrete indices, so every 20 updates a discrete
215
- proposal pass runs (`gradient_indices.py`, method
216
- `expert_parallel_output_gradient_prefix_backtracking_v1`):
217
-
218
- 1. Backprop the **output** loss to the reconstructed weight of one expert
219
- projection.
220
- 2. Propose alternative `(a,b)` index pairs near a small gradient step, capped at
221
- `max_fraction=0.001` of groups, `trust_ratio=0.01`, `target_ratio=0.03`; the
222
- current pair is always in the candidate set (changing nothing stays legal).
223
- 3. **Prefix backtracking acceptance:** try the top-k proposals with
224
- k∈{full, ¼, 1/16, 1}, accept only if the **training** SSE strictly drops
225
- **and** a separate **check batch** SSE does not increase (beyond 1e-7 slack);
226
- otherwise revert. Proposal deltas are broadcast sparsely (only routed rows).
227
-
228
- The transcripts flag this as the method's main theoretical weakness versus
229
- published PV-Tuning: codebooks are tuned for output accuracy while index
230
- proposals originally came from weight-reconstruction Hessians — the
231
- output-gradient proposal above was added to close that gap.
232
-
233
- ### 6.2 Cold arithmetic emulation
234
-
235
- Both training and evaluation emulate the serving numerics (`activation.py`,
236
- pinned to serving revision `b1380cf7…`): **four FP4 activation planes**
237
- (`fp4_planes4_fp16_boundaries_v1`) and **FP16 SwiGLU boundaries** (FP32 GEMM,
238
- FP16 SiLU/product), with straight-through estimators for gradients. This
239
- emulates the quantization boundaries only — **native SM120 MMA accumulation and
240
- TP reduction are not bit-exact** and were not qualified.
241
-
242
- ---
243
-
244
- ## 7. Acceptance gates
245
-
246
- Publication of a layer requires all of:
247
-
248
- - **Validation non-regression:** `0 ≤ final_reference_rel ≤ initial·(1+1e-6)`.
249
- - **Development-audit non-regression:** the held-out audit split (never used for
250
- updates or checkpoint selection) must satisfy
251
- `audit_reference_rel ≤ initial_audit·(1+1e-6)`; otherwise the layer keeps its
252
- initialization. There is no absolute error floor — the rule is purely
253
- final ≤ initial.
254
- - **Serialized-replay parity:** decoded/reloaded forward matches the tuned
255
- forward to <2e-6, and remote file hashes are re-verified after upload.
256
-
257
- The audit is honestly labeled **development data, not an untouched final test**;
258
- it is a split of the same calibration tokens the initialization already saw. Two
259
- notable per-layer decisions recorded in the card: **layer 30** retains its
260
- audit-qualified update-15 checkpoint after its validation-best update-20 failed
261
- the development audit; **layer 3** retains a separately qualified lower-LR
262
- refinement.
263
-
264
- ---
265
-
266
- ## 8. Calibration corpus
267
-
268
- The text corpus (`build_calib_v31.py` + `reasoning_slice.py`) is deterministic
269
- (seed 42), GLM-tokenized, with the following domain mix by tokens:
270
-
271
- | Share | Domain | Sources |
272
- |---|---|---|
273
- | ~32% | code | local vLLM sources, m-a-p/CodeFeedback, jtatman/python-code-500k |
274
- | ~20% | tool-calling / agentic | generated GLM chat-template sessions |
275
- | ~15% | reasoning | OpenR1-Math-220k + dolphin-r1 `<think>…</think>` → answer |
276
- | ~13% | instruction chat | tatsu-lab/alpaca + coding chat |
277
- | ~10% | medical | MedQA textbook continuation + medical Q&A |
278
- | ~10% | prose | vLLM docs markdown + databricks-dolly-15k |
279
-
280
- The **reasoning slice was added specifically** because the earlier calib-v3 mix
281
- starved the cold 2-bit experts of reasoning-termination behavior: 87% of its
282
- `<think>` blocks were empty. The slice emits multi-turn conversations where an
283
- **earlier** assistant turn carries a real `<think>…</think>` that terminates and
284
- hands off to an answer, followed by a trailing user turn, so the think-close
285
- token **`</think>` (id 154842)** renders in-stream rather than as a trailing
286
- generation prompt. Held-out prompts are explicitly excluded (last 25 vLLM docs,
287
- last 3 MedQA files, vLLM code beyond index 400 of the seed-42 shuffle), and
288
- 17 exact-match prompts were purged from the training set.
289
-
290
- The capture/training code also supports **up-weighting boundary rows** (rows
291
- whose next token is a boundary such as `</think>` / `<|endoftext|>`) via a
292
- `row_weight` term whose sum normalizes the loss, and an optional matched-mixed
293
- validation set; the published card describes reasoning-based validation.
294
-
295
- ---
296
-
297
- ## 9. Results captured during the build
298
-
299
- All errors below are **held-out relative L2**, either over the whole transformer
300
- block (including residual) or over the cold-experts sum, as noted. They are
301
- local reconstruction errors, **not** token-accuracy or perplexity.
302
-
303
- - **Pilot smoke test** (layer 3, one step): full-block error 0.005497 → 0.005456
304
- (~0.75%), 304/388 expert-projection proposals accepted; exported weights
305
- reproduced the retained result.
306
- - **Layer 3** (200-step pilot): full-block reference error **0.005497 →
307
- 0.005271 (−4.12%)**, best at step 200.
308
- - **Layer 4** (200-step pilot): full-block reference error **0.015760 →
309
- 0.015596 (−1.05%)**.
310
- - **reference vs same_input** (layer 4, matched): reference 0.21% better on
311
- validation, 0.08% worse on development audit.
312
- - **Faster fork parity:** the optimized fork (larger microbatch, specialized
313
- embedding-lookup backward, sparse proposal broadcast) matched the reference
314
- fork's validation and audit **exactly** on layers 3 and 4, with all propagated
315
- BF16 outputs equal; training loops fell from ~1525 s → ~337 s (layer 3) and
316
- ~1226 s → ~303 s (layer 4).
317
- - **Early-layer stability example** (layer 18, full-corpus): initial
318
- 0.017205, step-5 0.017176, step-69 0.017183 — differences ≤0.17%, treated as
319
- a signal to watch LR/noise rather than proof of convergence.
320
-
321
- Older cold-only checkpoints reported ~0.2211→0.2070 (layer 3) and 0.2364→0.2320,
322
- but those measured **cold-expert output only on different captures** and are
323
- **not comparable** to the full-block errors above.
324
-
325
- **No comparison against the FP8 donor or an AQLM variant on perplexity / KLD /
326
- top-1 / wikitext was recorded** in the reviewed transcripts, and no full-model
327
- task evaluation was performed. Those remain open.
328
-
329
- ---
330
-
331
- ## 10. Honest limitations
332
-
333
- - Full-model quality and native **SM120** execution were **not** evaluated; the
334
- arithmetic emulation covers FP4/FP16 boundaries only, not MMA/TP bit-exactness.
335
- - The development audit is a split of calibration data, not an independent test.
336
- - The reference target is the more OOD objective; its generalization advantage
337
- over same_input was small and mixed on the one matched layer tested.
338
- - Gates enforce local non-regression, not any absolute quality bar.
339
-
340
- *Prepared from the GLM-5.3 ARVQ build transcripts (2026-09-14 and 2026-09-16)
341
- and the `btx53/arvq88` fitting/serving code. Where a number could not be sourced
342
- it is stated as unknown rather than estimated.*
 
1
+ # ARVQ-v2 methodology correction
2
 
3
+ The previous file was copied from v1 and described the wrong target, scale dtype, corpus and schedule. The current methodology is documented in [reproduce/README.md](reproduce/README.md), with exact stage order in [RUNBOOK.md](reproduce/RUNBOOK.md) and settings in [recipe.json](reproduce/recipe.json).
 
 
 
 
 
4
 
5
+ Key facts: same-input target; original FP8 expert teacher; FP4 per-expert books; FP16 block scales;15,007,754-token cleaned corpus;50× next-boundary loss weighting; reassignment every10; decay by25 or earlier; patience10. Validation selection is unweighted target-relative L2. Most layers stop early. All75 layer receipts remain available. Historical contradictory metadata is archived under reproduce/historical_metadata.
6
 
7
+ The bundled reference packer/decoder now handles both uint8 E4M3 and actual FP16 scales, tested on CPU. No model weights were changed by this correction.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
README.md CHANGED
@@ -9,26 +9,15 @@ Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation.
9
  Other layers retain their previously published weights; see pv_progress.json.
10
  Each layer's two tensor files and reports are replaced together in one commit.
11
 
12
- Training draws sequentially from 18,001,846 text tokens at context 1024. Fixed validation and
13
- development-audit sets each contain 16,384 tokens. Adam trains FP4-constrained
14
- per-expert books and FP8-constrained per-block scales. The effective batch is
15
- 262,144 tokens, accumulated in four 65,536-token passes. Layers 4–26 use 69
16
- updates: book/scale LR .048/.032 through update 45, then .012/.008. From layer27,
17
- 69 updates is the maximum: three validation checks more than 0.1% worse than
18
- best trigger an earlier LR reduction; 15 updates without improvement after
19
- the reduction permit stopping. Layer30 retains audit-qualified update15 after
20
- its validation-best update20 failed the development audit. Layer 3
21
- retains its separately qualified lower-LR refinement. Output-gradient index
22
- reassignment occurs every 20 updates; validation every five updates retains
23
- the best checkpoint. Audit non-regression and export replay gate publication.
24
-
25
- Student inputs reflect retained tuned upstream layers. Fixed reference targets
26
- use original FP8-source routed experts and the donor backbone. A rolling cache
27
- carries both trajectories. Cold arithmetic emulates FP4 activation planes and
28
- FP16 boundaries; native SM120 kernel execution has not yet been smoke-tested
29
- (the evaluation below decodes the published weights to BF16 in software).
30
- The audit set is historical development data, not an untouched final test.
31
- Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
32
 
33
  ## Full-model evaluation (2026-09-22)
34
 
@@ -46,8 +35,7 @@ four held-out corpora):
46
  In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
47
  ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
48
  shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
49
- calibration-domain bias of the ARVQ calibration mix rather than a fitting
50
- regression; treat general-English quality accordingly.
51
 
52
  **Repo history note (2026-09-22):** the hot tier, `aqlm_layer_books` and
53
  `cold_assignment.json` were re-exported to match the cold_manifests allocation.
@@ -78,3 +66,9 @@ format version 4 to match. Earlier branches cannot load this repo:
78
  and predates the v4 format string; the `glm52-sm120` main branch and the v3
79
  `arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
80
  is for the v2r repo (quarter-scale residual decode), not this one.
 
 
 
 
 
 
 
9
  Other layers retain their previously published weights; see pv_progress.json.
10
  Each layer's two tensor files and reports are replaced together in one commit.
11
 
12
+ The production recipe uses **same-input targets**, **FP16 block scales**, and a **50× boundary loss weight**. It differs from v1; previous versions of this README accidentally retained the v1 training description.
13
+
14
+ The cleaned v3.2 training stream contains **15,007,754 tokens**, context1024, effective batch262,144 and microbatch65,536. It removes empty think blocks and frames documents with canonical GLM delimiters. Reasoning validation and development audit captures contain16,384 tokens each; audit is WikiText/OOD and was used during development. The maximum one-pass budget is58updates, with adaptive early stopping: LR .048/.032,4× decay by update25 or earlier after three worsening validation checks, then patience10. Reassignment runs every10updates and validation every5. Published receipts show67/75 layers selecting update5; a full corpus is available, but most layers do **not** consume an entire pass.
15
+
16
+ Boundary weighting is50 on the **position preceding** token154842 (`</think>`) or154820 (`<|endoftext|>`),1 elsewhere. It multiplies squared output residuals once and normalizes by summed weights; it is not2500×. Window-final positions have weight1. Validation metrics and discrete reassignment checks remain unweighted. Books stay on the FP4 grid; learned FP16 block scales serialize directly. The cold index rate is2bpw, or2.125bpw including block scales, plus codebook/global overhead.
17
+
18
+ Student inputs are propagated through the selected upstream hybrid layers. The target is the original FP8-source routed experts evaluated on **those same student inputs**, with frozen hot contribution subtracted. This is not v1's independent original-reference trajectory objective. Two trajectories may still be carried for diagnostic comparison. Reference labels in some historical reports are generic/stale; the reports' `target: same_input` and fitting code determine the actual objective.
19
+
20
+ Hot experts, backbone and vision components are inherited; the original bootstrap hot-allocation mismatch was corrected as noted below. Native runtime correctness is a separate qualification from software decoding.
 
 
 
 
 
 
 
 
 
 
 
21
 
22
  ## Full-model evaluation (2026-09-22)
23
 
 
35
  In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
36
  ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
37
  shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
38
+ possible calibration-domain mismatch. These results do not isolate corpus bias from quantization format, optimization, or trajectory effects; treat general-English quality accordingly.
 
39
 
40
  **Repo history note (2026-09-22):** the hot tier, `aqlm_layer_books` and
41
  `cold_assignment.json` were re-exported to match the cold_manifests allocation.
 
66
  and predates the v4 format string; the `glm52-sm120` main branch and the v3
67
  `arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
68
  is for the v2r repo (quarter-scale residual decode), not this one.
69
+
70
+ ## Reproduce this release
71
+
72
+ The [reproducibility guide](reproduce/README.md), [stage-by-stage runbook](reproduce/RUNBOOK.md), [recipe](reproduce/recipe.json), [data artifact hashes](reproduce/data_artifacts.json), and [source tree](reproduce/source) now accompany this checkpoint. The package contains202 recovered source files spanning corpus preparation, capture, fitting, PV, export and evaluation. Run `python reproduce/verify.py` for CPU checks.
73
+
74
+ **Scope:** the2026-09-26 audit adds code/documentation and corrects metadata; it does not change model weights. Exact historical retraining is not yet guaranteed: some original local corpus inputs, streamed-dataset revisions and environment locks were not preserved. Recovered artifacts, missing inputs and changed-source qualifications are explicitly documented rather than concealed.
arvq88_reference.py CHANGED
@@ -20,7 +20,66 @@ def pack_codebooks(c0,c1):
20
  levels=torch.tensor(LEVELS,device=cb.device);n=(cb[...,None]-levels).abs().argmin(-1)
21
  if not torch.equal(levels[n],cb):raise ValueError('Codebooks must be exactly on the FP4 grid')
22
  return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
23
- def decode(t,layer,proj,e,N=None,K=None):
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24
  pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
25
  default=(4096,6144) if proj=='w13' else (6144,2048)
26
  N,K=default if N is None else (N,K)
@@ -28,16 +87,27 @@ def decode(t,layer,proj,e,N=None,K=None):
28
  if cb.ndim==2:cb=cb[e]
29
  if cb.shape!=(512,):raise ValueError('Invalid packed codebook shape')
30
  values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
31
- scales=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128).view(torch.float8_e4m3fn).float()
32
- return ((values[a.long()]+values[256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
 
 
 
 
33
  def sha(p):
34
  h=hashlib.sha256()
35
  with open(p,'rb') as f:
36
  for chunk in iter(lambda:f.read(8<<20),b''):h.update(chunk)
37
  return h.hexdigest()
38
- def export_layer(source,dest,L):
39
  from safetensors.torch import save_file,load_file
40
  source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
 
 
 
 
 
 
 
41
  original=json.loads((source/'arvq-manifest.json').read_text());files={}
42
  for proj,tag in [('w13','gateup'),('w2','down')]:
43
  d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
@@ -45,20 +115,28 @@ def export_layer(source,dest,L):
45
  if expert and d['c0'].shape!=(E,256,8):raise ValueError('Codebook expert count mismatch')
46
  if proj=='w13':scope='expert' if expert else 'layer'
47
  elif scope!=('expert' if expert else 'layer'):raise ValueError('Mixed projection codebook scopes')
 
 
 
 
 
48
  t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
49
  pre+'codebooks':pack_codebooks(d['c0'],d['c1']),
50
- pre+'scales':d['s'].to(torch.float8_e4m3fn).view(torch.uint8).reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
51
  pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
52
  for e in range(E):
53
  a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
54
  for e in sorted({0,E//2,E-1}):
55
  c0,c1=(d['c0'][e],d['c1'][e]) if expert else (d['c0'],d['c1'])
56
- ref=((c0[d['a'][e].long()]+c1[d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
57
- assert torch.equal(ref,decode(t,L,proj,e,N,K))
 
 
58
  name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
59
- save_file(t,str(tmp),metadata={'format':'pt','arvq_format':'rvq256_256x8_expert' if expert else 'rvq256_256x8','codebook_scope':scope,'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires per-expert v3 loader/kernel' if expert else 'requires 8+8 loader/kernel'})
60
  tmp.replace(dest/name);del t,d
61
  loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64);del loaded
62
  files[proj]={'file':name,'sha256':sha(dest/name)}
63
- m={**original,'format':'rvq256_256x8_expert' if scope=='expert' else 'rvq256_256x8','version':3 if scope=='expert' else 2,'codebook_scope':scope,'bits':2.0625,'codebook_sizes':[256,256],'index_bits':[8,8],'words_per_tile':64,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True}
 
64
  (dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
 
20
  levels=torch.tensor(LEVELS,device=cb.device);n=(cb[...,None]-levels).abs().argmin(-1)
21
  if not torch.equal(levels[n],cb):raise ValueError('Codebooks must be exactly on the FP4 grid')
22
  return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
23
+ def pack_codebooks_mb(cb,factors):
24
+ """mcbook16: cb [E,256+B*256,8] in EFFECTIVE units. Rows are stored as plain-grid
25
+ nibbles; decode multiplies residual book m by factors[m] (exact powers of two)."""
26
+ B=len(factors)
27
+ if cb.ndim!=3 or cb.shape[1]!=256+B*256 or cb.shape[2]!=8:raise ValueError('Invalid mcbook codebook shape')
28
+ f=torch.cat([torch.ones(256),torch.tensor(factors).float().repeat_interleave(256)])[None,:,None]
29
+ grid=cb/f
30
+ levels=torch.tensor(LEVELS,device=cb.device);n=(grid[...,None]-levels).abs().argmin(-1)
31
+ if not torch.equal(levels[n]*f,cb):raise ValueError('Codebooks must be exactly on their per-book grids')
32
+ return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
33
+ def decode_mb(t,layer,proj,e,N,K):
34
+ pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
35
+ a,b=unpack(t[pre+'packed'][e],N,K)
36
+ cb=t[pre+'codebooks'][e].long();factors=t[pre+'book_factors'].float();B=len(factors)
37
+ if cb.shape!=(256+B*256,):raise ValueError('Invalid mcbook packed codebook shape')
38
+ values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
39
+ values=torch.cat([values[:256],values[256:]*factors.repeat_interleave(256)[:,None]])
40
+ sel=t[pre+'selectors'][e].long()
41
+ m=sel.repeat_interleave(16,0).repeat_interleave(8,1)
42
+ stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
43
+ if stored.dtype==torch.float16:scales=stored.float()
44
+ elif stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
45
+ else:raise ValueError('Unsupported packed block scale dtype')
46
+ return ((values[a.long()]+values[256+m*256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
47
+ def export_layer_mb(source,dest,L):
48
+ """mcbook16 export (format version 5): 16 residual books, per-tile 4-bit selectors,
49
+ fp16 block scales; packed index stream and scales unchanged from v4."""
50
+ from safetensors.torch import save_file,load_file
51
+ source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
52
+ original=json.loads((source/'arvq-manifest.json').read_text());files={};first_factors=None
53
+ for proj,tag in [('w13','gateup'),('w2','down')]:
54
+ d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
55
+ if d.get('scale_dtype')!='fp16':raise ValueError('mcbook export requires fp16 block scales')
56
+ if not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
57
+ factors=[float(f) for f in d['book_factor']]
58
+ if first_factors is None:first_factors=factors
59
+ elif factors!=first_factors:raise ValueError('Mixed projection book factors')
60
+ sel=d['selector']
61
+ if sel.shape!=(E,N//16,K//64) or int(sel.max())>=len(factors):raise ValueError('Invalid selector tensor')
62
+ t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
63
+ pre+'codebooks':pack_codebooks_mb(d['cb'],factors),
64
+ pre+'selectors':sel.to(torch.uint8).contiguous(),
65
+ pre+'book_factors':torch.tensor(factors,dtype=torch.float32),
66
+ pre+'scales':d['s'].half().reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
67
+ pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
68
+ for e in range(E):
69
+ a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
70
+ for e in sorted({0,E//2,E-1}):
71
+ m=d['selector'][e].long().repeat_interleave(16,0).repeat_interleave(8,1)
72
+ cb=d['cb'][e]
73
+ ref=((cb[d['a'][e].long()]+cb[256+m*256+d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
74
+ assert torch.equal(ref,decode_mb(t,L,proj,e,N,K))
75
+ name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
76
+ save_file(t,str(tmp),metadata={'format':'pt','arvq_format':'rvq256_mb16_256x8_expert_fp16block','codebook_scope':'expert','residual_books':str(len(factors)),'book_factors':json.dumps(factors),'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires mcbook16 v5 loader/kernel (vllm-glm52-sm120 branch experiment/arvq-mcbook16)'})
77
+ tmp.replace(dest/name)
78
+ loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64) and loaded[pre+'selectors'].shape==(E,N//16,K//64);del loaded
79
+ files[proj]={'file':name,'sha256':sha(dest/name)};del t,d
80
+ m={**original,'format':'rvq256_mb16_256x8_expert_fp16block','version':5,'residual_scale_shift':0,'block_scale_dtype':'float16','codebook_scope':'expert','bits':2.125+4/1024,'codebook_sizes':[256,16*256],'index_bits':[8,8],'selector_bits_per_tile':4,'words_per_tile':64,'book_factors':first_factors,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True,'requires_mcbook16_loader':True}
81
+ (dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
82
+ def decode(t,layer,proj,e,N=None,K=None,residual_scale_shift=0):
83
  pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
84
  default=(4096,6144) if proj=='w13' else (6144,2048)
85
  N,K=default if N is None else (N,K)
 
87
  if cb.ndim==2:cb=cb[e]
88
  if cb.shape!=(512,):raise ValueError('Invalid packed codebook shape')
89
  values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
90
+ stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
91
+ if stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
92
+ elif stored.dtype==torch.float16:scales=stored.float()
93
+ else:raise ValueError('Unsupported packed block scale dtype')
94
+ f=2.0**(-int(residual_scale_shift))
95
+ return ((values[a.long()]+f*values[256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
96
  def sha(p):
97
  h=hashlib.sha256()
98
  with open(p,'rb') as f:
99
  for chunk in iter(lambda:f.read(8<<20),b''):h.update(chunk)
100
  return h.hexdigest()
101
+ def export_layer(source,dest,L,residual_scale_shift=0):
102
  from safetensors.torch import save_file,load_file
103
  source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
104
+ probe=torch.load(source/'w13.pt',map_location='cpu',weights_only=True,mmap=True)['w13']
105
+ if 'selector' in probe:
106
+ del probe
107
+ if int(residual_scale_shift):raise ValueError('mcbook exports are shift-0 (per-book factors)')
108
+ return export_layer_mb(source,dest,L)
109
+ del probe
110
+ rss=int(residual_scale_shift);f=2.0**(-rss)
111
  original=json.loads((source/'arvq-manifest.json').read_text());files={}
112
  for proj,tag in [('w13','gateup'),('w2','down')]:
113
  d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
 
115
  if expert and d['c0'].shape!=(E,256,8):raise ValueError('Codebook expert count mismatch')
116
  if proj=='w13':scope='expert' if expert else 'layer'
117
  elif scope!=('expert' if expert else 'layer'):raise ValueError('Mixed projection codebook scopes')
118
+ fp16=d.get('scale_dtype','fp8_e4m3')=='fp16'
119
+ if proj=='w13':fp16_blocks=fp16
120
+ elif fp16_blocks!=fp16:raise ValueError('Mixed projection scale precision')
121
+ if fp16 and not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
122
+ stored_scales=d['s'].half() if fp16 else d['s'].to(torch.float8_e4m3fn).view(torch.uint8)
123
  t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
124
  pre+'codebooks':pack_codebooks(d['c0'],d['c1']),
125
+ pre+'scales':stored_scales.reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
126
  pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
127
  for e in range(E):
128
  a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
129
  for e in sorted({0,E//2,E-1}):
130
  c0,c1=(d['c0'][e],d['c1'][e]) if expert else (d['c0'],d['c1'])
131
+ ref=((c0[d['a'][e].long()]+f*c1[d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
132
+ assert torch.equal(ref,decode(t,L,proj,e,N,K,rss))
133
+ base_fmt='rvq256_256x8_expert_fp16block' if fp16 else ('rvq256_256x8_expert' if expert else 'rvq256_256x8')
134
+ arvq_fmt=base_fmt+(f'_rs{int(2**rss)}' if rss else '')
135
  name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
136
+ save_file(t,str(tmp),metadata={'format':'pt','arvq_format':arvq_fmt,'residual_scale_shift':str(rss),'codebook_scope':scope,'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires FP16 block-scale v4 loader/kernel' if fp16 else ('requires per-expert v3 loader/kernel' if expert else 'requires 8+8 loader/kernel')})
137
  tmp.replace(dest/name);del t,d
138
  loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64);del loaded
139
  files[proj]={'file':name,'sha256':sha(dest/name)}
140
+ base_mfmt='rvq256_256x8_expert_fp16block' if fp16_blocks else ('rvq256_256x8_expert' if scope=='expert' else 'rvq256_256x8')
141
+ m={**original,'format':base_mfmt+(f'_rs{int(2**rss)}' if rss else ''),'version':4 if fp16_blocks else (3 if scope=='expert' else 2),'residual_scale_shift':rss,'block_scale_dtype':'float16' if fp16_blocks else 'float8_e4m3fn','codebook_scope':scope,'bits':2.125 if fp16_blocks else 2.0625,'codebook_sizes':[256,256],'index_bits':[8,8],'words_per_tile':64,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True}
142
  (dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
calibration_corpus.json CHANGED
@@ -1,24 +1,17 @@
1
  {
2
- "source": "/tmp/glm53-unc-traces",
3
- "reasoning_rollouts": true,
4
- "counts": {
5
- "train": {
6
- "problems": 856,
7
- "tokens": 3048596
8
- },
9
- "validation": {
10
- "problems": 89,
11
- "tokens": 345875
12
- },
13
- "audit": {
14
- "problems": 104,
15
- "tokens": 322277
16
- }
17
- },
18
- "partitions": "disjoint prompt hashes; separately captured",
19
- "limitations": "Text-only captures in 2048-token windows; source datasets overlap earlier calibration. Teacher answers are not correctness-verified.",
20
- "initial_fit_partition": "train only",
21
- "capture_teacher": "RadixArk/GLM-5.3-NVFP4",
22
- "rollout_teacher": "unc NVFP4 teacher",
23
- "same_token_corpus_as": "/tmp/glm53-kernel-aware-pv/full"
24
- }
 
1
  {
2
+ "stage": "current recipe provenance correction",
3
+ "variant": "ARVQ-v2",
4
+ "full_pv_training_tokens": 15007754,
5
+ "sequence_length": 1024,
6
+ "validation_rows": 16384,
7
+ "audit_rows": 16384,
8
+ "boundary_boost": 50,
9
+ "boundary_token_ids": [
10
+ 154842,
11
+ 154820
12
+ ],
13
+ "data_pipeline": "reproduce/README.md",
14
+ "data_hashes": "reproduce/data_artifacts.json",
15
+ "historical_component_metadata": "reproduce/historical_metadata/calibration_corpus.json",
16
+ "exact_retraining_guaranteed": false
17
+ }
 
 
 
 
 
 
 
pv_progress.json CHANGED
@@ -77,111 +77,28 @@
77
  77
78
  ],
79
  "total_layers": 75,
80
- "recipe": "full_corpus_sequential_lr_decay",
81
- "training_tokens": 18001846,
82
  "batch_tokens": 262144,
83
  "microbatch_tokens": 65536,
84
- "max_updates_per_layer": 69,
85
  "adaptive_schedule": {
86
- "from_layer": 27,
87
  "relative_worsening": 0.001,
88
  "checks": 3,
89
- "patience_updates": 15,
90
- "max_updates": 69,
91
- "selection": "existing fixed reasoning validation; retain best checkpoint",
92
- "description": "Drop LR4x after three consecutive validation checks >0.1% worse than best, or after45 at latest. After reduction stop following15 updates without a new best."
93
  },
94
- "lr_decay_after": 45,
95
  "lr_decay_factor": 0.25,
96
- "layer3_exception": "45 high-LR updates plus retained lower-LR refinement checkpoint",
97
- "remaining_layers": "retain their previously published weights",
98
- "previous_campaign_progress": {
99
- "layers_pv_complete": [
100
- 3,
101
- 4,
102
- 5,
103
- 6,
104
- 7,
105
- 8
106
- ],
107
- "total_layers": 75,
108
- "recipe": "sequential_reference_gradient_pv_v1",
109
- "legacy_pv_layers": [
110
- 9,
111
- 10,
112
- 11,
113
- 12,
114
- 13,
115
- 14
116
- ],
117
- "initial_fit_layers": [
118
- 15,
119
- 16,
120
- 17,
121
- 18,
122
- 19,
123
- 20,
124
- 21,
125
- 22,
126
- 23,
127
- 24,
128
- 25,
129
- 26,
130
- 27,
131
- 28,
132
- 29,
133
- 30,
134
- 31,
135
- 32,
136
- 33,
137
- 34,
138
- 35,
139
- 36,
140
- 37,
141
- 38,
142
- 39,
143
- 40,
144
- 41,
145
- 42,
146
- 43,
147
- 44,
148
- 45,
149
- 46,
150
- 47,
151
- 48,
152
- 49,
153
- 50,
154
- 51,
155
- 52,
156
- 53,
157
- 54,
158
- 55,
159
- 56,
160
- 57,
161
- 58,
162
- 59,
163
- 60,
164
- 61,
165
- 62,
166
- 63,
167
- 64,
168
- 65,
169
- 66,
170
- 67,
171
- 68,
172
- 69,
173
- 70,
174
- 71,
175
- 72,
176
- 73,
177
- 74,
178
- 75,
179
- 76,
180
- 77
181
- ],
182
- "format": "rvq256_256x8_expert",
183
- "full_model_quality": "pending"
184
- },
185
- "format": "rvq256_256x8_expert",
186
- "full_model_quality": "pending"
187
- }
 
77
  77
78
  ],
79
  "total_layers": 75,
80
+ "recipe": "full_corpus_sequential_same_input_fp16block",
81
+ "training_tokens": 15007754,
82
  "batch_tokens": 262144,
83
  "microbatch_tokens": 65536,
84
+ "max_updates_per_layer": 58,
85
  "adaptive_schedule": {
 
86
  "relative_worsening": 0.001,
87
  "checks": 3,
88
+ "patience_updates": 10
 
 
 
89
  },
90
+ "lr_decay_after": 25,
91
  "lr_decay_factor": 0.25,
92
+ "layer3_exception": "v2 layer3 uses same LR recipe; selected step50 after58 updates, per its published report",
93
+ "format": "rvq256_256x8_expert_fp16block",
94
+ "full_model_quality": "See README full-model evaluation; software-dequant evaluation, not a guarantee of native runtime parity",
95
+ "target": "same_input",
96
+ "boundary_boost": 50.0,
97
+ "boundary_token_ids": [
98
+ 154842,
99
+ 154820
100
+ ],
101
+ "reassign_every": 10,
102
+ "block_scale_dtype": "float16",
103
+ "metadata_corrected_on": "2026-09-26"
104
+ }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
reproduce/AUDIT.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Audit evidence and corrections, 2026-09-26
2
+
3
+ Evidence precedence: actual checkpoint tensor headers and75 per-layer PV reports; saved campaign config; executable capture/loss/packing source; local RUNLOG53 chronology and Claude build conversation; model-card prose. Private transcripts are not uploaded.
4
+
5
+ Confirmed v2 defects corrected: copied v1 README and METHODOLOGY; incorrect18,001,846-token claim (actual15,007,754); FP8 vs actual FP16 scales; reference vs same-input target; decay45/patience15/reassign20 vs25/10/10; missing50× boundary loss rule; stale pv_progress; root reference decoder reinterpreting FP16 bytes as E4M3. Header samples show U8 in v1 and F16 in v2. CPU fixtures verify corrected decode in both formats.
6
+
7
+ v1 methodology now distinguishes ARVQ output-benefit allocation from the old AQLM allocation. Both ARVQ corpus metadata files now distinguish earlier/component provenance from final PV streams. AQLM notes clarify sampled activation PV, global75/25 row mixture,65536-entry shared layer/projection books and absence of boundary weighting.
8
+
9
+ v2 WikiText degradation is measured; attributing it solely to corpus bias was too strong and has been corrected. Quantization and optimization effects were not isolated by a controlled ablation.
10
+
11
+ Historical per-layer `checkpoint_selection` strings can be stale even when `target` and recorded metrics are correct; original reports are retained, not retroactively rewritten. Old publisher scripts are archived and can regenerate stale prose; use the corrected root documentation and recipe files as the current documentation contract.
12
+
13
+ Missing provenance is explicitly listed. This package is a practical reconstruction/runbook and source release, not certification of byte-identical historical retraining.
reproduce/README.md ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Reproducing the GLM-5.3 hybrid releases
2
+
3
+ This package was reconstructed on 2026-09-26 from the local quantization source, saved configurations, per-layer reports, checkpoint tensor headers, and Claude build history. The original private conversations are not redistributed. `checkpoint_before_audit.json` pins the weights/reports audited; the present documentation-only revision does not retrain or replace weights.
4
+
5
+ ## Reproducibility status
6
+
7
+ The fitting, capture, data-building, evaluation, packing and verification source is included in `source/` with SHA-256 checksums. This is a **recovered working-tree snapshot**, containing subsequent fixes and experiments, not a falsely reconstructed original Git commit. `recipe.json` selects the intended recipe; do not infer it from experimental filenames. Data hashes and current environment versions are included. Existing model weights and original per-layer reports remain the reference artifacts.
8
+
9
+ A bit-for-bit re-training claim is **not justified**: original streamed dataset revisions, local corpus inputs and the complete original dependency lockfile were not all retained. Some training paths were overwritten by later campaigns. The original AQLM `/data` activation and sample archives are no longer present at their recorded paths; the v1 18,001,846-token flat stream has not been recovered. These gaps must be resolved or a rerun must be labeled a new reproduction, with measured differences. Neither a fixed random seed nor the current package versions restores missing historical inputs.
10
+
11
+ ## Package layout
12
+
13
+ - `source/tools/`: text corpus builders, source loading/NVFP4 emulation, capture, AQLM fitting, PV and validation.
14
+ - `source/tools/capture53/`: GLM5.3 wrappers, multimodal corpus/capture, allocation, checkpoint assembly, MTP conversion, vision graft and full-model PPL/KL evaluation.
15
+ - `source/btx53/`: FP8 source ingest, raw-basis ARVQ initializer, activation emulation, coupled sequential PV, capture/propagation, REAP allocation and export tooling; `arvq88/perf/` contains the full-corpus controller and its gates.
16
+ - `recipe.json`, `layer_recipe_summary.json`: recovered variant-specific settings and original layer exceptions.
17
+ - `data_artifacts.json`: SHA-256, sizes/shapes and token-boundary counts for recovered token files; paths are historical identities, not downloadable links.
18
+ - `historical_metadata/`: original conflicting metadata retained as historical evidence, not the current recipe.
19
+ - `prepare_workspace.py`: creates a new relocatable copy; does not overwrite the archive or execute a training job.
20
+ - `verify.py`: CPU-only source checksum, boundary-weight semantics and FP8/FP16 reference-decoder checks.
21
+
22
+ ## Environment and paths
23
+
24
+ Training/capture used eight A100s with software dequantization/emulation. This is not proof of native SM120 kernel equivalence. `environment_observed.json` records the current audit environment, **not** the original training environment. `requirements.in` lists main Python dependencies, not an exact historical lock. Use a CUDA-compatible PyTorch build and a Transformers version supporting the donor tokenizer; install additional optional dependencies demanded by the chosen multimodal backend. Data-source scripts record their imports. Native serving requires the custom serving fork described on the model card.
25
+
26
+ ```bash
27
+ python -m venv .venv
28
+ . .venv/bin/activate
29
+ pip install -r reproduce/requirements.in
30
+ python reproduce/verify.py
31
+ python reproduce/prepare_workspace.py --dest /absolute/path/glm53-reproduction
32
+ cd /absolute/path/glm53-reproduction
33
+ # Use the same activated interpreter for all following commands.
34
+ ```
35
+
36
+ The preparer copies source and maps legacy `/data` and `/tmp` paths into the destination, and the old checkout path into this source tree. It writes the transformed source hashes and path map. Review `recipe.json` and scripts before GPU work: donor, tokenizer, corpus, source-cache and baseline locations must exist. Historical scripts that spawn `ROOT/.venv/bin/python` need that interpreter available (the preparer creates the link). Old publication/watchdog scripts are archival; **do not run them against the original public repos**. Export to a new directory and use your own destination repo after verification.
37
+
38
+ Acquire donor/base/vision snapshots using pinned revisions in `historical_metadata/source.json`, `backbone_sources.json`, available model-index provenance, and the audited checkpoint metadata. Where a revision is missing or marked `legacy_cache_revision_unverified`, record that gap rather than substituting today's HEAD as if it were historical. The source FP8 checkpoint and NVFP4 donor are distinct teachers. AQLM uses the NVFP4 donor; ARVQ cold fitting uses original FP8 expert weights. `btx53/tools/prefetch_cold_fp8.py`, `arvq88/inputs.py`, `local_source.py` and `tools/aqlm_quantize.py` implement source fetching/dequantization.
39
+
40
+ ## Data pipelines and non-portable inputs
41
+
42
+ ### AQLM text calibration (v3)
43
+
44
+ `tools/build_calib_v3.py`: seed42, target15M tokens, approximately40% code /25% agentic /15% instruction /10% medical /10% prose. Streams CodeFeedback, python-code-dataset-500k, Alpaca and Dolly; also uses local vLLM sources/docs, local tools, synthetic shell/agent sessions and MedQA material. See the builder for exact generator order and per-category stopping. It excludes reserved vLLM/MedQA files, but many dataset revisions were not pinned. The original `collect_expert_stats_v2.py` also reads host logs, diffs and shell outputs: these **host inputs must be frozen and reviewed** to reproduce the original corpus; they are not included here. Do not publish freshly collected private logs merely to recreate this pipeline.
45
+
46
+ The initial capture uses nonoverlapping2048-token windows. `stream_capture53.py stats` records routing/salience and sampled activations; it is not a claim that later AQLM PV trains on every15M token. AQLM convergence and PV use the retained activation sample.
47
+
48
+ ### Multimodal branch
49
+
50
+ `capture53/build_calib_mm.py`: seed42,1200 image/caption samples,300 each from medical, natural images, OCR/documents and screenshots. Primary candidates: eltorio/ROCOv2-radiology, lmms-lab/COCO-Caption2017, naver-clova-ix/synthdog-en and HuggingFaceM4/WebSight v0.2; explicit fallbacks live in the script. Retain generated `samples.jsonl`, selected dataset/config/split, dataset revision, sample IDs and checksums. Images are resized448×448 and processed with the checkpoint KimiK25 processor into256 vision tokens. Source licenses differ (ROCOv2 includes a noncommercial/share-alike restriction); model MIT metadata does not relicense the datasets. No image or caption archive is redistributed here.
51
+
52
+ `capture_mm53.py` splices the GLM5.2V tower/projector outputs and masks padding. `solve_r3.py` allocates using .75*per-layer-normalized text REAP + .25*per-layer-normalized multimodal salience. Mixed activation files concatenate32768 text rows and10922 multimodal rows. Thus75/25 is a **global row mixture**, not a guarantee that each expert's routed Hessian contains exactly25% visual energy/rows.
53
+
54
+ ### ARVQ v1 full-corpus stream
55
+
56
+ The published75 layer reports describe18,001,846 tokens (17,580 windows including the short tail), context1024. Historical metadata's3,048,596-token reasoning corpus identifies a component/earlier capture, not the complete final stream. The exact original concatenated stream is not recovered. Do not replace it silently with the current15M v32 file or label v32 a faithful v1 data reproduction. Capture reads `train_token_file`; supply a recovered checksum-verified v1 stream to reproduce that campaign. Use original layer reports for early-stop and recovery exceptions.
57
+
58
+ ### ARVQ v2 v3.2 data rebuild
59
+
60
+ `tools/build_calib_v32.py` actual executable budget is22% code /15% agentic /30% reasoning /13% instruction /10% medical /10% prose, seed42, target15M; actual complete documents produce15,007,754 tokens. Its old opening docstring predates the30% reasoning increase; the budget constants and this runbook are authoritative. Medical data switches to MedRAG textbooks when old local MedQA is unavailable. Reasoning sources include OpenR1-Math-220k and dolphin-r1. Reasoning is shortened to3500 characters, answers1200. Dataset availability, streaming order and local files can affect exact regeneration.
61
+
62
+ `clean_think` removes empty/whitespace-only think blocks, trailing generation-only prompts and unmatched think tags while retaining content. `frame_ids` adds canonical `[gMASK]<sop>` prefix and `<|endoftext|>` suffix only when absent. Use donor tokenization; IDs154822/154824/154820 correspond to those delimiters. Concatenate `shard_*.npy` in sorted order into the configured `train.npy`, then compare length and SHA to `data_artifacts.json`.
63
+
64
+ `tools/build_trace_refit_v32.py` builds three eval-capture sources: training reference from the first1M calibration tokens; held-out reasoning with RNG2025 skipping9000 reasoning documents; and OOD WikiText test text. The capturer uses RNG1234 to select64/16/16 nonoverlapping1024-token sequences, yielding65536 training-probe and16384 validation/audit rows. The procedural skip is not an independent content-deduplication guarantee; the audit split was repeatedly inspected during development.
65
+
66
+ ## Boundary weighting: exact semantics
67
+
68
+ For v2, within each1024-token sequence, set `w[t]=50` when the **next** token is154842 (`</think>`) or154820 (`<|endoftext|>`); otherwise1. The last row is1 because the implementation does not look across the window boundary. The same rule is in standalone capture and fused handoff. It weights the hidden output responsible for predicting a boundary, not the boundary token's own output.
69
+
70
+ The PV data term is `sum_t w[t] * ||cold_prediction[t] - required[t]||² / (D * sum_t w[t])`, with D the captured target-energy normalizer. The residual is squared **before** multiplying by50; the multiplier is50, not2500. The small anchor regularizer is separate. Reported validation/audit relative-L2 and boundary/bulk summaries are **unweighted**; checkpoint selection uses validation `target_rel`. Discrete index-reassignment proposal/check objectives are also unweighted in the recovered implementation. The50× multiplier is therefore a continuous PV-loss rule, not a blanket multiplier on all quantization/evaluation stages.
71
+
72
+ AQLM and ARVQ v1 have no identified50× boundary term in their production recipes. Do not retroactively attribute the v2 change to them. Both the corpus cleanup and weighting were motivated by boundary preservation; no causal guarantee of shorter generation follows from this alone.
73
+
74
+ ## Verification and acceptance
75
+
76
+ Before fitting: verify corpus hashes/counts, donor identity, tokenizer special IDs, cold/hot partition coverage and target dimensions. Freeze the assignment. Initial-fit output must cover every cold expert. After PV: retain the selected checkpoint, check finite metrics and audit/non-regression policy, then pack and reload for index and decoded-weight parity. Propagate the selected weights and recompute the next layer's inputs; do not mix independently repaired layers without re-evaluation. Save scripts/config hashes, training receipts, all per-layer metrics and environment lockfiles for the new run.
77
+
78
+ `tools/capture53/eval3_kld.py` / `eval_hybrid53.py` evaluate teacher-forced PPL/KL. `btx53/arvq88/pack.py` verifies packed indices and sample decoded weights; the root `arvq88_reference.py` supports actual uint8-E4M3 and float16 scale storage. v3 cold storage is2+8/128=2.0625 bpw before codebook overhead; v4 is2+16/128=2.125 bpw before overhead. AQLM's2bits likewise describes index rate, excluding its65536×8 codebook and scales. Verify actual bytes if comparing compression budgets.
79
+
80
+ This audit ran lightweight CPU verification; it did not rerun a multi-day quantization or certify native GPU serving. Historical PPL/task results on the model card retain their original scope.
81
+
82
+ ## Variant-specific commands
83
+
84
+ Read [RUNBOOK.md](RUNBOOK.md) for this repository’s stage order and recipe.
85
+
86
+ ## External code and optional modules
87
+
88
+ `external_sources.json` records observed upstream commits. Install b12x at its recorded revision only for the optional BTX comparison modules; these are not part of the ARVQ production recipe. The vLLM checkout is also an input to local code/prose generation and must be placed at `source/vllm` (or the relocated workspace’s `vllm`). Its current commit is not proof of the original dirty training checkout. Dataset IDs/licenses appear in the builders; preserve the resolved revisions and generated sample manifests on a new run. See THIRD_PARTY.md.
reproduce/RUNBOOK.md ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ARVQ-v2 runbook
2
+
3
+ Run the source from the prepared workspace. Preserve all75 original layer reports and the original cold allocation. `recipe.json` is the recovered v2 configuration.
4
+
5
+ 1. Prepare source experts from `zai-org/GLM-5.3` and the `RadixArk/GLM-5.3-NVFP4` donor. The published `source.json` pins a source revision but also marks legacy-cache revision verification unavailable; retain that qualification. Source dequantization uses block128×128 E4M3. Prepare the vision extension and capture inputs using the included capture tools.
6
+ 2. Run `tools/build_calib_v32.py`, concatenate sorted shards into `train.npy`, and run `tools/build_trace_refit_v32.py`. Confirm15,007,754 training tokens and the recorded hashes.
7
+ 3. Generate initial ARVQ fits and the ARVQ output-benefit allocation. Historical v2 launch (replace paths):
8
+
9
+ ```bash
10
+ python btx53/build_baseline.py --pv-steps 0 --workdir BASELINE \
11
+ --calib-tokens CALIB_SHARDS --calib-token-count 15000000 \
12
+ --src-cache FP8_EXPERT_CACHE --nvfp4-donor DONOR \
13
+ --dst baseline-only --gpus 0,1,2,3,4,5,6,7 --codebook-scope expert
14
+ ```
15
+
16
+ `--codebook-scope expert` is required. `arvq88/encoder.py` and `fit.py` implement raw-basis Hessian/LDLQ fitting; `arvq_reap.py` blends text/MM error-reduction benefit75/25 and keeps5750 hot experts. To reproduce the published assignment rather than resolve it anew, use the published `cold_assignment.json` and verify against every cold manifest. Never substitute the original AQLM allocation merely because the hot count matches.
17
+
18
+ 4. Place `recipe.json` as `WORK/config.json`, configure source/cache/corpus/baseline paths, and launch:
19
+
20
+ ```bash
21
+ python btx53/arvq88/perf/full_pipeline.py --work WORK --first 3 --last 77
22
+ ```
23
+
24
+ The controller captures full training rows plus fixed eval rows, runs coupled expert-parallel PV, gates retained outputs, and propagates the chosen student state. Fused handoff requires its bitwise replay qualification. The source includes later optimizations; preserve receipts/identity hashes and a fresh WORK if changing the recipe.
25
+
26
+ 5. v2: same-input source target, FP16 block scales,50× boundary weighting, LR .048/.032 with4× decay by update25 or earlier, patience10, index reassignment every10, evaluation every5. 58updates would cover the full15,007,754-token corpus; adaptive stopping often uses fewer. The published reports show67layers selecting update5. Do not describe this as one complete pass for every layer.
27
+ 6. Export through `arvq88/pack.py` / checkpoint assembly to a **new** destination. Run packed-index and decoded-weight parity, hot/cold consistency, full-model PPL/KL and native serving tests. The historical publishers reference public production repos and are provided for provenance, not as a safe rerun destination. Fixed metadata in this audit is the reference; historical publisher templates can reintroduce stale descriptions.
28
+
29
+ A v2 model-wide v4 marker requires FP16-scale cold layers. Do not deploy a partially replaced v3/v4 mixture. The old bootstrap also had a hot-tier mismatch, corrected at f53a8dcc; use a later audited snapshot.
reproduce/THIRD_PARTY.md ADDED
@@ -0,0 +1 @@
 
 
1
+ Source files retain their original notices. External checkouts and dataset material have their own licenses; the model card MIT tag does not override those. See external_sources.json and the dataset builder source for origins. The b12x/quant-toolkit/vLLM repositories are referenced rather than bulk vendored. No new license is asserted over third-party code or datasets by this audit.
reproduce/checkpoint_before_audit.json ADDED
@@ -0,0 +1,100 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "repo": "jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid",
3
+ "revision": "cbde0aed09e981e68bfefc827cb22d5fbdb81cfb",
4
+ "files": [
5
+ "METHODOLOGY.md",
6
+ "README.md",
7
+ "arvq88_reference.py",
8
+ "backbone_sources.json",
9
+ "btx_hybrid_manifest.json",
10
+ "build_provenance.json",
11
+ "calibration_corpus.json",
12
+ "cold_assignment.json",
13
+ "config.json",
14
+ "gates_report.json",
15
+ "generation_config.json",
16
+ "kimi_k25_processor.py",
17
+ "kimi_k25_vision_processing.py",
18
+ "media_utils.py",
19
+ "model.safetensors.index.json",
20
+ "preprocessor_config.json",
21
+ "pv_layers/layer_003.json",
22
+ "pv_layers/layer_004.json",
23
+ "pv_layers/layer_005.json",
24
+ "pv_layers/layer_006.json",
25
+ "pv_layers/layer_007.json",
26
+ "pv_layers/layer_008.json",
27
+ "pv_layers/layer_009.json",
28
+ "pv_layers/layer_010.json",
29
+ "pv_layers/layer_011.json",
30
+ "pv_layers/layer_012.json",
31
+ "pv_layers/layer_013.json",
32
+ "pv_layers/layer_014.json",
33
+ "pv_layers/layer_015.json",
34
+ "pv_layers/layer_016.json",
35
+ "pv_layers/layer_017.json",
36
+ "pv_layers/layer_018.json",
37
+ "pv_layers/layer_019.json",
38
+ "pv_layers/layer_020.json",
39
+ "pv_layers/layer_021.json",
40
+ "pv_layers/layer_022.json",
41
+ "pv_layers/layer_023.json",
42
+ "pv_layers/layer_024.json",
43
+ "pv_layers/layer_025.json",
44
+ "pv_layers/layer_026.json",
45
+ "pv_layers/layer_027.json",
46
+ "pv_layers/layer_028.json",
47
+ "pv_layers/layer_029.json",
48
+ "pv_layers/layer_030.json",
49
+ "pv_layers/layer_031.json",
50
+ "pv_layers/layer_032.json",
51
+ "pv_layers/layer_033.json",
52
+ "pv_layers/layer_034.json",
53
+ "pv_layers/layer_035.json",
54
+ "pv_layers/layer_036.json",
55
+ "pv_layers/layer_037.json",
56
+ "pv_layers/layer_038.json",
57
+ "pv_layers/layer_039.json",
58
+ "pv_layers/layer_040.json",
59
+ "pv_layers/layer_041.json",
60
+ "pv_layers/layer_042.json",
61
+ "pv_layers/layer_043.json",
62
+ "pv_layers/layer_044.json",
63
+ "pv_layers/layer_045.json",
64
+ "pv_layers/layer_046.json",
65
+ "pv_layers/layer_047.json",
66
+ "pv_layers/layer_048.json",
67
+ "pv_layers/layer_049.json",
68
+ "pv_layers/layer_050.json",
69
+ "pv_layers/layer_051.json",
70
+ "pv_layers/layer_052.json",
71
+ "pv_layers/layer_053.json",
72
+ "pv_layers/layer_054.json",
73
+ "pv_layers/layer_055.json",
74
+ "pv_layers/layer_056.json",
75
+ "pv_layers/layer_057.json",
76
+ "pv_layers/layer_058.json",
77
+ "pv_layers/layer_059.json",
78
+ "pv_layers/layer_060.json",
79
+ "pv_layers/layer_061.json",
80
+ "pv_layers/layer_062.json",
81
+ "pv_layers/layer_063.json",
82
+ "pv_layers/layer_064.json",
83
+ "pv_layers/layer_065.json",
84
+ "pv_layers/layer_066.json",
85
+ "pv_layers/layer_067.json",
86
+ "pv_layers/layer_068.json",
87
+ "pv_layers/layer_069.json",
88
+ "pv_layers/layer_070.json",
89
+ "pv_layers/layer_071.json",
90
+ "pv_layers/layer_072.json",
91
+ "pv_layers/layer_073.json",
92
+ "pv_layers/layer_074.json",
93
+ "pv_layers/layer_075.json",
94
+ "pv_layers/layer_076.json",
95
+ "pv_layers/layer_077.json",
96
+ "pv_progress.json",
97
+ "source.json",
98
+ "tokenizer_config.json"
99
+ ]
100
+ }
reproduce/data_artifacts.json ADDED
@@ -0,0 +1,333 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "note": "Recovered local artifact hashes; current paths may have been reused. v32 hashes belong to v2. No raw texts, local logs, private chats or image bytes are included. v1 original 18,001,846-token flat stream is not recovered by this manifest.",
3
+ "files": [
4
+ {
5
+ "path": "/tmp/glm52-calib-v3/shard_00000.npy",
6
+ "sha256": "84dba15d4fec9e222ab68d2196a926dfd60123a7741530b3cbc02a84e1f20cb8",
7
+ "bytes": 4000128,
8
+ "shape": [
9
+ 1000000
10
+ ],
11
+ "dtype": "uint32"
12
+ },
13
+ {
14
+ "path": "/tmp/glm52-calib-v3/shard_00001.npy",
15
+ "sha256": "25310686c241a8ddfd5463fa240bf0a9549885772cc0d0ff75d15138f00c1ea0",
16
+ "bytes": 4000128,
17
+ "shape": [
18
+ 1000000
19
+ ],
20
+ "dtype": "uint32"
21
+ },
22
+ {
23
+ "path": "/tmp/glm52-calib-v3/shard_00002.npy",
24
+ "sha256": "cbcfa8c283d8e894e83559198c52e65a95455ceee9bb5061c9cce90ec53478b9",
25
+ "bytes": 4000128,
26
+ "shape": [
27
+ 1000000
28
+ ],
29
+ "dtype": "uint32"
30
+ },
31
+ {
32
+ "path": "/tmp/glm52-calib-v3/shard_00003.npy",
33
+ "sha256": "b64418fda40e375922baa5755b8c01ef7b4b14b8d40c1fcd2ebd9c80e1cec381",
34
+ "bytes": 4000128,
35
+ "shape": [
36
+ 1000000
37
+ ],
38
+ "dtype": "uint32"
39
+ },
40
+ {
41
+ "path": "/tmp/glm52-calib-v3/shard_00004.npy",
42
+ "sha256": "b42bd5638c244923766898100cb0e75458919a3a127d42789c925d81ebe59699",
43
+ "bytes": 4000128,
44
+ "shape": [
45
+ 1000000
46
+ ],
47
+ "dtype": "uint32"
48
+ },
49
+ {
50
+ "path": "/tmp/glm52-calib-v3/shard_00005.npy",
51
+ "sha256": "a672f7431ddefb1b5233fc7e53e6f0fc09a36395231c5c2c8d8a7d4056fe74eb",
52
+ "bytes": 4000128,
53
+ "shape": [
54
+ 1000000
55
+ ],
56
+ "dtype": "uint32"
57
+ },
58
+ {
59
+ "path": "/tmp/glm52-calib-v3/shard_00006.npy",
60
+ "sha256": "e6dcafae0e4018eb7d99b7517366f05375570d934476ad0456b4e4448c890719",
61
+ "bytes": 4000128,
62
+ "shape": [
63
+ 1000000
64
+ ],
65
+ "dtype": "uint32"
66
+ },
67
+ {
68
+ "path": "/tmp/glm52-calib-v3/shard_00007.npy",
69
+ "sha256": "eb22593102c11b48a182852989f17399ca32b3bc77196b73b22cb3e3d58190f9",
70
+ "bytes": 4000128,
71
+ "shape": [
72
+ 1000000
73
+ ],
74
+ "dtype": "uint32"
75
+ },
76
+ {
77
+ "path": "/tmp/glm52-calib-v3/shard_00008.npy",
78
+ "sha256": "610bfc9c5dcaf030eabfeca2c4150f409391102634f9dff003a58ca32bb5b0ed",
79
+ "bytes": 4000128,
80
+ "shape": [
81
+ 1000000
82
+ ],
83
+ "dtype": "uint32"
84
+ },
85
+ {
86
+ "path": "/tmp/glm52-calib-v3/shard_00009.npy",
87
+ "sha256": "0464899366bb7429800be3f61486af04b871abde381a28c644a44b14793ff9a8",
88
+ "bytes": 4000128,
89
+ "shape": [
90
+ 1000000
91
+ ],
92
+ "dtype": "uint32"
93
+ },
94
+ {
95
+ "path": "/tmp/glm52-calib-v3/shard_00010.npy",
96
+ "sha256": "9691d0cfce1ec93e5ace299c4045e9c1bb2e53adcf713adccbd1336ebe9287a7",
97
+ "bytes": 4000128,
98
+ "shape": [
99
+ 1000000
100
+ ],
101
+ "dtype": "uint32"
102
+ },
103
+ {
104
+ "path": "/tmp/glm52-calib-v3/shard_00011.npy",
105
+ "sha256": "d48e3552bcbf2251ecbd98febc390167d15514f265d0cf4bab2169a173266438",
106
+ "bytes": 4000128,
107
+ "shape": [
108
+ 1000000
109
+ ],
110
+ "dtype": "uint32"
111
+ },
112
+ {
113
+ "path": "/tmp/glm52-calib-v3/shard_00012.npy",
114
+ "sha256": "255fa27435f0948f0af499e474810aa877951c68c84655afa5a7376c47592c04",
115
+ "bytes": 4000128,
116
+ "shape": [
117
+ 1000000
118
+ ],
119
+ "dtype": "uint32"
120
+ },
121
+ {
122
+ "path": "/tmp/glm52-calib-v3/shard_00013.npy",
123
+ "sha256": "94220e009467d6b645902d2bdd429aa29544a4eb9fa5d4f9b24bff823b8e0496",
124
+ "bytes": 4000128,
125
+ "shape": [
126
+ 1000000
127
+ ],
128
+ "dtype": "uint32"
129
+ },
130
+ {
131
+ "path": "/tmp/glm52-calib-v3/shard_00014.npy",
132
+ "sha256": "37d7b9604e62a12eef2c293ed899f6bb2a8032418e5c5b7c578c82ff29e5bf0c",
133
+ "bytes": 4000128,
134
+ "shape": [
135
+ 1000000
136
+ ],
137
+ "dtype": "uint32"
138
+ },
139
+ {
140
+ "path": "/tmp/glm52-calib-v3/shard_00015.npy",
141
+ "sha256": "a7a3a29377ef4436d6ec5185d5f45f2c7e615fbb093499e59a3651d78ef0181d",
142
+ "bytes": 26120,
143
+ "shape": [
144
+ 6498
145
+ ],
146
+ "dtype": "uint32"
147
+ },
148
+ {
149
+ "path": "/tmp/glm52-calib-v32/shard_00000.npy",
150
+ "sha256": "587d01f48f9d39719661004f51d288bed7ccba5808c2a86068ad83a72e8e9130",
151
+ "bytes": 4000128,
152
+ "shape": [
153
+ 1000000
154
+ ],
155
+ "dtype": "uint32"
156
+ },
157
+ {
158
+ "path": "/tmp/glm52-calib-v32/shard_00001.npy",
159
+ "sha256": "b34946701d15f586a7c7513ab37aa43c7bdbf1e9a51448a90ddea81850df8011",
160
+ "bytes": 4000128,
161
+ "shape": [
162
+ 1000000
163
+ ],
164
+ "dtype": "uint32"
165
+ },
166
+ {
167
+ "path": "/tmp/glm52-calib-v32/shard_00002.npy",
168
+ "sha256": "42b98201ecbfd0830fc333c6691ec67daaa6b82287b8d33a8ac2670dca14d842",
169
+ "bytes": 4000128,
170
+ "shape": [
171
+ 1000000
172
+ ],
173
+ "dtype": "uint32"
174
+ },
175
+ {
176
+ "path": "/tmp/glm52-calib-v32/shard_00003.npy",
177
+ "sha256": "87c619b34e746d40ab1b32827317aca92a5583d7a9a010c6199ccb63aeb52f8f",
178
+ "bytes": 4000128,
179
+ "shape": [
180
+ 1000000
181
+ ],
182
+ "dtype": "uint32"
183
+ },
184
+ {
185
+ "path": "/tmp/glm52-calib-v32/shard_00004.npy",
186
+ "sha256": "9ea10b0107c7afbb82256c5e00ddc001f5b0288d9f35e69712e88ffd20d70b28",
187
+ "bytes": 4000128,
188
+ "shape": [
189
+ 1000000
190
+ ],
191
+ "dtype": "uint32"
192
+ },
193
+ {
194
+ "path": "/tmp/glm52-calib-v32/shard_00005.npy",
195
+ "sha256": "5941cd2282b9c9eeeffcb500044a20c8bf5cb27d03ca3e0783491d956141af56",
196
+ "bytes": 4000128,
197
+ "shape": [
198
+ 1000000
199
+ ],
200
+ "dtype": "uint32"
201
+ },
202
+ {
203
+ "path": "/tmp/glm52-calib-v32/shard_00006.npy",
204
+ "sha256": "3a83b4ae5266f229e4431193594462f38f82357f0ecb20bb5bf2b79c1703e1ab",
205
+ "bytes": 4000128,
206
+ "shape": [
207
+ 1000000
208
+ ],
209
+ "dtype": "uint32"
210
+ },
211
+ {
212
+ "path": "/tmp/glm52-calib-v32/shard_00007.npy",
213
+ "sha256": "564afecdb45368874c7b371ed4627188b14d7a84ab3a8e2f26c26433fa228548",
214
+ "bytes": 4000128,
215
+ "shape": [
216
+ 1000000
217
+ ],
218
+ "dtype": "uint32"
219
+ },
220
+ {
221
+ "path": "/tmp/glm52-calib-v32/shard_00008.npy",
222
+ "sha256": "0a56d27bcab2ad9149e138c85b513016c97cc0950d8911ae49ed83e12c6f6c8c",
223
+ "bytes": 4000128,
224
+ "shape": [
225
+ 1000000
226
+ ],
227
+ "dtype": "uint32"
228
+ },
229
+ {
230
+ "path": "/tmp/glm52-calib-v32/shard_00009.npy",
231
+ "sha256": "7fe80d4b868c2d7f1d0a162a7fd80782f7d69f8f50ec6976f6a75c1a24dc2aca",
232
+ "bytes": 4000128,
233
+ "shape": [
234
+ 1000000
235
+ ],
236
+ "dtype": "uint32"
237
+ },
238
+ {
239
+ "path": "/tmp/glm52-calib-v32/shard_00010.npy",
240
+ "sha256": "61d886db2533ec343ca240ccdd859118476bd37dd59ba37cfcb0dfb56c6d5f57",
241
+ "bytes": 4000128,
242
+ "shape": [
243
+ 1000000
244
+ ],
245
+ "dtype": "uint32"
246
+ },
247
+ {
248
+ "path": "/tmp/glm52-calib-v32/shard_00011.npy",
249
+ "sha256": "d7e4af049ce7bf3f0186973d73f66990403f365deac2b1ba1c26bd04be917bfb",
250
+ "bytes": 4000128,
251
+ "shape": [
252
+ 1000000
253
+ ],
254
+ "dtype": "uint32"
255
+ },
256
+ {
257
+ "path": "/tmp/glm52-calib-v32/shard_00012.npy",
258
+ "sha256": "23bbeb34c4ac0f753192f1d8d17d3def7ba5cf6449149549bfe1ff9c0ef404e7",
259
+ "bytes": 4000128,
260
+ "shape": [
261
+ 1000000
262
+ ],
263
+ "dtype": "uint32"
264
+ },
265
+ {
266
+ "path": "/tmp/glm52-calib-v32/shard_00013.npy",
267
+ "sha256": "36f1aba2e26747a9227f7d60edde31d4e18b64f402578b42ccd1bbd16258414a",
268
+ "bytes": 4000128,
269
+ "shape": [
270
+ 1000000
271
+ ],
272
+ "dtype": "uint32"
273
+ },
274
+ {
275
+ "path": "/tmp/glm52-calib-v32/shard_00014.npy",
276
+ "sha256": "4d86d75d69d69b86fae7eba6aaf31153d9a043620bb94daee285925b4ce8a353",
277
+ "bytes": 4000128,
278
+ "shape": [
279
+ 1000000
280
+ ],
281
+ "dtype": "uint32"
282
+ },
283
+ {
284
+ "path": "/tmp/glm52-calib-v32/shard_00015.npy",
285
+ "sha256": "6a5fc10154611f997bc9957ced9c43f37c14372289a5c89d850a5f405358c5b9",
286
+ "bytes": 31144,
287
+ "shape": [
288
+ 7754
289
+ ],
290
+ "dtype": "uint32"
291
+ },
292
+ {
293
+ "path": "/tmp/glm53-vision-trace-refit/audit/tokens/shard_000.npy",
294
+ "sha256": "3f9fead556891373590e7017016833d7395674bdcf60c052dc663fb3190932bc",
295
+ "bytes": 1160816,
296
+ "shape": [
297
+ 290172
298
+ ],
299
+ "dtype": "uint32"
300
+ },
301
+ {
302
+ "path": "/tmp/glm53-vision-trace-refit/train/tokens/shard_000.npy",
303
+ "sha256": "587d01f48f9d39719661004f51d288bed7ccba5808c2a86068ad83a72e8e9130",
304
+ "bytes": 4000128,
305
+ "shape": [
306
+ 1000000
307
+ ],
308
+ "dtype": "uint32"
309
+ },
310
+ {
311
+ "path": "/tmp/glm53-vision-trace-refit/validation/tokens/shard_000.npy",
312
+ "sha256": "9f943a73f4a9eed1f8f817b34653f242f212b5df108d66dfeea4e7f0a5357f47",
313
+ "bytes": 1605812,
314
+ "shape": [
315
+ 401421
316
+ ],
317
+ "dtype": "uint32"
318
+ },
319
+ {
320
+ "path": "/tmp/glm53-layer3-full-corpus/train.npy",
321
+ "sha256": "26812c90e966c66bbe33b0e0c8bf9572906deb65eda7ce022f73c58f04eda29a",
322
+ "shape": [
323
+ 15007754
324
+ ],
325
+ "dtype": "uint32",
326
+ "boundary_counts": {
327
+ "154820": 18854,
328
+ "154841": 3099,
329
+ "154842": 3099
330
+ }
331
+ }
332
+ ]
333
+ }
reproduce/environment_observed.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "captured_on": "2026-09-26",
3
+ "versions": {
4
+ "torch": "2.11.0+cu130",
5
+ "numpy": "2.2.6",
6
+ "transformers": "5.13.0",
7
+ "datasets": "5.0.0",
8
+ "safetensors": "0.8.0",
9
+ "huggingface_hub": "1.22.0",
10
+ "Pillow": "12.3.0",
11
+ "scipy": "1.18.0",
12
+ "triton": "3.6.0"
13
+ },
14
+ "historical_environment_exactly_recovered": false,
15
+ "note": "Current audit environment, not an original training lockfile. Pin a tested CUDA/PyTorch/Transformers stack before rerunning."
16
+ }
reproduce/external_sources.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scope": "Observed checkout commits, not proof of original dataset/training source state",
3
+ "vllm": {
4
+ "url": "https://github.com/vllm-project/vllm.git",
5
+ "revision": "f32e2837ffc41ad369fff3ba8b05461ce5b53b55",
6
+ "use": "optional serving/code-corpus checkout; local dirty diff not reproduced"
7
+ },
8
+ "b12x": {
9
+ "url": "https://github.com/local-inference-lab/b12x",
10
+ "revision": "9043b448622764a598969518d413b3fd8b3c0c07",
11
+ "use": "optional BTX comparison modules, not ARVQ production codec"
12
+ },
13
+ "quant-toolkit": {
14
+ "url": "https://github.com/local-inference-lab/quant-toolkit",
15
+ "revision": "8bdb1016e52dd15f91e42d35c09096c0f31325f3",
16
+ "use": "historical external toolkit; not vendored"
17
+ }
18
+ }
reproduce/historical_metadata/METHODOLOGY.md ADDED
@@ -0,0 +1,342 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ARVQ cold-expert quantization: methodology
2
+
3
+ How the cold Mixture-of-Experts weights in this repository
4
+ (`GLM-5.3-Vision-NVFP4-ARVQ-hybrid`) were quantized and tuned. This document
5
+ reconstructs the actual pipeline and tuning decisions from the build transcripts
6
+ and the fitting/serving code; numbers are cited from those sources. It is a
7
+ methodology record, not a quality claim: full-model task quality and native
8
+ SM120 execution were **not** evaluated for this checkpoint.
9
+
10
+ ---
11
+
12
+ ## 1. What this model is (the hybrid)
13
+
14
+ GLM-5.3 is a large MoE. Each MoE block routes each token to its top-8 of 256
15
+ routed experts (plus a shared expert). This checkpoint is a **hybrid** in which
16
+ each expert is stored in one of two ways:
17
+
18
+ - **Hot experts** — kept in the NVFP4 donor format (unchanged).
19
+ - **Cold experts** — re-quantized to ~2 bits with **ARVQ** (additive residual
20
+ vector quantization), the subject of this document.
21
+
22
+ The hot/cold split is fixed by a REAP allocation that keeps the **5,750** most
23
+ important experts hot; the remaining **13,450** experts across the model are
24
+ cold (per the 09-14 build transcript, which verified "13,450 cold-expert source
25
+ files"). The allocation is 75% text-REAP / 25% multimodal-salience weighted and
26
+ is **not** modified by this campaign.
27
+
28
+ Only the **75 MoE blocks, layers 3–77**, are touched. Layers 0–2 are dense and
29
+ stay frozen. Everything except the cold experts is inherited unchanged:
30
+ attention, backbone, shared experts, hot NVFP4 experts, BF16 MTP, and the
31
+ vision components.
32
+
33
+ Cold experts are quantized per (layer, projection). The two projections are the
34
+ fused gate/up `w13` (N=4096, K=6144) and the down projection `w2` (N=6144,
35
+ K=2048).
36
+
37
+ ---
38
+
39
+ ## 2. ARVQ representation and bit budget
40
+
41
+ Each cold expert weight group of **8 contiguous columns** is represented as the
42
+ sum of two codebook atoms times a scale:
43
+
44
+ ```
45
+ w_group = global * block_scale * (c0[a] + c1[b])
46
+ ```
47
+
48
+ - **Two codebooks** `c0`, `c1`, each **256 entries × 8 dims** (512 codewords per
49
+ pair). This is the "rvq256_256x8" format. Codebooks are **per-expert** in this
50
+ v3 checkpoint (`codebook_scope = expert`, packed format
51
+ `rvq256_256x8_expert`, version 3).
52
+ - **Indices** `a`, `b` are one uint8 each per 8-weight group → **16 index bits
53
+ per group = 2.0 bits/weight**.
54
+ - **Codebook atoms are constrained to the FP4 grid** `{0, ±0.5, ±1, ±1.5, ±2,
55
+ ±3, ±4, ±6}` (project-to-FP4 after every update). The FP4 grid on the atoms is
56
+ why no incoherence/Hadamard rotation is used: the target space cannot be
57
+ rescaled without leaving the grid (see `encoder.py` docstring).
58
+ - **Block scales** are stored as **`float8_e4m3fn`**, one scale per **[row,
59
+ 128-column block]** (48 blocks/row for `w13`, 16 blocks/row for `w2`). At use
60
+ time each block scale is `repeat_interleave(16)` to cover the 16 eight-wide
61
+ groups in a 128-column block. Adding the FP8 scale (8 bits / 128 weights =
62
+ 0.0625 bit/weight) gives the reported **2.0625 bits/weight** plus the
63
+ amortized per-expert codebooks.
64
+ - A single fp32 `global` scalar per (layer, projection) sits above the block
65
+ scales so the index/scale layout is unchanged from earlier versions.
66
+
67
+ Packing (`pack.py`) writes MMA-fragment index tiles (64 uint32 words/tile) and
68
+ verifies, per layer, that every expert's indices round-trip exactly and that a
69
+ sample of decoded weights matches bit-for-bit before publishing.
70
+
71
+ ---
72
+
73
+ ## 3. Initial fit (Hessian-aware, no PV yet)
74
+
75
+ Before any gradient tuning, each (layer, projection) gets an
76
+ activation-Hessian-tuned initialization (`encoder.py`, `fit.py`):
77
+
78
+ 1. **Per-expert raw-basis Hessian.** `w13` uses `H = E[x xᵀ]` over routed
79
+ tokens; `w2` uses `H = E[m mᵀ]` where `m = silu(x·Wgᵀ)·(x·Wuᵀ)` is the BF16
80
+ SwiGLU intermediate. Experts with <32 routed tokens fall back to all captured
81
+ tokens. Activations are stored/computed in **float32**; Hessians are half.
82
+ 2. **Escalating-damp Cholesky** of `H⁻¹`: damping starts at `1e-3 · mean(diag H)`
83
+ and multiplies by 4× (up to 12 attempts) until the factorization succeeds.
84
+ 3. **Codebook EM** on a Hessian-importance-weighted subsample (20,000 groups/
85
+ expert) of normalized 8-dim groups: k-means warm start → 8 alternating
86
+ refine iterations, FP4-projecting the atoms after every M-step.
87
+ 4. **Per-expert LDLQ/GPTQ error-feedback column sweep** with the fixed
88
+ codebooks (col_block=128, 2 sweep passes, 1 inner refine): plain-L2 inner
89
+ assignment, off-diagonal Hessian carries cross-group error feedback.
90
+ 5. **Least-squares refit of the E4M3 block scales** given the chosen codes.
91
+
92
+ An earlier decision, recorded 09-14: a **tuned Hadamard rotation was tested and
93
+ rejected** — a matched 48-expert down-projection refit gave 0.372 (H128 rotated)
94
+ versus **0.191 unrotated**, so the unrotated raw-basis fit was adopted.
95
+
96
+ Nonfinite guards are raised as errors throughout (the fit/tuner refuse to
97
+ proceed on NaN/Inf rather than silently scrubbing).
98
+
99
+ ---
100
+
101
+ ## 4. The PV objective: `--target reference`
102
+
103
+ The cold experts are then output-tuned per layer. The objective for this
104
+ checkpoint is **`--target reference`** (`sequential_pv_full_corpus.py`).
105
+
106
+ For each captured token position the tuner forms:
107
+
108
+ - `frozen` — the block output from everything that is **not** a cold expert on
109
+ the **student** input (retained tuned upstream layers → attention → shared
110
+ expert → hot NVFP4 experts). This is fixed.
111
+ - `reference` — the **original FP8 reference block output**: the unchanged donor
112
+ backbone with the **original FP8 routed experts** (all 256), evaluated on the
113
+ **original reference trajectory** for that token position.
114
+ - `required = reference − frozen` — the residual the cold experts must supply.
115
+
116
+ The loss is the relative squared error of the cold-expert sum against `required`:
117
+
118
+ ```
119
+ loss = || student_cold(x_student) − required ||² / (denominator · Σw)
120
+ ```
121
+
122
+ where `denominator` is the per-row mean energy of `reference − frozen` over the
123
+ corpus (so the scale is comparable across layers). Selection metric is
124
+ `reference_rel = ||(frozen + cold) − reference|| / ||reference||`, measured over
125
+ the **whole transformer block including the residual**, not the cold-experts sum
126
+ alone.
127
+
128
+ **Why reference (and its known weakness).** The transcripts weighed
129
+ `reference` against a `same_input` target (match the block's output on the
130
+ *student's own* drifted input). The rationale for reference (09-16): "Matching
131
+ the original reference trajectory is more directly aligned with preserving the
132
+ model's behavior." The acknowledged risk: because each layer receives a
133
+ **different (student) input** than the original, the cold experts may lack the
134
+ capacity to reproduce the reference output and could "learn corrections that
135
+ work on training examples but fail on unseen ones" — i.e. reference-target is
136
+ the more out-of-distribution objective. It was chosen only after a matched
137
+ layer-4 A/B:
138
+
139
+ | Layer-4 metric | reference vs same_input |
140
+ |---|---|
141
+ | validation reference error | **0.21% better** with reference |
142
+ | development-audit error | 0.08% **worse** with reference |
143
+
144
+ A small mixed difference; reference was retained for behavioral alignment. Two
145
+ important design consequences: the target is the **complete block output**
146
+ (including the residual), so matching only the MoE contribution cannot leave the
147
+ incoming residual error untouched; and the loss compares the **routing-weighted
148
+ sum of all cold experts** jointly (gate/up and down together).
149
+
150
+ ---
151
+
152
+ ## 5. Sequential per-layer pipeline with trajectory coupling
153
+
154
+ Layers are processed **sequentially, 3 → 77** (`full_reference.py` /
155
+ `sequential_capture.py`):
156
+
157
+ 1. **Capture** (8-way sequence-parallel). For layer L, the student input is the
158
+ **retained tuned output of block L-1** (`student_outputs.pt`), while the
159
+ reference target re-runs the **donor backbone + original FP8 experts** on the
160
+ original reference hidden state. Both trajectories are carried forward in a
161
+ rolling cache; layer 3 is the control (no preceding ARVQ layer, so student =
162
+ reference input). This cross-layer coupling means each layer is tuned against
163
+ the exact drifted inputs it will see at serving time.
164
+ 2. **Fit / PV tune** the cold experts (Section 6).
165
+ 3. **Gate** (Section 7).
166
+ 4. **Export + serialized replay** — decode to the packed format and require the
167
+ reloaded forward to match the tuned forward to <2e-6 before use.
168
+ 5. **Per-layer HF publish** — `publish_reference.py` re-assembles cold slots
169
+ across the 8 expert shards, checks slot coverage/allocation/global metadata,
170
+ runs the exact index round-trip and sampled decoded-weight parity, then
171
+ **atomically replaces that layer's two tensor files + reports in one commit**
172
+ with parent-commit protection; remote hashes are re-verified.
173
+
174
+ Parallelism: cold experts are **sharded across 8 GPUs** (each rank owns distinct
175
+ experts); routed outputs are summed with an all-reduce so a single coupled-output
176
+ objective is optimized without a dense-model replica.
177
+
178
+ ---
179
+
180
+ ## 6. PV tuning numerics (the published full-corpus campaign)
181
+
182
+ Per the model card, all 75 layers were replaced by a **full-corpus sequential
183
+ PV campaign** drawing sequentially from **18,001,846 training tokens** at
184
+ context **1024** (53,248 legacy tokens were removed around 17 exact matches to
185
+ held-out prompts). Fixed **validation** and **development-audit** sets each hold
186
+ **16,384 tokens**.
187
+
188
+ - **Optimizer:** Adam. Two trainable parameter groups — FP4-constrained
189
+ per-expert **codebooks** and FP8-constrained per-block **scales**. (The
190
+ faster fork trains per-row scale deltas instead; the published full-corpus
191
+ fork trains the per-128-block log-scales.)
192
+ - **Effective batch:** 262,144 tokens, accumulated as **four 65,536-token
193
+ microbatch passes** (256 whole 1024-token sequences), one shuffled
194
+ no-replacement pass over the corpus.
195
+ - **Learning-rate schedule:** book/scale LR **0.048 / 0.032** through update 45,
196
+ then dropped by ×0.25 to **0.012 / 0.008** (`lr_decay_after=45`,
197
+ `lr_decay_factor=0.25`). (Note: these campaign LRs are higher than the
198
+ in-repo code defaults of 0.003/0.002, consistent with the much larger
199
+ 262k-token batch.)
200
+ - **Update budget:** layers 4–26 use a fixed **69 updates**. From layer 27, 69
201
+ is the maximum with an **adaptive early stop**: three validation checks more
202
+ than 0.1% worse than best trigger an earlier LR reduction; 15 updates without
203
+ improvement after that reduction permit stopping.
204
+ - **Validation every 5 updates**, retaining the best checkpoint.
205
+ - **Index reassignment every 20 updates** (Section 6.1).
206
+ - **Regularization:** codebook drift-from-anchor penalty + scale drift penalty,
207
+ weighted 0.01; grad-norm clip 1.0; codebook atoms clamped to [-6, 6];
208
+ log-scales clamped to within ±0.35 of their initial value.
209
+ - **Coverage gate:** experts with fewer than **256 routed training rows** are
210
+ frozen (gradient masked) and keep their initial books/scales.
211
+
212
+ ### 6.1 Output-gradient index reassignment
213
+
214
+ Continuous Adam cannot move the discrete indices, so every 20 updates a discrete
215
+ proposal pass runs (`gradient_indices.py`, method
216
+ `expert_parallel_output_gradient_prefix_backtracking_v1`):
217
+
218
+ 1. Backprop the **output** loss to the reconstructed weight of one expert
219
+ projection.
220
+ 2. Propose alternative `(a,b)` index pairs near a small gradient step, capped at
221
+ `max_fraction=0.001` of groups, `trust_ratio=0.01`, `target_ratio=0.03`; the
222
+ current pair is always in the candidate set (changing nothing stays legal).
223
+ 3. **Prefix backtracking acceptance:** try the top-k proposals with
224
+ k∈{full, ¼, 1/16, 1}, accept only if the **training** SSE strictly drops
225
+ **and** a separate **check batch** SSE does not increase (beyond 1e-7 slack);
226
+ otherwise revert. Proposal deltas are broadcast sparsely (only routed rows).
227
+
228
+ The transcripts flag this as the method's main theoretical weakness versus
229
+ published PV-Tuning: codebooks are tuned for output accuracy while index
230
+ proposals originally came from weight-reconstruction Hessians — the
231
+ output-gradient proposal above was added to close that gap.
232
+
233
+ ### 6.2 Cold arithmetic emulation
234
+
235
+ Both training and evaluation emulate the serving numerics (`activation.py`,
236
+ pinned to serving revision `b1380cf7…`): **four FP4 activation planes**
237
+ (`fp4_planes4_fp16_boundaries_v1`) and **FP16 SwiGLU boundaries** (FP32 GEMM,
238
+ FP16 SiLU/product), with straight-through estimators for gradients. This
239
+ emulates the quantization boundaries only — **native SM120 MMA accumulation and
240
+ TP reduction are not bit-exact** and were not qualified.
241
+
242
+ ---
243
+
244
+ ## 7. Acceptance gates
245
+
246
+ Publication of a layer requires all of:
247
+
248
+ - **Validation non-regression:** `0 ≤ final_reference_rel ≤ initial·(1+1e-6)`.
249
+ - **Development-audit non-regression:** the held-out audit split (never used for
250
+ updates or checkpoint selection) must satisfy
251
+ `audit_reference_rel ≤ initial_audit·(1+1e-6)`; otherwise the layer keeps its
252
+ initialization. There is no absolute error floor — the rule is purely
253
+ final ≤ initial.
254
+ - **Serialized-replay parity:** decoded/reloaded forward matches the tuned
255
+ forward to <2e-6, and remote file hashes are re-verified after upload.
256
+
257
+ The audit is honestly labeled **development data, not an untouched final test**;
258
+ it is a split of the same calibration tokens the initialization already saw. Two
259
+ notable per-layer decisions recorded in the card: **layer 30** retains its
260
+ audit-qualified update-15 checkpoint after its validation-best update-20 failed
261
+ the development audit; **layer 3** retains a separately qualified lower-LR
262
+ refinement.
263
+
264
+ ---
265
+
266
+ ## 8. Calibration corpus
267
+
268
+ The text corpus (`build_calib_v31.py` + `reasoning_slice.py`) is deterministic
269
+ (seed 42), GLM-tokenized, with the following domain mix by tokens:
270
+
271
+ | Share | Domain | Sources |
272
+ |---|---|---|
273
+ | ~32% | code | local vLLM sources, m-a-p/CodeFeedback, jtatman/python-code-500k |
274
+ | ~20% | tool-calling / agentic | generated GLM chat-template sessions |
275
+ | ~15% | reasoning | OpenR1-Math-220k + dolphin-r1 `<think>…</think>` → answer |
276
+ | ~13% | instruction chat | tatsu-lab/alpaca + coding chat |
277
+ | ~10% | medical | MedQA textbook continuation + medical Q&A |
278
+ | ~10% | prose | vLLM docs markdown + databricks-dolly-15k |
279
+
280
+ The **reasoning slice was added specifically** because the earlier calib-v3 mix
281
+ starved the cold 2-bit experts of reasoning-termination behavior: 87% of its
282
+ `<think>` blocks were empty. The slice emits multi-turn conversations where an
283
+ **earlier** assistant turn carries a real `<think>…</think>` that terminates and
284
+ hands off to an answer, followed by a trailing user turn, so the think-close
285
+ token **`</think>` (id 154842)** renders in-stream rather than as a trailing
286
+ generation prompt. Held-out prompts are explicitly excluded (last 25 vLLM docs,
287
+ last 3 MedQA files, vLLM code beyond index 400 of the seed-42 shuffle), and
288
+ 17 exact-match prompts were purged from the training set.
289
+
290
+ The capture/training code also supports **up-weighting boundary rows** (rows
291
+ whose next token is a boundary such as `</think>` / `<|endoftext|>`) via a
292
+ `row_weight` term whose sum normalizes the loss, and an optional matched-mixed
293
+ validation set; the published card describes reasoning-based validation.
294
+
295
+ ---
296
+
297
+ ## 9. Results captured during the build
298
+
299
+ All errors below are **held-out relative L2**, either over the whole transformer
300
+ block (including residual) or over the cold-experts sum, as noted. They are
301
+ local reconstruction errors, **not** token-accuracy or perplexity.
302
+
303
+ - **Pilot smoke test** (layer 3, one step): full-block error 0.005497 → 0.005456
304
+ (~0.75%), 304/388 expert-projection proposals accepted; exported weights
305
+ reproduced the retained result.
306
+ - **Layer 3** (200-step pilot): full-block reference error **0.005497 →
307
+ 0.005271 (−4.12%)**, best at step 200.
308
+ - **Layer 4** (200-step pilot): full-block reference error **0.015760 →
309
+ 0.015596 (−1.05%)**.
310
+ - **reference vs same_input** (layer 4, matched): reference 0.21% better on
311
+ validation, 0.08% worse on development audit.
312
+ - **Faster fork parity:** the optimized fork (larger microbatch, specialized
313
+ embedding-lookup backward, sparse proposal broadcast) matched the reference
314
+ fork's validation and audit **exactly** on layers 3 and 4, with all propagated
315
+ BF16 outputs equal; training loops fell from ~1525 s → ~337 s (layer 3) and
316
+ ~1226 s → ~303 s (layer 4).
317
+ - **Early-layer stability example** (layer 18, full-corpus): initial
318
+ 0.017205, step-5 0.017176, step-69 0.017183 — differences ≤0.17%, treated as
319
+ a signal to watch LR/noise rather than proof of convergence.
320
+
321
+ Older cold-only checkpoints reported ~0.2211→0.2070 (layer 3) and 0.2364→0.2320,
322
+ but those measured **cold-expert output only on different captures** and are
323
+ **not comparable** to the full-block errors above.
324
+
325
+ **No comparison against the FP8 donor or an AQLM variant on perplexity / KLD /
326
+ top-1 / wikitext was recorded** in the reviewed transcripts, and no full-model
327
+ task evaluation was performed. Those remain open.
328
+
329
+ ---
330
+
331
+ ## 10. Honest limitations
332
+
333
+ - Full-model quality and native **SM120** execution were **not** evaluated; the
334
+ arithmetic emulation covers FP4/FP16 boundaries only, not MMA/TP bit-exactness.
335
+ - The development audit is a split of calibration data, not an independent test.
336
+ - The reference target is the more OOD objective; its generalization advantage
337
+ over same_input was small and mixed on the one matched layer tested.
338
+ - Gates enforce local non-regression, not any absolute quality bar.
339
+
340
+ *Prepared from the GLM-5.3 ARVQ build transcripts (2026-09-14 and 2026-09-16)
341
+ and the `btx53/arvq88` fitting/serving code. Where a number could not be sourced
342
+ it is stated as unknown rather than estimated.*
reproduce/historical_metadata/backbone_sources.json ADDED
The diff for this file is too large to render. See raw diff
 
reproduce/historical_metadata/build_provenance.json ADDED
@@ -0,0 +1,178 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage": "per-expert initial fit; alternating PV in progress",
3
+ "allocation": {
4
+ "mode": "ARVQ output-benefit allocation",
5
+ "text_weight": 0.75,
6
+ "mm_weight": 0.25,
7
+ "hot_count": 5750,
8
+ "floor": 8,
9
+ "cap": 176,
10
+ "text_partition": "train only",
11
+ "metric": "positive routing-weighted squared-error reduction from retaining NVFP4 over fitted ARVQ",
12
+ "scores_sha256": {
13
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_3.npz": "99e689edad7a7f026759c9c66e682b499ca3263e6d19ad54f59c1f1dccdb3ea9",
14
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_4.npz": "7813ddafdb23f7c4802bf431a2d063f7afc9f3463356079bc97ef8289e30d823",
15
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_5.npz": "551940cef83bd52f2e032bd738a88fe293e74efc3e918cfdabf025e111c0d586",
16
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_6.npz": "1949e2540a0d92169ae9e5374789443d026b0868860149e785342eee19b521b5",
17
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_7.npz": "cb1d4f6597057c32d2d8136700ea357b86a12ccb881d4bb4a613033c6ea0a34e",
18
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_8.npz": "c2d4cfb7641a3a9fe96abffa76ab59774eb37cc629b058b3f9bebb1f98bbb09e",
19
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_9.npz": "c82bff242306c5b27acac0bdd8530e9ad603c1e00642855face4a9aefd03f97a",
20
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_10.npz": "59fcb8922894f0155704577296cd8752323d1f880a51e801a5b5b1c969647aa0",
21
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_11.npz": "45b94e85267b83b94203eddca795a21de68fb359a55dd0e304082597a8cb2cee",
22
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_12.npz": "a277261ee1c2b60dd45de874e0e7e4819356480af132fcb23dfff411a2775077",
23
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_13.npz": "911c1ba89ea9807704759cad0522532127329259b19ba3ade47ce95c1a5c4e49",
24
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_14.npz": "e672d24e525c3c077ff2a9a2617b6435e0d3d97398cd2c901f5298bd4a0b7197",
25
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_15.npz": "e163205f5eb51d0c8796df2070fde3bf80caf3bdf166311336669f53a695c437",
26
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_16.npz": "9cda2a0e36f882a5daa485c698b9d1eb5ff3aba7ffeac81f7706bed09214f6c6",
27
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_17.npz": "0f085320e0e1f09bc52183d619a8d91535bdb33b5d56c3de05f212b4337bc2a7",
28
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_18.npz": "627d542b236da70a5dd2ed098d21cee75dc4cb38db6105369f947361dbe7448a",
29
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_19.npz": "3613a6fa1e852c110ee05ad901a4c89e1ba659c4782b88988bc935a3a69406bf",
30
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_20.npz": "7be6a86b425bb811436bdf0f1e4a9a55a0d64aa8a2a7df163923a1898e88e3d9",
31
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_21.npz": "4abd00a72ff0e76af4cf708caace639e945bce378168ee47892db2645af0de81",
32
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_22.npz": "92deb3e93534084dd0f43a0f3ab6a11bde3156b5a6218128f03ad58c31e5da08",
33
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_23.npz": "6866eee649260c556ad663a9f2b198109d1a3aa65dff69a26ef9a33ebfe90b17",
34
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_24.npz": "9d2a097d7039b8016e50fe48ef7db9175bd6af7652fb8f02d2845e0608c085c8",
35
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_25.npz": "9675f20e29ebc295606eff124e7eea8e512123d3a8ec37774fac5db06797a5c6",
36
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_26.npz": "db25e41b5b130d39bcb6d508f719eaa76164b71813fa2bbdba61a249653f9fe3",
37
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_27.npz": "07d874eae8001ca5fc0e976ba7199afa6e83f54925606b0a8e2ea3836920dbcd",
38
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_28.npz": "b83eafbc8d8cf70d62c805c37cc3b6b6a2554f2236cd858a84d44696a45a526b",
39
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_29.npz": "084317cbf70cdb24996a55bda7423b32eb2ee807209d73a0baba7fcad1307cd7",
40
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_30.npz": "68a6b57cae5acdb0d01fd498ff324ae4e3b01a3c9e3121d7814b1d93fc53cb78",
41
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_31.npz": "b2f421260f518195814fd1f42b10dbf59060dfd08ad726e1afd0de2f6b5f7442",
42
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_32.npz": "88f73b9073cf9b4532aac041fad881f9df4392fea7b367101e361e643183a01a",
43
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_33.npz": "cce07637bfafcc225057a4c7e3c304c167965ca28b551e7da0f0d13f41b00505",
44
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_34.npz": "0251a6abee66ab96080a23ff1b545ad933841525a8782d3de3dac5b472a681bb",
45
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_35.npz": "1dc1d919c3fb1de2ae8ec4765e5576e53de6049d43608d3f2e2c5eb687b01b0b",
46
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_36.npz": "43815d9f766ec63f51268edd90433bf6331327766888a1a089cf7bf31c9a28b6",
47
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_37.npz": "dc7a1f5b296bc6105a8ca969f558e2f1261238889abfc73326b50ef83b2c0216",
48
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_38.npz": "2fb9251355a449817aef26486305cc41edace820a91d6113a2f3a0b8f05258a8",
49
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_39.npz": "f69d1d2da6fd0b3a277e0d3810be2eff1e95050742e26013a30226cb2ba86b50",
50
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_40.npz": "c59aeceb845ee137ebf890bf5789d181198272838d1f794bf2c96ae31502e940",
51
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_41.npz": "d962c915b2b672f58bd5f1bfa65e806f09dd5a842a5986c1ca4ac03422620130",
52
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_42.npz": "acbdc6f83401fac52ce2a5a5e4d174be609842ffba5f1a3e219e9faae913c55a",
53
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_43.npz": "7ffa833b873294a9f6f1f005644b1101fda459b10ecdb5375eb4c85f7cf6ad5b",
54
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_44.npz": "906e81150c5f462b68750f74966df5bb93527a30f3b2cb0ecc410da5c9859840",
55
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_45.npz": "df9f6938632d32af48ceef8013694d0a92bd7208f4335fb06dc74ee2cd38ba12",
56
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_46.npz": "9fd39f469414d71da762d10b7671e5e0bdcbf087808804e2c868b73193702065",
57
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_47.npz": "9d2eb792b6553fa078c5cecac1ac53db461e5f2b7bc3b971246f6b8bc7846477",
58
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_48.npz": "49b108964ddb44d54b4d2375e82772d14275e94ab43da17855953717367ff583",
59
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_49.npz": "6775fdc746a2b11a443645140a86d1a11dcb389df71289fbef3fe300dec787e2",
60
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_50.npz": "f6d604b98a1c161b023f42205abf6d1eeecc9119871a8e920623908af0784fc9",
61
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_51.npz": "0aac9c4e66efaf5c0924107749e82a1e19e70a0b8e31db46434acf38353b8312",
62
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_52.npz": "7601958d27deb50c0bbf21a6ebb79c20859ca08c962b6f751f972b0f3453bbe6",
63
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_53.npz": "7bf959678dfc05ec4ef9f7b063d74a3de80e3cf757a199ecf083988ca55d41b7",
64
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_54.npz": "53097dbffee6e800d22551ed5ea068fe5bae3bad836813e2288f0667612bb6e6",
65
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_55.npz": "1ec4c9979cc55150e3e3163f8bf52f5d4042e351c15429bd11acc8bc7a77d273",
66
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_56.npz": "1675552ccfc8929bd29d7eed04e18e387222858beadf170e9652e17b27966766",
67
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_57.npz": "7e7029ec0de486aad04720d1059b69d74358c59432a33aa16e526041f4506609",
68
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_58.npz": "813756873a97c71e40fdff33313809723d599dbef2b2ebf36af0d56118b4e62d",
69
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_59.npz": "fcaadf941258b398bd6bf6a96f7865fccfe9008bf4ac1b93df8896065f761383",
70
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_60.npz": "e73c14eb000ec4c2fdadb4078011f12ab44bfd470bf690c22c3e0ffdbe89d11a",
71
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_61.npz": "a4c10b888ee7879054c08e4de97540e786f6c44c8823f3b7988dfb0fbe1b8927",
72
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_62.npz": "b73347956fed5d783082fe8dd76027372ec50db12683447b180e9a053f9abe61",
73
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_63.npz": "fe9211a06376cee25e813131672c6101c4564a7c27b5820eafe54b3efa0623e1",
74
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_64.npz": "348ecb20769be5b1bec0cb31157f513b92f911ef36a8439bef7d45b31321f4d9",
75
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_65.npz": "2d37ba213e5e4c5e598aedb928c7f50f16205665253d9fc04cad569aa01384db",
76
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_66.npz": "28615fbdc3c2f0b2458d38a52f05e041dd02501b3d18448ba449b4626a334cd6",
77
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_67.npz": "9e5a6100b7acb907c99f80ce1cf89f2e46e392c7ca48bb000eda2c126ae1413a",
78
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_68.npz": "dc0fdd42f3332b73508943c6f9ed0c4166ca0065c2b444823eccbd37a451be6f",
79
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_69.npz": "4a2b3139ba9f6cd9327285f6ca2417eeacbca49c6e96d8dc525f80e1a7d6eed0",
80
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_70.npz": "6c03c2284504c6b89cd7ea6af120b6236691dee58ecb2ab1680be19e379068e0",
81
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_71.npz": "c51e64b3d7d9347bb8fa91a6f4b74631983b7ded564f291b44ba496ff198622c",
82
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_72.npz": "298c23118c8377d56d7a025cfc59ca782246c1269f0a0b47cdc44f8f53778dc6",
83
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_73.npz": "0e2aaa5468da7aa0f48a0007eb7fdbed01abe6fdcca5f62ccc8c0d0febccd582",
84
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_74.npz": "7c33c7cb1a136dd9f438ab36b83007554227ecbb2488fdf883d545d1c680cc94",
85
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_75.npz": "0c71e68caa217700531d8611731f67fb5842e53df15778bbb480472028d564f6",
86
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_76.npz": "9dc6ec2ca2dd15912b5505a62ea3b296be8a174bc3c69eeeaa0d6ba316dcaf67",
87
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/text/layer_77.npz": "a2c9a43c8566867ede5e47f37b72c2623adc075edb44dd43508037d9a137bd06",
88
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_3.npz": "2e102e3892277e8808818ffa8373389b9e02e3ce3e1b280c8511dbf0e47353be",
89
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_4.npz": "c007140d4edbd361da7e15c6bd74bb49e10786ccddc49f0c343592c3939d5519",
90
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_5.npz": "99e460fca675b7c78122c054e9293b5bfc745d43a4cd57625c4b53dd6fb7dc09",
91
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_6.npz": "f1d5673cb87462244ddf87c79feee9f38887bcf3443092335b3b38c7ee350e0d",
92
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_7.npz": "b7e29194ef43c192fedd848312f2fd33687075be7271fbd5e8e882c600a5e4b7",
93
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_8.npz": "25b9e46d9c422178b825f2d20cf23a1adb09470dea7454e6b70000b927b58f98",
94
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_9.npz": "3757ebe69d33ab85557c141c3a37e5c800fe2c2ee2c966d2cdef3a6e689a1420",
95
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_10.npz": "07ab61f067a9fff3d4a1c9bb5039c745417d5bf305d824be85ddd6b114b9829f",
96
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_11.npz": "269c82abbd5c60440a169962d051b285670e86beef6f60f84a23ddb260e2419f",
97
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_12.npz": "0e2326f560e5e0b3b7441d8e6584720fefc42d6209bbd0ffc5c4af3679b5f542",
98
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_13.npz": "4e8c4f5e081cbe193d0468ad98842f4904ddceef140d61363eadce54bb9d8ce9",
99
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_14.npz": "16ccc9716e778239850f0d8c7f827ab3eefbd4673b775d1bdd17673141550973",
100
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_15.npz": "199256730bd069357ace4ff4a9528f4a4a2b316e5d299071d2b8751a05a0f026",
101
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_16.npz": "4631ef874d195778a08af4f0d03a2e6a6a2d1ed6c830b7016c758a0687a1a235",
102
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_17.npz": "66fbd376505e0f6f921e3f483253f8e733378f8df909733c8fc65194b6eeb092",
103
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_18.npz": "ee3134b30d8287684cdf95d1eb44da997b90c3ef4c681ac998a770109b35adbf",
104
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_19.npz": "611b57f18f927ae09479a29d77425fafba33514d0f15ff69dae366733c91e1d3",
105
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_20.npz": "b3d431eae9b9c3aa5a36c60975378bbec4112377cc858b5ccf0d0946778a6079",
106
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_21.npz": "cefc699b376b5b43c14769884cc481c22f4a284e168c1fa88feeb9614588f505",
107
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_22.npz": "7966b47efd7a838d8f87c2f5697320648c6f0bcdd5ffe7762cef03ff2ffbfa95",
108
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_23.npz": "9b55935db844c788fc59417a1c053722fd5685ad6ec4060a192a0a1fc1582e4f",
109
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_24.npz": "65c1134ad6a7a38dabea537c0b19899126aa8de849306c1385c14a3c09a22d37",
110
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_25.npz": "038eab98b035a9ee810e1c7bb08ad89d227bcb87012bf88164d3dea0ad3bc1cd",
111
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_26.npz": "8e31e623cf17a166d7ee056ba20b5c6f19a693e039c9910a851c0dc3a9cc983b",
112
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_27.npz": "1306ea155b6746f73a4463cb8be7474bb7f80b3870daead9930e37d0a8fa12cf",
113
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_28.npz": "3cd4ce023f5f11199b423dda6484bbb7377617b086e2e0bf5dc57afd29e7f566",
114
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_29.npz": "c7ee73af65346bc85e3c908825198faf4acc4987778f204df0821f4b642c7990",
115
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_30.npz": "81607884d312e7062274e230fec71dffc2a8a9f073ad8479d3ba19ea3df66899",
116
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_31.npz": "514aa850fc3099da75b19b74be73e944d7b58cb3577b4a221f58bef4169d9f75",
117
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_32.npz": "63b2e2a7e689a5efe641536527594ce26be5c807a32acc09a81aa6709150ccdb",
118
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_33.npz": "acdf16dc4552aef9184c5b687a9b7e0a4ea5edda0fa9ae496eb5f4f6bec12db1",
119
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_34.npz": "9c70b6174cf4a7a02a6300b14bd730fc51f6d1520dff274b8a3d6588368624e8",
120
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_35.npz": "25a82c26d0da013eb14be7d9b9f1811eeb9768fa45687e4a71c96c9b9752d9e6",
121
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_36.npz": "aeae4e169f803c42af7030b1395e1db2459fffaf363d4729075302df46d07f26",
122
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_37.npz": "c3787ea553f8bc3d239269f7e329e6f39d02b4ec63d06022405286dc640fbf2e",
123
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_38.npz": "1a486331012e24ff643d0dfec19f2d71bbf29e5de7df3706d7e16da33541ffca",
124
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_39.npz": "811e7c07267a7c80dbbcd51462fa01bae403341761ee9d89fce1ec67d32eec31",
125
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_40.npz": "712935236730a5e4142b5cf4da8dca2ba78416a358d79001b2787ee2a0f5b552",
126
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_41.npz": "e5fcd78bedc942ba734c28e16127c63fb02280ee279a8aee70143e2ec6d16c6e",
127
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_42.npz": "324b4909e5317f0217290669a78b1ef6f21db35665f13dbac7a78650a166a9b5",
128
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_43.npz": "b990182a89192ae58b4cfc82b3dd530bdfefbea01dd517a8d361c28c47980370",
129
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_44.npz": "cdddf30edd971309d21935c342fc353d0fe7b7a45a12c404cf8b7ed5d7a7dece",
130
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_45.npz": "525ef223869db6a843bca47df8fec0d8cf7f08f21d3626f8af03f4a33f0e896a",
131
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_46.npz": "8c4001f106b27f2646c61809d701a29245ee9997b0e291a7487f091e2651a7e3",
132
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_47.npz": "ce82107d4c1575654f41c9bc0bb070aade6d316c5a3bfd73bf84a75e9713fdeb",
133
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_48.npz": "d0a12442d24126677dd652b204f6910aff45d31ed76d368ce0efaff95943069f",
134
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_49.npz": "1566580cf72cf4ffc51c7cf5d9238bb09e92236512e92a7ae4a52fd887d5ac2e",
135
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_50.npz": "55be42b2dee8fecee914584711e405a48bd1be061aee8f8840ef43ed9a8ffe5b",
136
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_51.npz": "eca0da0b8d7522bc39c587cbd4c3f2ee88e7fc7f27370ff1d1b57c2e96d45565",
137
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_52.npz": "53491b6e515ddd86a49a3c4b5dcd5d00cd9687a0eabc595248a0cd28a85b2c3f",
138
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_53.npz": "561e99733a55b073842560ee9847b6b7dcd798d7bc769f3504ee373613708877",
139
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_54.npz": "3d0d33114524ff3a30e83b4d4f5c80aa44b031fd0e4da7fc776eb63fe89ff322",
140
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_55.npz": "3166fea4ab9810e08c4c05e5960c3fb338be355a0b9a583d9dd2a2e121c64982",
141
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_56.npz": "f4123b5dd6c2f722d029139b6f82919ee3c34de8d67326d819c8a6ca65931ff3",
142
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_57.npz": "e32c77a50be665998bd0682742535bd10b6c2478894f42779d141581b09788c0",
143
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_58.npz": "69430a65fc86a0166ccc9193746aa61fb4682ae65983944f8304907b34509bbd",
144
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_59.npz": "c35ebeaf949917077fb132e219558381c9638ef4de3819b95f289918cc4f9c3f",
145
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_60.npz": "6dc83319d29beb864c6f36d92a11ec898870e6fdf173917f0c9e73f6b18cac5c",
146
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_61.npz": "69c7b7440c70f7c4383000d1a798b5ef7eb9d3d9a7c6f28c49bd124f484f5fb5",
147
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_62.npz": "3bf69359b4ad241cb4e6dfdc00205fd637e1f7f959e35ec40eca480f6c41f80b",
148
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_63.npz": "280ad144e82b822b7ebf60af0d367dfcd261531bbccebd5791152aa75ddcf237",
149
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_64.npz": "1303b6bb6d7fad77472424cbc3696cb3368cc9e8e3f442dba0f362885b283a49",
150
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_65.npz": "f396c0ed74ed729685ff8681b2d2029acde74c28db559386b4cf15fff7cb8769",
151
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_66.npz": "b4d92dc502dc505d873f079044a09b8c2c0c310e95bdb2602a8cd572603c9953",
152
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_67.npz": "aa4765dd3124747d07ea1e06d353c12bdf2cf3f67777af71556c81d51e047a48",
153
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_68.npz": "f8df75fb71d9459a84789008ce0b0dc5ee8f0cafce7dad057e6fd6b544da9131",
154
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_69.npz": "339f59f847d5d98d52dd986eada9092a04036f82c4e570af4db66b7c51eb2000",
155
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_70.npz": "1156e221583b3435eb14c4116ef2629e89f84a8a0979253d88556d57424afec0",
156
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_71.npz": "473f3146f0ba8dbd340b44665954354d4d8fac6b5acb9ebcc952cf3f298b345c",
157
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_72.npz": "58179fe2da6eed5b63eac4710ef8bc54ceff4ff9ba2d8108942643585d21b744",
158
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_73.npz": "73fcb0af07e5a5accaa1d165446302cf22c595ffacd45015a460524d8e9cd9b1",
159
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_74.npz": "6e50287f298b566b96d44fe9961ebc7c9f0010c81b035fc4ff7456833ba86b46",
160
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_75.npz": "1a219f6efabaf773a6268b08f82716dc6177d257c728f70f81b0694d01d65cdf",
161
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_76.npz": "5aade66c63087e535fcd2c52bd6d8e075f439185a1745a8c8644bd4133a02eaf",
162
+ "/tmp/glm53-vision-expert-full/arvq_reap/scores/mm/layer_77.npz": "88e5ec9012bd07e07ef84d675db34d5fccfc45361ecf0ccbef2aca0a2a7e8bda"
163
+ },
164
+ "limitation": "per-expert additive error proxy; excludes cross-expert covariance and end-to-end model quality",
165
+ "candidate_codebook_scope": "expert"
166
+ },
167
+ "hub_revisions": {
168
+ "zai-org/GLM-5.3": "aca966e4e02791568aa6a4ced368624b3d897f42",
169
+ "RadixArk/GLM-5.3-NVFP4": "11af4cba759e6559eda70358a5778bd1bddddd78",
170
+ "jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m": "2b883d28bb9dd13a9511e2bd45a8ad1cbacbad74"
171
+ },
172
+ "cold_source": "zai-org/GLM-5.3",
173
+ "source_nonexpert_mtp": "nvfp4_donor",
174
+ "source_vision": "base_hybrid",
175
+ "initial_fit_partition": "train only",
176
+ "pv_reassign_every": 40,
177
+ "format": "rvq256_256x8_expert"
178
+ }
reproduce/historical_metadata/calibration_corpus.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source": "/tmp/glm53-unc-traces",
3
+ "reasoning_rollouts": true,
4
+ "counts": {
5
+ "train": {
6
+ "problems": 856,
7
+ "tokens": 3048596
8
+ },
9
+ "validation": {
10
+ "problems": 89,
11
+ "tokens": 345875
12
+ },
13
+ "audit": {
14
+ "problems": 104,
15
+ "tokens": 322277
16
+ }
17
+ },
18
+ "partitions": "disjoint prompt hashes; separately captured",
19
+ "limitations": "Text-only captures in 2048-token windows; source datasets overlap earlier calibration. Teacher answers are not correctness-verified.",
20
+ "initial_fit_partition": "train only",
21
+ "capture_teacher": "RadixArk/GLM-5.3-NVFP4",
22
+ "rollout_teacher": "unc NVFP4 teacher",
23
+ "same_token_corpus_as": "/tmp/glm53-kernel-aware-pv/full"
24
+ }
reproduce/historical_metadata/cold_assignment.json ADDED
The diff for this file is too large to render. See raw diff
 
reproduce/historical_metadata/pv_progress.json ADDED
@@ -0,0 +1,187 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "layers_pv_complete": [
3
+ 3,
4
+ 4,
5
+ 5,
6
+ 6,
7
+ 7,
8
+ 8,
9
+ 9,
10
+ 10,
11
+ 11,
12
+ 12,
13
+ 13,
14
+ 14,
15
+ 15,
16
+ 16,
17
+ 17,
18
+ 18,
19
+ 19,
20
+ 20,
21
+ 21,
22
+ 22,
23
+ 23,
24
+ 24,
25
+ 25,
26
+ 26,
27
+ 27,
28
+ 28,
29
+ 29,
30
+ 30,
31
+ 31,
32
+ 32,
33
+ 33,
34
+ 34,
35
+ 35,
36
+ 36,
37
+ 37,
38
+ 38,
39
+ 39,
40
+ 40,
41
+ 41,
42
+ 42,
43
+ 43,
44
+ 44,
45
+ 45,
46
+ 46,
47
+ 47,
48
+ 48,
49
+ 49,
50
+ 50,
51
+ 51,
52
+ 52,
53
+ 53,
54
+ 54,
55
+ 55,
56
+ 56,
57
+ 57,
58
+ 58,
59
+ 59,
60
+ 60,
61
+ 61,
62
+ 62,
63
+ 63,
64
+ 64,
65
+ 65,
66
+ 66,
67
+ 67,
68
+ 68,
69
+ 69,
70
+ 70,
71
+ 71,
72
+ 72,
73
+ 73,
74
+ 74,
75
+ 75,
76
+ 76,
77
+ 77
78
+ ],
79
+ "total_layers": 75,
80
+ "recipe": "full_corpus_sequential_lr_decay",
81
+ "training_tokens": 18001846,
82
+ "batch_tokens": 262144,
83
+ "microbatch_tokens": 65536,
84
+ "max_updates_per_layer": 69,
85
+ "adaptive_schedule": {
86
+ "from_layer": 27,
87
+ "relative_worsening": 0.001,
88
+ "checks": 3,
89
+ "patience_updates": 15,
90
+ "max_updates": 69,
91
+ "selection": "existing fixed reasoning validation; retain best checkpoint",
92
+ "description": "Drop LR4x after three consecutive validation checks >0.1% worse than best, or after45 at latest. After reduction stop following15 updates without a new best."
93
+ },
94
+ "lr_decay_after": 45,
95
+ "lr_decay_factor": 0.25,
96
+ "layer3_exception": "45 high-LR updates plus retained lower-LR refinement checkpoint",
97
+ "remaining_layers": "retain their previously published weights",
98
+ "previous_campaign_progress": {
99
+ "layers_pv_complete": [
100
+ 3,
101
+ 4,
102
+ 5,
103
+ 6,
104
+ 7,
105
+ 8
106
+ ],
107
+ "total_layers": 75,
108
+ "recipe": "sequential_reference_gradient_pv_v1",
109
+ "legacy_pv_layers": [
110
+ 9,
111
+ 10,
112
+ 11,
113
+ 12,
114
+ 13,
115
+ 14
116
+ ],
117
+ "initial_fit_layers": [
118
+ 15,
119
+ 16,
120
+ 17,
121
+ 18,
122
+ 19,
123
+ 20,
124
+ 21,
125
+ 22,
126
+ 23,
127
+ 24,
128
+ 25,
129
+ 26,
130
+ 27,
131
+ 28,
132
+ 29,
133
+ 30,
134
+ 31,
135
+ 32,
136
+ 33,
137
+ 34,
138
+ 35,
139
+ 36,
140
+ 37,
141
+ 38,
142
+ 39,
143
+ 40,
144
+ 41,
145
+ 42,
146
+ 43,
147
+ 44,
148
+ 45,
149
+ 46,
150
+ 47,
151
+ 48,
152
+ 49,
153
+ 50,
154
+ 51,
155
+ 52,
156
+ 53,
157
+ 54,
158
+ 55,
159
+ 56,
160
+ 57,
161
+ 58,
162
+ 59,
163
+ 60,
164
+ 61,
165
+ 62,
166
+ 63,
167
+ 64,
168
+ 65,
169
+ 66,
170
+ 67,
171
+ 68,
172
+ 69,
173
+ 70,
174
+ 71,
175
+ 72,
176
+ 73,
177
+ 74,
178
+ 75,
179
+ 76,
180
+ 77
181
+ ],
182
+ "format": "rvq256_256x8_expert",
183
+ "full_model_quality": "pending"
184
+ },
185
+ "format": "rvq256_256x8_expert",
186
+ "full_model_quality": "pending"
187
+ }
reproduce/historical_metadata/source.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "repo": "zai-org/GLM-5.3",
3
+ "revision": "aca966e4e02791568aa6a4ced368624b3d897f42",
4
+ "meta": {
5
+ "kind": "block_fp8",
6
+ "block": [
7
+ 128,
8
+ 128
9
+ ],
10
+ "fmt": "e4m3"
11
+ },
12
+ "cache": "/tmp/glm53-fp8-cold",
13
+ "legacy_cache_revision_unverified": true
14
+ }
reproduce/layer_recipe_summary.json ADDED
@@ -0,0 +1,1727 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "layer": 3,
4
+ "target": "same_input",
5
+ "block_scale_dtype": "fp16",
6
+ "training_sequence_count": 14657,
7
+ "training_tokens_seen": 15007754,
8
+ "training_passes": 1.0,
9
+ "best_step": 50,
10
+ "codebook_lr": 0.048,
11
+ "scale_lr": 0.032,
12
+ "lr_decay_after": 25,
13
+ "adaptive_schedule": {
14
+ "best": 0.004300016159961927,
15
+ "drop_after": 20,
16
+ "relative_worsening": 0.001,
17
+ "checks": 3,
18
+ "patience_updates": 10,
19
+ "bad_checks": 3,
20
+ "last_best": 50,
21
+ "early_drop": true
22
+ },
23
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
24
+ },
25
+ {
26
+ "layer": 4,
27
+ "target": "same_input",
28
+ "block_scale_dtype": "fp16",
29
+ "training_sequence_count": 14657,
30
+ "training_tokens_seen": 7864320,
31
+ "training_passes": 0.5240171180844249,
32
+ "best_step": 5,
33
+ "codebook_lr": 0.048,
34
+ "scale_lr": 0.032,
35
+ "lr_decay_after": 25,
36
+ "adaptive_schedule": {
37
+ "best": 0.011536361290655902,
38
+ "drop_after": 20,
39
+ "relative_worsening": 0.001,
40
+ "checks": 3,
41
+ "patience_updates": 10,
42
+ "bad_checks": 3,
43
+ "last_best": 5,
44
+ "early_drop": true
45
+ },
46
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
47
+ },
48
+ {
49
+ "layer": 5,
50
+ "target": "same_input",
51
+ "block_scale_dtype": "fp16",
52
+ "training_sequence_count": 14657,
53
+ "training_tokens_seen": 7864320,
54
+ "training_passes": 0.5240171180844249,
55
+ "best_step": 5,
56
+ "codebook_lr": 0.048,
57
+ "scale_lr": 0.032,
58
+ "lr_decay_after": 25,
59
+ "adaptive_schedule": {
60
+ "best": 0.017510809110214375,
61
+ "drop_after": 20,
62
+ "relative_worsening": 0.001,
63
+ "checks": 3,
64
+ "patience_updates": 10,
65
+ "bad_checks": 3,
66
+ "last_best": 5,
67
+ "early_drop": true
68
+ },
69
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
70
+ },
71
+ {
72
+ "layer": 6,
73
+ "target": "same_input",
74
+ "block_scale_dtype": "fp16",
75
+ "training_sequence_count": 14657,
76
+ "training_tokens_seen": 7864320,
77
+ "training_passes": 0.5240171180844249,
78
+ "best_step": 5,
79
+ "codebook_lr": 0.048,
80
+ "scale_lr": 0.032,
81
+ "lr_decay_after": 25,
82
+ "adaptive_schedule": {
83
+ "best": 0.018967027905769453,
84
+ "drop_after": 20,
85
+ "relative_worsening": 0.001,
86
+ "checks": 3,
87
+ "patience_updates": 10,
88
+ "bad_checks": 3,
89
+ "last_best": 5,
90
+ "early_drop": true
91
+ },
92
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
93
+ },
94
+ {
95
+ "layer": 7,
96
+ "target": "same_input",
97
+ "block_scale_dtype": "fp16",
98
+ "training_sequence_count": 14657,
99
+ "training_tokens_seen": 7864320,
100
+ "training_passes": 0.5240171180844249,
101
+ "best_step": 5,
102
+ "codebook_lr": 0.048,
103
+ "scale_lr": 0.032,
104
+ "lr_decay_after": 25,
105
+ "adaptive_schedule": {
106
+ "best": 0.02562102057711342,
107
+ "drop_after": 20,
108
+ "relative_worsening": 0.001,
109
+ "checks": 3,
110
+ "patience_updates": 10,
111
+ "bad_checks": 3,
112
+ "last_best": 5,
113
+ "early_drop": true
114
+ },
115
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
116
+ },
117
+ {
118
+ "layer": 8,
119
+ "target": "same_input",
120
+ "block_scale_dtype": "fp16",
121
+ "training_sequence_count": 14657,
122
+ "training_tokens_seen": 13107200,
123
+ "training_passes": 0.8733618634740414,
124
+ "best_step": 40,
125
+ "codebook_lr": 0.048,
126
+ "scale_lr": 0.032,
127
+ "lr_decay_after": 25,
128
+ "adaptive_schedule": {
129
+ "best": 0.022436852159198863,
130
+ "drop_after": 25,
131
+ "relative_worsening": 0.001,
132
+ "checks": 3,
133
+ "patience_updates": 10,
134
+ "bad_checks": 0,
135
+ "last_best": 40,
136
+ "early_drop": false
137
+ },
138
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
139
+ },
140
+ {
141
+ "layer": 9,
142
+ "target": "same_input",
143
+ "block_scale_dtype": "fp16",
144
+ "training_sequence_count": 14657,
145
+ "training_tokens_seen": 15007754,
146
+ "training_passes": 1.0,
147
+ "best_step": 58,
148
+ "codebook_lr": 0.048,
149
+ "scale_lr": 0.032,
150
+ "lr_decay_after": 25,
151
+ "adaptive_schedule": {
152
+ "best": 0.0017143418153641949,
153
+ "drop_after": 25,
154
+ "relative_worsening": 0.001,
155
+ "checks": 3,
156
+ "patience_updates": 10,
157
+ "bad_checks": 0,
158
+ "last_best": 58,
159
+ "early_drop": false
160
+ },
161
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
162
+ },
163
+ {
164
+ "layer": 10,
165
+ "target": "same_input",
166
+ "block_scale_dtype": "fp16",
167
+ "training_sequence_count": 14657,
168
+ "training_tokens_seen": 7864320,
169
+ "training_passes": 0.5240171180844249,
170
+ "best_step": 5,
171
+ "codebook_lr": 0.048,
172
+ "scale_lr": 0.032,
173
+ "lr_decay_after": 25,
174
+ "adaptive_schedule": {
175
+ "best": 0.0013590989109528494,
176
+ "drop_after": 20,
177
+ "relative_worsening": 0.001,
178
+ "checks": 3,
179
+ "patience_updates": 10,
180
+ "bad_checks": 3,
181
+ "last_best": 5,
182
+ "early_drop": true
183
+ },
184
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
185
+ },
186
+ {
187
+ "layer": 11,
188
+ "target": "same_input",
189
+ "block_scale_dtype": "fp16",
190
+ "training_sequence_count": 14657,
191
+ "training_tokens_seen": 7864320,
192
+ "training_passes": 0.5240171180844249,
193
+ "best_step": 5,
194
+ "codebook_lr": 0.048,
195
+ "scale_lr": 0.032,
196
+ "lr_decay_after": 25,
197
+ "adaptive_schedule": {
198
+ "best": 0.0013980529941903324,
199
+ "drop_after": 20,
200
+ "relative_worsening": 0.001,
201
+ "checks": 3,
202
+ "patience_updates": 10,
203
+ "bad_checks": 3,
204
+ "last_best": 5,
205
+ "early_drop": true
206
+ },
207
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
208
+ },
209
+ {
210
+ "layer": 12,
211
+ "target": "same_input",
212
+ "block_scale_dtype": "fp16",
213
+ "training_sequence_count": 14657,
214
+ "training_tokens_seen": 7864320,
215
+ "training_passes": 0.5240171180844249,
216
+ "best_step": 5,
217
+ "codebook_lr": 0.048,
218
+ "scale_lr": 0.032,
219
+ "lr_decay_after": 25,
220
+ "adaptive_schedule": {
221
+ "best": 0.001638401368651591,
222
+ "drop_after": 20,
223
+ "relative_worsening": 0.001,
224
+ "checks": 3,
225
+ "patience_updates": 10,
226
+ "bad_checks": 3,
227
+ "last_best": 5,
228
+ "early_drop": true
229
+ },
230
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
231
+ },
232
+ {
233
+ "layer": 13,
234
+ "target": "same_input",
235
+ "block_scale_dtype": "fp16",
236
+ "training_sequence_count": 14657,
237
+ "training_tokens_seen": 7864320,
238
+ "training_passes": 0.5240171180844249,
239
+ "best_step": 5,
240
+ "codebook_lr": 0.048,
241
+ "scale_lr": 0.032,
242
+ "lr_decay_after": 25,
243
+ "adaptive_schedule": {
244
+ "best": 0.00271562211069347,
245
+ "drop_after": 20,
246
+ "relative_worsening": 0.001,
247
+ "checks": 3,
248
+ "patience_updates": 10,
249
+ "bad_checks": 3,
250
+ "last_best": 5,
251
+ "early_drop": true
252
+ },
253
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
254
+ },
255
+ {
256
+ "layer": 14,
257
+ "target": "same_input",
258
+ "block_scale_dtype": "fp16",
259
+ "training_sequence_count": 14657,
260
+ "training_tokens_seen": 7864320,
261
+ "training_passes": 0.5240171180844249,
262
+ "best_step": 5,
263
+ "codebook_lr": 0.048,
264
+ "scale_lr": 0.032,
265
+ "lr_decay_after": 25,
266
+ "adaptive_schedule": {
267
+ "best": 0.0030091782715398136,
268
+ "drop_after": 20,
269
+ "relative_worsening": 0.001,
270
+ "checks": 3,
271
+ "patience_updates": 10,
272
+ "bad_checks": 3,
273
+ "last_best": 5,
274
+ "early_drop": true
275
+ },
276
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
277
+ },
278
+ {
279
+ "layer": 15,
280
+ "target": "same_input",
281
+ "block_scale_dtype": "fp16",
282
+ "training_sequence_count": 14657,
283
+ "training_tokens_seen": 7864320,
284
+ "training_passes": 0.5240171180844249,
285
+ "best_step": 5,
286
+ "codebook_lr": 0.048,
287
+ "scale_lr": 0.032,
288
+ "lr_decay_after": 25,
289
+ "adaptive_schedule": {
290
+ "best": 0.003168485440774852,
291
+ "drop_after": 20,
292
+ "relative_worsening": 0.001,
293
+ "checks": 3,
294
+ "patience_updates": 10,
295
+ "bad_checks": 3,
296
+ "last_best": 5,
297
+ "early_drop": true
298
+ },
299
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
300
+ },
301
+ {
302
+ "layer": 16,
303
+ "target": "same_input",
304
+ "block_scale_dtype": "fp16",
305
+ "training_sequence_count": 14657,
306
+ "training_tokens_seen": 7864320,
307
+ "training_passes": 0.5240171180844249,
308
+ "best_step": 5,
309
+ "codebook_lr": 0.048,
310
+ "scale_lr": 0.032,
311
+ "lr_decay_after": 25,
312
+ "adaptive_schedule": {
313
+ "best": 0.004101365709508885,
314
+ "drop_after": 20,
315
+ "relative_worsening": 0.001,
316
+ "checks": 3,
317
+ "patience_updates": 10,
318
+ "bad_checks": 3,
319
+ "last_best": 5,
320
+ "early_drop": true
321
+ },
322
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
323
+ },
324
+ {
325
+ "layer": 17,
326
+ "target": "same_input",
327
+ "block_scale_dtype": "fp16",
328
+ "training_sequence_count": 14657,
329
+ "training_tokens_seen": 7864320,
330
+ "training_passes": 0.5240171180844249,
331
+ "best_step": 5,
332
+ "codebook_lr": 0.048,
333
+ "scale_lr": 0.032,
334
+ "lr_decay_after": 25,
335
+ "adaptive_schedule": {
336
+ "best": 0.003969997484537619,
337
+ "drop_after": 20,
338
+ "relative_worsening": 0.001,
339
+ "checks": 3,
340
+ "patience_updates": 10,
341
+ "bad_checks": 3,
342
+ "last_best": 5,
343
+ "early_drop": true
344
+ },
345
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
346
+ },
347
+ {
348
+ "layer": 18,
349
+ "target": "same_input",
350
+ "block_scale_dtype": "fp16",
351
+ "training_sequence_count": 14657,
352
+ "training_tokens_seen": 7864320,
353
+ "training_passes": 0.5240171180844249,
354
+ "best_step": 5,
355
+ "codebook_lr": 0.048,
356
+ "scale_lr": 0.032,
357
+ "lr_decay_after": 25,
358
+ "adaptive_schedule": {
359
+ "best": 0.00536306292142981,
360
+ "drop_after": 20,
361
+ "relative_worsening": 0.001,
362
+ "checks": 3,
363
+ "patience_updates": 10,
364
+ "bad_checks": 3,
365
+ "last_best": 5,
366
+ "early_drop": true
367
+ },
368
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
369
+ },
370
+ {
371
+ "layer": 19,
372
+ "target": "same_input",
373
+ "block_scale_dtype": "fp16",
374
+ "training_sequence_count": 14657,
375
+ "training_tokens_seen": 7864320,
376
+ "training_passes": 0.5240171180844249,
377
+ "best_step": 5,
378
+ "codebook_lr": 0.048,
379
+ "scale_lr": 0.032,
380
+ "lr_decay_after": 25,
381
+ "adaptive_schedule": {
382
+ "best": 0.005900326847720345,
383
+ "drop_after": 20,
384
+ "relative_worsening": 0.001,
385
+ "checks": 3,
386
+ "patience_updates": 10,
387
+ "bad_checks": 3,
388
+ "last_best": 5,
389
+ "early_drop": true
390
+ },
391
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
392
+ },
393
+ {
394
+ "layer": 20,
395
+ "target": "same_input",
396
+ "block_scale_dtype": "fp16",
397
+ "training_sequence_count": 14657,
398
+ "training_tokens_seen": 7864320,
399
+ "training_passes": 0.5240171180844249,
400
+ "best_step": 5,
401
+ "codebook_lr": 0.048,
402
+ "scale_lr": 0.032,
403
+ "lr_decay_after": 25,
404
+ "adaptive_schedule": {
405
+ "best": 0.00827687275281147,
406
+ "drop_after": 20,
407
+ "relative_worsening": 0.001,
408
+ "checks": 3,
409
+ "patience_updates": 10,
410
+ "bad_checks": 3,
411
+ "last_best": 5,
412
+ "early_drop": true
413
+ },
414
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
415
+ },
416
+ {
417
+ "layer": 21,
418
+ "target": "same_input",
419
+ "block_scale_dtype": "fp16",
420
+ "training_sequence_count": 14657,
421
+ "training_tokens_seen": 7864320,
422
+ "training_passes": 0.5240171180844249,
423
+ "best_step": 5,
424
+ "codebook_lr": 0.048,
425
+ "scale_lr": 0.032,
426
+ "lr_decay_after": 25,
427
+ "adaptive_schedule": {
428
+ "best": 0.007670614413287488,
429
+ "drop_after": 20,
430
+ "relative_worsening": 0.001,
431
+ "checks": 3,
432
+ "patience_updates": 10,
433
+ "bad_checks": 3,
434
+ "last_best": 5,
435
+ "early_drop": true
436
+ },
437
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
438
+ },
439
+ {
440
+ "layer": 22,
441
+ "target": "same_input",
442
+ "block_scale_dtype": "fp16",
443
+ "training_sequence_count": 14657,
444
+ "training_tokens_seen": 7864320,
445
+ "training_passes": 0.5240171180844249,
446
+ "best_step": 5,
447
+ "codebook_lr": 0.048,
448
+ "scale_lr": 0.032,
449
+ "lr_decay_after": 25,
450
+ "adaptive_schedule": {
451
+ "best": 0.009540958133708182,
452
+ "drop_after": 20,
453
+ "relative_worsening": 0.001,
454
+ "checks": 3,
455
+ "patience_updates": 10,
456
+ "bad_checks": 3,
457
+ "last_best": 5,
458
+ "early_drop": true
459
+ },
460
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
461
+ },
462
+ {
463
+ "layer": 23,
464
+ "target": "same_input",
465
+ "block_scale_dtype": "fp16",
466
+ "training_sequence_count": 14657,
467
+ "training_tokens_seen": 7864320,
468
+ "training_passes": 0.5240171180844249,
469
+ "best_step": 5,
470
+ "codebook_lr": 0.048,
471
+ "scale_lr": 0.032,
472
+ "lr_decay_after": 25,
473
+ "adaptive_schedule": {
474
+ "best": 0.011060020056539258,
475
+ "drop_after": 20,
476
+ "relative_worsening": 0.001,
477
+ "checks": 3,
478
+ "patience_updates": 10,
479
+ "bad_checks": 3,
480
+ "last_best": 5,
481
+ "early_drop": true
482
+ },
483
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
484
+ },
485
+ {
486
+ "layer": 24,
487
+ "target": "same_input",
488
+ "block_scale_dtype": "fp16",
489
+ "training_sequence_count": 14657,
490
+ "training_tokens_seen": 7864320,
491
+ "training_passes": 0.5240171180844249,
492
+ "best_step": 5,
493
+ "codebook_lr": 0.048,
494
+ "scale_lr": 0.032,
495
+ "lr_decay_after": 25,
496
+ "adaptive_schedule": {
497
+ "best": 0.009994059305569054,
498
+ "drop_after": 20,
499
+ "relative_worsening": 0.001,
500
+ "checks": 3,
501
+ "patience_updates": 10,
502
+ "bad_checks": 3,
503
+ "last_best": 5,
504
+ "early_drop": true
505
+ },
506
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
507
+ },
508
+ {
509
+ "layer": 25,
510
+ "target": "same_input",
511
+ "block_scale_dtype": "fp16",
512
+ "training_sequence_count": 14657,
513
+ "training_tokens_seen": 7864320,
514
+ "training_passes": 0.5240171180844249,
515
+ "best_step": 5,
516
+ "codebook_lr": 0.048,
517
+ "scale_lr": 0.032,
518
+ "lr_decay_after": 25,
519
+ "adaptive_schedule": {
520
+ "best": 0.010544123217574992,
521
+ "drop_after": 20,
522
+ "relative_worsening": 0.001,
523
+ "checks": 3,
524
+ "patience_updates": 10,
525
+ "bad_checks": 3,
526
+ "last_best": 5,
527
+ "early_drop": true
528
+ },
529
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
530
+ },
531
+ {
532
+ "layer": 26,
533
+ "target": "same_input",
534
+ "block_scale_dtype": "fp16",
535
+ "training_sequence_count": 14657,
536
+ "training_tokens_seen": 7864320,
537
+ "training_passes": 0.5240171180844249,
538
+ "best_step": 5,
539
+ "codebook_lr": 0.048,
540
+ "scale_lr": 0.032,
541
+ "lr_decay_after": 25,
542
+ "adaptive_schedule": {
543
+ "best": 0.012252411564001522,
544
+ "drop_after": 20,
545
+ "relative_worsening": 0.001,
546
+ "checks": 3,
547
+ "patience_updates": 10,
548
+ "bad_checks": 3,
549
+ "last_best": 5,
550
+ "early_drop": true
551
+ },
552
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
553
+ },
554
+ {
555
+ "layer": 27,
556
+ "target": "same_input",
557
+ "block_scale_dtype": "fp16",
558
+ "training_sequence_count": 14657,
559
+ "training_tokens_seen": 7864320,
560
+ "training_passes": 0.5240171180844249,
561
+ "best_step": 5,
562
+ "codebook_lr": 0.048,
563
+ "scale_lr": 0.032,
564
+ "lr_decay_after": 25,
565
+ "adaptive_schedule": {
566
+ "best": 0.017041586186703876,
567
+ "drop_after": 20,
568
+ "relative_worsening": 0.001,
569
+ "checks": 3,
570
+ "patience_updates": 10,
571
+ "bad_checks": 3,
572
+ "last_best": 5,
573
+ "early_drop": true
574
+ },
575
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
576
+ },
577
+ {
578
+ "layer": 28,
579
+ "target": "same_input",
580
+ "block_scale_dtype": "fp16",
581
+ "training_sequence_count": 14657,
582
+ "training_tokens_seen": 9175040,
583
+ "training_passes": 0.611353304431829,
584
+ "best_step": 25,
585
+ "codebook_lr": 0.048,
586
+ "scale_lr": 0.032,
587
+ "lr_decay_after": 25,
588
+ "adaptive_schedule": {
589
+ "best": 0.018901359302169598,
590
+ "drop_after": 20,
591
+ "relative_worsening": 0.001,
592
+ "checks": 3,
593
+ "patience_updates": 10,
594
+ "bad_checks": 3,
595
+ "last_best": 25,
596
+ "early_drop": true
597
+ },
598
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
599
+ },
600
+ {
601
+ "layer": 29,
602
+ "target": "same_input",
603
+ "block_scale_dtype": "fp16",
604
+ "training_sequence_count": 14657,
605
+ "training_tokens_seen": 7864320,
606
+ "training_passes": 0.5240171180844249,
607
+ "best_step": 5,
608
+ "codebook_lr": 0.048,
609
+ "scale_lr": 0.032,
610
+ "lr_decay_after": 25,
611
+ "adaptive_schedule": {
612
+ "best": 0.01730259021141979,
613
+ "drop_after": 20,
614
+ "relative_worsening": 0.001,
615
+ "checks": 3,
616
+ "patience_updates": 10,
617
+ "bad_checks": 3,
618
+ "last_best": 5,
619
+ "early_drop": true
620
+ },
621
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
622
+ },
623
+ {
624
+ "layer": 30,
625
+ "target": "same_input",
626
+ "block_scale_dtype": "fp16",
627
+ "training_sequence_count": 14657,
628
+ "training_tokens_seen": 9175040,
629
+ "training_passes": 0.611353304431829,
630
+ "best_step": 25,
631
+ "codebook_lr": 0.048,
632
+ "scale_lr": 0.032,
633
+ "lr_decay_after": 25,
634
+ "adaptive_schedule": {
635
+ "best": 0.01953328825434068,
636
+ "drop_after": 20,
637
+ "relative_worsening": 0.001,
638
+ "checks": 3,
639
+ "patience_updates": 10,
640
+ "bad_checks": 3,
641
+ "last_best": 25,
642
+ "early_drop": true
643
+ },
644
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
645
+ },
646
+ {
647
+ "layer": 31,
648
+ "target": "same_input",
649
+ "block_scale_dtype": "fp16",
650
+ "training_sequence_count": 14657,
651
+ "training_tokens_seen": 7864320,
652
+ "training_passes": 0.5240171180844249,
653
+ "best_step": 5,
654
+ "codebook_lr": 0.048,
655
+ "scale_lr": 0.032,
656
+ "lr_decay_after": 25,
657
+ "adaptive_schedule": {
658
+ "best": 0.0185387897556427,
659
+ "drop_after": 20,
660
+ "relative_worsening": 0.001,
661
+ "checks": 3,
662
+ "patience_updates": 10,
663
+ "bad_checks": 3,
664
+ "last_best": 5,
665
+ "early_drop": true
666
+ },
667
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
668
+ },
669
+ {
670
+ "layer": 32,
671
+ "target": "same_input",
672
+ "block_scale_dtype": "fp16",
673
+ "training_sequence_count": 14657,
674
+ "training_tokens_seen": 7864320,
675
+ "training_passes": 0.5240171180844249,
676
+ "best_step": 5,
677
+ "codebook_lr": 0.048,
678
+ "scale_lr": 0.032,
679
+ "lr_decay_after": 25,
680
+ "adaptive_schedule": {
681
+ "best": 0.02538125707815895,
682
+ "drop_after": 20,
683
+ "relative_worsening": 0.001,
684
+ "checks": 3,
685
+ "patience_updates": 10,
686
+ "bad_checks": 3,
687
+ "last_best": 5,
688
+ "early_drop": true
689
+ },
690
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
691
+ },
692
+ {
693
+ "layer": 33,
694
+ "target": "same_input",
695
+ "block_scale_dtype": "fp16",
696
+ "training_sequence_count": 14657,
697
+ "training_tokens_seen": 7864320,
698
+ "training_passes": 0.5240171180844249,
699
+ "best_step": 5,
700
+ "codebook_lr": 0.048,
701
+ "scale_lr": 0.032,
702
+ "lr_decay_after": 25,
703
+ "adaptive_schedule": {
704
+ "best": 0.02889075275987553,
705
+ "drop_after": 20,
706
+ "relative_worsening": 0.001,
707
+ "checks": 3,
708
+ "patience_updates": 10,
709
+ "bad_checks": 3,
710
+ "last_best": 5,
711
+ "early_drop": true
712
+ },
713
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
714
+ },
715
+ {
716
+ "layer": 34,
717
+ "target": "same_input",
718
+ "block_scale_dtype": "fp16",
719
+ "training_sequence_count": 14657,
720
+ "training_tokens_seen": 7864320,
721
+ "training_passes": 0.5240171180844249,
722
+ "best_step": 5,
723
+ "codebook_lr": 0.048,
724
+ "scale_lr": 0.032,
725
+ "lr_decay_after": 25,
726
+ "adaptive_schedule": {
727
+ "best": 0.03261215982052183,
728
+ "drop_after": 20,
729
+ "relative_worsening": 0.001,
730
+ "checks": 3,
731
+ "patience_updates": 10,
732
+ "bad_checks": 3,
733
+ "last_best": 5,
734
+ "early_drop": true
735
+ },
736
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
737
+ },
738
+ {
739
+ "layer": 35,
740
+ "target": "same_input",
741
+ "block_scale_dtype": "fp16",
742
+ "training_sequence_count": 14657,
743
+ "training_tokens_seen": 7864320,
744
+ "training_passes": 0.5240171180844249,
745
+ "best_step": 5,
746
+ "codebook_lr": 0.048,
747
+ "scale_lr": 0.032,
748
+ "lr_decay_after": 25,
749
+ "adaptive_schedule": {
750
+ "best": 0.02849563835939984,
751
+ "drop_after": 20,
752
+ "relative_worsening": 0.001,
753
+ "checks": 3,
754
+ "patience_updates": 10,
755
+ "bad_checks": 3,
756
+ "last_best": 5,
757
+ "early_drop": true
758
+ },
759
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
760
+ },
761
+ {
762
+ "layer": 36,
763
+ "target": "same_input",
764
+ "block_scale_dtype": "fp16",
765
+ "training_sequence_count": 14657,
766
+ "training_tokens_seen": 7864320,
767
+ "training_passes": 0.5240171180844249,
768
+ "best_step": 5,
769
+ "codebook_lr": 0.048,
770
+ "scale_lr": 0.032,
771
+ "lr_decay_after": 25,
772
+ "adaptive_schedule": {
773
+ "best": 0.033030121256077974,
774
+ "drop_after": 20,
775
+ "relative_worsening": 0.001,
776
+ "checks": 3,
777
+ "patience_updates": 10,
778
+ "bad_checks": 3,
779
+ "last_best": 5,
780
+ "early_drop": true
781
+ },
782
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
783
+ },
784
+ {
785
+ "layer": 37,
786
+ "target": "same_input",
787
+ "block_scale_dtype": "fp16",
788
+ "training_sequence_count": 14657,
789
+ "training_tokens_seen": 7864320,
790
+ "training_passes": 0.5240171180844249,
791
+ "best_step": 5,
792
+ "codebook_lr": 0.048,
793
+ "scale_lr": 0.032,
794
+ "lr_decay_after": 25,
795
+ "adaptive_schedule": {
796
+ "best": 0.035430789227037796,
797
+ "drop_after": 20,
798
+ "relative_worsening": 0.001,
799
+ "checks": 3,
800
+ "patience_updates": 10,
801
+ "bad_checks": 3,
802
+ "last_best": 5,
803
+ "early_drop": true
804
+ },
805
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
806
+ },
807
+ {
808
+ "layer": 38,
809
+ "target": "same_input",
810
+ "block_scale_dtype": "fp16",
811
+ "training_sequence_count": 14657,
812
+ "training_tokens_seen": 7864320,
813
+ "training_passes": 0.5240171180844249,
814
+ "best_step": 5,
815
+ "codebook_lr": 0.048,
816
+ "scale_lr": 0.032,
817
+ "lr_decay_after": 25,
818
+ "adaptive_schedule": {
819
+ "best": 0.03938840653001622,
820
+ "drop_after": 20,
821
+ "relative_worsening": 0.001,
822
+ "checks": 3,
823
+ "patience_updates": 10,
824
+ "bad_checks": 3,
825
+ "last_best": 5,
826
+ "early_drop": true
827
+ },
828
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
829
+ },
830
+ {
831
+ "layer": 39,
832
+ "target": "same_input",
833
+ "block_scale_dtype": "fp16",
834
+ "training_sequence_count": 14657,
835
+ "training_tokens_seen": 7864320,
836
+ "training_passes": 0.5240171180844249,
837
+ "best_step": 5,
838
+ "codebook_lr": 0.048,
839
+ "scale_lr": 0.032,
840
+ "lr_decay_after": 25,
841
+ "adaptive_schedule": {
842
+ "best": 0.04198908163931699,
843
+ "drop_after": 20,
844
+ "relative_worsening": 0.001,
845
+ "checks": 3,
846
+ "patience_updates": 10,
847
+ "bad_checks": 3,
848
+ "last_best": 5,
849
+ "early_drop": true
850
+ },
851
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
852
+ },
853
+ {
854
+ "layer": 40,
855
+ "target": "same_input",
856
+ "block_scale_dtype": "fp16",
857
+ "training_sequence_count": 14657,
858
+ "training_tokens_seen": 7864320,
859
+ "training_passes": 0.5240171180844249,
860
+ "best_step": 5,
861
+ "codebook_lr": 0.048,
862
+ "scale_lr": 0.032,
863
+ "lr_decay_after": 25,
864
+ "adaptive_schedule": {
865
+ "best": 0.040921155033096596,
866
+ "drop_after": 20,
867
+ "relative_worsening": 0.001,
868
+ "checks": 3,
869
+ "patience_updates": 10,
870
+ "bad_checks": 3,
871
+ "last_best": 5,
872
+ "early_drop": true
873
+ },
874
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
875
+ },
876
+ {
877
+ "layer": 41,
878
+ "target": "same_input",
879
+ "block_scale_dtype": "fp16",
880
+ "training_sequence_count": 14657,
881
+ "training_tokens_seen": 7864320,
882
+ "training_passes": 0.5240171180844249,
883
+ "best_step": 5,
884
+ "codebook_lr": 0.048,
885
+ "scale_lr": 0.032,
886
+ "lr_decay_after": 25,
887
+ "adaptive_schedule": {
888
+ "best": 0.04285828934035643,
889
+ "drop_after": 20,
890
+ "relative_worsening": 0.001,
891
+ "checks": 3,
892
+ "patience_updates": 10,
893
+ "bad_checks": 3,
894
+ "last_best": 5,
895
+ "early_drop": true
896
+ },
897
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
898
+ },
899
+ {
900
+ "layer": 42,
901
+ "target": "same_input",
902
+ "block_scale_dtype": "fp16",
903
+ "training_sequence_count": 14657,
904
+ "training_tokens_seen": 7864320,
905
+ "training_passes": 0.5240171180844249,
906
+ "best_step": 5,
907
+ "codebook_lr": 0.048,
908
+ "scale_lr": 0.032,
909
+ "lr_decay_after": 25,
910
+ "adaptive_schedule": {
911
+ "best": 0.04590353936917018,
912
+ "drop_after": 20,
913
+ "relative_worsening": 0.001,
914
+ "checks": 3,
915
+ "patience_updates": 10,
916
+ "bad_checks": 3,
917
+ "last_best": 5,
918
+ "early_drop": true
919
+ },
920
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
921
+ },
922
+ {
923
+ "layer": 43,
924
+ "target": "same_input",
925
+ "block_scale_dtype": "fp16",
926
+ "training_sequence_count": 14657,
927
+ "training_tokens_seen": 7864320,
928
+ "training_passes": 0.5240171180844249,
929
+ "best_step": 5,
930
+ "codebook_lr": 0.048,
931
+ "scale_lr": 0.032,
932
+ "lr_decay_after": 25,
933
+ "adaptive_schedule": {
934
+ "best": 0.0445766345634676,
935
+ "drop_after": 20,
936
+ "relative_worsening": 0.001,
937
+ "checks": 3,
938
+ "patience_updates": 10,
939
+ "bad_checks": 3,
940
+ "last_best": 5,
941
+ "early_drop": true
942
+ },
943
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
944
+ },
945
+ {
946
+ "layer": 44,
947
+ "target": "same_input",
948
+ "block_scale_dtype": "fp16",
949
+ "training_sequence_count": 14657,
950
+ "training_tokens_seen": 7864320,
951
+ "training_passes": 0.5240171180844249,
952
+ "best_step": 5,
953
+ "codebook_lr": 0.048,
954
+ "scale_lr": 0.032,
955
+ "lr_decay_after": 25,
956
+ "adaptive_schedule": {
957
+ "best": 0.04553051376240955,
958
+ "drop_after": 20,
959
+ "relative_worsening": 0.001,
960
+ "checks": 3,
961
+ "patience_updates": 10,
962
+ "bad_checks": 3,
963
+ "last_best": 5,
964
+ "early_drop": true
965
+ },
966
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
967
+ },
968
+ {
969
+ "layer": 45,
970
+ "target": "same_input",
971
+ "block_scale_dtype": "fp16",
972
+ "training_sequence_count": 14657,
973
+ "training_tokens_seen": 7864320,
974
+ "training_passes": 0.5240171180844249,
975
+ "best_step": 5,
976
+ "codebook_lr": 0.048,
977
+ "scale_lr": 0.032,
978
+ "lr_decay_after": 25,
979
+ "adaptive_schedule": {
980
+ "best": 0.04488142473736064,
981
+ "drop_after": 20,
982
+ "relative_worsening": 0.001,
983
+ "checks": 3,
984
+ "patience_updates": 10,
985
+ "bad_checks": 3,
986
+ "last_best": 5,
987
+ "early_drop": true
988
+ },
989
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
990
+ },
991
+ {
992
+ "layer": 46,
993
+ "target": "same_input",
994
+ "block_scale_dtype": "fp16",
995
+ "training_sequence_count": 14657,
996
+ "training_tokens_seen": 7864320,
997
+ "training_passes": 0.5240171180844249,
998
+ "best_step": 5,
999
+ "codebook_lr": 0.048,
1000
+ "scale_lr": 0.032,
1001
+ "lr_decay_after": 25,
1002
+ "adaptive_schedule": {
1003
+ "best": 0.053625666510220674,
1004
+ "drop_after": 20,
1005
+ "relative_worsening": 0.001,
1006
+ "checks": 3,
1007
+ "patience_updates": 10,
1008
+ "bad_checks": 3,
1009
+ "last_best": 5,
1010
+ "early_drop": true
1011
+ },
1012
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1013
+ },
1014
+ {
1015
+ "layer": 47,
1016
+ "target": "same_input",
1017
+ "block_scale_dtype": "fp16",
1018
+ "training_sequence_count": 14657,
1019
+ "training_tokens_seen": 7864320,
1020
+ "training_passes": 0.5240171180844249,
1021
+ "best_step": 5,
1022
+ "codebook_lr": 0.048,
1023
+ "scale_lr": 0.032,
1024
+ "lr_decay_after": 25,
1025
+ "adaptive_schedule": {
1026
+ "best": 0.05097487104130676,
1027
+ "drop_after": 20,
1028
+ "relative_worsening": 0.001,
1029
+ "checks": 3,
1030
+ "patience_updates": 10,
1031
+ "bad_checks": 3,
1032
+ "last_best": 5,
1033
+ "early_drop": true
1034
+ },
1035
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1036
+ },
1037
+ {
1038
+ "layer": 48,
1039
+ "target": "same_input",
1040
+ "block_scale_dtype": "fp16",
1041
+ "training_sequence_count": 14657,
1042
+ "training_tokens_seen": 7864320,
1043
+ "training_passes": 0.5240171180844249,
1044
+ "best_step": 5,
1045
+ "codebook_lr": 0.048,
1046
+ "scale_lr": 0.032,
1047
+ "lr_decay_after": 25,
1048
+ "adaptive_schedule": {
1049
+ "best": 0.05253228264238492,
1050
+ "drop_after": 20,
1051
+ "relative_worsening": 0.001,
1052
+ "checks": 3,
1053
+ "patience_updates": 10,
1054
+ "bad_checks": 3,
1055
+ "last_best": 5,
1056
+ "early_drop": true
1057
+ },
1058
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1059
+ },
1060
+ {
1061
+ "layer": 49,
1062
+ "target": "same_input",
1063
+ "block_scale_dtype": "fp16",
1064
+ "training_sequence_count": 14657,
1065
+ "training_tokens_seen": 7864320,
1066
+ "training_passes": 0.5240171180844249,
1067
+ "best_step": 5,
1068
+ "codebook_lr": 0.048,
1069
+ "scale_lr": 0.032,
1070
+ "lr_decay_after": 25,
1071
+ "adaptive_schedule": {
1072
+ "best": 0.054531234896217015,
1073
+ "drop_after": 20,
1074
+ "relative_worsening": 0.001,
1075
+ "checks": 3,
1076
+ "patience_updates": 10,
1077
+ "bad_checks": 3,
1078
+ "last_best": 5,
1079
+ "early_drop": true
1080
+ },
1081
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1082
+ },
1083
+ {
1084
+ "layer": 50,
1085
+ "target": "same_input",
1086
+ "block_scale_dtype": "fp16",
1087
+ "training_sequence_count": 14657,
1088
+ "training_tokens_seen": 7864320,
1089
+ "training_passes": 0.5240171180844249,
1090
+ "best_step": 5,
1091
+ "codebook_lr": 0.048,
1092
+ "scale_lr": 0.032,
1093
+ "lr_decay_after": 25,
1094
+ "adaptive_schedule": {
1095
+ "best": 0.05778833536756532,
1096
+ "drop_after": 20,
1097
+ "relative_worsening": 0.001,
1098
+ "checks": 3,
1099
+ "patience_updates": 10,
1100
+ "bad_checks": 3,
1101
+ "last_best": 5,
1102
+ "early_drop": true
1103
+ },
1104
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1105
+ },
1106
+ {
1107
+ "layer": 51,
1108
+ "target": "same_input",
1109
+ "block_scale_dtype": "fp16",
1110
+ "training_sequence_count": 14657,
1111
+ "training_tokens_seen": 7864320,
1112
+ "training_passes": 0.5240171180844249,
1113
+ "best_step": 5,
1114
+ "codebook_lr": 0.048,
1115
+ "scale_lr": 0.032,
1116
+ "lr_decay_after": 25,
1117
+ "adaptive_schedule": {
1118
+ "best": 0.046839461506473105,
1119
+ "drop_after": 20,
1120
+ "relative_worsening": 0.001,
1121
+ "checks": 3,
1122
+ "patience_updates": 10,
1123
+ "bad_checks": 3,
1124
+ "last_best": 5,
1125
+ "early_drop": true
1126
+ },
1127
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1128
+ },
1129
+ {
1130
+ "layer": 52,
1131
+ "target": "same_input",
1132
+ "block_scale_dtype": "fp16",
1133
+ "training_sequence_count": 14657,
1134
+ "training_tokens_seen": 7864320,
1135
+ "training_passes": 0.5240171180844249,
1136
+ "best_step": 5,
1137
+ "codebook_lr": 0.048,
1138
+ "scale_lr": 0.032,
1139
+ "lr_decay_after": 25,
1140
+ "adaptive_schedule": {
1141
+ "best": 0.05742379963086141,
1142
+ "drop_after": 20,
1143
+ "relative_worsening": 0.001,
1144
+ "checks": 3,
1145
+ "patience_updates": 10,
1146
+ "bad_checks": 3,
1147
+ "last_best": 5,
1148
+ "early_drop": true
1149
+ },
1150
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1151
+ },
1152
+ {
1153
+ "layer": 53,
1154
+ "target": "same_input",
1155
+ "block_scale_dtype": "fp16",
1156
+ "training_sequence_count": 14657,
1157
+ "training_tokens_seen": 7864320,
1158
+ "training_passes": 0.5240171180844249,
1159
+ "best_step": 5,
1160
+ "codebook_lr": 0.048,
1161
+ "scale_lr": 0.032,
1162
+ "lr_decay_after": 25,
1163
+ "adaptive_schedule": {
1164
+ "best": 0.05401293698903637,
1165
+ "drop_after": 20,
1166
+ "relative_worsening": 0.001,
1167
+ "checks": 3,
1168
+ "patience_updates": 10,
1169
+ "bad_checks": 3,
1170
+ "last_best": 5,
1171
+ "early_drop": true
1172
+ },
1173
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1174
+ },
1175
+ {
1176
+ "layer": 54,
1177
+ "target": "same_input",
1178
+ "block_scale_dtype": "fp16",
1179
+ "training_sequence_count": 14657,
1180
+ "training_tokens_seen": 7864320,
1181
+ "training_passes": 0.5240171180844249,
1182
+ "best_step": 5,
1183
+ "codebook_lr": 0.048,
1184
+ "scale_lr": 0.032,
1185
+ "lr_decay_after": 25,
1186
+ "adaptive_schedule": {
1187
+ "best": 0.05698656719338924,
1188
+ "drop_after": 20,
1189
+ "relative_worsening": 0.001,
1190
+ "checks": 3,
1191
+ "patience_updates": 10,
1192
+ "bad_checks": 3,
1193
+ "last_best": 5,
1194
+ "early_drop": true
1195
+ },
1196
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1197
+ },
1198
+ {
1199
+ "layer": 55,
1200
+ "target": "same_input",
1201
+ "block_scale_dtype": "fp16",
1202
+ "training_sequence_count": 14657,
1203
+ "training_tokens_seen": 7864320,
1204
+ "training_passes": 0.5240171180844249,
1205
+ "best_step": 5,
1206
+ "codebook_lr": 0.048,
1207
+ "scale_lr": 0.032,
1208
+ "lr_decay_after": 25,
1209
+ "adaptive_schedule": {
1210
+ "best": 0.05936456247157864,
1211
+ "drop_after": 20,
1212
+ "relative_worsening": 0.001,
1213
+ "checks": 3,
1214
+ "patience_updates": 10,
1215
+ "bad_checks": 3,
1216
+ "last_best": 5,
1217
+ "early_drop": true
1218
+ },
1219
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1220
+ },
1221
+ {
1222
+ "layer": 56,
1223
+ "target": "same_input",
1224
+ "block_scale_dtype": "fp16",
1225
+ "training_sequence_count": 14657,
1226
+ "training_tokens_seen": 7864320,
1227
+ "training_passes": 0.5240171180844249,
1228
+ "best_step": 5,
1229
+ "codebook_lr": 0.048,
1230
+ "scale_lr": 0.032,
1231
+ "lr_decay_after": 25,
1232
+ "adaptive_schedule": {
1233
+ "best": 0.0527071627420072,
1234
+ "drop_after": 20,
1235
+ "relative_worsening": 0.001,
1236
+ "checks": 3,
1237
+ "patience_updates": 10,
1238
+ "bad_checks": 3,
1239
+ "last_best": 5,
1240
+ "early_drop": true
1241
+ },
1242
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1243
+ },
1244
+ {
1245
+ "layer": 57,
1246
+ "target": "same_input",
1247
+ "block_scale_dtype": "fp16",
1248
+ "training_sequence_count": 14657,
1249
+ "training_tokens_seen": 7864320,
1250
+ "training_passes": 0.5240171180844249,
1251
+ "best_step": 5,
1252
+ "codebook_lr": 0.048,
1253
+ "scale_lr": 0.032,
1254
+ "lr_decay_after": 25,
1255
+ "adaptive_schedule": {
1256
+ "best": 0.04895510529481134,
1257
+ "drop_after": 20,
1258
+ "relative_worsening": 0.001,
1259
+ "checks": 3,
1260
+ "patience_updates": 10,
1261
+ "bad_checks": 3,
1262
+ "last_best": 5,
1263
+ "early_drop": true
1264
+ },
1265
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1266
+ },
1267
+ {
1268
+ "layer": 58,
1269
+ "target": "same_input",
1270
+ "block_scale_dtype": "fp16",
1271
+ "training_sequence_count": 14657,
1272
+ "training_tokens_seen": 7864320,
1273
+ "training_passes": 0.5240171180844249,
1274
+ "best_step": 5,
1275
+ "codebook_lr": 0.048,
1276
+ "scale_lr": 0.032,
1277
+ "lr_decay_after": 25,
1278
+ "adaptive_schedule": {
1279
+ "best": 0.05838005265713665,
1280
+ "drop_after": 20,
1281
+ "relative_worsening": 0.001,
1282
+ "checks": 3,
1283
+ "patience_updates": 10,
1284
+ "bad_checks": 3,
1285
+ "last_best": 5,
1286
+ "early_drop": true
1287
+ },
1288
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1289
+ },
1290
+ {
1291
+ "layer": 59,
1292
+ "target": "same_input",
1293
+ "block_scale_dtype": "fp16",
1294
+ "training_sequence_count": 14657,
1295
+ "training_tokens_seen": 7864320,
1296
+ "training_passes": 0.5240171180844249,
1297
+ "best_step": 5,
1298
+ "codebook_lr": 0.048,
1299
+ "scale_lr": 0.032,
1300
+ "lr_decay_after": 25,
1301
+ "adaptive_schedule": {
1302
+ "best": 0.06076521669703373,
1303
+ "drop_after": 20,
1304
+ "relative_worsening": 0.001,
1305
+ "checks": 3,
1306
+ "patience_updates": 10,
1307
+ "bad_checks": 3,
1308
+ "last_best": 5,
1309
+ "early_drop": true
1310
+ },
1311
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1312
+ },
1313
+ {
1314
+ "layer": 60,
1315
+ "target": "same_input",
1316
+ "block_scale_dtype": "fp16",
1317
+ "training_sequence_count": 14657,
1318
+ "training_tokens_seen": 7864320,
1319
+ "training_passes": 0.5240171180844249,
1320
+ "best_step": 5,
1321
+ "codebook_lr": 0.048,
1322
+ "scale_lr": 0.032,
1323
+ "lr_decay_after": 25,
1324
+ "adaptive_schedule": {
1325
+ "best": 0.05275984301503296,
1326
+ "drop_after": 20,
1327
+ "relative_worsening": 0.001,
1328
+ "checks": 3,
1329
+ "patience_updates": 10,
1330
+ "bad_checks": 3,
1331
+ "last_best": 5,
1332
+ "early_drop": true
1333
+ },
1334
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1335
+ },
1336
+ {
1337
+ "layer": 61,
1338
+ "target": "same_input",
1339
+ "block_scale_dtype": "fp16",
1340
+ "training_sequence_count": 14657,
1341
+ "training_tokens_seen": 7864320,
1342
+ "training_passes": 0.5240171180844249,
1343
+ "best_step": 5,
1344
+ "codebook_lr": 0.048,
1345
+ "scale_lr": 0.032,
1346
+ "lr_decay_after": 25,
1347
+ "adaptive_schedule": {
1348
+ "best": 0.05175780703646056,
1349
+ "drop_after": 20,
1350
+ "relative_worsening": 0.001,
1351
+ "checks": 3,
1352
+ "patience_updates": 10,
1353
+ "bad_checks": 3,
1354
+ "last_best": 5,
1355
+ "early_drop": true
1356
+ },
1357
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1358
+ },
1359
+ {
1360
+ "layer": 62,
1361
+ "target": "same_input",
1362
+ "block_scale_dtype": "fp16",
1363
+ "training_sequence_count": 14657,
1364
+ "training_tokens_seen": 7864320,
1365
+ "training_passes": 0.5240171180844249,
1366
+ "best_step": 5,
1367
+ "codebook_lr": 0.048,
1368
+ "scale_lr": 0.032,
1369
+ "lr_decay_after": 25,
1370
+ "adaptive_schedule": {
1371
+ "best": 0.05643135330733986,
1372
+ "drop_after": 20,
1373
+ "relative_worsening": 0.001,
1374
+ "checks": 3,
1375
+ "patience_updates": 10,
1376
+ "bad_checks": 3,
1377
+ "last_best": 5,
1378
+ "early_drop": true
1379
+ },
1380
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1381
+ },
1382
+ {
1383
+ "layer": 63,
1384
+ "target": "same_input",
1385
+ "block_scale_dtype": "fp16",
1386
+ "training_sequence_count": 14657,
1387
+ "training_tokens_seen": 7864320,
1388
+ "training_passes": 0.5240171180844249,
1389
+ "best_step": 5,
1390
+ "codebook_lr": 0.048,
1391
+ "scale_lr": 0.032,
1392
+ "lr_decay_after": 25,
1393
+ "adaptive_schedule": {
1394
+ "best": 0.0563437113825566,
1395
+ "drop_after": 20,
1396
+ "relative_worsening": 0.001,
1397
+ "checks": 3,
1398
+ "patience_updates": 10,
1399
+ "bad_checks": 3,
1400
+ "last_best": 5,
1401
+ "early_drop": true
1402
+ },
1403
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1404
+ },
1405
+ {
1406
+ "layer": 64,
1407
+ "target": "same_input",
1408
+ "block_scale_dtype": "fp16",
1409
+ "training_sequence_count": 14657,
1410
+ "training_tokens_seen": 7864320,
1411
+ "training_passes": 0.5240171180844249,
1412
+ "best_step": 5,
1413
+ "codebook_lr": 0.048,
1414
+ "scale_lr": 0.032,
1415
+ "lr_decay_after": 25,
1416
+ "adaptive_schedule": {
1417
+ "best": 0.048833171244225114,
1418
+ "drop_after": 20,
1419
+ "relative_worsening": 0.001,
1420
+ "checks": 3,
1421
+ "patience_updates": 10,
1422
+ "bad_checks": 3,
1423
+ "last_best": 5,
1424
+ "early_drop": true
1425
+ },
1426
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1427
+ },
1428
+ {
1429
+ "layer": 65,
1430
+ "target": "same_input",
1431
+ "block_scale_dtype": "fp16",
1432
+ "training_sequence_count": 14657,
1433
+ "training_tokens_seen": 7864320,
1434
+ "training_passes": 0.5240171180844249,
1435
+ "best_step": 5,
1436
+ "codebook_lr": 0.048,
1437
+ "scale_lr": 0.032,
1438
+ "lr_decay_after": 25,
1439
+ "adaptive_schedule": {
1440
+ "best": 0.060005061741192876,
1441
+ "drop_after": 20,
1442
+ "relative_worsening": 0.001,
1443
+ "checks": 3,
1444
+ "patience_updates": 10,
1445
+ "bad_checks": 3,
1446
+ "last_best": 5,
1447
+ "early_drop": true
1448
+ },
1449
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1450
+ },
1451
+ {
1452
+ "layer": 66,
1453
+ "target": "same_input",
1454
+ "block_scale_dtype": "fp16",
1455
+ "training_sequence_count": 14657,
1456
+ "training_tokens_seen": 1310720,
1457
+ "training_passes": 0.08733618634740414,
1458
+ "best_step": 5,
1459
+ "codebook_lr": 0.048,
1460
+ "scale_lr": 0.032,
1461
+ "lr_decay_after": 25,
1462
+ "adaptive_schedule": {
1463
+ "best": 0.05761269836156414,
1464
+ "drop_after": 25,
1465
+ "relative_worsening": 0.001,
1466
+ "checks": 3,
1467
+ "patience_updates": 10,
1468
+ "bad_checks": 0,
1469
+ "last_best": 5,
1470
+ "early_drop": false
1471
+ },
1472
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1473
+ },
1474
+ {
1475
+ "layer": 67,
1476
+ "target": "same_input",
1477
+ "block_scale_dtype": "fp16",
1478
+ "training_sequence_count": 14657,
1479
+ "training_tokens_seen": 7864320,
1480
+ "training_passes": 0.5240171180844249,
1481
+ "best_step": 5,
1482
+ "codebook_lr": 0.048,
1483
+ "scale_lr": 0.032,
1484
+ "lr_decay_after": 25,
1485
+ "adaptive_schedule": {
1486
+ "best": 0.05480543365221943,
1487
+ "drop_after": 20,
1488
+ "relative_worsening": 0.001,
1489
+ "checks": 3,
1490
+ "patience_updates": 10,
1491
+ "bad_checks": 3,
1492
+ "last_best": 5,
1493
+ "early_drop": true
1494
+ },
1495
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1496
+ },
1497
+ {
1498
+ "layer": 68,
1499
+ "target": "same_input",
1500
+ "block_scale_dtype": "fp16",
1501
+ "training_sequence_count": 14657,
1502
+ "training_tokens_seen": 7864320,
1503
+ "training_passes": 0.5240171180844249,
1504
+ "best_step": 5,
1505
+ "codebook_lr": 0.048,
1506
+ "scale_lr": 0.032,
1507
+ "lr_decay_after": 25,
1508
+ "adaptive_schedule": {
1509
+ "best": 0.062059367848489574,
1510
+ "drop_after": 20,
1511
+ "relative_worsening": 0.001,
1512
+ "checks": 3,
1513
+ "patience_updates": 10,
1514
+ "bad_checks": 3,
1515
+ "last_best": 5,
1516
+ "early_drop": true
1517
+ },
1518
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1519
+ },
1520
+ {
1521
+ "layer": 69,
1522
+ "target": "same_input",
1523
+ "block_scale_dtype": "fp16",
1524
+ "training_sequence_count": 14657,
1525
+ "training_tokens_seen": 10485760,
1526
+ "training_passes": 0.6986894907792331,
1527
+ "best_step": 30,
1528
+ "codebook_lr": 0.048,
1529
+ "scale_lr": 0.032,
1530
+ "lr_decay_after": 25,
1531
+ "adaptive_schedule": {
1532
+ "best": 0.056006381621796365,
1533
+ "drop_after": 20,
1534
+ "relative_worsening": 0.001,
1535
+ "checks": 3,
1536
+ "patience_updates": 10,
1537
+ "bad_checks": 3,
1538
+ "last_best": 30,
1539
+ "early_drop": true
1540
+ },
1541
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1542
+ },
1543
+ {
1544
+ "layer": 70,
1545
+ "target": "same_input",
1546
+ "block_scale_dtype": "fp16",
1547
+ "training_sequence_count": 14657,
1548
+ "training_tokens_seen": 7864320,
1549
+ "training_passes": 0.5240171180844249,
1550
+ "best_step": 5,
1551
+ "codebook_lr": 0.048,
1552
+ "scale_lr": 0.032,
1553
+ "lr_decay_after": 25,
1554
+ "adaptive_schedule": {
1555
+ "best": 0.0733148360257097,
1556
+ "drop_after": 20,
1557
+ "relative_worsening": 0.001,
1558
+ "checks": 3,
1559
+ "patience_updates": 10,
1560
+ "bad_checks": 3,
1561
+ "last_best": 5,
1562
+ "early_drop": true
1563
+ },
1564
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1565
+ },
1566
+ {
1567
+ "layer": 71,
1568
+ "target": "same_input",
1569
+ "block_scale_dtype": "fp16",
1570
+ "training_sequence_count": 14657,
1571
+ "training_tokens_seen": 7864320,
1572
+ "training_passes": 0.5240171180844249,
1573
+ "best_step": 5,
1574
+ "codebook_lr": 0.048,
1575
+ "scale_lr": 0.032,
1576
+ "lr_decay_after": 25,
1577
+ "adaptive_schedule": {
1578
+ "best": 0.06478566802319799,
1579
+ "drop_after": 20,
1580
+ "relative_worsening": 0.001,
1581
+ "checks": 3,
1582
+ "patience_updates": 10,
1583
+ "bad_checks": 3,
1584
+ "last_best": 5,
1585
+ "early_drop": true
1586
+ },
1587
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1588
+ },
1589
+ {
1590
+ "layer": 72,
1591
+ "target": "same_input",
1592
+ "block_scale_dtype": "fp16",
1593
+ "training_sequence_count": 14657,
1594
+ "training_tokens_seen": 7864320,
1595
+ "training_passes": 0.5240171180844249,
1596
+ "best_step": 5,
1597
+ "codebook_lr": 0.048,
1598
+ "scale_lr": 0.032,
1599
+ "lr_decay_after": 25,
1600
+ "adaptive_schedule": {
1601
+ "best": 0.05970967496138404,
1602
+ "drop_after": 20,
1603
+ "relative_worsening": 0.001,
1604
+ "checks": 3,
1605
+ "patience_updates": 10,
1606
+ "bad_checks": 3,
1607
+ "last_best": 5,
1608
+ "early_drop": true
1609
+ },
1610
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1611
+ },
1612
+ {
1613
+ "layer": 73,
1614
+ "target": "same_input",
1615
+ "block_scale_dtype": "fp16",
1616
+ "training_sequence_count": 14657,
1617
+ "training_tokens_seen": 7864320,
1618
+ "training_passes": 0.5240171180844249,
1619
+ "best_step": 5,
1620
+ "codebook_lr": 0.048,
1621
+ "scale_lr": 0.032,
1622
+ "lr_decay_after": 25,
1623
+ "adaptive_schedule": {
1624
+ "best": 0.0664773336538204,
1625
+ "drop_after": 20,
1626
+ "relative_worsening": 0.001,
1627
+ "checks": 3,
1628
+ "patience_updates": 10,
1629
+ "bad_checks": 3,
1630
+ "last_best": 5,
1631
+ "early_drop": true
1632
+ },
1633
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1634
+ },
1635
+ {
1636
+ "layer": 74,
1637
+ "target": "same_input",
1638
+ "block_scale_dtype": "fp16",
1639
+ "training_sequence_count": 14657,
1640
+ "training_tokens_seen": 7864320,
1641
+ "training_passes": 0.5240171180844249,
1642
+ "best_step": 5,
1643
+ "codebook_lr": 0.048,
1644
+ "scale_lr": 0.032,
1645
+ "lr_decay_after": 25,
1646
+ "adaptive_schedule": {
1647
+ "best": 0.05450847628729126,
1648
+ "drop_after": 20,
1649
+ "relative_worsening": 0.001,
1650
+ "checks": 3,
1651
+ "patience_updates": 10,
1652
+ "bad_checks": 3,
1653
+ "last_best": 5,
1654
+ "early_drop": true
1655
+ },
1656
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1657
+ },
1658
+ {
1659
+ "layer": 75,
1660
+ "target": "same_input",
1661
+ "block_scale_dtype": "fp16",
1662
+ "training_sequence_count": 14657,
1663
+ "training_tokens_seen": 9175040,
1664
+ "training_passes": 0.611353304431829,
1665
+ "best_step": 25,
1666
+ "codebook_lr": 0.048,
1667
+ "scale_lr": 0.032,
1668
+ "lr_decay_after": 25,
1669
+ "adaptive_schedule": {
1670
+ "best": 0.04963358343974084,
1671
+ "drop_after": 20,
1672
+ "relative_worsening": 0.001,
1673
+ "checks": 3,
1674
+ "patience_updates": 10,
1675
+ "bad_checks": 3,
1676
+ "last_best": 25,
1677
+ "early_drop": true
1678
+ },
1679
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1680
+ },
1681
+ {
1682
+ "layer": 76,
1683
+ "target": "same_input",
1684
+ "block_scale_dtype": "fp16",
1685
+ "training_sequence_count": 14657,
1686
+ "training_tokens_seen": 7864320,
1687
+ "training_passes": 0.5240171180844249,
1688
+ "best_step": 5,
1689
+ "codebook_lr": 0.048,
1690
+ "scale_lr": 0.032,
1691
+ "lr_decay_after": 25,
1692
+ "adaptive_schedule": {
1693
+ "best": 0.05268687462948389,
1694
+ "drop_after": 20,
1695
+ "relative_worsening": 0.001,
1696
+ "checks": 3,
1697
+ "patience_updates": 10,
1698
+ "bad_checks": 3,
1699
+ "last_best": 5,
1700
+ "early_drop": true
1701
+ },
1702
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1703
+ },
1704
+ {
1705
+ "layer": 77,
1706
+ "target": "same_input",
1707
+ "block_scale_dtype": "fp16",
1708
+ "training_sequence_count": 14657,
1709
+ "training_tokens_seen": 14417920,
1710
+ "training_passes": 0.9606980498214457,
1711
+ "best_step": 45,
1712
+ "codebook_lr": 0.048,
1713
+ "scale_lr": 0.032,
1714
+ "lr_decay_after": 25,
1715
+ "adaptive_schedule": {
1716
+ "best": 0.03900097511967364,
1717
+ "drop_after": 25,
1718
+ "relative_worsening": 0.001,
1719
+ "checks": 3,
1720
+ "patience_updates": 10,
1721
+ "bad_checks": 0,
1722
+ "last_best": 45,
1723
+ "early_drop": false
1724
+ },
1725
+ "checkpoint_selection": "held-out original reference trajectory for both arms"
1726
+ }
1727
+ ]
reproduce/package_manifest.json ADDED
@@ -0,0 +1,247 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "files": {
3
+ "arvq88_reference.py": "b0606bf8ae91bab655ed5b26bb9b7437f0e28718301fdb1b86f475b1b31924e1",
4
+ "METHODOLOGY.md": "45ad59ea38585133a8005e169898a4dcd7ae56b9a4ce8ff3ba7e9f1ca00e9253",
5
+ "pv_progress.json": "583c9ebe0f26dfd4ffee434f1a7abc66af61a149b4f818ab5ccc4d5d78e7e3a9",
6
+ "README.md": "6d219086224a5c4fdc8085d04bae3292747673408d52275fa7ce17561c0a9c4b",
7
+ "calibration_corpus.json": "43517a1e4a3580b4da7c628019d0ad5d0a610aa469cfddcafc05f86bff47a770",
8
+ "reproduce/source_manifest.json": "eadfaac3a4552c341d64293b25f702b66358ba29450bbdba0e9b80f473ac9819",
9
+ "reproduce/environment_observed.json": "e9ae2da2456329999caf5ff468067e3c47caabd754bd80ae27b74f3c6c6f16e3",
10
+ "reproduce/requirements.in": "3f58e110a3d1816aef49af4fe3e44e2ff36e4bffd8c3187a099d556ef690d671",
11
+ "reproduce/data_artifacts.json": "bf1e037960df3435a829526d7145149036553f6e2ebe514d35c308a37e9a4126",
12
+ "reproduce/checkpoint_before_audit.json": "a8822199bb0af460349ba8c930e7abdbdf23096ba68276f429068c94b325fcce",
13
+ "reproduce/layer_recipe_summary.json": "1bdb40cd40eb3a80e9b0fc4e7b162be0143cd42cfbdcc592170e24e9f918a68e",
14
+ "reproduce/recipe.json": "4a6e876f0320909178001af53b086063a1951191e2ddfc26e7c696c76cdbc524",
15
+ "reproduce/README.md": "007935e4deb8d402f7d43bcd2dc7e2d2d27cf78dfde39fad35a58b4eae96126e",
16
+ "reproduce/RUNBOOK.md": "d13c303ef1c80eba3bc0aea0df43a263b042a5cd3fd76500e88432fe561b6407",
17
+ "reproduce/prepare_workspace.py": "cc679deae776de7beacece214c4868e51e35bc4a61a7a2725a4d05d1bb8324e5",
18
+ "reproduce/verify.py": "418df7c6dc54de146f0ff465618d42e69f8489eecc7743abc9464f8e15fa96e2",
19
+ "reproduce/test_roundtrip.py": "fdce4358d869c1fb5fe01a7a9a878d6a29710af65779bbf0637be52eefd84735",
20
+ "reproduce/external_sources.json": "42d948405368409e11c646a313f055e417271f4a828a907743d86a3094bba82a",
21
+ "reproduce/THIRD_PARTY.md": "b26dc2e38b00f1b7e17881535567ac8e811725abbdc709342f67f2eb08b93ac5",
22
+ "reproduce/tensor_header_sample.json": "7f804b22fa108706614d8fb7d12302eae7e4c537c746299c7ca00dc5670f4916",
23
+ "reproduce/AUDIT.md": "ae0e0e3447269352176307811ed4d07c4b9b91525f4d06724e9840cce3af1c8b",
24
+ "reproduce/historical_metadata/source.json": "c797ff4755b4ef9bf4b1fe8143ea7cfd55e6b52383fe11debaf4a23bf4cb1ffc",
25
+ "reproduce/historical_metadata/backbone_sources.json": "b9fc2ca4ed0da88b4061aff2d6cf53d59afae80a1d291a1ac8952f9a2063cd51",
26
+ "reproduce/historical_metadata/cold_assignment.json": "2bc853d8c282e2e7b7a3cb6fe390bda59d912d046063c6ee235cfa551014f695",
27
+ "reproduce/historical_metadata/calibration_corpus.json": "5c3e0f16c03632cf987d4e6cc4af731dd51abd7f12e18366285dbe0766150743",
28
+ "reproduce/historical_metadata/build_provenance.json": "d137642397e2c8cab6da61f1a7975453b5bf9d010ba1217a2d2b683d9cf75dbd",
29
+ "reproduce/historical_metadata/pv_progress.json": "ebec214f80ea0b16365f1aa819f6723a9461f5bf14a80f977f705025140ac6bc",
30
+ "reproduce/historical_metadata/METHODOLOGY.md": "43591534ca9a1e8db9015fd05368177a37efb1400b5d2ed699e2844e17485805",
31
+ "reproduce/source/tools/gguf_remote.py": "8befc8a77f2e7964c9bca0c5732a919b58dc5598fe33b00d4708bf0d8547232a",
32
+ "reproduce/source/tools/budget.py": "562dbf3a01822eb500f92d46f9710c987ba48bd2faf4fe1f0a12474ed1cf08cb",
33
+ "reproduce/source/tools/make_manifests.py": "b1c92a95a3e49bc5d27207e915a402da6a8713facbae1e09c00b5d1a932a4740",
34
+ "reproduce/source/tools/range_download.py": "6b8462c4af6372b5899bdae608537a76e5a470fc4938bc3d67ea5de7f16e9581",
35
+ "reproduce/source/tools/validate_serve.py": "e328e06a20246d58b0225ddb0d6cc00fa7450623c7ebb464da41cb58f29c8ea6",
36
+ "reproduce/source/tools/build_checkpoint.py": "af00a9e5c1a1b96c302bf7ea7eb7f5f9b5f32060de2c186085067b2bd5868ee5",
37
+ "reproduce/source/tools/patch_checkpoint.py": "cb5bdf47be506c0293b96055b83cf38f0ca807414aa0b9c640c16c2186246a3b",
38
+ "reproduce/source/tools/strip_stale.py": "3684269f45eeb2ea1b6698c80774a4a74cf616dea4832004be884ef1d5483b8e",
39
+ "reproduce/source/tools/collect_expert_stats.py": "979d22afbaf18e920a576857e6aedadbb74a120fdea593f2609a2e813342d16f",
40
+ "reproduce/source/tools/build_checkpoint_v3.py": "e0dc5a53b242a4a7d54ba7412798792bfc75f95ccd31632f7e589949a27500c2",
41
+ "reproduce/source/tools/solve_assignment.py": "f95d133a3dd96a797bec1a589e49e405056ffe9a525174994b95ec0a9059bba2",
42
+ "reproduce/source/tools/collect_expert_stats_v2.py": "7837bca536acd8513b2a7f6f127460b2e490bb615eabcc81fa06675d4915e1c4",
43
+ "reproduce/source/tools/patch_checkpoint_v3.py": "8e0479ec80071339dcd1ca06a9de4cbc12e661750142c037c24ee3f51df5a60e",
44
+ "reproduce/source/tools/solve_assignment_2tier.py": "c0de546bfe8ec4698eb8c0ff5c73a7f270513f217670ed44481cf7913cec3430",
45
+ "reproduce/source/tools/make_hot_manifest.py": "310b088822bbcf423bc4fc67ccd1033b0c8bfbb761f778f9b13f95ec303822d5",
46
+ "reproduce/source/tools/build_checkpoint_v4.py": "46624935e36542442b15b65d4f27d7cb4fb4ee40759ec59a9823b69785c7d4d6",
47
+ "reproduce/source/tools/build_checkpoint_v5.py": "b72c276e1977b67ab31f7e269dc1c82529d0431e45df1c1676cca378db19fd29",
48
+ "reproduce/source/tools/make_hot_manifest2.py": "a7b8ddc26edf59956413e2ba51297b694a9edc33797fd726596798d249637dd6",
49
+ "reproduce/source/tools/build_checkpoint_v6.py": "bd707467d8d8ff0738eddc4ecb324125b7b6c094fab6532addacca9e7cd26725",
50
+ "reproduce/source/tools/run_rtx6000.sh": "6ab7fdcffa0fd1f4768134d0f98d670259935fafcef9e478204579b6376ac478",
51
+ "reproduce/source/tools/capture_golden.py": "353997aca54a26b622d831dc6224199bf9c73bfcdb31be61a9abdb4bfc6bb100",
52
+ "reproduce/source/tools/make_kernel_vectors.py": "633dcde4ca640a19754baddaa98eea065ee8c45af80c68c075f4cb9ed542f395",
53
+ "reproduce/source/tools/verify_sm120.py": "51f4f52f28e6424bc094273055c6ad62d138d70780c7632edfba0415524d5433",
54
+ "reproduce/source/tools/capture_acts.py": "907206bce4a843da2c5d3d3c153dc09c0787a70bae62fd388664fd2d114da632",
55
+ "reproduce/source/tools/aqlm_converge.py": "9cb4c093a6c8dfd7639b46bd68f98c7b113d39666642a4c117748776cd625cd5",
56
+ "reproduce/source/tools/CONVERGER.md": "3394b8437e2eb6038dfd8d7159768634dd4c87137847293a38dfb2b252472894",
57
+ "reproduce/source/tools/score_experts_reap.py": "21e675cc5ad0b2727ccedd5695341541a98a8575fb4ce9797ae8110a9d175b23",
58
+ "reproduce/source/tools/build_checkpoint_v7.py": "54e4838a08d6b6c8b6e4e9e8cc83c7636b46a1f4511e85140c72dd4280c2d911",
59
+ "reproduce/source/tools/build_calib_v3.py": "38f51b1d4ada52631d35c6afab1983b3118e4dadcf1a67c5f7ff3ede30ac499c",
60
+ "reproduce/source/tools/phase15_fit.py": "b2c4fe789e9fc355abcbdecb9a89fe53922f58564cb53af381ff27fb1abb6346",
61
+ "reproduce/source/tools/build_checkpoint_v8.py": "c30f8265ea2c6a982dcb6e2b432b28f34c4d765e1e55ad06ad42a275992b11b7",
62
+ "reproduce/source/tools/capture_acts_v2.py": "15ee0237b58c0e1cc68dbfcf832dd7e741fdfb0270678ea30bd1819749d01b8c",
63
+ "reproduce/source/tools/aqlm_full.py": "37b3757632e3be9e9e88ebb7145769c96ca956835593e1c039a5b36d79b84c0d",
64
+ "reproduce/source/tools/make_bf16_manifest.py": "7716ffabdd110ad606fe69992d1bc9380d7f4713c1f2dc52c1163e28f1b7be61",
65
+ "reproduce/source/tools/phase21_pipeline.py": "45fa1b0a34863d5a0112f4d8ced733b4a284dc6a2fe3b4d74308cbb64411b8ad",
66
+ "reproduce/source/tools/build_heldout_xl.py": "a331c916257890166729bb6b4de470c93a735f06194fd740a58e23f25233bd3a",
67
+ "reproduce/source/tools/bf16_stream.py": "67b7cdcb13c621f7c53d60d1c4d2d3aef6f64044a7c5e29e19ae6b42955360c4",
68
+ "reproduce/source/tools/pv_tune.py": "878ee5749eca7cc88ca3591a0b5d1f5d9c2f33e1baf584f7a1ae9631517f5038",
69
+ "reproduce/source/tools/capture_acts_fresh.py": "e46b3148902f8cee8a82a145ccfeac9a3c21332773b9f8b0e37ffe96b328d9b7",
70
+ "reproduce/source/tools/build_checkpoint_v9.py": "3f4e177466d57661d3f4f5fa75af22a11d89b46401c4e53a0cea970a82d46131",
71
+ "reproduce/source/tools/convert_fp8_attn.py": "20ed2e8f8398b1098e54ce953b8234f1484a75e4bd5508bbe2bc4ddd98d4af77",
72
+ "reproduce/source/tools/fetch_nvfp4_experts.py": "79c8f59d25880eaf288b2403a95c67b95df48144b8ae5a283cb55ff0a65bebaa",
73
+ "reproduce/source/tools/build_checkpoint_v10.py": "d84b3a38b74a7f6bef2913a03e710eac13a6d434a1a2adc96863b06de45863f8",
74
+ "reproduce/source/tools/encode_onto_codebook.py": "f9e3343a6cd5d95da2a48f79431cedb70b7c5c311dfd1392e398f8c425425626",
75
+ "reproduce/source/tools/run_b200_tp8.sh": "a71e0d82b6a3e9cc485f87098c8cb8f9fb86e0755d0bf3845521452f781dffda",
76
+ "reproduce/source/tools/aqlm_quantize.py": "da0cea94862d7703c028a139bc15fbc41d719390c09633303a1b678e6fcd31cd",
77
+ "reproduce/source/tools/build_calib_v31.py": "b87f2d9415380e28cfba2a0ce75f876ee48e37bb6e865c4a02d43e2300698fa7",
78
+ "reproduce/source/tools/reasoning_slice.py": "6878c1f6162cb97a98ee11679f97e39f616c0e61818312670bd8a4ffb5639f23",
79
+ "reproduce/source/tools/build_calib_v32.py": "6b369e36e210351fbdead6754170a83c6410e9692292e83a29926e50d23d1b08",
80
+ "reproduce/source/tools/build_trace_refit_v32.py": "0d4bdfa5345c0051bf03a7bcf9798922cb4cccdf3e3a379526e541e2af4f3ab9",
81
+ "reproduce/source/btx53/pilot_expert.py": "fb09fd073d8168d3b563ed5affb184be22bf6954b15d721fb039a5dd1fb40399",
82
+ "reproduce/source/btx53/retry_prefetch.sh": "9c27f3c724b803b284bc79dcd8a5d2fba6d145c3e4f7be43aebbf9f09b981da9",
83
+ "reproduce/source/btx53/pipeline.py": "091bcbe282ac930930704865bb02b2738103c5f42abbb1d2d1e8aed6d2849f24",
84
+ "reproduce/source/btx53/build_baseline.py": "0fd3b364f5ed470354723342ae766e8695b1804d7a50495adfe68941cfaf9344",
85
+ "reproduce/source/btx53/build_v2_seed_checkpoint.sh": "d0b4445045ee7900f4f423a54e2c1adcf4634756ba056f73b227b7bd228f7a56",
86
+ "reproduce/source/btx53/arvq_gate_loop.sh": "c7c8c9527efbafd16d9e11a604e51e505fa8fbbff6079edb9bfbdcc2476837fd",
87
+ "reproduce/source/btx53/arvq_gate.sh": "530ce8c4f6755a164862cfd1fce5c839e57d523288401bed0ab2ee25b26e1f65",
88
+ "reproduce/source/btx53/arvq_watchdog.sh": "415cb80615b9d6d0e2918cbfad978c4f312d49b271f37db633eb7e51b978ad60",
89
+ "reproduce/source/tools/sanity/sc1_schema.py": "78bff17dea376fb982ce8e58e77ce34772b3891b35e0acde5c70ba9a9e3f5cc1",
90
+ "reproduce/source/tools/sanity/sc2_dequant_stats.py": "f8440c5a4982482a5d5efb10d0b9e390c30bcace65d809ae8da0e4be8c7abc78",
91
+ "reproduce/source/tools/sanity/sc6_ppl.py": "7cafb6d5c6d9222e7582241c1533bfc6389e55a35079e1c6fee1235d3c7882a1",
92
+ "reproduce/source/tools/sanity/bench_small.py": "75bb8bec5854d8415aaa7b806ada089af29be7cc36283c6fbbe14a7be65f7647",
93
+ "reproduce/source/tools/evalsets/build_heldout_53.py": "e913a8b1223044c86665769497bf551be8da10a88e1ae94f64f97c4970619883",
94
+ "reproduce/source/tools/evalsets/build_neutral_sets.py": "21ec003f850b4397345cbd30ef320feb8c7db7f85ecc73c9fc17f22c824c087e",
95
+ "reproduce/source/tools/vision/fetch_tensor.py": "40b10e1b5e37096b23663884f0f3d7c9dc7326c2b6a09918e9f53beaf879934f",
96
+ "reproduce/source/tools/vision/compare_embeds.py": "ba890a855820b3f9cbc1c4e042da183ac1c521b3705921f917439d06b7a708dd",
97
+ "reproduce/source/tools/vision/compare_tower.py": "2cf6db6c278107837ec16239a4c2dcbe3a93a7cab99d9c592bf599e9d4d685d5",
98
+ "reproduce/source/tools/capture53/launch_stats.sh": "d4f4453b0ffb48cb7fcbd441bd1cac1e6af132c982cdd4eeb276bd91c10b0d01",
99
+ "reproduce/source/tools/capture53/solve_assignment53.py": "5cbdb5a27c8b34b194b04083a120d6e0f9bde64171a19d2629ace3349c4e77e8",
100
+ "reproduce/source/tools/capture53/mtp_requant53.py": "ce982688427c49e3f04ddd1e18e89b6dfa5ee9dcbf454fb9c2e184ded72e6ee6",
101
+ "reproduce/source/tools/capture53/vision_probe53.py": "b6a772b02b808bf2f72920b5f2678cf09210e91d6210b9e7c501c2691c88c7cc",
102
+ "reproduce/source/tools/capture53/score_reap53.py": "c711b388826f3fac6a0684d16772e98c54bbc29b35876513ff98feef79a45e58",
103
+ "reproduce/source/tools/capture53/pv_tune53.py": "9f2c246411878d6af6902c4b7074fc997a70e3f9064e6a235796da2630d03fc7",
104
+ "reproduce/source/tools/capture53/eval_hybrid53.py": "dd291b3ff47935773ae02b6c22d574065db5af54bc779d3605e9757eb8f98047",
105
+ "reproduce/source/tools/capture53/solve_blend53.py": "bda421790c9e42292cc4d6c1a023d0d8b60a2777262f9165e54a378eeed86420",
106
+ "reproduce/source/tools/capture53/build_checkpoint53.py": "f4ef159add9890d34580e62f153b9e8db72e9f8411396477da249c45e4510d40",
107
+ "reproduce/source/tools/capture53/solve_reap53.py": "29d35e7abbe23985e1adbeeb439015884f8523ba36bd602cecb817a5c9fd0308",
108
+ "reproduce/source/tools/capture53/build_calib_mm.py": "cbc0364753cef10607327fa3126caa825207882a05d0315d23f52dcc9a03862c",
109
+ "reproduce/source/tools/capture53/graft_vision53.py": "04fbca9c6bec0c935dad247ac2f06ee3df583ba5053c46235b7ad7b94504b85d",
110
+ "reproduce/source/tools/capture53/aqlm_converge53.py": "b3ccd5521e0ad3a82d38e56a6ce070625c2a99528d38cd932ac57aadc64a1af0",
111
+ "reproduce/source/tools/capture53/solve_r3.py": "5979122eab06cbcc1568f0942896f8d056773378064bf24551dc3c1ccb9b6c09",
112
+ "reproduce/source/tools/capture53/capture_mm53.py": "9201ddd9cbfa89092016f4f51b7859a6cb20d25ab13c575f0bb270d7b786bc4c",
113
+ "reproduce/source/tools/capture53/stream_capture53.py": "9ca7120673abb2ac0b4fbeaf25811dc36ca25da9d7791cceb65ae53d935d8a08",
114
+ "reproduce/source/tools/capture53/prepare_unc_trace_model.py": "bb08ea3058856ec31b595f2cf4aa24eb541cbf1ee2fb7f77f2008cc1e2d954af",
115
+ "reproduce/source/tools/capture53/eval3_kld.py": "0c270ef4822f47cfbc92721456297e74a4ef115608a6d25d7eaba1d91e55cfe7",
116
+ "reproduce/source/btx53/arvq88/__init__.py": "0bf1139931af3c4357fa592ee2f2abc669bc2ce17826b4ab9cb4f593f773d018",
117
+ "reproduce/source/btx53/arvq88/inputs.py": "4a8ebeef4ad6a07e47b7d8e28474dd2ea48a192f755cd14abada993c052305a0",
118
+ "reproduce/source/btx53/arvq88/fit.py": "7288bbe580b307dbc399b161ee5899ea0cab9d57e57d40224ebce74152575497",
119
+ "reproduce/source/btx53/arvq88/checkpoint.py": "4d1194767fb46102085732eb114e524bf052653ad82f80367a655ba1155b7b90",
120
+ "reproduce/source/btx53/arvq88/cli.py": "452f32eb2eaa6fde3a8270c79599e185e36ab11e6993901cc7f94f56d5903840",
121
+ "reproduce/source/btx53/arvq88/gate_worker.py": "05199a15d687bf943ba0e8b9f8d16958a195a5874aff9d1b42ae00bbebda697f",
122
+ "reproduce/source/btx53/arvq88/reap_worker.py": "ac39da41fb132a026d666c892e300948b7606492610816068b0223848b14049d",
123
+ "reproduce/source/btx53/arvq88/reap.py": "9e69dbfcf48987b53c8694aa8c32f29e04bea270481a7c98a98b0a2e0a277704",
124
+ "reproduce/source/btx53/arvq88/pv_campaign.py": "95a483dedc27f61fb767cb8020471a04fd5cfa05078f816efbfc6fdb9cbd15cb",
125
+ "reproduce/source/btx53/arvq88/validation.py": "829c6a4ece05c189653a18dfe9429f3a48cb29c54e373fcd398706e3439a47ce",
126
+ "reproduce/source/btx53/arvq88/promote_initial.py": "27b7600139e9f5bf30ce0b3db0cb587284c16a79e77baa1cd3049a03b0e04dc0",
127
+ "reproduce/source/btx53/arvq88/local_source.py": "0c57365308afcf75bf4b485a1e86fd1ca169ce7c4ddb999f80f2c3f938bb1bd9",
128
+ "reproduce/source/btx53/arvq88/backbone.py": "88e88930ea9ecec52259f5ab648915dc9b3a4d8144b302dadfbeb736c9dcc45b",
129
+ "reproduce/source/btx53/arvq88/activation.py": "cf13e2170085459b97907a270abb32d694af3ce69f723b332119a62be00e6641",
130
+ "reproduce/source/btx53/arvq88/alternating.py": "19e38a8230dba40bc022180ed4234b5a60f53765d36a417677e35711e9f0c80c",
131
+ "reproduce/source/btx53/arvq88/arvq_reap.py": "c39faf937462d85b2ede04f19f14c2cc8ef875c6788a1dab2445a03834882cfd",
132
+ "reproduce/source/btx53/arvq88/jobs.py": "4b9021e5763ba15b45cf4c0cf76d6f62f705ac175b076cdae380032a2d9ac59d",
133
+ "reproduce/source/btx53/arvq88/experiment.py": "65272954df9ef44f88467c05f5ce5f12c8045b38765f67b773353890d7877de4",
134
+ "reproduce/source/btx53/arvq88/incremental_publish.py": "70526567046728bd2905d059bf475f6a1d4d74d4242f7288cd64091e9137f910",
135
+ "reproduce/source/btx53/arvq88/gradient_indices.py": "09bd66d77575deec66d1732677355c406cb9b228abfc5002e79cbae5adb72e83",
136
+ "reproduce/source/btx53/arvq88/sequential_pv.py": "807609be6d53c94ed578ae91f81e74d3be60451d436d6d5f1b5053004ea821f0",
137
+ "reproduce/source/btx53/arvq88/sequential_pilot.py": "030f3bd1d83ea2f9bcd97ee8514d01dc69d269e7d7b7f87159da667b46b76710",
138
+ "reproduce/source/btx53/arvq88/encoder.py": "6c20c56fc7a308724ab4faf4174c4e8e32cba5a43c3fccd3f3e6d4da1f0b583e",
139
+ "reproduce/source/btx53/arvq88/sequential_capture.py": "bcbb6cb450243411e064624c70b1ec460e052680eb7f9958f418b25577d9e671",
140
+ "reproduce/source/btx53/arvq88/final_publish.py": "e7dcb2859de2e08a5c2bea7b21b9aae61e96bb2361b28ba4270cc145fb0da188",
141
+ "reproduce/source/btx53/arvq88/pv.py": "82f3f559e44b5db2480e041fa45d983bf1e6afc2d74a68c0511af0d174f9ed68",
142
+ "reproduce/source/btx53/arvq88/pack.py": "b0606bf8ae91bab655ed5b26bb9b7437f0e28718301fdb1b86f475b1b31924e1",
143
+ "reproduce/source/btx53/arvqprep/__init__.py": "921c16de3c268ed12dcac8329852f47fb6875d926a2ea665a03e69c81059b23d",
144
+ "reproduce/source/btx53/arvqprep/pack.py": "ce459847f07a558377872e0f1f83620d9f26d9eec7ec94d08247052876bab656",
145
+ "reproduce/source/btx53/arvqprep/encoder.py": "ace5d22c886994a0327e8320c6674f71c1caab779d215e3d05295d1f490bfec1",
146
+ "reproduce/source/btx53/arvqprep/smoke.py": "fb445c69e081b1c23acaf03feef234e215c61cceaf1b6d0b4c47aa63f9032a34",
147
+ "reproduce/source/btx53/arvqprep/container.py": "3f5ab2a065a1a44826cd0183ac39ba1d016ac87be655ed3b9502b1c3137ae7ad",
148
+ "reproduce/source/btx53/arvqprep/workflow.py": "778d6782e95cf81ea7d7c44f8f508aab65706e64a18cc3d15adc57eb61e5e095",
149
+ "reproduce/source/btx53/arvqprep/publish.py": "3249526e87beba17e69612d034cc400c57cf88a09eea3b3622eb8e4f43b868ce",
150
+ "reproduce/source/btx53/arvqprep/pv_workflow.py": "5a57e8ae4e76f938f6cda3b245321c4812e5e7c16e52b3976adcfa96b3cb9187",
151
+ "reproduce/source/btx53/btxprep/core.py": "b4fad5c192efa2e9e0088427ad8e90dacac59b9220479809b81ca826556c8348",
152
+ "reproduce/source/btx53/btxprep/transforms.py": "2f0bda3c93e39836ae6bca82deceb876e87b54797dd696c194a99041f6bd6d1c",
153
+ "reproduce/source/btx53/btxprep/container.py": "f260e70e4e6fda4c20b8cb85469620ebe67d343983776c370b0bfe7bdb2f967e",
154
+ "reproduce/source/btx53/btxprep/__init__.py": "f8b773e399d3d5b859d05a85347d8f6a91eed2808447c934d917c1de39d43a40",
155
+ "reproduce/source/btx53/btxprep/encoder.py": "ac8745d1ba1590817a16bc4ed48acc9af980a7103ad22e32aeb179b48b5813f2",
156
+ "reproduce/source/btx53/tools/remote_st.py": "4a27ac143fff1cc8cdf6fb349ac58bba1d22933c77c562779bb111b727f0031e",
157
+ "reproduce/source/btx53/tools/ingest.py": "e05630716d805903f1bfe658512da2d3f63abbd280a3af87ca9baadb5becd52c",
158
+ "reproduce/source/btx53/tools/upload_btx.py": "72bc32e8f2c2755e9ec5c713d6226bd7021265d1cddaf1e927b3682b508e5cba",
159
+ "reproduce/source/btx53/tools/filter_shards.py": "749d7b6f8a5b6ba24442508ce66e4108086b837e89c20234c3bf4939b27ada7d",
160
+ "reproduce/source/btx53/tools/fetch_hyb_kinds.py": "aaa37fd535d26b3772c4e467c55502c810ba39b73ca998078c487aeb1a7be846",
161
+ "reproduce/source/btx53/tools/gates_btx.py": "bcd2d8b5e4ef9e3779bfbec625d7af5b0727d954502db63a46147de882de170c",
162
+ "reproduce/source/btx53/tools/upload_btx_full.py": "ee8b46dd631629e15d1f1bb86c0b06d4ad22e573c674e129e58c04b5e7e1dc8d",
163
+ "reproduce/source/btx53/tools/launch_btx_workers.sh": "4e8484172fe4dea0b48390c72959d81bc4d25ac05265187b9bb90df9c5bc592a",
164
+ "reproduce/source/btx53/tools/upload_hybrid.py": "5b3e4be4ccf076d569ea5a48f0e79d6168236d4d6eccf84053a4bf183a4cc2f6",
165
+ "reproduce/source/btx53/tools/gates_arvq.py": "ce6b09045cc8f1c9cea7647083e133a52885b284c09879ccf4b9090d58d27453",
166
+ "reproduce/source/btx53/tools/launch_arvq_workers.sh": "78b013b3c8f5e2750fc6a3b7e88d58ddc5a61aafdea0d51eb666d0d3dba2000c",
167
+ "reproduce/source/btx53/tools/arvq_driver.sh": "13d72054ed28f5e112b98ddd7db0cdb7de00e0a4ead5cbd54b03bc0386860cc8",
168
+ "reproduce/source/btx53/tools/pv_arvq.py": "82dd57c204c057eaacd17aefcd37faa1d6b90ab7f6931aa2872b3acf65822858",
169
+ "reproduce/source/btx53/tools/prefetch_cold_fp8.py": "5825b1069ee10681516d69a703699abae4ac46011f6e71d6537acfece4345921",
170
+ "reproduce/source/btx53/arvq88/tests/test_workflow.py": "d9bc78d2a898a9c8dab96ea2a79dcf183cf182fb535992ff106ff452c8961979",
171
+ "reproduce/source/btx53/arvq88/tests/test_final_publish.py": "6c2b104ec40284415763fea5005510df0b63055bf99f081345e8248de0f7eb24",
172
+ "reproduce/source/btx53/arvq88/tests/test_validation.py": "f563a433c66412085a191deb8c693865632aedf1de19948567c710d11c3606b7",
173
+ "reproduce/source/btx53/arvq88/tests/test_local_source.py": "544513082fe0c666ade7ed00a837aeeeb915c52fb1325362a488e1c8c3a050ed",
174
+ "reproduce/source/btx53/arvq88/tests/test_donor.py": "1948eca1225198bed8964ff9d56e5f2b98c86847537a108670cf59260a20d76e",
175
+ "reproduce/source/btx53/arvq88/tests/test_backbone.py": "58f3875440e45b9bfd5be915452a7f1a18cd36688909f737d5830e1a5062fa85",
176
+ "reproduce/source/btx53/arvq88/tests/test_activation.py": "eff65a3e37b8b4b77b8fbb8bd979f6ca5012680a9643716f257f4e8985662944",
177
+ "reproduce/source/btx53/arvq88/tests/test_expert_books.py": "5942045e2e09a0a50ea1bae852a2f9a411b8dfd949e78e8e67151b8e119ec33f",
178
+ "reproduce/source/btx53/arvq88/tests/test_incremental_publish.py": "d0296210c666b4c556dcc6139500a1a0a36646186ee7a019d0ca2d41b9a80368",
179
+ "reproduce/source/btx53/arvq88/tests/test_gradient_indices.py": "6f1527d21b73c882500bb8ff229c54a023abe89d43ad23c6c3176122c188347e",
180
+ "reproduce/source/btx53/arvq88/tests/test_sequential_pv.py": "7d3c65cf41648b1b564a1939bcbee3d495ec4bd9d0af7028a5d788b052fadd18",
181
+ "reproduce/source/btx53/arvq88/docs/per_expert_handoff.md": "554425132e253e53a7890ba6f270c2cbfcec4e7115921db2983036c5b6d20664",
182
+ "reproduce/source/btx53/arvq88/docs/expert_experiments.md": "6d78a3ee627255f13e2be7995bf21df4c36f9703b192bc5020a5c1c7a7a7336e",
183
+ "reproduce/source/btx53/arvq88/docs/gradient_pv_recipe.md": "77686dc9d94c3a4c5559024598dbbcafe7ca55f17858e7305fb2ad069a4f274f",
184
+ "reproduce/source/btx53/arvq88/docs/sequential_pilot.md": "079342a060d5fc384e00b724ca3e1189c169ef4b7c590df12b1739c1dceba6d5",
185
+ "reproduce/source/btx53/arvq88/perf/graph_expert.py": "9d663b49ef7bb23db843f78068f146e2c178f23137360acfaa1aa544c6942d73",
186
+ "reproduce/source/btx53/arvq88/perf/bench_graph_expert.py": "9b6b5759376fb771caed63e908e256d9b187fec93e0d6bcb6a16a16c97cdb42e",
187
+ "reproduce/source/btx53/arvq88/perf/sparse_delta.py": "52565e3bdaa477d1676fffaf10bdffd10e1ad2d1b4275ded2cdd1d210bd0f81e",
188
+ "reproduce/source/btx53/arvq88/perf/test_perf.py": "3f9221722a3d0aa3360ba02ca5823612162aefb27d084f0d4a3bacf4bac45ef9",
189
+ "reproduce/source/btx53/arvq88/perf/README.md": "ebeeb89781cd162e1711837bab91a9077df496a876bc99ab6d6a0fe53779c21b",
190
+ "reproduce/source/btx53/arvq88/perf/rerun_pilot.py": "691665c2f6d482d8518f7153124814f8333a810d44c8c03269756ef2bbf8199f",
191
+ "reproduce/source/btx53/arvq88/perf/full_reference.py": "cabbd41c1be1f2aa77e3b298b5958a79aed88cc73263ec358ef2c64a9d4846b5",
192
+ "reproduce/source/btx53/arvq88/perf/sequential_pv_repair.py": "74ece4df20d6573ae8bf752736adb12217968098c49e6f5d3bfbb7ebc1d3b3f3",
193
+ "reproduce/source/btx53/arvq88/perf/test_repair.py": "78112f17972eadf45c7f48f578b8c75170c70e7fe0aa75f63bf861c6161f952c",
194
+ "reproduce/source/btx53/arvq88/perf/sequential_pv_cached.py": "9c629c8083fe5cc2e760e1d6afc3dc1f60c1320ac30b4dcd95db35dc93d8b163",
195
+ "reproduce/source/btx53/arvq88/perf/batch_benchmark.py": "c78c5608221c051b93817b373c94c7ca844416851f72d3905f4e231272776829",
196
+ "reproduce/source/btx53/arvq88/perf/capture_full_layer3.py": "3359edb2956d41de364c32e209275c1d87d047bbfe48ee84f29d76d9ee535a1b",
197
+ "reproduce/source/btx53/arvq88/perf/test_full_corpus.py": "98491ab3943d70b4ed2536ecbd72bd163561002d6fba61124f829adcbcf5bf1c",
198
+ "reproduce/source/btx53/arvq88/perf/async_shards.py": "b536e99bf5ea3badd4291f3fbc9b22a0594d334e9b551f3146a71288991d7529",
199
+ "reproduce/source/btx53/arvq88/perf/test_async_shards.py": "ecee760fa718fa0542a8650bf8a5b7d1b7ccade05b6682f0660401cac4208f34",
200
+ "reproduce/source/btx53/arvq88/perf/FULL_CORPUS_PIPELINE.md": "2c842e5833b64c1de55b9b3d33756ddb5cc6dfaef808ccd652cbc37f07280040",
201
+ "reproduce/source/btx53/arvq88/perf/distributed_batch.py": "ff4fcfdc9c74bc46f06bb51a126d2764aa4d983da5b9eebc06885081330fa39e",
202
+ "reproduce/source/btx53/arvq88/perf/test_distributed_batch.py": "36bc25e59d2ec9045751409ccb74a2f0464a2de502fa5adb4affff5fec3a56cc",
203
+ "reproduce/source/btx53/arvq88/perf/check_recapture.py": "9b49570cf9dca3d42ee9262822b7992b2d1ddf6025838f4fbe3d9ff5ab090593",
204
+ "reproduce/source/btx53/arvq88/perf/publish_full_corpus.py": "a361a0fe959a214dc8f7550c90248697b46f1d1fa3aeba1f149503723617b0b7",
205
+ "reproduce/source/btx53/arvq88/perf/capture_transfer.py": "aa6f8702cbaaf3ebcc9dad65053c422f3ec7f1e1d9c4ac3fc60834769c38b386",
206
+ "reproduce/source/btx53/arvq88/perf/test_capture_transfer.py": "c33c0ffda4ca906930760cd100d7e0d283d32eeba1fbebd47af34bffa680abfa",
207
+ "reproduce/source/btx53/arvq88/perf/monitor_output_error.py": "cce2dd878e8e164e8b3a296cc70699cc8eb6e3225bac3f9b8c5171e063af40cb",
208
+ "reproduce/source/btx53/arvq88/perf/capture_matched_validation.py": "5d0b9d1892923acbd066f0f88030221db39aecee6ca62e8a1f094a29833a473d",
209
+ "reproduce/source/btx53/arvq88/perf/adaptive_schedule.py": "504f54400931422d706544029ca1c3673224f33f77d7801c89dda1d4ba53dd30",
210
+ "reproduce/source/btx53/arvq88/perf/test_adaptive_schedule.py": "456bc945a66c6ff1893f86a972a63b2efb44ae3a386e09e02df23673780cc537",
211
+ "reproduce/source/btx53/arvq88/perf/campaign_watchdog.py": "548d78758937c79d17601772f2c5424a8bbf5d9952144b6fbbde396a8de4e499",
212
+ "reproduce/source/btx53/arvq88/perf/row_scale_pilot.py": "a84aec6f9aa2824aaa9f00e8d53c19b5f5deca76ddd4c0da69230bebe95c0181",
213
+ "reproduce/source/btx53/arvq88/perf/four_layer_row_pilot.py": "de817289405b9f2a4eb34e28b315e779dcadadfafc401c1e28d54c75b5a0acf7",
214
+ "reproduce/source/btx53/arvq88/perf/row_scale_combined_eval.py": "6e2084b496058983fc4b143142ec3ebcb9703bed5801bb7bb0a599017976e17f",
215
+ "reproduce/source/btx53/arvq88/perf/FP16_BLOCK_SCALE_HANDOFF.md": "9304e796d7fbff92157737b9aa6cf2a6e29e66d16656409af4ab06c5f18e8cfa",
216
+ "reproduce/source/btx53/arvq88/perf/sequential_pv_same_input.py": "c6fb2c7a2f565c45f326837da2546ebd3fd01d2496353539d3a3303c9c370f13",
217
+ "reproduce/source/btx53/arvq88/perf/full_corpus.py": "db9e8cb6b3d7ea22fc8490fc355fd9421dfc7429c600431cc37f52017f9811c2",
218
+ "reproduce/source/btx53/arvq88/perf/full_pipeline.py": "c35476f87e9c2231cc1acc0059e17f91d0c51f04a99d464423054d54b4d793a4",
219
+ "reproduce/source/btx53/arvq88/perf/fused_handoff.py": "6c62bf5606ed48fa8a1a9d27a5f694f353f44c7a02f651475522f2e431a5c37a",
220
+ "reproduce/source/btx53/arvq88/perf/fused_evaluation_capture.py": "10ef6f00129bcea2915b41febdb08ba2f942c4f132e61ca4a774d8f913bf8daa",
221
+ "reproduce/source/btx53/arvq88/perf/capture_full_sequential.py": "ee221434a21fb4818f9ab20925ea66c31a396c8b3557dda1a45174fabd284927",
222
+ "reproduce/source/btx53/arvq88/perf/publish_v2_streaming.py": "07ea5e0fb78b462590f422895e05c823208da79d929980e3a6b0dde071f69395",
223
+ "reproduce/source/btx53/arvq88/perf/sequential_pv_fast.py": "a9b04ba27ebfd00f10bfbf96791c816858cef27d560086610e3bf6906fde9356",
224
+ "reproduce/source/btx53/arvq88/perf/propagate_full.py": "efb52d2784d1fca48beb3404c462255756c3fe499f68a418baa8a790d2018bf9",
225
+ "reproduce/source/btx53/arvq88/perf/baseline_propagation.py": "a0169695c1993f9c57da6c078e6e22aaefd404e1108a39a1989ceb1142d1cf35",
226
+ "reproduce/source/btx53/arvq88/perf/sequential_pv_full_corpus.py": "297df5e8e276e2caaae2af049b4dc0bf2e6185e0a9844bfdf705272aa4f96c3c",
227
+ "reproduce/source/btx53/arvq88/perf/publish_same_input.py": "bc75d242557ad1cac3b2cdc316b5760a501da5b221f7ed1f439a39d09b3de978",
228
+ "reproduce/source/btx53/arvq88/perf/publish_reference.py": "204d86bd3393bd91b5b43a5f7eb8f5c40d2cb3811962b772ab5cfffca1f63bab",
229
+ "reproduce/source/btx53/arvq88/perf/sequential_pv_mcbook.py": "f60b89d8325f4afd07349980c9190f78caf7ca0c05494805b04a1032de8fc749",
230
+ "reproduce/source/btx53/arvqprep/tests/test_workflow.py": "9250f47f793f13098ea48ce9ec51baec71ff65c410b3beac4609f32f8decefeb",
231
+ "reproduce/source/btx53/btxprep/tests/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
232
+ "reproduce/source/btx53/btxprep/tests/test_roundtrip.py": "fdce4358d869c1fb5fe01a7a9a878d6a29710af65779bbf0637be52eefd84735"
233
+ },
234
+ "weights_changed": false,
235
+ "validated": "2026-09-26",
236
+ "checks": [
237
+ "source SHA256",
238
+ "all Python parses",
239
+ "CPU FP8/FP16 decode fixtures",
240
+ "50x next-token boundary semantics",
241
+ "relocated workspace smoke"
242
+ ],
243
+ "not_validated": [
244
+ "full historical retraining",
245
+ "native SM120 execution"
246
+ ]
247
+ }
reproduce/prepare_workspace.py ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Materialize a fresh portable source copy; no fitting or publication is run."""
2
+ import argparse,json,shutil,sys,re,hashlib
3
+ from pathlib import Path
4
+ p=argparse.ArgumentParser();p.add_argument('--dest',type=Path,required=True);a=p.parse_args();dest=a.dest.resolve();root=Path(__file__).resolve().parent
5
+ if dest.exists():raise SystemExit('Destination must not exist; refusing to modify an existing workspace')
6
+ shutil.copytree(root/'source',dest)
7
+ mapping={'/home/coder/git/glm52':str(dest),'/data':str(dest/'data'),'/tmp':str(dest/'tmp')}
8
+ pattern="(?:"+"|".join(re.escape(k) for k in mapping)+")(?=/|$|[\\s\"' ])"
9
+ changed={}
10
+ for f in dest.rglob('*'):
11
+ if not f.is_file() or f.suffix not in ('.py','.sh','.md'):continue
12
+ s=f.read_text();t=s
13
+ # Interpreter locations must resolve to the activated environment.
14
+ t=re.sub(r'/tmp/(?:venv[^/]*|[^/]+/venv)/bin/python(?:3)?','__GLM53_PYTHON__',t)
15
+ t=re.sub(pattern,lambda m:mapping[m.group()],t).replace('__GLM53_PYTHON__',sys.executable)
16
+ if t!=s:f.write_text(t);changed[str(f.relative_to(dest))]=hashlib.sha256(f.read_bytes()).hexdigest()
17
+ (dest/'data').mkdir(exist_ok=True);(dest/'tmp').mkdir(exist_ok=True);(dest/'.venv/bin').mkdir(parents=True,exist_ok=True);(dest/'.venv/bin/python').symlink_to(sys.executable)
18
+ cfg=json.loads((root/'recipe.json').read_text())
19
+ def relocate(v):
20
+ if isinstance(v,str):
21
+ return re.sub(pattern,lambda m:mapping[m.group()],v)
22
+ if isinstance(v,list):return [relocate(x) for x in v]
23
+ if isinstance(v,dict):return {k:relocate(x) for k,x in v.items()}
24
+ return v
25
+ (dest/'recipe.json').write_text(json.dumps(relocate(cfg),indent=2));(dest/'relocation.json').write_text(json.dumps({'mapping':mapping,'python':sys.executable,'changed_sha256':changed},indent=2))
26
+ print(f'Prepared {dest}. Populate donor/data/cache paths, inspect recipe.json, then follow RUNBOOK.md. No job launched.')
reproduce/recipe.json ADDED
@@ -0,0 +1,116 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "baseline": "/tmp/glm53-vision-expert-full",
3
+ "corpus": "/tmp/glm53-vision-trace-refit",
4
+ "source_cache": "/tmp/glm53-fp8-cold",
5
+ "sequence_length": 1024,
6
+ "sequences": [
7
+ 64,
8
+ 16,
9
+ 16
10
+ ],
11
+ "world_size": 8,
12
+ "record_initial_audit": true,
13
+ "train_token_file": "/tmp/glm53-layer3-full-corpus/train.npy",
14
+ "batch_tokens": 262144,
15
+ "microbatch_tokens": 65536,
16
+ "codebook_lr": 0.048,
17
+ "scale_lr": 0.032,
18
+ "reference_only": false,
19
+ "target": "same_input",
20
+ "block_scale_dtype": "fp16",
21
+ "codebook_dtype": "fp4_grid",
22
+ "boundary_boost": 50.0,
23
+ "boundary_token_ids": [
24
+ 154842,
25
+ 154820
26
+ ],
27
+ "reassign_every": 10,
28
+ "reassign_max_fraction": 0.005,
29
+ "reassign_trust_ratio": 0.02,
30
+ "reassign_target_ratio": 0.1,
31
+ "lr_decay_after": 25,
32
+ "lr_decay_factor": 0.25,
33
+ "adaptive_schedule": {
34
+ "relative_worsening": 0.001,
35
+ "checks": 3,
36
+ "patience_updates": 10
37
+ },
38
+ "fuse_next_capture": true,
39
+ "layers": [
40
+ 3,
41
+ 4,
42
+ 5,
43
+ 6,
44
+ 7,
45
+ 8,
46
+ 9,
47
+ 10,
48
+ 11,
49
+ 12,
50
+ 13,
51
+ 14,
52
+ 15,
53
+ 16,
54
+ 17,
55
+ 18,
56
+ 19,
57
+ 20,
58
+ 21,
59
+ 22,
60
+ 23,
61
+ 24,
62
+ 25,
63
+ 26,
64
+ 27,
65
+ 28,
66
+ 29,
67
+ 30,
68
+ 31,
69
+ 32,
70
+ 33,
71
+ 34,
72
+ 35,
73
+ 36,
74
+ 37,
75
+ 38,
76
+ 39,
77
+ 40,
78
+ 41,
79
+ 42,
80
+ 43,
81
+ 44,
82
+ 45,
83
+ 46,
84
+ 47,
85
+ 48,
86
+ 49,
87
+ 50,
88
+ 51,
89
+ 52,
90
+ 53,
91
+ 54,
92
+ 55,
93
+ 56,
94
+ 57,
95
+ 58,
96
+ 59,
97
+ 60,
98
+ 61,
99
+ 62,
100
+ 63,
101
+ 64,
102
+ 65,
103
+ 66,
104
+ 67,
105
+ 68,
106
+ 69,
107
+ 70,
108
+ 71,
109
+ 72,
110
+ 73,
111
+ 74,
112
+ 75,
113
+ 76,
114
+ 77
115
+ ]
116
+ }
reproduce/requirements.in ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ torch
2
+ numpy
3
+ transformers
4
+ datasets
5
+ safetensors
6
+ huggingface_hub
7
+ accelerate
8
+ Pillow
9
+ scipy
10
+ triton
reproduce/source/btx53/arvq88/__init__.py ADDED
@@ -0,0 +1 @@
 
 
1
+ """GLM-5.3 256+256 FP4 additive-vector quantization workflow."""
reproduce/source/btx53/arvq88/activation.py ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Serving activation-plane reference, pinned to b1380cf74d170b69a8708f1b1287cc09a6eb7df4.
2
+
3
+ Matches hybrid.cu pack_planes and the default FP16 SwiGLU path. This emulates
4
+ quantization boundaries, not SM120 MMA instruction accumulation bit-for-bit.
5
+ """
6
+ import torch
7
+ REFERENCE_REVISION='b1380cf74d170b69a8708f1b1287cc09a6eb7df4'
8
+ ARITHMETIC='fp4_planes4_fp16_boundaries_v1'
9
+
10
+ def fp16_ste(x):
11
+ rounded=x.to(torch.float16).float()
12
+ return x+(rounded-x).detach() if x.requires_grad else rounded
13
+
14
+ def planes(x):
15
+ """Return FP4 codes, FP8 scale bytes, and decoded *unweighted* planes."""
16
+ if x.shape[-1]%16:raise ValueError('Activation input dimension must be divisible by 16')
17
+ v=x.to(torch.float16).float().reshape(*x.shape[:-1],-1,16)
18
+ if not torch.isfinite(v).all():raise ValueError('Nonfinite FP16 serving activation')
19
+ cuts=torch.tensor([.25,.75,1.25,1.75,2.5,3.5,5.],device=x.device)
20
+ table=torch.tensor([0.,.5,1.,1.5,2.,3.,4.,6.],device=x.device)
21
+ codes=[];scales=[];decoded=[]
22
+ for _ in range(4):
23
+ peak=v.abs().amax(-1,keepdim=True)
24
+ exponent=torch.ceil(torch.log2((peak/6).clamp_min(2**-20))).clamp(-6,8)
25
+ scale=torch.exp2(exponent)
26
+ # CUDA compares strictly > each midpoint: ties go toward zero, not ties-to-even.
27
+ q=(v.abs().div(scale).unsqueeze(-1)>cuts).sum(-1)
28
+ dec=table[q]*scale*torch.where(v<0,-1.,1.)
29
+ codes.append((q|((v<0).long()<<3)).to(torch.uint8).reshape(x.shape))
30
+ scales.append(((exponent.squeeze(-1).long()+7)<<3).to(torch.uint8))
31
+ decoded.append(dec.reshape(x.shape))
32
+ v=(v-dec)*16
33
+ return torch.stack(codes),torch.stack(scales),torch.stack(decoded)
34
+
35
+ def activation_ste(x):
36
+ with torch.no_grad():
37
+ _,_,p=planes(x)
38
+ result=p[0]+p[1]/16+p[2]/256+p[3]/4096
39
+ return x+(result-x).detach() if x.requires_grad else result
40
+
41
+ def swiglu_ste(gu):
42
+ # Serving: FP32 projection -> FP16; SiLU FP16 result; FP16 product.
43
+ gu=fp16_ste(gu);gate,up=gu.chunk(2,-1)
44
+ return fp16_ste(fp16_ste(torch.nn.functional.silu(gate))*up)
reproduce/source/btx53/arvq88/alternating.py ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Discrete LDLQ proposals interleaved with continuous output matching.
2
+
3
+ This is a bounded alternating optimizer, not a reproduction of the PV-Tuning
4
+ paper. Proposals use training activation Hessians and are retained only if the
5
+ training routed-output objective decreases. Validation still selects checkpoints.
6
+ """
7
+ import math
8
+ from pathlib import Path
9
+ import torch
10
+ from .encoder import cholesky_inv_upper, hessian_fc1, sweep_expert
11
+ from .activation import ARITHMETIC, activation_ste, swiglu_ste
12
+
13
+
14
+ def accept_candidate(before, after):
15
+ return math.isfinite(after) and after < before - 1e-8
16
+
17
+
18
+ @torch.no_grad()
19
+ def reassign_indices(p13, p2, x, ids, train, cold, args, meta, evaluate_train):
20
+ from ingest import load_expert_bf16
21
+ before = evaluate_train()
22
+ # Codes must participate in both validation-best snapshots and rollback.
23
+ old = [(p.codes_a.clone(), p.codes_b.clone()) for p in (p13, p2)]
24
+ changed = total = 0
25
+ try:
26
+ for e, eid in enumerate(cold):
27
+ rows = train[(ids[train] == eid).any(1)]
28
+ if len(rows) < 32:
29
+ continue # No held-out-row fallback for rare experts.
30
+ z = x[rows[:args.reassign_hessian_tokens]]
31
+ wg, wu, wd = [w.to(x.device).float() for w in load_expert_bf16(
32
+ str(Path(args.src_cache)/f'layer_{args.layer}'/f'expert_{eid}.safetensors'), meta)]
33
+ z = activation_ste(z) if args.arithmetic == ARITHMETIC else z
34
+ # Down sees the current quantized gate/up, not teacher intermediates.
35
+ gu = z @ p13.weight(e).T
36
+ if args.arithmetic == ARITHMETIC:
37
+ down_x = activation_ste(swiglu_ste(gu))
38
+ else:
39
+ g, u = gu.chunk(2, -1)
40
+ down_x = torch.nn.functional.silu(g) * u
41
+ for p, W, inp in ((p13, torch.cat([wg, wu]), z), (p2, wd, down_x)):
42
+ H = hessian_fc1(inp)
43
+ hv = cholesky_inv_upper(H)
44
+ cb, scale, glob = p.quantized(e)
45
+ a, b, _ = sweep_expert(W, hv, cb[:256], cb[256:], scale,
46
+ float(glob), col_block=128, refine=2)
47
+ changed += int(((a != p.codes_a[e]) | (b != p.codes_b[e])).sum())
48
+ total += a.numel()
49
+ p.codes_a[e].copy_(a); p.codes_b[e].copy_(b)
50
+ del H, hv, a, b
51
+ del wg, wu, wd, gu, down_x, z
52
+ after = evaluate_train()
53
+ accepted = accept_candidate(before, after)
54
+ if not accepted:
55
+ for p, (a, b) in zip((p13, p2), old):
56
+ p.codes_a.copy_(a); p.codes_b.copy_(b)
57
+ return {'method': 'LDLQ training-Hessian proposal with routed-training-output acceptance',
58
+ 'training_before': before, 'training_candidate': after,
59
+ 'accepted': accepted, 'proposed_changed_groups': changed,
60
+ 'groups_considered': total}
61
+ except BaseException:
62
+ for p, (a, b) in zip((p13, p2), old):
63
+ p.codes_a.copy_(a); p.codes_b.copy_(b)
64
+ raise
reproduce/source/btx53/arvq88/arvq_reap.py ADDED
@@ -0,0 +1,174 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Score actual ARVQ vs NVFP4 expert outputs on training captures.
2
+
3
+ Uses fitted ARVQ candidates for ALL 256 experts, including currently hot experts.
4
+ The score is the positive reduction in routing-weighted squared output error
5
+ against original FP8 when an expert is kept NVFP4 instead of ARVQ. Cross-expert
6
+ error covariance and downstream model quality remain outside this proxy.
7
+ """
8
+ import argparse
9
+ import json
10
+ import sys
11
+ from pathlib import Path
12
+ import numpy as np
13
+ import torch
14
+
15
+ ROOT = Path(__file__).resolve().parents[1]
16
+ sys.path[:0] = [str(ROOT), str(ROOT/'tools'), str(ROOT.parent/'tools')]
17
+ from arvq88.activation import activation_ste, swiglu_ste, ARITHMETIC
18
+ from arvq88.encoder import EncodedLayerProj, reconstruct
19
+ from arvq88.inputs import write
20
+ from arvq88.pack import sha
21
+ from arvq88.reap import blend_allocate
22
+
23
+
24
+ def hot_benefit(reference, cold_output, hot_output, gates):
25
+ if reference.shape != cold_output.shape or reference.shape != hot_output.shape:
26
+ raise ValueError('Expert outputs must have matching shapes')
27
+ weight = gates.double().square()
28
+ cold = (cold_output.double()-reference.double()).square().sum(-1)
29
+ hot = (hot_output.double()-reference.double()).square().sum(-1)
30
+ return float((weight*(cold-hot)).sum())
31
+
32
+
33
+ def score_output(x, w13, w2, quantized):
34
+ if quantized:
35
+ return activation_ste(swiglu_ste(activation_ste(x) @ w13.T)) @ w2.T
36
+ g,u=(x @ w13.T).chunk(2,-1)
37
+ return (torch.nn.functional.silu(g)*u) @ w2.T
38
+
39
+
40
+ @torch.no_grad()
41
+ def score_layer(args):
42
+ from ingest import SourceMeta, load_expert_bf16
43
+ from aqlm_quantize import ShardReader, dequant_nvfp4, FP4_LUT
44
+ torch.set_num_threads(2)
45
+ dev=args.device; L=args.layer; source=Path(args.fit_dir)/f'layer_{L:05d}'
46
+ manifest=json.loads((source/'arvq-manifest.json').read_text())
47
+ if manifest['cold_expert_ids']!=list(range(256)):
48
+ raise ValueError('ARVQ allocation scoring requires candidates for all 256 experts')
49
+ store={}
50
+ for name in ('w13','w2'):
51
+ store.update(torch.load(source/f'{name}.pt',map_location='cpu',weights_only=True))
52
+ enc={k:EncodedLayerProj(**store[k]) for k in ('w13','w2')}
53
+ captures=torch.load(args.capture,map_location='cpu',weights_only=True)
54
+ if 'pv_split' in captures:
55
+ mask=captures['pv_split']==0
56
+ captures={k:v[mask] for k,v in captures.items() if k in ('x','topk_ids','topk_weights')}
57
+ x=captures['x'].to(dev).float();ids=captures['topk_ids'].to(dev);g=captures['topk_weights'].to(dev).float()
58
+ meta=SourceMeta(**json.loads(Path(args.source_meta).read_text())['meta'])
59
+ reader=ShardReader(args.donor);lut=torch.tensor(FP4_LUT,device=dev,dtype=torch.float32)
60
+ benefits=np.zeros(256); counts=np.zeros(256,dtype=np.int64)
61
+ for e in range(256):
62
+ rows=(ids==e).any(1).nonzero().flatten()
63
+ counts[e]=len(rows)
64
+ if not len(rows):continue
65
+ cold13=reconstruct(enc['w13'],e,dev);cold2=reconstruct(enc['w2'],e,dev)
66
+ wg,wu,wd=[w.to(dev).float() for w in load_expert_bf16(str(Path(args.src_cache)/f'layer_{L}'/f'expert_{e}.safetensors'),meta)]
67
+ pre=f'model.layers.{L}.mlp.experts.{e}'
68
+ hot=[dequant_nvfp4(reader,pre+'.'+proj,dev,lut).float() for proj in ('gate_proj','up_proj','down_proj')]
69
+ hot13=torch.cat(hot[:2]);hot2=hot[2];ref13=torch.cat([wg,wu])
70
+ for batch in rows.split(args.batch):
71
+ z=x[batch];weight=(g[batch]*(ids[batch]==e)).sum(1)
72
+ ref=score_output(z,ref13,wd,False)
73
+ cold=score_output(z,cold13,cold2,True)
74
+ hot_out=score_output(z,hot13,hot2,True)
75
+ benefits[e]+=hot_benefit(ref,cold,hot_out,weight)
76
+ del cold13,cold2,wg,wu,wd,hot,hot13,hot2,ref13
77
+ out=Path(args.out);out.parent.mkdir(parents=True,exist_ok=True)
78
+ np.savez(out,scores=benefits.clip(min=0),signed_hot_benefit=benefits,routed_rows=counts,layer=L)
79
+ write(out.with_suffix('.json'),{'layer':L,'metric':'sum routing_weight^2 * (ARVQ squared error - NVFP4 squared error), clipped at zero',
80
+ 'reference':'original block-FP8 weights','arithmetic':ARITHMETIC,'capture_sha256':sha(args.capture),
81
+ 'candidate_manifest_sha256':sha(source/'arvq-manifest.json'),'scope':'all 256 experts; training tokens only',
82
+ 'limitation':'FP32 GEMMs with serving rounding emulation; no native MMA or cross-expert covariance'})
83
+ print('SCORED',L,args.capture,flush=True)
84
+
85
+
86
+ def allocate_scores(text_files,mm_files,hot_count=5750):
87
+ if len(text_files)!=75 or len(mm_files)!=75:raise ValueError('Need both modalities for all 75 layers')
88
+ def read(files):
89
+ scores=[]
90
+ for L,p in zip(range(3,78),files):
91
+ d=np.load(p)
92
+ if int(d['layer'])!=L:raise ValueError('Wrong score layer order')
93
+ scores.append(d['scores'])
94
+ return np.stack(scores)
95
+ text,mm=read(text_files),read(mm_files)
96
+ result=blend_allocate(text,mm,hot_count)
97
+ result['provenance']={'mode':'ARVQ output-benefit allocation','text_weight':.75,'mm_weight':.25,
98
+ 'hot_count':hot_count,'floor':8,'cap':176,'text_partition':'train only',
99
+ 'metric':'positive routing-weighted squared-error reduction from retaining NVFP4 over fitted ARVQ',
100
+ 'scores_sha256':{str(p):sha(p) for p in [*text_files,*mm_files]},
101
+ 'limitation':'per-expert additive error proxy; excludes cross-expert covariance and end-to-end model quality'}
102
+ return result
103
+
104
+
105
+
106
+
107
+ def recompute(args, hub, work):
108
+ """Fit all expert candidates, measure both modalities, then spend hot budget."""
109
+ from types import SimpleNamespace
110
+ from arvq88.inputs import ensure_source, LAYERS
111
+ from arvq88.jobs import run_jobs
112
+ if not getattr(args,'mm_acts',None):
113
+ raise ValueError('ARVQ REAP requires --mm-acts with matching donor multimodal captures')
114
+ mm=Path(args.mm_acts)
115
+ if not all((mm/f'acts_layer{L}.pt').is_file() for L in LAYERS):
116
+ raise ValueError('Incomplete multimodal captures; refusing text-only ARVQ REAP')
117
+ prefit=work/'arvq_reap';prefit.mkdir(exist_ok=True)
118
+ a=SimpleNamespace(**vars(args));a.workdir=str(prefit)
119
+ # These are candidates, not a valid final 5750-hot allocation.
120
+ all_experts={'layers':{str(L):{'cold_5750':list(range(256)),'hot':[]} for L in LAYERS}}
121
+ write(prefit/'status.json',{'stage':'preparing_original_fp8_sources','experts':19200})
122
+ ensure_source(a,hub,prefit,all_experts)
123
+ write(prefit/'assignment.json',all_experts);write(prefit/'run_config.json',vars(a))
124
+ write(prefit/'status.json',{'stage':'fitting_all_expert_candidates','layers':75,'experts':19200})
125
+ run_jobs([(f'fit-{L}',lambda gpu,L=L:[sys.executable,'-u',str(ROOT/'arvq88/fit.py'),str(prefit),str(L),f'cuda:{gpu}']) for L in LAYERS],args.gpus,prefit/'logs')
126
+ jobs=[]
127
+ for modality,folder in [('text',Path(args.calib)),('mm',mm)]:
128
+ for L in LAYERS:
129
+ output=prefit/'scores'/modality/f'layer_{L}.npz'
130
+ jobs.append((f'score-{modality}-{L}',lambda gpu,L=L,folder=folder,output=output:[
131
+ sys.executable,'-u',str(ROOT/'arvq88/arvq_reap.py'),
132
+ '--fit-dir',str(prefit/'initial'),'--capture',str(folder/f'acts_layer{L}.pt'),
133
+ '--src-cache',args.src_cache,'--source-meta',str(prefit/'source.json'),
134
+ '--donor',args.donor_cache,'--layer',str(L),'--device',f'cuda:{gpu}','--out',str(output)]))
135
+ write(prefit/'status.json',{'stage':'scoring_text_and_multimodal','jobs':150})
136
+ run_jobs(jobs,args.gpus,prefit/'logs')
137
+ result=allocate_scores([prefit/'scores/text'/f'layer_{L}.npz' for L in LAYERS],
138
+ [prefit/'scores/mm'/f'layer_{L}.npz' for L in LAYERS],args.hot_count)
139
+ result['provenance']['candidate_codebook_scope']=args.codebook_scope
140
+ write(prefit/'status.json',{'stage':'allocation_complete','hot':args.hot_count,'cold':19200-args.hot_count})
141
+ return result
142
+
143
+
144
+ def select_candidates(prefit,work,allocation):
145
+ """Keep exactly the fitted candidates used to calculate hot-slot benefit."""
146
+ from arvq88.pack import export_layer
147
+ prefit,work=Path(prefit),Path(work)
148
+ for L in range(3,78):
149
+ source=prefit/'initial'/f'layer_{L:05d}';dest=work/'initial'/f'layer_{L:05d}'
150
+ dest.mkdir(parents=True,exist_ok=True)
151
+ manifest=json.loads((source/'arvq-manifest.json').read_text())
152
+ if manifest['cold_expert_ids']!=list(range(256)):raise ValueError('Candidates must cover all experts')
153
+ cold=allocation['layers'][str(L)]['cold_5750'];files={}
154
+ for proj in ['w13','w2']:
155
+ d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj]
156
+ for key in ['a','b','s']:d[key]=d[key][cold]
157
+ if d['c0'].ndim==3:
158
+ for key in ['c0','c1']:d[key]=d[key][cold]
159
+ path=dest/f'{proj}.pt';tmp=path.with_suffix('.tmp');torch.save({proj:d},tmp);tmp.replace(path)
160
+ files[proj]={'file':path.name,'sha256':sha(path)}
161
+ write(dest/'arvq-manifest.json',{**manifest,'cold_expert_ids':cold,'files':files,
162
+ 'selection':'Exact subset of all-expert ARVQ candidates scored for allocation',
163
+ 'candidate_manifest_sha256':sha(source/'arvq-manifest.json')})
164
+ export_layer(dest,work/'cold',L)
165
+
166
+
167
+ if __name__=='__main__':
168
+ ap=argparse.ArgumentParser(description=__doc__)
169
+ ap.add_argument('--fit-dir',required=True);ap.add_argument('--capture',required=True)
170
+ ap.add_argument('--src-cache',required=True);ap.add_argument('--source-meta',required=True)
171
+ ap.add_argument('--donor',required=True);ap.add_argument('--layer',type=int,required=True)
172
+ ap.add_argument('--out',required=True);ap.add_argument('--device',default='cuda:0')
173
+ ap.add_argument('--batch',type=int,default=256)
174
+ score_layer(ap.parse_args())
reproduce/source/btx53/arvq88/backbone.py ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Explicit tensor provenance: language/MTP from donor, vision from base."""
2
+ import copy,json,re,os
3
+ from pathlib import Path
4
+
5
+ def vision_tensor(name):
6
+ return name.startswith(('vision_tower.','mm_projector.'))
7
+ def routed_tensor(name):
8
+ m=re.match(r'model.layers\.(\d+)\.mlp\.experts\.',name)
9
+ return bool(m and 3<=int(m[1])<=77)
10
+
11
+ def copy_backbone(base,donor,dest,save):
12
+ from safetensors import safe_open
13
+ from .pack import sha
14
+ provenance={}
15
+ for root,label,choose in [(base,'default-vision',vision_tensor),(donor,'nvfp4-donor',lambda n:not vision_tensor(n) and not routed_tensor(n))]:
16
+ index=json.loads((root/'model.safetensors.index.json').read_text())['weight_map']
17
+ for i,filename in enumerate(sorted(set(index.values()))):
18
+ names=sorted(n for n,f in index.items() if f==filename and choose(n))
19
+ if not names:continue
20
+ target=f'{label}-{i:05d}.safetensors'
21
+ with safe_open(str(root/filename),framework='pt') as stream:
22
+ save(target,{n:stream.get_tensor(n) for n in names})
23
+ # save_file preserves dtype and values; gate the resulting immutable shard identity.
24
+ provenance[target]={'source':str(root/filename),'tensor_names':names,'sha256':sha(dest/target),'role':label}
25
+ expected={n for n in json.loads((base/'model.safetensors.index.json').read_text())['weight_map'] if not routed_tensor(n) and not vision_tensor(n) and not n.endswith(('.weight_scale','.weight_scale_2','.input_scale','.input_scale_2'))}
26
+ actual={n for f in provenance.values() if f['role']=='nvfp4-donor' for n in f['tensor_names']}
27
+ if expected-actual:raise ValueError(f'Donor missing language/MTP tensors: {sorted(expected-actual)[:5]}')
28
+ return provenance
29
+
30
+ def hybrid_config(base,donor):
31
+ config=copy.deepcopy(base);text=copy.deepcopy(donor.get('text_config',donor))
32
+ qc=copy.deepcopy(base.get('quantization_config',{}))
33
+ donor_qc=text.pop('quantization_config',{})
34
+ if donor_qc.get('quant_method')=='modelopt':qc['nvfp4']=donor_qc
35
+ elif donor_qc.get('nvfp4'):qc['nvfp4']=donor_qc['nvfp4']
36
+ else:raise ValueError('Expected NVFP4 donor quantization config')
37
+ if 'text_config' in config:
38
+ config['text_config']=text
39
+ for name in ['pad_token_id','eos_token_id','tie_word_embeddings','dtype']:
40
+ if name in text:config[name]=text[name]
41
+ else:config=text
42
+ config['quantization_config']=copy.deepcopy(qc)
43
+ config.get('text_config',config)['quantization_config']=copy.deepcopy(qc)
44
+ return config
reproduce/source/btx53/arvq88/checkpoint.py ADDED
@@ -0,0 +1,168 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Assemble/gate/publish an explicitly versioned 8+8 hybrid checkpoint."""
2
+ import json,re,shutil,os
3
+ from pathlib import Path
4
+ from .inputs import LAYERS,complete_snapshot,write,token
5
+ from .pack import sha
6
+
7
+ def build(args,hub,work,allocation):
8
+ # Refresh the card renderer after long PV runs so reviewed provenance text is current.
9
+ import importlib,sys
10
+ if 'arvq88.pv_campaign' in sys.modules:importlib.reload(sys.modules['arvq88.pv_campaign'])
11
+ import torch
12
+ from safetensors import safe_open
13
+ from safetensors.torch import save_file
14
+ base=Path(args.local_hot)
15
+ if not complete_snapshot(base):base=hub.snapshot(args.base_hybrid,work/'base-hybrid')
16
+ dest=work/'checkpoint'
17
+ if dest.resolve()==base.resolve():raise ValueError('Output checkpoint cannot overwrite the local hot donor')
18
+ dest.mkdir(exist_ok=True)
19
+ manifests=[json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text()) for L in LAYERS]
20
+ formats={(m['format'],m['version']) for m in manifests}
21
+ if len(formats)!=1 or not formats<= {('rvq256_256x8',2),('rvq256_256x8_expert',3)}:
22
+ raise ValueError('Mixed or unsupported cold formats')
23
+ cold_format,cold_version=next(iter(formats))
24
+ index=json.loads((base/'model.safetensors.index.json').read_text())['weight_map'];wm={};total=0
25
+ def save(name,tensors):
26
+ nonlocal total
27
+ if not tensors:return
28
+ save_file(tensors,str(dest/name));wm.update({k:name for k in tensors});total+=sum(t.numel()*t.element_size() for t in tensors.values())
29
+ from .backbone import copy_backbone,hybrid_config
30
+ donor=Path(args.donor_cache)
31
+ if not complete_snapshot(donor):donor=hub.snapshot(args.nvfp4_donor,donor)
32
+ if dest.resolve()==donor.resolve():raise ValueError('Output cannot overwrite donor')
33
+ # A stale assembled checkpoint must not retain weights from the previous provenance rule.
34
+ marker=dest/'backbone_sources.json'
35
+ if any(dest.glob('*.safetensors')) and not marker.exists():raise ValueError('Use a fresh checkpoint directory for donor-backbone assembly')
36
+ provenance=copy_backbone(base,donor,dest,save)
37
+ write(marker,provenance)
38
+ dwm=json.loads((donor/'model.safetensors.index.json').read_text())['weight_map']
39
+ def donor_get(name):
40
+ nonlocal dwm,donor
41
+ if dwm is None:
42
+ if not complete_snapshot(donor):donor=hub.snapshot(args.nvfp4_donor,donor)
43
+ dwm=json.loads((donor/'model.safetensors.index.json').read_text())['weight_map']
44
+ with safe_open(str(donor/dwm[name]),framework='pt') as f:return f.get_tensor(name)
45
+ for L in LAYERS:
46
+ pre=f'model.layers.{L}.mlp.experts.';row=allocation['layers'][str(L)];cold=row['cold_5750'];hot=row['hot']
47
+ with safe_open(str(base/index[pre+'hyb_kind']),framework='pt') as f:old=f.get_tensor(pre+'hyb_kind')
48
+ tensors={};expected=torch.zeros(256,dtype=torch.int8);expected[cold]=2;tensors[pre+'hyb_kind']=expected
49
+ suffixes=['nvfp4_w13_packed','nvfp4_w13_bscale','nvfp4_w13_scale2','nvfp4_w2_packed','nvfp4_w2_bscale','nvfp4_w2_scale2']
50
+ if torch.equal(old,expected) and args.nvfp4_donor=='RadixArk/GLM-5.3-NVFP4':
51
+ for suffix in suffixes:
52
+ name=pre+suffix
53
+ with safe_open(str(base/index[name]),framework='pt') as f:tensors[name]=f.get_tensor(name)
54
+ else:
55
+ parts={k:[] for k in suffixes}
56
+ for e in hot:
57
+ ep=pre+str(e)+'.'
58
+ for suffix,out_suffix in [('weight','packed'),('weight_scale','bscale')]:
59
+ gu=torch.cat([donor_get(ep+p+'.'+suffix) for p in ['gate_proj','up_proj']]).view(torch.uint8)
60
+ dn=donor_get(ep+'down_proj.'+suffix).view(torch.uint8)
61
+ parts['nvfp4_w13_'+out_suffix].append(gu);parts['nvfp4_w2_'+out_suffix].append(dn)
62
+ parts['nvfp4_w13_scale2'].append(torch.stack([donor_get(ep+p+'.weight_scale_2').float().reshape(()) for p in ['gate_proj','up_proj']]))
63
+ parts['nvfp4_w2_scale2'].append(donor_get(ep+'down_proj.weight_scale_2').float().reshape(1))
64
+ tensors.update({pre+k:torch.stack(v) for k,v in parts.items()})
65
+ save(f'hot-layer-{L:03d}.safetensors',tensors)
66
+ manifest=json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text())
67
+ if manifest['cold_expert_ids']!=cold:raise ValueError('Cold fit/assignment mismatch')
68
+ for f in manifest['files'].values():
69
+ src=work/'cold'/f['file']
70
+ if sha(src)!=f['sha256']:raise ValueError('Cold shard identity mismatch')
71
+ target=dest/src.name
72
+ if not target.exists():os.link(src,target)
73
+ elif sha(target)!=f['sha256']:raise ValueError('Stale assembled cold shard')
74
+ with safe_open(str(target),framework='pt') as stream:
75
+ for name in stream.keys():
76
+ t=stream.get_tensor(name);wm[name]=src.name;total+=t.numel()*t.element_size()
77
+ for p in base.iterdir():
78
+ if p.is_file() and p.suffix in ['.json','.py','.jinja'] and p.name not in ['config.json','model.safetensors.index.json','backbone_sources.json'] and 'report' not in p.name and 'provenance' not in p.name:shutil.copy2(p,dest/p.name)
79
+ # Keep the vision wrapper/processor; language config and tokenizer follow donor.
80
+ for p in donor.iterdir():
81
+ if p.is_file() and (p.name.startswith(('tokenizer','special_tokens','added_tokens','chat_template','generation_config')) or p.name.endswith('.model')):shutil.copy2(p,dest/p.name)
82
+ config=hybrid_config(json.loads((base/'config.json').read_text()),json.loads((donor/'config.json').read_text()))
83
+ books={str(L):{'n_nvfp4':len(allocation['layers'][str(L)]['hot']),'n_base':0,'n_cold':len(allocation['layers'][str(L)]['cold_5750'])} for L in LAYERS}
84
+ for cfg in [config,config.get('text_config',config)]:
85
+ qc=cfg.setdefault('quantization_config',{});qc['quant_method']='nvfp4_arvq_hybrid'
86
+ qc['arvq']={'format':cold_format,'version':cold_version,'activation_planes':4,'weight_scale_group':128,
87
+ 'codebook_scope':'expert' if cold_version==3 else 'layer','codebook_sizes':[256,256]}
88
+ qc['aqlm_layer_books']=books
89
+ write(dest/'config.json',config);write(dest/'model.safetensors.index.json',{'metadata':{'total_size':total},'weight_map':wm})
90
+ shutil.copy2(work/'assignment.json',dest/'cold_assignment.json');shutil.copy2(work/'source.json',dest/'source.json')
91
+ shutil.copytree(work/'cold',dest/'cold_manifests',dirs_exist_ok=True,ignore=shutil.ignore_patterns('*.safetensors'))
92
+ return dest
93
+
94
+ def gates(work,layers=None,device="cpu",report_path=None):
95
+ import torch
96
+ from safetensors import safe_open
97
+ from .pack import decode
98
+ torch.set_num_threads(2);layers=LAYERS if layers is None else layers;ck=work/'checkpoint';wm=json.loads((ck/'model.safetensors.index.json').read_text())['weight_map'];errors=[];hashes={}
99
+ if any(k.endswith(('.w13_codes','.w2c_codes','.w2m_codes')) for k in wm):errors.append('Old AQLM indices leaked')
100
+ for L in layers:
101
+ m=json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text());E=len(m['cold_expert_ids'])
102
+ pre=f'model.layers.{L}.mlp.experts.'
103
+ with safe_open(str(ck/wm[pre+'hyb_kind']),framework='pt') as stream:
104
+ kind=stream.get_tensor(pre+'hyb_kind')
105
+ if (kind==2).nonzero().flatten().tolist()!=m['cold_expert_ids']:errors.append('Hot/cold assignment mismatch')
106
+ for suffix,shape in [('nvfp4_w13_packed',(256-E,4096,3072)),('nvfp4_w2_packed',(256-E,6144,1024)),('nvfp4_w13_bscale',(256-E,4096,384)),('nvfp4_w2_bscale',(256-E,6144,128)),('nvfp4_w13_scale2',(256-E,2)),('nvfp4_w2_scale2',(256-E,1))]:
107
+ if tuple(stream.get_slice(pre+suffix).get_shape())!=shape:errors.append(f'Hot shape mismatch {L} {suffix}')
108
+ for proj in ['w13','w2']:
109
+ f=m['files'][proj];path=ck/f['file'];digest=sha(path);hashes[f['file']]=digest
110
+ if digest!=f['sha256']:errors.append(f'SHA mismatch {path.name}')
111
+ with safe_open(str(path),framework='pt') as stream:
112
+ t={k:stream.get_tensor(k) for k in stream.keys()}
113
+ N,K=(4096,6144) if proj=='w13' else (6144,2048);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
114
+ cbshape=(E,512) if m.get('version')==3 else (512,)
115
+ config=json.loads((ck/'config.json').read_text())
116
+ qconfig=config.get('text_config',config)['quantization_config']['arvq']
117
+ if qconfig['format']!=m['format'] or qconfig['version']!=m['version']:errors.append(f'Format metadata mismatch {L}')
118
+ for suffix,shape,dtype in [('packed',(E,N//16,K//64,64),torch.uint32),('codebooks',cbshape,torch.uint32),('scales',(E,N//16,K//128,16),torch.uint8),('global',(1,),torch.float32)]:
119
+ if t[pre+suffix].shape!=shape or t[pre+suffix].dtype!=dtype:errors.append(f'Shape/dtype {L} {suffix}')
120
+ if wm.get(pre+suffix)!=f['file']:errors.append('Index mismatch')
121
+ scales=t[pre+'scales'].view(torch.float8_e4m3fn).float()
122
+ if not torch.isfinite(scales).all() or (scales<0).any() or not torch.isfinite(t[pre+'global']).all() or not (t[pre+'global']>0).all():errors.append(f'Invalid scales {L} {proj}')
123
+ if not torch.isfinite(decode({k:v.to(device) for k,v in t.items()},L,proj,0)).all():errors.append('Nonfinite decoded sample')
124
+ del t,scales
125
+ backbone=json.loads((ck/'backbone_sources.json').read_text())
126
+ for name,entry in backbone.items():
127
+ if sha(ck/name)!=entry['sha256']:errors.append(f'Backbone SHA mismatch {name}')
128
+ if any(wm.get(n)!=name for n in entry['tensor_names']):errors.append(f'Backbone index mismatch {name}')
129
+ report={'passed':not errors,'layers':len(layers),'layer_ids':layers,'errors':errors,'cold_file_sha256':hashes,'scope':'Initial-fit export has exhaustive index roundtrips and sampled weight parity; checkpoint hashes, shapes, scales and decoded finiteness checked. No SM120 serving or model quality gate.'}
130
+ report['metadata_sha256']={n:sha(ck/n) for n in ['config.json','model.safetensors.index.json','backbone_sources.json']}
131
+ write(report_path or ck/'gates_report.json',report)
132
+ if errors:raise RuntimeError(f'8+8 gates failed: {errors[:3]}')
133
+ return report
134
+
135
+ def parallel_gates(work,gpus):
136
+ import subprocess,sys
137
+ from .inputs import ROOT
138
+ folder=work/'gate_parts';folder.mkdir(exist_ok=True);jobs=[]
139
+ for i,gpu in enumerate(gpus):
140
+ layers=LAYERS[i::len(gpus)]
141
+ if not layers:continue
142
+ result=folder/f'gpu{gpu}.json'
143
+ with open(folder/f'gpu{gpu}.log','a') as log:
144
+ p=subprocess.Popen([sys.executable,str(ROOT/'arvq88/gate_worker.py'),str(work),','.join(map(str,layers)),f'cuda:{gpu}',str(result)],stdout=log,stderr=subprocess.STDOUT,start_new_session=True)
145
+ jobs.append((p,result))
146
+ codes=[p.wait() for p,_ in jobs]
147
+ if any(codes):raise RuntimeError(f'Gate workers failed: {codes}; inspect {folder}')
148
+ parts=[json.loads(r.read_text()) for _,r in jobs];ids=[l for p in parts for l in p['layer_ids']]
149
+ if sorted(ids)!=LAYERS:raise RuntimeError('Gate layer coverage mismatch')
150
+ if any(p['metadata_sha256']!=parts[0]['metadata_sha256'] for p in parts):raise RuntimeError('Checkpoint metadata changed during gates')
151
+ report={**parts[0],'passed':all(p['passed'] for p in parts),'layers':75,'layer_ids':LAYERS,
152
+ 'errors':[e for p in parts for e in p['errors']],
153
+ 'cold_file_sha256':{k:v for p in parts for k,v in p['cold_file_sha256'].items()},'workers':len(parts)}
154
+ write(work/'checkpoint/gates_report.json',report)
155
+ if not report['passed']:raise RuntimeError('Gates did not pass')
156
+ return report
157
+
158
+ def publish(args,hub,work):
159
+ ck=work/'checkpoint';report=json.loads((ck/'gates_report.json').read_text())
160
+ if not report['passed'] or len(report['cold_file_sha256'])!=150:raise RuntimeError('Incomplete initial checkpoint gates')
161
+ for name,digest in {**report['cold_file_sha256'],**report['metadata_sha256']}.items():
162
+ if sha(ck/name)!=digest:raise RuntimeError('Weights changed after gates')
163
+ parent=hub.api.model_info(args.dst).sha
164
+ commit=hub.api.upload_folder(repo_id=args.dst,folder_path=str(ck),revision='main',parent_commit=parent,commit_message='GLM-5.3 NVFP4 + ARVQ 8+8 activation-calibrated initial fit')
165
+ remote={s.rfilename:s for s in hub.api.model_info(args.dst,revision=commit.oid,files_metadata=True).siblings}
166
+ for name,digest in report['cold_file_sha256'].items():
167
+ if remote[name].lfs.sha256!=digest:raise RuntimeError('Remote cold-weight hash mismatch')
168
+ write(work/'upload.json',{'repo':args.dst,'revision':commit.oid,'cold_files_verified':150});return commit.commit_url
reproduce/source/btx53/arvq88/cli.py ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Local-or-Hub GLM-5.3 to NVFP4/ARVQ 8+8 with audited PV tuning."""
2
+ import argparse,fcntl,json,os,subprocess,sys,time,traceback,shutil
3
+ from pathlib import Path
4
+ from .inputs import ROOT,LAYERS,Hub,assignment,ensure_captures,ensure_source,write
5
+
6
+ def parser():
7
+ p=argparse.ArgumentParser(description=__doc__)
8
+ p.add_argument('--arvq88',action='store_true',help=argparse.SUPPRESS)
9
+ p.add_argument('--src',default='zai-org/GLM-5.3',help='Hugging Face repo ID or local model directory');p.add_argument('--dst',required=True)
10
+ p.add_argument('--workdir',default='/tmp/glm53-arvq88');p.add_argument('--gpus',default='0,1,2,3,4,5,6,7')
11
+ p.add_argument('--assign',default=str(ROOT/'cold_assignment.json'));p.add_argument('--reap-scores')
12
+ p.add_argument('--mm-stats',help='Cached multimodal salience NPZ for r3 allocation')
13
+ p.add_argument('--mm-calib',default='/tmp/glm53-calib-mm')
14
+ p.add_argument('--vision-dir',default='/tmp/glm53-vision')
15
+ p.add_argument('--reap-parts',help='Historical all-expert 1x16 AQLM warm starts, allocation proxy only')
16
+ p.add_argument('--reap-metric',choices=['aqlm','arvq'],default='aqlm',help='ARVQ recomputes hot allocation from fitted ARVQ vs NVFP4 output error on both modalities')
17
+ p.add_argument('--mm-acts',help='Matching donor multimodal activation directory for ARVQ allocation scoring')
18
+ p.add_argument('--hot-count',type=int,default=5750);p.add_argument('--calib',default='/tmp/glm53-capture/acts')
19
+ p.add_argument('--calib-tokens',default='/tmp/glm52-calib-v3');p.add_argument('--calib-token-count',type=int,default=15_000_000)
20
+ p.add_argument('--src-cache',default='/tmp/glm53-fp8-cold')
21
+ p.add_argument('--nvfp4-donor',default='RadixArk/GLM-5.3-NVFP4',help='NVFP4 Hugging Face repo ID or local sharded checkpoint directory; download cache is automatic')
22
+ p.add_argument('--base-hybrid',default='jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m')
23
+ p.add_argument('--local-hot',default='/tmp/glm53-hybrid-pipeline/checkpoint')
24
+ p.add_argument('--download-workers',type=int,default=8);p.add_argument('--threads',type=int,default=2)
25
+ p.add_argument('--col-block',type=int,default=128);p.add_argument('--cb-iters',type=int,default=8)
26
+ p.add_argument('--sweep-passes',type=int,default=2);p.add_argument('--refine',type=int,default=1)
27
+ p.add_argument('--codebook-scope',choices=['layer','expert'],default='layer')
28
+ p.add_argument('--pv-reassign-every',type=int,default=0,help='Alternate index proposals every N PV steps; 0 freezes indices')
29
+ from .activation import ARITHMETIC
30
+ p.add_argument('--pv-arithmetic',choices=['fp32',ARITHMETIC],default=ARITHMETIC,help='PV forward arithmetic: serving activation emulation or legacy FP32')
31
+ p.add_argument('--pv-steps',type=int,default=200,help='PV steps per layer; 0 produces initial fits only')
32
+ p.add_argument('--clean-repo',action='store_true',help='After verified final PV upload, remove obsolete repository files')
33
+ p.add_argument('--detach',action='store_true');p.add_argument('--do-upload',action='store_true')
34
+ p.add_argument('--plan',action='store_true',help='Show resolved workflow without downloading, fitting or uploading')
35
+ return p
36
+
37
+ def resolve_donor(args,work):
38
+ from .local_source import local_directory,fingerprint
39
+ from .inputs import complete_snapshot
40
+ local=local_directory(args.nvfp4_donor)
41
+ args.donor_identity=None
42
+ if local:
43
+ if not complete_snapshot(local) or not (local/'config.json').is_file():
44
+ raise ValueError('Local NVFP4 donor requires config.json, model.safetensors.index.json and all indexed shards')
45
+ args.nvfp4_donor=str(local);args.donor_cache=str(local)
46
+ args.donor_identity=fingerprint(local)
47
+ else:
48
+ import hashlib
49
+ legacy=Path('/tmp/glm53-dl/nvfp4-full')
50
+ args.donor_cache=str(legacy if args.nvfp4_donor=='RadixArk/GLM-5.3-NVFP4' and complete_snapshot(legacy)
51
+ else work/'donors'/hashlib.sha256(args.nvfp4_donor.encode()).hexdigest()[:16])
52
+
53
+ def main():
54
+ args=parser().parse_args()
55
+ from .local_source import local_directory,fingerprint
56
+ local=local_directory(args.src)
57
+ if local:
58
+ args.src=str(local);fingerprint(local) # Validate config/index/shards before capture or GPU work.
59
+ args.gpus=[int(x) for x in args.gpus.split(',')]
60
+ if not args.gpus or len(args.gpus)!=len(set(args.gpus)) or min(args.gpus)<0:raise ValueError('GPUs must be a unique nonnegative list')
61
+ if args.pv_steps!=0 and args.pv_steps<40:raise ValueError('PV steps must be 0 or at least 40')
62
+ if args.pv_reassign_every<0:raise ValueError('Invalid index reassignment interval')
63
+ if args.reap_metric=='arvq' and not args.mm_acts:raise ValueError('ARVQ scoring requires --mm-acts')
64
+ if args.clean_repo and args.pv_steps==0:raise ValueError('Repository cleanup requires final PV weights')
65
+ if args.threads<1 or args.download_workers<1 or args.calib_token_count<2048:raise ValueError('Invalid CPU/download/calibration count')
66
+ if args.col_block<=0 or args.col_block%8 or 2048%args.col_block or args.cb_iters<1 or args.sweep_passes<1 or args.refine<0:raise ValueError('Invalid encoder settings')
67
+ work=Path(args.workdir);work.mkdir(parents=True,exist_ok=True)
68
+ resolve_donor(args,work)
69
+ # Do not silently reuse this machine's original-source artifacts for another input.
70
+ if args.src!='zai-org/GLM-5.3':
71
+ for name,default,replacement in [('src_cache','/tmp/glm53-fp8-cold',work/'source'),('calib','/tmp/glm53-capture/acts',work/'capture/acts'),('calib_tokens','/tmp/glm52-calib-v3',work/'calib-tokens'),('assign',str(ROOT/'cold_assignment.json'),work/'recomputed-assignment.json')]:
72
+ if getattr(args,name)==default:setattr(args,name,str(replacement))
73
+ if args.nvfp4_donor!='RadixArk/GLM-5.3-NVFP4':
74
+ if args.calib=='/tmp/glm53-capture/acts':args.calib=str(work/'capture/acts')
75
+ if args.assign==str(ROOT/'cold_assignment.json'):args.assign=str(work/'recomputed-assignment.json')
76
+ if args.base_hybrid!='jarrelscy/GLM-5.3-Vision-NVFP4-AQLM-hybrid-1m' and args.local_hot=='/tmp/glm53-hybrid-pipeline/checkpoint':args.local_hot=str(work/'base-hybrid')
77
+ if args.plan:
78
+ print(json.dumps({'format':'rvq256_256x8_expert' if args.codebook_scope=='expert' else 'rvq256_256x8','stage':'PV-tuned' if args.pv_steps else 'initial fit (no PV)','args':vars(args),'steps':['reuse/download text calibration','reuse/capture NVFP4-teacher activations','reuse allocation or scores, otherwise recalculate 75% text REAP + 25% multimodal salience allocation','reuse/fetch original cold source tensors','fit 8+8 with per-layer cache identities',*(['PV with train/validation/audit splits, early stopping and audit fallback'] if args.pv_steps else []),'assemble donor hot/nonexpert/MTP + ARVQ cold + default vision','gate packed format and hashes','upload main only with --do-upload'],'fork_changes':False},indent=2));return
79
+ compat='/tmp/compat13/usr/local/cuda-13.2/compat'
80
+ if Path(compat).exists() and compat not in os.environ.get('LD_LIBRARY_PATH','').split(':'):
81
+ os.execvpe(sys.executable,[sys.executable]+sys.argv,{**os.environ,'LD_LIBRARY_PATH':compat+':'+os.environ.get('LD_LIBRARY_PATH','')})
82
+ if args.detach:
83
+ with open(work/'driver.log','a') as log:
84
+ child=subprocess.Popen([sys.executable,'-u']+[a for a in sys.argv if a!='--detach'],stdin=subprocess.DEVNULL,stdout=log,stderr=subprocess.STDOUT,start_new_session=True)
85
+ print(f'Detached 8+8 pipeline PID {child.pid}; {work}/driver.log');return
86
+ lock=open(work/'.workflow.lock','a');fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB)
87
+ import torch
88
+ torch.set_num_threads(args.threads)
89
+ if not torch.cuda.is_available():raise RuntimeError('CUDA required for 8+8 pipeline')
90
+ # Never overbook another user's or an existing detached worker's GPU.
91
+ result=subprocess.run(['nvidia-smi','--query-compute-apps=gpu_uuid,pid','--format=csv,noheader'],capture_output=True,text=True,check=True)
92
+ gpuinfo=subprocess.run(['nvidia-smi','--query-gpu=index,uuid','--format=csv,noheader'],capture_output=True,text=True,check=True)
93
+ selected={line.split(',')[1].strip() for line in gpuinfo.stdout.splitlines() if int(line.split(',')[0]) in args.gpus}
94
+ if any(line.split(',')[0].strip() in selected for line in result.stdout.splitlines() if line.strip()):raise RuntimeError('Requested GPUs have active compute jobs; choose free --gpus or wait. No processes were stopped.')
95
+ try:
96
+ binding={'src':args.src,'nvfp4_donor':args.nvfp4_donor,'base_hybrid':args.base_hybrid,'format':'rvq256_256x8_expert' if args.codebook_scope=='expert' else 'rvq256_256x8',
97
+ 'reap_metric':args.reap_metric,'pv_reassign_every':args.pv_reassign_every}
98
+ if args.donor_identity:binding['local_donor_identity']=args.donor_identity
99
+ if (work/'binding.json').exists() and json.loads((work/'binding.json').read_text())!=binding:raise ValueError('Workdir belongs to another input/donor/format')
100
+ write(work/'binding.json',binding);hub=Hub(work/'hfcache')
101
+ ensure_captures(args,hub,work)
102
+ allocation=assignment(args,hub,work)
103
+ ensure_source(args,hub,work,allocation)
104
+ write(work/'run_config.json',vars(args));pending=list(LAYERS);running={}
105
+ if args.reap_metric=='arvq':
106
+ from .arvq_reap import select_candidates
107
+ select_candidates(work/'arvq_reap',work,allocation);pending=[]
108
+ while pending or running:
109
+ for gpu in args.gpus:
110
+ if gpu in running or not pending:continue
111
+ L=pending.pop(0)
112
+ with open(work/f'fit-layer{L}.log','a') as log:
113
+ child=subprocess.Popen([sys.executable,'-u',str(ROOT/'arvq88/fit.py'),str(work),str(L),f'cuda:{gpu}'],stdout=log,stderr=subprocess.STDOUT,stdin=subprocess.DEVNULL,start_new_session=True)
114
+ running[gpu]=(L,child);print('FIT',L,'GPU',gpu,'PID',child.pid,flush=True)
115
+ time.sleep(15)
116
+ for gpu,(L,child) in list(running.items()):
117
+ rc=child.poll()
118
+ if rc is None:continue
119
+ if rc:raise RuntimeError(f'Fit layer {L} failed; see fit-layer{L}.log')
120
+ del running[gpu];print('FIT DONE',L,flush=True)
121
+ if args.pv_steps:
122
+ from .pv_campaign import run as run_pv
123
+ run_pv(args,work,work/'initial')
124
+ from .checkpoint import build,parallel_gates,publish
125
+ ck=build(args,hub,work,allocation);parallel_gates(work,args.gpus)
126
+ write(ck/'build_provenance.json',{'inputs':binding,'hub_revisions':hub.revisions,'allocation':allocation['provenance'],'capture_dir':args.calib,'calib_tokens':args.calib_tokens,'stage':'PV-tuned' if args.pv_steps else 'initial fit, no PV','format':binding['format'],'source_nonexpert_mtp':'nvfp4_donor','source_vision':'base_hybrid','capture_teacher':'nvfp4_donor','allocation_recipe':allocation['provenance']})
127
+ (ck/'README.md').write_text(f'''---
128
+ base_model: {args.src}
129
+ tags: [arvq, nvfp4, experimental]
130
+ ---
131
+ # GLM-5.3 NVFP4 / ARVQ 8+8 hybrid
132
+
133
+ Cold experts use two shared FP4 dictionaries of 256 entries, 8-dimensional groups,
134
+ 16 index bits/group and FP8 scales/128 weights (2.0625 bpw plus codebooks).
135
+ This is the activation-calibrated initial Hessian fit, **before PV**.
136
+ Hot experts use the selected NVFP4 donor. Nonexpert and MTP tensors come from the NVFP4 donor; vision comes from the
137
+ base hybrid. See build_provenance.json for every donor and allocation method.
138
+
139
+ Codebook scope: {args.codebook_scope}. Requires a loader/kernel supporting
140
+ {binding['format']} version {3 if args.codebook_scope=='expert' else 2}: 512 codewords per pair,
141
+ 64 uint32 index words per tile and 8+8 index extraction. The prior 8+7 loader is
142
+ incompatible. This pipeline does not modify the serving fork.
143
+ Offline packing/structure checks passed; serving and model-level quality are untested.
144
+ ''')
145
+ shutil.copy2(ROOT/'arvq88/pack.py',ck/'arvq88_reference.py')
146
+ if args.pv_steps:
147
+ from .pv_campaign import write_card
148
+ write_card(args,work)
149
+ from .final_publish import publish
150
+ write(work/'complete.json',{'checkpoint':str(ck),'uploaded':False})
151
+ if args.do_upload:print('UPLOADED',publish(args,hub,work),flush=True)
152
+ else:print('Checkpoint ready:',ck,'; add --do-upload to publish',flush=True)
153
+ except BaseException:
154
+ (work/'run.failed').write_text(traceback.format_exc());raise
reproduce/source/btx53/arvq88/docs/expert_experiments.md ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Per-expert books, alternating updates, and ARVQ allocation
2
+
3
+ These changes are experimental. Production v2 defaults remain shared books and
4
+ fixed indices. No new weights are uploaded by the ablation runner.
5
+
6
+ The serving-agent prompt is in `per_expert_handoff.md`. Version 3 changes only
7
+ the packed codebook tensor from `[512]` to `[E,512]` per layer/projection.
8
+ The paired index tensor, block scales and single projection global scale retain
9
+ the v2 layout. Expert slots follow the ordered cold-expert manifest.
10
+
11
+ `pv.py --reassign-every 40` interleaves continuous updates with LDLQ discrete
12
+ proposals computed from training activation Hessians. Proposals are accepted
13
+ only if a fixed training subset improves in routed cold-output error. This is
14
+ an alternating optimizer, not a reproduction of the PV-Tuning paper. Best-state
15
+ snapshots include code indices, so validation selection and audit fallback
16
+ restore complete models, not just codebooks. Rare experts with fewer than 32
17
+ training rows are skipped during discrete proposals rather than using held-out
18
+ rows. This can limit improvement for rare experts.
19
+
20
+ `fit.py` accepts `codebook_scope: expert` in run_config.json and learns separate
21
+ books for each expert with a common projection global scale. Fitting filters
22
+ merged captures to `pv_split == 0` when labels are present.
23
+
24
+ Run the controlled pilot:
25
+
26
+ ```bash
27
+ LD_LIBRARY_PATH=/tmp/compat13/usr/local/cuda-13.2/compat \
28
+ .venv/bin/python -u btx53/arvq88/experiment.py \
29
+ --base /tmp/glm53-vision-trace-refit \
30
+ --workdir /tmp/glm53-arvq-expert-ablation --layers 3,40,70
31
+ ```
32
+
33
+ This waits for all selected GPUs to be free, fits independent books, then runs
34
+ shared-fixed/shared-alternating/expert-fixed/expert-alternating variants with
35
+ identical source, split captures, expert selection and continuous step budgets.
36
+ Alternating variants have additional discrete-search work; this is not an
37
+ iso-compute benchmark. Outputs are comparison.json and packed/<variant>/ shards.
38
+ The original shared fit is reused as a baseline; the per-expert fit is fresh.
39
+ No production promotion is automatic. Native serving parity and task evaluation
40
+ are owned by the serving agent. The reused audit corpus is not a new independent
41
+ benchmark; do not tune repeatedly against audit results.
42
+
43
+ ARVQ allocation is available through these new pipeline arguments:
44
+
45
+ ```bash
46
+ .venv/bin/python btx53/pipeline.py \
47
+ --src zai-org/GLM-5.3 --dst jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-hybrid \
48
+ --workdir /tmp/glm53-arvq-expert-full \
49
+ --calib /tmp/glm53-vision-trace-refit/trace_capture/acts \
50
+ --calib-tokens /tmp/glm53-vision-trace-refit/train/tokens \
51
+ --codebook-scope expert --pv-reassign-every 40 --reap-metric arvq \
52
+ --mm-acts /tmp/glm53-vision-trace-refit/capture-mm-rlianv4b/acts_mm
53
+ ```
54
+
55
+ Run a full campaign only after the pilot and serving tests warrant it; the above
56
+ command is documented, not launched. It fits candidate ARVQ weights for all 256
57
+ experts in all 75 layers, fetches missing original FP8 experts, scores candidates
58
+ on both text and multimodal captures, then selects exactly 5750 hot experts using
59
+ the 75/25 layer-normalized blend, floor 8, cap 176. The final initial weights are
60
+ an exact subset of those candidates; allocation is followed by alternating PV.
61
+
62
+ For each expert the new score is:
63
+
64
+ max(0, sum_t g[t,e]^2 * (||ARVQ_e(x_t)-FP8_e(x_t)||^2
65
+ - ||NVFP4_e(x_t)-FP8_e(x_t)||^2))
66
+
67
+ Both quantized candidates use four-plane / FP16-boundary emulation, not native
68
+ SM120 MMA. The score measures the local benefit of retaining the NVFP4 expert,
69
+ with no AQLM error proxy. It omits cross-expert error covariance, routing changes,
70
+ and downstream error accumulation. Both modalities are required; an empty or
71
+ nonpositive modality layer fails instead of silently using text-only allocation.
72
+
73
+ The initial implementation does not use full-model cross-entropy/distillation
74
+ and does not establish end-to-end quality. The old corpus captures use fixed
75
+ teacher routing and 2048-token windows. Independent reasoning/task evaluations
76
+ remain necessary even if all layer audits pass.
reproduce/source/btx53/arvq88/docs/gradient_pv_recipe.md ADDED
@@ -0,0 +1,180 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Proposed Vision ARVQ v3 tuning recipe
2
+
3
+ Prepared 18 September 2026 (Melbourne). Status: plan, not a production restart.
4
+ The prior full-run driver and publisher were stopped at the user's request.
5
+ Initial fits, REAP allocation, local results, and published weights are preserved.
6
+
7
+ ## Findings verified against code and logs
8
+
9
+ - Claude transcript `d7bce691-1871-46a6-af7a-61eae60194fc.jsonl`,
10
+ 2026-09-01T22:31:07Z: AQLM converge optimized Hessian-weighted weight error;
11
+ subsequent PV optimized layer output MSE with fixed indices.
12
+ - `tools/capture53/pv_tune53.py` inherits `tools/pv_tune.py`: per-layer,
13
+ expert-coordinate gradient updates; cached routing; NVFP4-dequant teacher;
14
+ validation best restore. Its `reseed()` body is a no-op. This is not the
15
+ full-model algorithm in the PV-Tuning paper.
16
+ - `PLAN53.md` and `tools/capture53/eval_hybrid53.py`: the old pipeline also
17
+ performed full-model, layer-streamed perplexity evaluation. The old evaluator
18
+ decodes AQLM and needs an ARVQ v3 loader and arithmetic changes.
19
+ - Current ARVQ trains all codebooks and per-128-weight scales. Per expert:
20
+ 8,192 dictionary values + 294,912 scales; 13,450 cold experts expose
21
+ 3,966,566,400 scale values. Low storage precision is not immunity to overfit.
22
+ - Current merged capture: 98,304 rows/layer, 32,768 each for train/val/audit.
23
+ It contains x, topk_ids, topk_weights, pv_split, but no per-row prompt IDs.
24
+ - Current objective matches cold FP8 contributions only. The old AQLM code
25
+ explicitly subtracted frozen student-hot contributions from the full target.
26
+ With FP8 reference versus NVFP4 hot weights these are not interchangeable.
27
+ - New sparse gradient prototype passes the 42-test package suite, including
28
+ gradient direction, trust limit, training-row isolation and improvement tests.
29
+ A real layer-3/two-expert/two-step smoke completed packing/reload audit, but
30
+ rejected all 8 expert/projection proposals. This is not quality qualification.
31
+
32
+ ## Fixed scope and baselines
33
+
34
+ Preserve expert-specific books, v3 serialization, original FP8 weight source,
35
+ current 5,750/13,450 REAP allocation, and existing corpus. No refitting/re-download
36
+ is necessary for this optimizer change. Preserve all old PV artifacts, including
37
+ 12 previously published layers. Compare proposed replacements against both the
38
+ initial fit and the currently published layer. If both are candidates, select on
39
+ validation and reserve the new final audit for the frozen recipe.
40
+
41
+ This remains layer-local alternating optimization, not formal EM and not a
42
+ claim to reproduce full-model published PV-Tuning.
43
+
44
+ ## Phase 1: data and trustworthy objective
45
+
46
+ 1. Recapture or expand from the same training problems to an initial target of
47
+ 131,072 representative training rows/layer. Keep separate prompt-defined
48
+ validation; retain prompt ID, token position, domain, modality and capture
49
+ provenance per row. Use reservoir/stratified sampling across conversations,
50
+ not the first rows or unrestricted token-level resplitting.
51
+ 2. Log routed token counts AND distinct prompt counts for every expert. Tentative
52
+ minimum for expert-specific discrete tuning: 256 training routes from at least
53
+ 16 prompts; below this freeze indices and scale corrections, and initially
54
+ freeze its books too. Pilot coverage decides whether more capture is needed.
55
+ 3. Include multimodal TRAINING captures, not just multimodal REAP scores. Start
56
+ with the existing 75/25 text/MM mixture as a hypothesis, report both domains
57
+ separately, and do not treat previously allocation-used MM rows as held-out.
58
+ 4. Build full routed FP8 reference outputs, subtract frozen student-hot output,
59
+ then fit the cold sum to that residual. Report both full routed error and
60
+ cold-only error. Match actual hot arithmetic or clearly label the emulator.
61
+ 5. Normalize training MSE with a fixed training-set reference energy rather than
62
+ changing denominators per minibatch. Preserve natural routing weights;
63
+ expert-enriched samples require sampling corrections to avoid silently
64
+ changing the deployment objective.
65
+ 6. Audit forward arithmetic against the serving commit reported as b8db885d9,
66
+ including routing weights, intermediate rounding, scale layout, cold slots,
67
+ accumulation and TP reductions. A100 surrogate gradients are acceptable;
68
+ native forward parity needs the SM120 environment, not a claim of bit parity.
69
+
70
+ ## Phase 2: constrained continuous updates
71
+
72
+ - Start with books + regularized per-output-row multiplicative scale corrections.
73
+ Fold corrections into the existing E4M3 per-128 scales at export: no format or
74
+ kernel change. Freeze projection-global scales to remove avoidable scale
75
+ ambiguity. This exposes 10,240 row corrections per expert instead of 294,912
76
+ independent block-scale adjustments.
77
+ - Use FP32 latent parameters, FP4/E4M3 quantized forward and STE backward.
78
+ - Penalize drift from initial quantized weights/books and log-scale corrections;
79
+ use stronger shrinkage for poorly supported experts. Do not describe this as
80
+ a guarantee. The tiny prototype's uniform latent anchor is only scaffolding.
81
+ - Pilot conservative book learning rates around 3e-4 versus the prior 3e-3,
82
+ and row-log-scale LR around 1e-4, with warmup then decay. These are starting
83
+ hypotheses, not transferred AQLM hyperparameters: parameter units differ.
84
+ - Effective batch 1,024 through four 256-row microbatches; measure wall time.
85
+ Unfreeze individual block scales only if a later bounded ablation establishes
86
+ an independent-validation benefit. Capacity should be earned by evidence.
87
+
88
+ ## Phase 3: output-gradient discrete updates
89
+
90
+ - After a short continuous warmup, propose every 20 continuous updates initially.
91
+ - Compute gradients of the SAME routed-output objective with respect to one
92
+ expert's decoded gate/up or down matrix at a time, with other experts fixed.
93
+ The full routed residual retains interactions between experts.
94
+ - Precondition/damp the gradient and construct a temporary nearby target weight.
95
+ Search code pairs against this gradient-stepped target, NOT original FP8
96
+ matrices. Weight-space projection itself is legitimate: the target matters.
97
+ - Use beam search over both books (small beam e.g. 4, retain current pair),
98
+ shortlist by predicted improvement and consider at most 0.1% of groups per
99
+ expert/projection initially. Hard ceiling 1%; also bound decoded weight change.
100
+ - Backtrack the target step and/or number of changed groups when actual loss
101
+ worsens. Raw gradient direction plus a global norm constraint was insufficient
102
+ in the smoke test. No forced replacements and no all-layer all-or-nothing sweep.
103
+ - Accept each expert/projection separately using actual quantized forward loss
104
+ on proposal-training rows and a second rotating training check batch. Never
105
+ query validation to accept individual proposals. Both batches remain training.
106
+ - Exploit the down projection's linearity while gate/up is fixed: for a candidate
107
+ compute its routed output delta and exact squared-loss change
108
+ `2 * <residual, delta_y> + ||delta_y||^2`. This is a more faithful ranking
109
+ than weight distance. Batched changes still require a combined check because
110
+ their output deltas interact. Gate/up proposals require nonlinear forward checks.
111
+ - For gate/up changes recompute downstream activations. Following acceptance,
112
+ refresh gradients/residuals before the next projection. Restore codes and any
113
+ affected optimizer state on rollback; codebooks/scales are fixed in this step.
114
+ - Stream one expert/projection's dense gradient/temporary target. Do not allocate
115
+ persistent dense Adam moments or shadow weights for all cold experts: those
116
+ are impractical at this model size. Sparse/factorized state needs memory tests.
117
+
118
+ ## Phase 4: stopping, validation and final audit
119
+
120
+ - Report fixed training-probe and validation loss every 25-32 continuous steps,
121
+ by domain and at prompt level. Include accepted proposals, actual/predicted
122
+ improvement, quantized value changes, and coverage. Batch loss alone is not
123
+ evidence of a train/validation gap.
124
+ - First pilot uses 200 steps for a matched comparison. For the chosen full recipe
125
+ use a bounded 2-4 passes over the expanded training cache, rather than assuming
126
+ 200 steps is convergence. Always retain the best validation state.
127
+ - Do not stop for a short plateau. After >=2 passes, require sustained training
128
+ improvement AND validation deterioration over >=3 checks, exceeding uncertainty
129
+ from paired prompt-level resampling. Separate compute-budget exhaustion and
130
+ plateau from demonstrated overfitting. Small fixed percentage thresholds in the
131
+ prototype are provisional, not statistical proof.
132
+ - Existing audit data has already influenced several accepted runs. Build a new
133
+ disjoint final audit and do not inspect it during optimizer/recipe selection.
134
+ One final pass/fail; do not tune until that audit passes.
135
+
136
+ ## Phase 5: pilot, rollout and full-model verification
137
+
138
+ 1. Three full representative layers (3,40,70), baseline initial/current weights.
139
+ 2. Compare continuous-only against continuous + gradient-guided indices with
140
+ identical data, scale parameterization, regularization and step budget. This
141
+ isolates whether discrete optimization actually helps; no shared-book trials.
142
+ 3. Qualify on held-out error including text/MM slices, actual serialized forward,
143
+ proposal improvement, no malformed tensors, deterministic rollback and peak
144
+ memory/runtime. Rejection rate alone is neither success nor failure.
145
+ 4. Launch all 75 only after the pilot demonstrates a held-out benefit or useful
146
+ nonregression at realistic cost. Incremental uploads retain truthful recipe
147
+ provenance and current validation status; no stale old-run result reuse.
148
+ 5. Port the existing streamed evaluator to ARVQ v3. Compare NVFP4 donor, initial
149
+ per-expert checkpoint, and final candidate on identical held-out/neutral text
150
+ and multimodal checks. Weight-only decoded evaluation is a diagnostic; the
151
+ activation-quantized forward and native SM120 tests remain required.
152
+ 6. Measure teacher-input versus student-input activation/routing drift. If local
153
+ improvements fail to transfer, use one sequential refinement sweep: feed each
154
+ layer with already-quantized upstream student activations, pair with reference
155
+ trajectories and recompute appropriate routing/targets. Preserve held-out
156
+ prompt separation. Do not recapture the teacher and label it student capture.
157
+ 7. Consider short multi-layer reconstruction or full-model KL distillation only
158
+ if measured cross-layer drift justifies it and a memory/runtime prototype fits.
159
+ Independent layer tuning cannot be claimed to minimize end-to-end loss.
160
+
161
+ ## Code scope
162
+
163
+ - `gradient_indices.py`: damped/beam proposals, backtracking, two-batch acceptance,
164
+ memory-bounded state, exact per-expert rollback and diagnostics. Prototype exists.
165
+ - `pv.py`: target residual, restricted scale parameterization, fixed normalization,
166
+ gradient accumulation, regularization, uncertainty-aware stopping. Partial
167
+ prototype exists; not approved as the final recipe by its smoke result.
168
+ - New `pv_data.py` / capture extension: prompt provenance, coverage and multimodal
169
+ partitions, larger representative reservoirs, student-trajectory capture.
170
+ - `validation.py`: per-prompt paired metrics, separate overfit versus plateau,
171
+ fresh final audit and baseline comparisons.
172
+ - `pv_campaign.py`: recipe hash, unique workdir, pilot selection and true optimizer
173
+ resume. Restart from initial fits by default; keep old tuned artifacts intact.
174
+ - `incremental_publish.py` / `write_card`: truthful optimizer/coverage/stopping
175
+ provenance and mixed-revision status; no old Hessian method described as new PV.
176
+ - New `eval_arvq53.py`: adapt `tools/capture53/eval_hybrid53.py`, with explicit
177
+ weight-only versus activation-emulated modes and serving parity fixtures.
178
+
179
+ No new completion ETA until the selected full-layer pilot measures iteration,
180
+ proposal, validation and upload costs. All user-facing times: Melbourne.
reproduce/source/btx53/arvq88/docs/per_expert_handoff.md ADDED
@@ -0,0 +1,78 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Implement per-expert ARVQ 8+8 codebook support in the active GLM serving fork.
2
+
3
+ Coordinate with the fitting agent in /home/coder/git/glm52. That agent owns fitting,
4
+ alternating index updates, ARVQ-specific REAP scoring, and checkpoint export. You own
5
+ the serving loader/kernel and actual SM120 tests. Do not duplicate fitter changes.
6
+
7
+ Context
8
+ - Existing 8+8 checkpoints share two FP4 books per (layer, projection).
9
+ - New experiments give each cold expert its own pair of 256 x 8 FP4 books.
10
+ - gate/up (w13) share a pair; down (w2) has a separate pair.
11
+ - New tensor layout is an explicit version 3 format. Existing version 2 must keep
12
+ working unchanged. The stale /tmp/arvq-inspect2 checkout only supported 8+7 at
13
+ inspection time: work from the actual current 8+8 serving branch, not that clone.
14
+ - No new production checkpoint is ready yet. Start with synthetic parity fixtures.
15
+
16
+ Exact proposed format contract (fitter will export this)
17
+ config.quantization_config.arvq, also under text_config when present:
18
+ format: rvq256_256x8_expert
19
+ version: 3
20
+ codebook_scope: expert
21
+ codebook_sizes: [256, 256]
22
+ activation_planes: 4
23
+ weight_scale_group: 128
24
+
25
+ For each layer L, prefix model.layers.L.mlp.experts.arvq_{w13|w2}_:
26
+ packed: uint32 [E, N/16, K/64, 64] (identical to v2)
27
+ codebooks: uint32 [E, 512] (NEW expert axis)
28
+ scales: uint8 [E, N/16, K/128, 16] (identical to v2)
29
+ global: float32 [1] (identical to v2; not per expert)
30
+ E is the number of cold experts in this layer.
31
+ w13: N=4096, K=6144. w2: N=6144, K=2048.
32
+ Each book row packs 8 FP4 E2M1 values into one uint32: component j occupies
33
+ bits [4*j,4*j+3]. Codes 0..7 represent [0,.5,1,1.5,2,3,4,6], bit 3 is sign.
34
+ Rows 0..255 are book 0; rows 256..511 are book 1.
35
+ The representation is W[e,r,k] = global * scale[e,r,k//128] *
36
+ (C[e,a[e,r,k//8],k%8] + C[e,256+b[e,r,k//8],k%8]).
37
+ Index pairs retain a | (b << 8), packed in the existing v2 MMA-fragment order.
38
+ Use the existing v2 row/column packing, not a new natural-order interpretation.
39
+ See btx53/arvq88/pack.py for the exporter/reference decoder.
40
+
41
+ Critical expert mapping
42
+ - Book row e is the COLD SLOT, matching packed[e], scales[e], and the ordered
43
+ cold_expert_ids manifest. It is not the original global expert ID.
44
+ - TP shards retain the full local codebook pair for every locally represented
45
+ cold expert; books are not sliced along input/output axes. Existing N/K index
46
+ and scale sharding remains unchanged.
47
+ - EP/expert reordering must reorder codebooks with the other expert tensors.
48
+ - Preserve the shared [512] v2 codebook path. Reject incompatible format/shape
49
+ combinations rather than silently interpreting expert books as shared.
50
+
51
+ Kernel changes
52
+ - Select cb + cold_slot * 512 for v3; shared cb base for v2.
53
+ - Keep FP4 two-book MMA algebra and four activation planes unchanged.
54
+ - Audit all direct, LUT, fused, prefill/decode, and CUDA-graph paths. A lookup table
55
+ computed once from layer-shared books cannot be reused across different experts.
56
+ Make caches/LUT keys include the relevant expert or dispatch to a correct path.
57
+ - Account for any shared-memory sizing that assumed exactly one shared book pair.
58
+ - No additional per-weight index bits are required; actual throughput/cache effects
59
+ must be measured, not inferred from unchanged MMA count.
60
+
61
+ Required tests before claiming correctness
62
+ 1. Two experts with identical indices/scales and deliberately DIFFERENT books must
63
+ decode to different known outputs. Exercise permuted original expert IDs.
64
+ 2. Compare packed decode to an independent CPU oracle, including a,b=0 and 255,
65
+ negative FP4 values, extreme valid scales, both projections, and TP slicing.
66
+ 3. Duplicated per-expert books must agree with the existing shared-book path within
67
+ the same kernel's numerical tolerance; test full MoE outputs and routing weights.
68
+ 4. On SM120 compare gate/up outputs, FP16 SwiGLU, down outputs, and summed MoE
69
+ outputs to the FP4-plane reference on identical inputs. Report max abs and
70
+ relative L2 discrepancies for TP1 and actual deployment TP.
71
+ 5. Test both decode and prefill, graphs on/off, expert permutations, and LUT/direct
72
+ dispatch. Benchmark latency/throughput and memory against version 2.
73
+ 6. Run the actual failing model tests once the fitting agent supplies candidate
74
+ weights. Offline layer audit success is not evidence of end-to-end quality.
75
+
76
+ Scope: implement and test serving support; return commit/patch, commands, parity
77
+ results, and benchmark results. Do not upload, overwrite, or deploy a production
78
+ checkpoint. Report hardware-dependent checks you cannot run as untested.
reproduce/source/btx53/arvq88/docs/sequential_pilot.md ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Two-block reference-trajectory pilot
2
+
3
+ Work directory: `/tmp/glm53-sequential-pilot`.
4
+ Driver: `python btx53/arvq88/sequential_pilot.py --work /tmp/glm53-sequential-pilot`.
5
+ No Hugging Face publication is performed by this pilot.
6
+
7
+ The first routed layers are 3 and 4. Frozen donor layers 0–2 generate the prefix.
8
+ Both reference and student use this prefix and unchanged donor attention, router,
9
+ shared experts and norms. Reference routed experts in layers 3/4 come from the
10
+ original FP8 source. This is a controlled hybrid reference, not a claim to run
11
+ all-original FP8 nonexpert weights or native SM120 kernels.
12
+
13
+ Capture uses 64 training, 16 validation, 16 development-audit windows of 1,024
14
+ tokens, sampled deterministically from the existing prompt-separated corpus.
15
+ Within-split windows are not guaranteed to represent distinct prompts. The
16
+ historical audit split is development data, not a fresh untouched final test.
17
+ Attention is dense-equivalent at this context length (<= index_topk).
18
+
19
+ Layer 3 has identical same-input/reference-trajectory targets and is trained once.
20
+ Its retained serialized outputs become the student's layer-4 inputs. The
21
+ reference layer-3 outputs independently propagate into reference layer 4.
22
+ Layer 4 then receives two matched 200-step runs from identical initial weights:
23
+
24
+ - `same_input`: FP8 reference experts on the student's own block input.
25
+ - `reference`: the original reference trajectory's complete block output.
26
+
27
+ Both subtract the frozen student's residual/attention/shared/hot contributions
28
+ before fitting the cold sum. Loss includes the complete block residual mismatch.
29
+ The same fixed loss normalization, training minibatches, seed, optimizer settings,
30
+ validation split, allocation and initial codebooks are used in both arms.
31
+
32
+ Eight GPUs own disjoint cold-expert subsets within each layer. Their partial
33
+ routed outputs are summed, and each rank differentiates the same coupled loss
34
+ with respect to its local experts. No averaging of local expert targets replaces
35
+ this objective. Continuous updates use effective batch 1,024 in four microbatches.
36
+ Books use LR 3e-4, row-log-scale corrections 1e-4, and an initialization anchor.
37
+ Block scales and projection globals are frozen; row corrections fold into the
38
+ existing E4M3 scale layout. Experts with <256 routed training rows are frozen.
39
+
40
+ Every 20 steps, expert-local output gradients propose sparse code changes.
41
+ Per expert/projection proposals are generated concurrently; their output deltas
42
+ are then accepted in deterministic rank order against the up-to-date coupled
43
+ residual. Prefix backtracking tries smaller changes. Acceptance requires lower
44
+ error on the proposal-training batch and nonregression on a separate training
45
+ check batch. Validation never accepts individual proposals.
46
+
47
+ Quantized forward emulates FP4 cold books, E4M3 weight scales, four activation
48
+ planes and FP16 SwiGLU boundaries. Hot experts use NVFP4-dequantized weights with
49
+ float32 GEMMs. Cross-GPU sums and donor BF16 backbone are surrogates for native
50
+ serving numerics; record this limitation with results.
51
+
52
+ Validation every 25 steps retains the best state. Budget is 200 steps. Early
53
+ stopping is allowed only after >=150 steps and 3 consecutive checks where
54
+ training improves >=0.5% while validation worsens >=0.5% versus the validation
55
+ best. This is a conservative heuristic, not statistical proof of overfitting.
56
+
57
+ `comparison.json` compares both layer-4 candidates against the SAME reference
58
+ trajectory on held-out windows; includes a paired-window bootstrap diagnostic.
59
+ Output errors are full-block relative L2, so they are not numerically comparable
60
+ to earlier cold-contribution relative errors (~0.2–0.3).
61
+
62
+ Preflight: 8-GPU one-step layer-3 smoke completed; 304/388 proposals accepted,
63
+ 118,355 groups changed. Held-out full-block relative L2 0.0054971166 ->
64
+ 0.0054560048. Saved/reloaded forward matched. These are smoke measurements,
65
+ not results of the 200-step target comparison.
reproduce/source/btx53/arvq88/encoder.py ADDED
@@ -0,0 +1,321 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Activation-Hessian-tuned ARVQ fit for one (layer, projection).
2
+
3
+ Codebooks (c0[256,8], c1[256,8], FP4-constrained) are SHARED across all cold
4
+ experts of a (layer, projection), so this is a joint per-layer fit:
5
+
6
+ 1. per expert: raw-basis Hessian H_e, escalating-damp Cholesky hinv_e,
7
+ per-(row,128block) init scale s_e, global scalar (fixed).
8
+ 2. codebook EM on a Hessian-importance-weighted subsample of normalized
9
+ 8-dim weight groups (k-means warm start -> alternating refine, FP4 project
10
+ after every M-step). <- "fit codebooks to data" (a)
11
+ 3. per expert final assignment via a GPTQ/LDLQ error-feedback column sweep
12
+ with the fixed codebooks (plain-L2 inner assignment; cross-group feedback
13
+ carries the off-diagonal Hessian), then LS refit of the E4M3 scales.
14
+ <- "Hessian-aware code selection + error feedback" (b)
15
+
16
+ No incoherence rotation: Hessians and targets are in the raw weight basis
17
+ (the SM120 kernel feeds raw activations). Weighting is in the EM importance +
18
+ error feedback + scale LS, never a sqrt(diag(H)) whitening (the FP4 grid on the
19
+ codebook forbids rescaling the target space).
20
+ """
21
+ from __future__ import annotations
22
+
23
+ from dataclasses import dataclass
24
+
25
+ import torch
26
+ import torch.nn.functional as F
27
+
28
+ from arvqprep.pack import project_to_fp4
29
+
30
+
31
+ # --------------------------------------------------------------- Hessians
32
+ def hessian_fc1(x: torch.Tensor) -> torch.Tensor:
33
+ """gate/up share E[x x^T] over the H axis. x: [T, H] (routed tokens)."""
34
+ x = x.float()
35
+ return (x.t() @ x) / max(x.shape[0], 1)
36
+
37
+
38
+ def hessian_fc2(x: torch.Tensor, wg: torch.Tensor, wu: torch.Tensor) -> torch.Tensor:
39
+ """down E[m m^T], m = silu(x Wg^T) * (x Wu^T) over the I axis. x:[T,H]."""
40
+ x = x.float()
41
+ m = F.silu(x @ wg.float().t()) * (x @ wu.float().t()) # [T, I]
42
+ return (m.t() @ m) / max(x.shape[0], 1)
43
+
44
+
45
+ def cholesky_inv_upper(hess: torch.Tensor, damp_start: float = 1e-3) -> torch.Tensor:
46
+ """Escalating-damp upper Cholesky of H^-1 (same recipe as btxprep/encoder)."""
47
+ k = hess.shape[0]
48
+ mean_diag = hess.diagonal().mean().clamp(min=1e-12)
49
+ damp = damp_start
50
+ eye = torch.eye(k, device=hess.device, dtype=hess.dtype)
51
+ for _ in range(12):
52
+ try:
53
+ return torch.linalg.cholesky(
54
+ torch.linalg.inv(hess + damp * mean_diag * eye), upper=True)
55
+ except Exception: # noqa: BLE001
56
+ damp *= 4.0
57
+ raise RuntimeError("ARVQ Cholesky failed at maximum damp")
58
+
59
+
60
+ # ---------------------------------------------------------- codebook EM
61
+ def _nearest(x, c):
62
+ """argmin_k ||x - c_k||^2 (plain L2). x:[M,8], c:[K,8] -> [M]."""
63
+ return (c.square().sum(1)[None, :] - 2 * x @ c.t()).argmin(1)
64
+
65
+
66
+ def _wmeans(x, w, ids, count, prev):
67
+ """Weighted per-cluster mean of x (scalar sample weights w); keep prev if empty."""
68
+ sums = torch.zeros_like(prev)
69
+ sums.index_add_(0, ids, x * w[:, None])
70
+ wsum = torch.zeros(count, device=x.device, dtype=x.dtype)
71
+ wsum.index_add_(0, ids, w)
72
+ return torch.where(wsum[:, None] > 0, sums / wsum.clamp_min(1e-20)[:, None], prev)
73
+
74
+
75
+ def fit_codebooks(x, w, *, iters=8, seed=0):
76
+ """Weighted FP4 additive-RVQ codebook fit on a subsample.
77
+ x:[M,8] normalized group vectors, w:[M] importance weights.
78
+ Returns c0[256,8], c1[256,8] (FP4-on-grid), best weighted relative-L2."""
79
+ g = torch.Generator(device=x.device).manual_seed(seed)
80
+ perm = torch.randperm(x.shape[0], generator=g, device=x.device)
81
+ c0 = project_to_fp4(x[perm[:256]].clone())
82
+ i = _nearest(x, c0)
83
+ c1 = project_to_fp4(_kmeans(x - c0[i], 256, g))
84
+ j = _nearest(x - c0[i], c1)
85
+ wsum = (w * x.square().sum(1)).sum().clamp_min(1e-20)
86
+ best = None
87
+ for _ in range(iters):
88
+ i = _nearest(x - c1[j], c0)
89
+ c0 = project_to_fp4(_wmeans(x - c1[j], w, i, 256, c0))
90
+ j = _nearest(x - c0[i], c1)
91
+ c1 = project_to_fp4(_wmeans(x - c0[i], w, j, 256, c1))
92
+ err = ((w * (x - c0[i] - c1[j]).square().sum(1)).sum() / wsum).item()
93
+ if best is None or err < best[0]:
94
+ best = (err, c0.clone(), c1.clone())
95
+ err, c0, c1 = best
96
+ return c0, c1, err ** 0.5
97
+
98
+
99
+ def _kmeans(x, count, gen, iters=12):
100
+ c = x[torch.randperm(x.shape[0], generator=gen, device=x.device)[:count]].clone()
101
+ for _ in range(iters):
102
+ ids = _nearest(x, c)
103
+ sums = torch.zeros_like(c)
104
+ sums.index_add_(0, ids, x)
105
+ cnt = torch.bincount(ids, minlength=count).float().clamp_min(1)[:, None]
106
+ c = torch.where(torch.bincount(ids, minlength=count)[:, None] > 0, sums / cnt, c)
107
+ return c
108
+
109
+
110
+ # ----------------------------------------------------- per-expert sweep
111
+ def _assign(tn, c0, c0n, c1, c1n, refine):
112
+ """Plain-L2 additive assignment. tn:[N,8]. Returns a,b (long)."""
113
+ a = (c0n[None, :] - 2 * tn @ c0.t()).argmin(1)
114
+ r = tn - c0[a]
115
+ b = (c1n[None, :] - 2 * r @ c1.t()).argmin(1)
116
+ for _ in range(refine):
117
+ r = tn - c1[b]
118
+ a = (c0n[None, :] - 2 * r @ c0.t()).argmin(1)
119
+ r = tn - c0[a]
120
+ b = (c1n[None, :] - 2 * r @ c1.t()).argmin(1)
121
+ return a, b
122
+
123
+
124
+ def sweep_expert(W, hinv, c0, c1, s, glob, *, col_block, refine):
125
+ """LDLQ error-feedback assignment for one expert.
126
+
127
+ W:[N,K] fp32 target; hinv:[K,K] upper fp32; s:[N,K/128] fp32 (E4M3-valued)
128
+ scale; glob: scalar. Returns a,b uint8 [N,K/8] and reconstruction q [N,K]."""
129
+ N, K = W.shape
130
+ w = W.clone()
131
+ q = torch.empty_like(w)
132
+ a_full = torch.empty(N, K // 8, dtype=torch.uint8, device=W.device)
133
+ b_full = torch.empty(N, K // 8, dtype=torch.uint8, device=W.device)
134
+ c0n = c0.square().sum(1)
135
+ c1n = c1.square().sum(1)
136
+ for j0 in range(0, K, col_block):
137
+ j1 = min(j0 + col_block, K)
138
+ blk = w[:, j0:j1]
139
+ qhat = torch.empty_like(blk)
140
+ for gg in range((j1 - j0) // 8):
141
+ c = j0 + gg * 8
142
+ sden = (glob * s[:, c // 128]).clamp_min(1e-20)[:, None] # [N,1]
143
+ tn = blk[:, gg * 8:gg * 8 + 8] / sden
144
+ a, b = _assign(tn, c0, c0n, c1, c1n, refine)
145
+ qhat[:, gg * 8:gg * 8 + 8] = sden * (c0[a] + c1[b])
146
+ a_full[:, c // 8] = a.to(torch.uint8)
147
+ b_full[:, c // 8] = b.to(torch.uint8)
148
+ q[:, j0:j1] = qhat
149
+ err = (blk - qhat) @ torch.linalg.inv(hinv[j0:j1, j0:j1])
150
+ if j1 < K:
151
+ w[:, j1:] -= err @ hinv[j0:j1, j1:]
152
+ return a_full, b_full, q
153
+
154
+
155
+ def refit_scales(W, a, b, c0, c1, glob, hdiag, K):
156
+ """LS per-(row,128block) scale given codes; E4M3-quantized. Returns [N,K/128]."""
157
+ N = W.shape[0]
158
+ u = (c0[a.long()] + c1[b.long()]).reshape(N, K // 8, 8) * glob # unit recon
159
+ Wg = W.reshape(N, K // 8, 8)
160
+ hd = hdiag.reshape(K // 8, 8)[None] # [1,K/8,8]
161
+ num = (hd * Wg * u).reshape(N, K // 128, 128).sum(-1)
162
+ den = (hd * u * u).reshape(N, K // 128, 128).sum(-1).clamp_min(1e-20)
163
+ s = (num / den)
164
+ s = s.clamp_min(0).to(torch.float8_e4m3fn).float()
165
+ return s
166
+
167
+
168
+ @dataclass
169
+ class EncodedLayerProj:
170
+ c0: torch.Tensor # [256,8] fp32 (FP4 grid)
171
+ c1: torch.Tensor # [256,8] fp32; [E,256,8] for expert scope
172
+ glob: float
173
+ a: torch.Tensor # [E,N,K/8] uint8
174
+ b: torch.Tensor # [E,N,K/8] uint8
175
+ s: torch.Tensor # [E,N,K/128] fp32 (E4M3-valued)
176
+ N: int
177
+ K: int
178
+ cb_rel_l2: float
179
+ recon_rel_fro: float # weight-space, Hessian-weighted, over experts
180
+ scale_dtype: str = "fp8_e4m3"
181
+ codebook_dtype: str = "fp4_grid" # 'fp4_grid' (v2/v3) or 'fp16' (free atoms)
182
+
183
+
184
+ def _e4m3(x):
185
+ return x.to(torch.float8_e4m3fn).float()
186
+
187
+
188
+ def fit_layer_projection(W, H, *, col_block=128, cb_iters=8, sweep_passes=2,
189
+ refine=1, subsample_per_expert=20000, seed=0,
190
+ device="cuda", verbose=False, global_scale=None,
191
+ codebook_scope="layer"):
192
+ """W: list of E weight targets [N,K] (fp16/fp32, CPU or GPU).
193
+ H: list of E Hessians [K,K] (fp32). Returns EncodedLayerProj."""
194
+ dev = torch.device(device)
195
+ E = len(W)
196
+ N, K = W[0].shape
197
+ gen = torch.Generator(device=dev).manual_seed(seed)
198
+ if codebook_scope not in ('layer','expert'):
199
+ raise ValueError('Unknown codebook scope')
200
+ if codebook_scope == 'expert':
201
+ # Retain one global scale per projection so the v2 index/scale layout
202
+ # and arithmetic stay unchanged. Only the dictionaries gain an axis.
203
+ if global_scale is None:
204
+ pool=[]
205
+ for weight in W:
206
+ rms=weight.to(dev).float().reshape(N,K//128,128).square().mean(-1).sqrt().flatten()
207
+ pick=torch.randperm(len(rms),generator=gen,device=dev)[:4096]
208
+ pool.append(rms[pick])
209
+ global_scale=float(torch.cat(pool).median().clamp_min(1e-8))
210
+ del pool,rms
211
+ results=[]
212
+ for e in range(E):
213
+ result=fit_layer_projection([W[e]],[H[e]],col_block=col_block,
214
+ cb_iters=cb_iters,sweep_passes=sweep_passes,refine=refine,
215
+ subsample_per_expert=subsample_per_expert,seed=seed+e,
216
+ device=device,verbose=verbose,global_scale=global_scale)
217
+ results.append(result)
218
+ return EncodedLayerProj(torch.stack([r.c0 for r in results]),
219
+ torch.stack([r.c1 for r in results]),float(global_scale),
220
+ torch.cat([r.a for r in results]),torch.cat([r.b for r in results]),
221
+ torch.cat([r.s for r in results]),N,K,
222
+ sum(r.cb_rel_l2 for r in results)/E,
223
+ sum(r.recon_rel_fro for r in results)/E)
224
+
225
+ # ---- Phase A: hinv, hdiag, init scale, global
226
+ hinv, hdiag, s = [], [], []
227
+ rms_pool = []
228
+ for e in range(E):
229
+ We = W[e].to(dev).float()
230
+ hd = H[e].to(dev).float().diagonal().clamp_min(1e-12)
231
+ hinv.append(cholesky_inv_upper(H[e].to(dev).float()).half())
232
+ hdiag.append(hd.half())
233
+ blk_rms = We.reshape(N, K // 128, 128).square().mean(-1).sqrt() # [N,K/128]
234
+ s.append(blk_rms) # tmp (pre-global)
235
+ rms_pool.append(blk_rms.flatten()[torch.randperm(N * (K // 128),
236
+ generator=gen, device=dev)[:4096]])
237
+ del We
238
+ glob = (float(torch.cat(rms_pool).median().clamp_min(1e-8).item())
239
+ if global_scale is None else float(global_scale))
240
+ if not __import__('math').isfinite(glob) or glob <= 0:
241
+ raise ValueError('Global scale must be finite and positive')
242
+ for e in range(E):
243
+ s[e] = _e4m3(s[e] / glob) # E4M3 scale
244
+
245
+ # ---- Phase B: importance-weighted subsample of normalized group vectors
246
+ xs, ws = [], []
247
+ for e in range(E):
248
+ We = W[e].to(dev).float().reshape(N, K // 8, 8)
249
+ sden = (glob * s[e]).clamp_min(1e-20) # [N,K/128]
250
+ sden = sden.repeat_interleave(16, dim=1) # [N,K/8]
251
+ tn = (We / sden[:, :, None]).reshape(-1, 8) # [N*K/8,8]
252
+ imp = (hdiag[e].float().reshape(K // 8, 8).mean(1)[None, :]
253
+ * (s[e].repeat_interleave(16, dim=1) ** 2)).reshape(-1) # [N*K/8]
254
+ m = tn.shape[0]
255
+ idx = torch.randperm(m, generator=gen, device=dev)[:subsample_per_expert]
256
+ xs.append(tn[idx])
257
+ ws.append(imp[idx])
258
+ del We, tn, imp
259
+ xs = torch.cat(xs)
260
+ ws = torch.cat(ws).clamp_min(1e-12)
261
+ c0, c1, cb_rel = fit_codebooks(xs, ws, iters=cb_iters, seed=seed)
262
+ if verbose:
263
+ print(f" codebook subsample rel-L2 {cb_rel:.4f} (M={xs.shape[0]})",
264
+ flush=True)
265
+ del xs, ws
266
+
267
+ # ---- Phase D: per-expert error-feedback sweep + scale refit
268
+ a_all = torch.empty(E, N, K // 8, dtype=torch.uint8)
269
+ b_all = torch.empty(E, N, K // 8, dtype=torch.uint8)
270
+ s_all = torch.empty(E, N, K // 128, dtype=torch.float32)
271
+ num_w = den_w = 0.0
272
+ for e in range(E):
273
+ We = W[e].to(dev).float()
274
+ hv = hinv[e].float()
275
+ hd = hdiag[e].float()
276
+ se = s[e]
277
+ a = b = q = None
278
+ for p in range(sweep_passes):
279
+ a, b, q = sweep_expert(We, hv, c0, c1, se, glob,
280
+ col_block=col_block, refine=refine)
281
+ if p + 1 < sweep_passes:
282
+ se = refit_scales(We, a, b, c0, c1, glob, hd, K)
283
+ se = refit_scales(We, a, b, c0, c1, glob, hd, K)
284
+ a, b, q = sweep_expert(We, hv, c0, c1, se, glob,
285
+ col_block=col_block, refine=refine)
286
+ a_all[e], b_all[e], s_all[e] = a.cpu(), b.cpu(), se.cpu()
287
+ d = (We - q)
288
+ # Half-stored Hessians can lose positive semidefiniteness. Evaluate
289
+ # against the stabilized positive metric used by the sweep instead:
290
+ # H_eff^-1 = U.T @ U, hence d H_eff d.T = ||d U^-1||^2.
291
+ num_w += positive_quadratic(d, hv)
292
+ den_w += positive_quadratic(We, hv)
293
+ del We, hv, q, d
294
+ hinv[e] = None
295
+ recon = (num_w / max(den_w, 1e-20)) ** 0.5
296
+ return EncodedLayerProj(c0.cpu(), c1.cpu(), glob, a_all, b_all, s_all,
297
+ N, K, cb_rel, recon)
298
+
299
+
300
+ def positive_quadratic(weight, inverse_cholesky):
301
+ """Nonnegative squared norm under the regularized fitting Hessian."""
302
+ z = torch.linalg.solve_triangular(inverse_cholesky.T, weight.T,
303
+ upper=False)
304
+ value = float(z.double().square().sum())
305
+ if not __import__('math').isfinite(value):
306
+ raise ValueError('Nonfinite regularized Hessian diagnostic')
307
+ return value
308
+
309
+
310
+ def reconstruct(enc: EncodedLayerProj, e: int, device="cpu") -> torch.Tensor:
311
+ """Decode one expert's weight [N,K] from stored codes/scale/codebooks."""
312
+ dev = torch.device(device)
313
+ c0, c1 = enc.c0.to(dev), enc.c1.to(dev)
314
+ if c0.ndim == 3:
315
+ c0, c1 = c0[e], c1[e]
316
+ a = enc.a[e].to(dev).long()
317
+ b = enc.b[e].to(dev).long()
318
+ s = enc.s[e].to(dev).repeat_interleave(16, dim=1) # [N,K/8]
319
+ cb = (c0[a] + c1[b]) # [N,K/8,8]
320
+ W = (enc.glob * s[:, :, None] * cb).reshape(enc.N, enc.K)
321
+ return W
reproduce/source/btx53/arvq88/experiment.py ADDED
@@ -0,0 +1,129 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Matched layer ablation for shared/expert books and fixed/alternating indices.
2
+
3
+ Waits for unoccupied GPUs; never stops another job. Exports local experimental
4
+ weights only. Full-model serving evaluation must precede a production campaign.
5
+ """
6
+ import argparse
7
+ import fcntl
8
+ import json
9
+ import os
10
+ import shutil
11
+ import subprocess
12
+ import sys
13
+ import time
14
+ import traceback
15
+ from pathlib import Path
16
+
17
+ ROOT=Path(__file__).resolve().parents[1]
18
+ sys.path[:0]=[str(ROOT),str(ROOT/'tools')]
19
+ from arvq88.inputs import write
20
+ from arvq88.jobs import run_jobs
21
+ from arvq88.pack import export_layer,sha
22
+ from arvq88.activation import ARITHMETIC
23
+
24
+
25
+ def summary(work,layers):
26
+ rows={}
27
+ for variant in ['shared_fixed','shared_alternating','expert_fixed','expert_alternating']:
28
+ rows[variant]={}
29
+ for L in layers:
30
+ p=work/variant/f'layer_{L:05d}'/'report.json'
31
+ if not p.exists():continue
32
+ r=json.loads(p.read_text())
33
+ rows[variant][str(L)]={k:r[k] for k in ['initial_output_rel','serialized_output_rel','initial_audit_rel','serialized_audit_rel','audit_passed','audit_candidate_accepted','steps_run','codebook_scope','indices_frozen','index_reassignment']}
34
+ write(work/'comparison.json',{'variants':rows,'scope':'Same cold allocation, corpus splits, source weights and PV budgets; isolated layer ablation. Not full-model task quality.',
35
+ 'allocation_next_step':'Use --reap-metric arvq --mm-acts with all-expert fitted candidates, after serving parity and task evaluation.'})
36
+
37
+
38
+ def main():
39
+ ap=argparse.ArgumentParser(description=__doc__)
40
+ ap.add_argument('--base',default='/tmp/glm53-vision-trace-refit')
41
+ ap.add_argument('--workdir',default='/tmp/glm53-arvq-expert-ablation')
42
+ ap.add_argument('--layers',default='3,40,70');ap.add_argument('--gpus',default='0,1,2,3,4,5,6,7')
43
+ ap.add_argument('--steps',type=int,default=200);ap.add_argument('--reassign-every',type=int,default=40)
44
+ ap.add_argument('--check',action='store_true')
45
+ a=ap.parse_args();base=Path(a.base);work=Path(a.workdir);work.mkdir(parents=True,exist_ok=True)
46
+ layers=[int(v) for v in a.layers.split(',')];gpus=[int(v) for v in a.gpus.split(',')]
47
+ if not layers or len(set(layers))!=len(layers) or any(L not in range(3,78) for L in layers):raise ValueError('Invalid layers')
48
+ if not gpus or len(set(gpus))!=len(gpus) or min(gpus)<0 or a.reassign_every<=0 or a.steps<40:raise ValueError('Invalid run budget')
49
+ config=json.loads((base/'config.json').read_text())
50
+ train=base/'train/capture/acts';captures=base/'trace_capture/acts'
51
+ for L in layers:
52
+ for p in [train/f'acts_layer{L}.pt',captures/f'acts_layer{L}.pt',base/'initial'/f'layer_{L:05d}/arvq-manifest.json']:
53
+ if not p.is_file():raise FileNotFoundError(p)
54
+ identity={'base':str(base),'assignment_sha256':sha(base/'assignment.json'),'source_sha256':sha(base/'source.json'),
55
+ 'layers':layers,'steps':a.steps,'reassign_every':a.reassign_every,'format':'rvq256_256x8_expert',
56
+ 'inputs_sha256':{str(p):sha(p) for L in layers for p in [train/f'acts_layer{L}.pt',captures/f'acts_layer{L}.pt',base/'initial'/f'layer_{L:05d}/arvq-manifest.json']},
57
+ 'code_sha256':{name:sha(ROOT/'arvq88'/name) for name in ['pv.py','alternating.py','encoder.py','pack.py','fit.py']}}
58
+ if (work/'binding.json').exists() and json.loads((work/'binding.json').read_text())!=identity:
59
+ raise ValueError('Experiment inputs or code changed; choose a fresh workdir')
60
+ if a.check:
61
+ print('PREFLIGHT OK',json.dumps(identity),flush=True);return
62
+ lock=open(work/'.experiment.lock','a');fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB)
63
+ write(work/'binding.json',identity)
64
+ if (work/'complete.json').exists():return
65
+ def status(stage,**kw):write(work/'status.json',{'stage':stage,'updated_utc':time.strftime('%Y-%m-%dT%H:%M:%SZ',time.gmtime()),**kw});print(stage,json.dumps(kw),flush=True)
66
+ try:
67
+ status('waiting_for_free_gpus',gpus=gpus,layers=layers)
68
+ while True:
69
+ rows=subprocess.check_output(['nvidia-smi','--query-gpu=index,memory.used','--format=csv,noheader,nounits'],text=True)
70
+ mem={int(i):int(m) for i,m in (r.split(',') for r in rows.splitlines())}
71
+ if any(g not in mem for g in gpus):raise ValueError('Selected GPU does not exist')
72
+ if all(mem[g]<1024 for g in gpus):break
73
+ time.sleep(30)
74
+ expert=work/'expert_initial';expert.mkdir(exist_ok=True)
75
+ for name in ['assignment.json','source.json']:shutil.copy2(base/name,expert/name)
76
+ config.update(workdir=str(expert),calib=str(train),codebook_scope='expert',do_upload=False)
77
+ write(expert/'run_config.json',config)
78
+ status('fitting_per_expert_books',layers=layers)
79
+ pending=[]
80
+ for L in layers:
81
+ folder=expert/'initial'/f'layer_{L:05d}'
82
+ path=folder/'arvq-manifest.json'
83
+ if path.exists():
84
+ m=json.loads(path.read_text());fit_id=m['source_identity']
85
+ expected=json.loads((expert/'assignment.json').read_text())['layers'][str(L)]['cold_5750']
86
+ if (m['cold_expert_ids']!=expected or fit_id['capture_sha256']!=sha(train/f'acts_layer{L}.pt')
87
+ or fit_id['source']!=json.loads((expert/'source.json').read_text())
88
+ or fit_id['settings'].get('codebook_scope')!='expert'):
89
+ raise ValueError('Completed fit inputs do not match experiment')
90
+ for entry in m['files'].values():
91
+ if sha(folder/entry['file'])!=entry['sha256']:raise ValueError('Completed fit weights changed')
92
+ print('REUSE VALIDATED FIT',L,flush=True)
93
+ else:pending.append(L)
94
+ run_jobs([(f'fit-{L}',lambda gpu,L=L:[sys.executable,'-u',str(ROOT/'arvq88/fit.py'),str(expert),str(L),f'cuda:{gpu}']) for L in pending],gpus,work/'logs')
95
+ jobs=[]
96
+ for variant in ['shared_fixed','shared_alternating','expert_fixed','expert_alternating']:
97
+ initial=expert/'initial' if variant.startswith('expert') else base/'initial'
98
+ reassign=a.reassign_every if variant.endswith('alternating') else 0
99
+ for L in layers:
100
+ out=work/variant/f'layer_{L:05d}'
101
+ if (out/'report.json').exists():continue
102
+ jobs.append((f'{variant}-{L}',lambda gpu,L=L,out=out,initial=initial,reassign=reassign:[
103
+ sys.executable,'-u',str(ROOT/'arvq88/pv.py'),'--arvq-dir',str(initial),
104
+ '--layer',str(L),'--experts','0','--steps',str(a.steps),'--device',f'cuda:{gpu}',
105
+ '--src-cache',config['src_cache'],'--source-meta',str(base/'source.json'),
106
+ '--calib',str(captures),'--codebook-lr','0.003','--scale-lr','0.002',
107
+ '--arithmetic',ARITHMETIC,'--reassign-every',str(reassign),'--out',str(out)]))
108
+ status('matched_pv_ablation',variants=4,layers=layers)
109
+ run_jobs(jobs,gpus,work/'logs')
110
+ status('exporting_experimental_weights')
111
+ for variant in ['shared_fixed','shared_alternating','expert_fixed','expert_alternating']:
112
+ initial=expert/'initial' if variant.startswith('expert') else base/'initial'
113
+ for L in layers:
114
+ out=work/variant/f'layer_{L:05d}';r=json.loads((out/'report.json').read_text())
115
+ manifest=json.loads((initial/f'layer_{L:05d}/arvq-manifest.json').read_text())
116
+ stage=work/'exports'/variant/f'layer_{L:05d}';stage.mkdir(parents=True,exist_ok=True)
117
+ for proj in ['w13','w2']:
118
+ target=stage/f'{proj}.pt'
119
+ if not target.exists():os.link(out/'offline_weights.pt',target)
120
+ write(stage/'arvq-manifest.json',{**manifest,'fit':variant,'pv_report':r})
121
+ export_layer(stage,work/'packed'/variant,L)
122
+ summary(work,layers)
123
+ write(work/'complete.json',{'uploaded':False,'comparison':str(work/'comparison.json'),'serving_validation':'pending'})
124
+ status('complete',uploaded=False)
125
+ except BaseException:
126
+ (work/'run.failed').write_text(traceback.format_exc());status('failed',report=str(work/'run.failed'));raise
127
+
128
+
129
+ if __name__=='__main__':main()
reproduce/source/btx53/arvq88/final_publish.py ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Publish final PV artifacts, verify them, then remove obsolete repo files."""
2
+ import hashlib,json
3
+ from pathlib import Path
4
+ from .pack import sha
5
+ from .inputs import write
6
+ from .validation import require_audit
7
+
8
+ def inventory(root):
9
+ result={}
10
+ for p in sorted(root.rglob('*')):
11
+ if not p.is_file():continue
12
+ name=str(p.relative_to(root))
13
+ # upload_folder ignores .cache/huggingface/ by default, so it never reaches the remote;
14
+ # including it here would make verify() fail on a file that was intentionally not uploaded.
15
+ if name.startswith('.cache/'):continue
16
+ size=p.stat().st_size
17
+ entry={'size':size,'sha256':sha(p)}
18
+ if size<10*1024*1024:
19
+ data=p.read_bytes();entry['blob_id']=hashlib.sha1(f'blob {len(data)}\0'.encode()+data).hexdigest()
20
+ result[name]=entry
21
+ return result
22
+
23
+ def verify(api,repo,revision,expected):
24
+ files={s.rfilename:s for s in api.model_info(repo,revision=revision,files_metadata=True).siblings}
25
+ for name,e in expected.items():
26
+ if name not in files:raise RuntimeError(f'Remote file missing: {name}')
27
+ s=files[name]
28
+ if s.size!=e['size']:raise RuntimeError(f'Remote size mismatch: {name}')
29
+ if s.lfs:
30
+ if s.lfs.sha256!=e['sha256']:raise RuntimeError(f'Remote SHA mismatch: {name}')
31
+ elif s.blob_id!=e.get('blob_id'):raise RuntimeError(f'Remote Git blob mismatch: {name}')
32
+ return files
33
+
34
+ def publish(args,hub,work):
35
+ ck=work/'checkpoint';pv=json.loads((work/'pv_report.json').read_text());gates=json.loads((ck/'gates_report.json').read_text())
36
+ if not pv['passed'] or set(pv['layers_completed'])!=set(range(3,78)) or not gates['passed']:raise RuntimeError('Final PV/gates incomplete')
37
+ if set(map(int,pv.get('layers',{})))!=set(range(3,78)):raise RuntimeError('Missing per-layer audits')
38
+ for report in pv['layers'].values():require_audit(report)
39
+ for name,digest in {**gates['cold_file_sha256'],**gates['metadata_sha256']}.items():
40
+ if sha(ck/name)!=digest:raise RuntimeError('Checkpoint changed after gates')
41
+ return publish_verified(args,hub,work,'complete GLM-5.3 NVFP4 / ARVQ 8+8 PV-tuned hybrid')
42
+
43
+ def publish_verified(args,hub,work,description):
44
+ from huggingface_hub import CommitOperationDelete
45
+ ck=work/'checkpoint'
46
+ expected=inventory(ck);write(work/'final_upload_inventory.json',expected)
47
+ from huggingface_hub.errors import RepositoryNotFoundError
48
+ try:before=hub.api.model_info(args.dst,files_metadata=True)
49
+ except RepositoryNotFoundError:
50
+ hub.api.create_repo(repo_id=args.dst,repo_type='model',exist_ok=True)
51
+ before=hub.api.model_info(args.dst,files_metadata=True)
52
+ write(work/'repo_before_final.json',{'revision':before.sha,'files':[s.rfilename for s in before.siblings]})
53
+ commit=hub.api.upload_folder(repo_id=args.dst,folder_path=str(ck),revision='main',parent_commit=before.sha,commit_message='Publish '+description)
54
+ files=verify(hub.api,args.dst,commit.oid,expected)
55
+ write(work/'final_upload_verified.json',{'revision':commit.oid,'files':len(expected)})
56
+ revision=commit.oid
57
+ if args.clean_repo:
58
+ obsolete=sorted(set(files)-set(expected)-{'.gitattributes'})
59
+ write(work/'repo_cleanup_plan.json',{'verified_revision':revision,'delete':obsolete,'keep':sorted(expected)})
60
+ if obsolete:
61
+ cleanup=hub.api.create_commit(repo_id=args.dst,revision='main',parent_commit=revision,
62
+ operations=[CommitOperationDelete(path_in_repo=name) for name in obsolete],
63
+ commit_message='Keep only current 8+8 checkpoint, metadata and validation reports')
64
+ revision=cleanup.oid
65
+ remaining=verify(hub.api,args.dst,revision,expected)
66
+ if set(remaining)-set(expected)-{'.gitattributes'}:raise RuntimeError('Obsolete files remain after cleanup')
67
+ write(work/'all.done',{'repo':args.dst,'revision':revision,'files_verified':len(expected),'clean_repo':args.clean_repo})
68
+ print('FINAL PUBLISHED AND VERIFIED',args.dst,revision,flush=True)
69
+ return revision
reproduce/source/btx53/arvq88/fit.py ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """One 8+8 initial-fit worker. Source and cache identity are pinned by the driver."""
2
+ import sys,json,time,fcntl
3
+ from pathlib import Path
4
+ ROOT=Path(__file__).resolve().parents[1];sys.path[:0]=[str(ROOT),str(ROOT/'tools')]
5
+ import torch
6
+ from arvq88 import encoder as enc
7
+ from arvq88.pack import export_layer
8
+ from arvq88.inputs import write
9
+ from arvqprep.workflow import sha256
10
+ from ingest import SourceMeta,load_expert_bf16
11
+
12
+ def fit(work,L,dev):
13
+ args=json.loads((work/'run_config.json').read_text());allocation=json.loads((work/'assignment.json').read_text())
14
+ cold=allocation['layers'][str(L)]['cold_5750'];base=work/'initial'/f'layer_{L:05d}';base.mkdir(parents=True,exist_ok=True)
15
+ lock=open(base/'.lock','a');fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB)
16
+ torch.set_num_threads(args['threads']);source=json.loads((work/'source.json').read_text());meta=SourceMeta(**source['meta'])
17
+ act_path=Path(args['calib'])/f'acts_layer{L}.pt'
18
+ settings={k:args[k] for k in ['cb_iters','sweep_passes','refine','col_block']}
19
+ settings['codebook_scope']=args.get('codebook_scope','layer')
20
+ fingerprint={'source':source,'capture_sha256':sha256(act_path),'cold_ids':cold,'encoder_sha256':sha256(Path(enc.__file__)),'settings':settings}
21
+ identity=base/'fit_identity.json'
22
+ if identity.exists() and json.loads(identity.read_text())!=fingerprint:raise ValueError('Cached initial fit inputs/settings changed')
23
+ if not identity.exists() and any(base.glob('*.pt')):raise ValueError('Unidentified fit cache; use import-cache command or fresh workdir')
24
+ write(identity,fingerprint)
25
+ if (base/'arvq-manifest.json').exists():
26
+ export_layer(base,work/'cold',L);return
27
+ acts=torch.load(act_path,weights_only=True,map_location='cpu');x=acts['x'].float();ids=acts['topk_ids'].long()
28
+ if 'pv_split' in acts:
29
+ train=acts['pv_split']==0;x=x[train];ids=ids[train]
30
+ if not len(x):raise ValueError('No training activations for initial fit')
31
+ files={};metrics={}
32
+ for proj in ['w13','w2']:
33
+ path=base/f'{proj}.pt';mp=base/f'{proj}.json'
34
+ if not path.exists() or not mp.exists():
35
+ W=[];H=[];start=time.time()
36
+ for eid in cold:
37
+ wg,wu,wd=[w.to(dev) for w in load_expert_bf16(str(Path(args['src_cache'])/f'layer_{L}'/f'expert_{eid}.safetensors'),meta)]
38
+ xr=x[(ids==eid).any(1)];xr=(x if len(xr)<32 else xr).to(dev)
39
+ H.append((enc.hessian_fc1(xr) if proj=='w13' else enc.hessian_fc2(xr,wg,wu)).half())
40
+ W.append((torch.cat([wg,wu]) if proj=='w13' else wd).half());del wg,wu,wd,xr
41
+ result=enc.fit_layer_projection(W,H,device=dev,verbose=True,seed=0,**fingerprint['settings'])
42
+ tmp=path.with_suffix('.tmp');torch.save({proj:vars(result)},tmp);tmp.replace(path)
43
+ write(mp,{'cb_rel_l2':result.cb_rel_l2,'hess_recon_rel_fro':result.recon_rel_fro,'diagnostic_metric':'regularized fitting Hessian; mean expert relative error for expert scope','seconds':time.time()-start})
44
+ del result,W,H;torch.cuda.empty_cache()
45
+ files[proj]={'file':path.name,'sha256':sha256(path)};metrics[proj]=json.loads(mp.read_text())
46
+ write(base/'arvq-manifest.json',{'layer':L,'cold_expert_ids':cold,'files':files,'per_proj':metrics,'no_rotation':True,'source_identity':fingerprint,'format':'rvq256_256x8'})
47
+ export_layer(base,work/'cold',L)
48
+ if __name__=='__main__':fit(Path(sys.argv[1]),int(sys.argv[2]),sys.argv[3])
reproduce/source/btx53/arvq88/gate_worker.py ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ import sys
2
+ from pathlib import Path
3
+ root=Path(__file__).resolve().parents[1];sys.path[:0]=[str(root),str(root/'tools')]
4
+ from arvq88.checkpoint import gates
5
+ gates(Path(sys.argv[1]),[int(x) for x in sys.argv[2].split(',')],sys.argv[3],Path(sys.argv[4]))
reproduce/source/btx53/arvq88/gradient_indices.py ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Sparse proximal code proposals driven by routed-output gradients.
2
+
3
+ This is a layer-local optimizer, not full-model published PV-Tuning. Each
4
+ expert/projection proposal is evaluated against actual training output loss.
5
+ """
6
+ import math
7
+ import torch
8
+ from .activation import ARITHMETIC, activation_ste, swiglu_ste
9
+
10
+ METHOD = 'routed_output_gradient_sparse_v1'
11
+
12
+ @torch.no_grad()
13
+ def propose_codes(p, e, grad, max_fraction=.01, trust_ratio=.01, target_ratio=.1):
14
+ """Search single-book alternatives near a gradient step; cap count and norm.
15
+
16
+ The candidate pool is the highest-gradient 4*budget groups. Search all 256
17
+ choices in either book, with the other book fixed. Score using the proximal
18
+ linearization g*delta + ||delta||^2/(2*eta). Original FP8 weights are never used.
19
+ """
20
+ w=p.weight(e).detach().reshape(-1,8);g=grad.detach().reshape(-1,8)
21
+ if not torch.isfinite(g).all():raise ValueError('Nonfinite index gradient')
22
+ gn=g.norm();wn=w.norm();budget=int(len(w)*max_fraction)
23
+ if budget<1 or float(gn)==0 or float(wn)==0:return None
24
+ eta=target_ratio*wn/gn
25
+ pool=min(len(w),4*budget)
26
+ selected=g.square().sum(1).topk(pool,sorted=False).indices
27
+ cb,s,glob=p.quantized(e);cb=cb.detach();factor=(s.detach().repeat_interleave(16,1)*glob.detach()).flatten()
28
+ old_a=p.codes_a[e].flatten();old_b=p.codes_b[e].flatten()
29
+ scores=[];aa=[];bb=[];norms=[]
30
+ for ix in selected.split(2048):
31
+ a=old_a[ix].long();b=old_b[ix].long();f=factor[ix,None]
32
+ current=w[ix];target=current-eta*g[ix]
33
+ # Distances need only [chunk,256], not [chunk,256,8].
34
+ best=[]
35
+ for book,other in ((cb[:256],cb[256+b]),(cb[256:],cb[a])):
36
+ residual=target-f*other
37
+ distances=f.square()*book.square().sum(1)[None,:]-2*f*(residual@book.T)
38
+ choice=distances.argmin(1)
39
+ candidate=f*(book[choice]+other);delta=candidate-current
40
+ cost=(g[ix]*delta).sum(1)+delta.square().sum(1)/(2*eta)
41
+ best.append((choice,cost,delta.square().sum(1)))
42
+ choose_a=best[0][1]<=best[1][1]
43
+ scores.append(torch.minimum(best[0][1],best[1][1]))
44
+ aa.append(torch.where(choose_a,best[0][0],a));bb.append(torch.where(choose_a,b,best[1][0]))
45
+ norms.append(torch.where(choose_a,best[0][2],best[1][2]))
46
+ score=torch.cat(scores);a=torch.cat(aa);b=torch.cat(bb);delta2=torch.cat(norms)
47
+ valid=((score<0)&(delta2>0)).nonzero().flatten()
48
+ if not len(valid):return None
49
+ order=valid[score[valid].argsort()[:budget]]
50
+ # Strict trust radius for the reconstructed projection, not latent values.
51
+ order=order[delta2[order].cumsum(0)<=(trust_ratio*wn).square()]
52
+ if not len(order):return None
53
+ ix=selected[order]
54
+ return {'positions':ix,'a':a[order].to(old_a.dtype),'b':b[order].to(old_b.dtype),
55
+ 'old_a':old_a[ix].clone(),'old_b':old_b[ix].clone(),
56
+ 'predicted_change':float(score[order].sum()),
57
+ 'relative_weight_change':float(delta2[order].sum().sqrt()/wn)}
58
+
59
+
60
+ def expert_output(z,w13,w2,arithmetic):
61
+ if arithmetic==ARITHMETIC:
62
+ gu=activation_ste(z)@w13.T
63
+ return activation_ste(swiglu_ste(gu))@w2.T
64
+ gu=z@w13.T;gate,up=gu.chunk(2,-1)
65
+ return (torch.nn.functional.silu(gate)*up)@w2.T
66
+
67
+
68
+ @torch.no_grad()
69
+ def reassign_indices(p13,p2,x,rows,cold,args,target,prediction,routes):
70
+ """Sequential expert/projection updates with exact routed residual checks.
71
+
72
+ Only the provided training rows enter gradients, candidate scoring or acceptance.
73
+ The routed residual includes every cold expert, so gates and cross-expert
74
+ error interactions are included. Validation/audit are inaccessible here.
75
+ """
76
+ ref=target[rows];residual=prediction(rows)-ref
77
+ denominator=ref.square().sum().clamp_min(1e-20)
78
+ before=float((residual.square().sum()/denominator).sqrt())
79
+ accepted=proposed=changed=skipped=0;details=[]
80
+ for e,eid in enumerate(cold):
81
+ weights=routes[e][rows];use=weights!=0
82
+ if int(use.sum())<args.reassign_min_rows:skipped+=1;continue
83
+ z=x[rows[use]];gate=weights[use,None]
84
+ for p,name in ((p13,'w13'),(p2,'w2')):
85
+ w13=p13.weight(e).detach();w2=p2.weight(e).detach()
86
+ with torch.enable_grad():
87
+ weight=w13 if name=='w13' else w2;weight.requires_grad_(True)
88
+ old_output=expert_output(z,w13,w2,args.arithmetic)
89
+ # At current weights this equals the actual coupled output loss.
90
+ routed=residual[use].detach()+gate*(old_output-old_output.detach())
91
+ loss=routed.square().sum()/denominator
92
+ grad,=torch.autograd.grad(loss,weight)
93
+ old_output=old_output.detach();w13=w13.detach();w2=w2.detach()
94
+ proposal=propose_codes(p,e,grad,args.reassign_max_fraction,
95
+ args.reassign_trust_ratio,args.reassign_target_ratio)
96
+ del grad,loss,routed,weight
97
+ if proposal is None:continue
98
+ proposed+=1;ix=proposal['positions']
99
+ a=p.codes_a[e].view(-1);b=p.codes_b[e].view(-1)
100
+ a[ix]=proposal['a'];b[ix]=proposal['b']
101
+ try:
102
+ new_output=expert_output(z,p13.weight(e).detach(),p2.weight(e).detach(),args.arithmetic)
103
+ candidate=residual[use]+gate*(new_output-old_output)
104
+ old_loss=float(residual[use].square().sum());new_loss=float(candidate.square().sum())
105
+ keep=math.isfinite(new_loss) and new_loss<old_loss-max(1e-12,old_loss*1e-7)
106
+ if keep:residual[use]=candidate;accepted+=1;changed+=len(ix)
107
+ else:a[ix]=proposal['old_a'];b[ix]=proposal['old_b']
108
+ except BaseException:
109
+ a[ix]=proposal['old_a'];b[ix]=proposal['old_b'];raise
110
+ details.append({'expert':eid,'projection':name,'accepted':keep,'changed_groups':len(ix),
111
+ 'relative_weight_change':proposal['relative_weight_change'],
112
+ 'training_before_local':old_loss,'training_candidate_local':new_loss})
113
+ # Recompute rather than trust accumulated residual arithmetic. Rollback is
114
+ # handled by the caller's full index snapshot if the net check regresses.
115
+ after=float(((prediction(rows)-ref).square().sum()/denominator).sqrt())
116
+ return {'method':METHOD,'training_before':before,'training_candidate':after,
117
+ 'accepted':math.isfinite(after) and after<before,
118
+ 'accepted_proposals':accepted,'proposals':proposed,'changed_groups':changed,
119
+ 'skipped_experts':skipped,'max_changed_fraction':args.reassign_max_fraction,
120
+ 'trust_ratio':args.reassign_trust_ratio,'details':details}
reproduce/source/btx53/arvq88/incremental_publish.py ADDED
@@ -0,0 +1,122 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Atomic initial v3 checkpoint publication, then audited layer replacements."""
2
+ import json
3
+ import os
4
+ import shutil
5
+ import time
6
+ from pathlib import Path
7
+ from .inputs import write
8
+ from .pack import sha
9
+ from .validation import require_audit
10
+ from .final_publish import inventory,verify
11
+
12
+
13
+ def progress_card(repo,completed):
14
+ return f'''---
15
+ library_name: vllm
16
+ tags: [arvq, nvfp4, experimental]
17
+ ---
18
+ # {repo.rsplit('/',1)[-1]}
19
+
20
+ Complete Vision NVFP4 / per-expert ARVQ 8+8 checkpoint, format
21
+ `rvq256_256x8_expert`, version 3. Requires the per-expert-book serving loader.
22
+ The new ARVQ-based 75% text / 25% multimodal allocation and all initial fits
23
+ are complete. **{len(completed)}/75 layers have been replaced by audited PV
24
+ weights.** Remaining layers retain their initial fits. See pv_progress.json.
25
+
26
+ Each cold expert has separate two-book FP4 dictionaries for gate/up and down.
27
+ PV emulates four FP4 activation planes and FP16 boundaries, with discrete
28
+ Hessian-aware index proposals every 40 steps. Validation selects the best
29
+ state; audit regression restores the initial fit. Reference source is original
30
+ GLM-5.3 FP8; activations use the saved NVFP4-teacher conversation corpus.
31
+ Hot experts, language backbone and BF16 MTP come from RadixArk NVFP4; vision
32
+ comes from the existing vision hybrid. Indices, cold slots and hot allocation
33
+ are coherent at every published revision. Each layer replacement is atomic.
34
+
35
+ Offline audits/packing checks are not full-model quality or native-SM120 parity
36
+ benchmarks. Independent-book serving quality validation remains pending.
37
+ '''
38
+
39
+
40
+ def commit_snapshot(hub,repo,root,parent,message,delete=()):
41
+ from huggingface_hub import CommitOperationAdd,CommitOperationDelete
42
+ expected=inventory(root)
43
+ operations=[CommitOperationAdd(path_in_repo=name,path_or_fileobj=str(Path(root)/name)) for name in expected]
44
+ # Upload LFS objects first, but expose all tensor/config changes in ONE commit.
45
+ hub.api.preupload_lfs_files(repo,operations,revision='main',num_threads=8)
46
+ result=hub.api.create_commit(repo_id=repo,revision='main',parent_commit=parent,
47
+ operations=operations+[CommitOperationDelete(path_in_repo=name) for name in delete],
48
+ commit_message=message,num_threads=8)
49
+ verify(hub.api,repo,result.oid,expected)
50
+ return result.oid
51
+
52
+
53
+ def publish_initial(args,hub,work):
54
+ ck=work/'checkpoint';gates=json.loads((ck/'gates_report.json').read_text())
55
+ if not gates['passed']:raise ValueError('Initial checkpoint gates failed')
56
+ for name,digest in {**gates['cold_file_sha256'],**gates['metadata_sha256']}.items():
57
+ if sha(ck/name)!=digest:raise ValueError('Checkpoint changed after initial gates')
58
+ remote=hub.api.model_info(args.dst,files_metadata=True)
59
+ local={str(p.relative_to(ck)) for p in ck.rglob('*') if p.is_file()}
60
+ obsolete=sorted({s.rfilename for s in remote.siblings}-local-{'.gitattributes'})
61
+ revision=commit_snapshot(hub,args.dst,ck,remote.sha,'Replace with complete per-expert ARVQ v3 initial fit; PV follows',obsolete)
62
+ state={'repo':args.dst,'revision':revision,'layers':[]}
63
+ write(work/'incremental_upload.json',state)
64
+ print('INITIAL CHECKPOINT UPLOADED AND VERIFIED',revision,flush=True)
65
+ return state
66
+
67
+
68
+ def layer_snapshot(work,L,completed,repo):
69
+ ck=work/'checkpoint'
70
+ try:m=json.loads((work/'cold'/f'layer-{L:03d}-manifest.json').read_text())
71
+ except (FileNotFoundError,json.JSONDecodeError):return None
72
+ r=m.get('pv_report')
73
+ if not r:return None
74
+ require_audit(r)
75
+ if not r.get('passed') or r.get('indices_frozen',True) or r.get('codebook_scope')!='expert':
76
+ raise ValueError('Unexpected PV recipe for incremental upload')
77
+ assignment=json.loads((work/'assignment.json').read_text())['layers'][str(L)]['cold_5750']
78
+ if m['cold_expert_ids']!=assignment or m.get('version')!=3:raise ValueError('Layer allocation/format mismatch')
79
+ stage=work/'upload_layers'/f'layer_{L:05d}';stage.mkdir(parents=True,exist_ok=True)
80
+ gates=json.loads((ck/'gates_report.json').read_text())
81
+ for entry in m['files'].values():
82
+ source=work/'cold'/entry['file']
83
+ if sha(source)!=entry['sha256']:raise ValueError('Packed layer hash mismatch')
84
+ dst=stage/entry['file']
85
+ if dst.exists():dst.unlink()
86
+ os.link(source,dst)
87
+ gates['cold_file_sha256'][entry['file']]=entry['sha256']
88
+ gates['scope']='Initial structural gates plus incremental packing/parity and PV audits; final full gates pending'
89
+ write(stage/'gates_report.json',gates)
90
+ write(stage/'cold_manifests'/f'layer-{L:03d}-manifest.json',m)
91
+ write(stage/'pv_layers'/f'layer_{L:03d}.json',r)
92
+ write(stage/'pv_progress.json',{'layers_pv_complete':sorted(completed),'total_layers':75,
93
+ 'remaining_layers':'initial fit','format':'rvq256_256x8_expert'})
94
+ (stage/'README.md').write_text(progress_card(repo,completed))
95
+ return stage
96
+
97
+
98
+ def publish_layers(args,hub,work):
99
+ statefile=work/'incremental_upload.json'
100
+ state=json.loads(statefile.read_text()) if statefile.exists() else publish_initial(args,hub,work)
101
+ done=set(state['layers'])
102
+ while len(done)<75:
103
+ if (work/'run.failed').exists():raise RuntimeError('Fit/PV campaign failed; completed uploads retained')
104
+ changed=False
105
+ for L in range(3,78):
106
+ if L in done:continue
107
+ stage=layer_snapshot(work,L,done|{L},args.dst)
108
+ if stage is None:continue
109
+ revision=commit_snapshot(hub,args.dst,stage,state['revision'],f'Replace layer {L} with audited per-expert alternating PV ({len(done)+1}/75)')
110
+ # Keep local assembly aligned; atomic replaces avoid mutating initial hardlinks.
111
+ for source in stage.rglob('*'):
112
+ if not source.is_file():continue
113
+ dst=work/'checkpoint'/source.relative_to(stage);dst.parent.mkdir(parents=True,exist_ok=True)
114
+ tmp=dst.with_suffix(dst.suffix+'.replace')
115
+ if tmp.exists():tmp.unlink()
116
+ if source.suffix=='.safetensors':os.link(source,tmp)
117
+ else:shutil.copy2(source,tmp)
118
+ tmp.replace(dst)
119
+ done.add(L);state.update(revision=revision,layers=sorted(done));write(statefile,state)
120
+ print('PV LAYER UPLOADED AND VERIFIED',L,len(done),'/75',revision,flush=True);changed=True
121
+ if not changed:time.sleep(20)
122
+ write(work/'incremental_pv.done.json',state)
reproduce/source/btx53/arvq88/inputs.py ADDED
@@ -0,0 +1,198 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Validated cache reuse, pinned Hub input fetch and REAP allocation."""
2
+ import json,os,shutil,sys,subprocess
3
+ from pathlib import Path
4
+ from concurrent.futures import ThreadPoolExecutor
5
+ import numpy as np
6
+ from .local_source import local_directory,model_index,fingerprint,LocalShard
7
+ from arvqprep.workflow import sha256
8
+ ROOT=Path(__file__).resolve().parents[1]
9
+ LAYERS=list(range(3,78))
10
+ def write(path,obj):
11
+ path=Path(path);path.parent.mkdir(parents=True,exist_ok=True)
12
+ tmp=path.with_suffix(path.suffix+'.tmp');tmp.write_text(json.dumps(obj,indent=2));tmp.replace(path)
13
+ def validate_assignment(obj,hot_count=5750):
14
+ layers=obj.get('layers',obj)
15
+ if set(map(int,layers))!=set(LAYERS):raise ValueError('Allocation must cover layers 3–77')
16
+ result={}
17
+ for L in LAYERS:
18
+ d=layers[str(L)];cold=d.get('cold_5750',d.get('cold'))
19
+ if cold is None or len(set(cold))!=len(cold) or any(type(e)!=int or not 0<=e<256 for e in cold):raise ValueError(f'Invalid cold IDs layer {L}')
20
+ hot=sorted(set(range(256))-set(cold))
21
+ if 'hot' in d and sorted(d['hot'])!=hot:raise ValueError('Overlapping/missing expert assignment')
22
+ result[str(L)]={'cold_5750':sorted(cold),'hot':hot}
23
+ if sum(len(v['hot']) for v in result.values())!=hot_count:raise ValueError('Allocation hot budget mismatch')
24
+ return {'layers':result}
25
+ def allocate(scores,hot_count=5750,floor=8,cap=176):
26
+ scores=np.asarray(scores,dtype=np.float64)
27
+ if scores.shape!=(75,256) or not np.isfinite(scores).all() or (scores<0).any():raise ValueError('Invalid REAP score array')
28
+ if not 75*floor<=hot_count<=75*cap:raise ValueError('Infeasible hot budget')
29
+ norm=scores/np.maximum(scores.sum(1,keepdims=True),1e-30)
30
+ hot=[set(np.argsort(-row,kind='stable')[:floor].tolist()) for row in norm];placed=75*floor
31
+ for pos in np.argsort(-norm.ravel(),kind='stable'):
32
+ if placed==hot_count:break
33
+ l,e=divmod(int(pos),256)
34
+ if len(hot[l])<cap and e not in hot[l]:hot[l].add(e);placed+=1
35
+ return validate_assignment({'layers':{str(L):{'cold':sorted(set(range(256))-hot[i])} for i,L in enumerate(LAYERS)}},hot_count)
36
+ def token():
37
+ p=Path.home()/'.cache/huggingface/token'
38
+ return p.read_text().strip() if p.exists() else None
39
+ class Hub:
40
+ def __init__(self,cache):
41
+ from huggingface_hub import HfApi
42
+ self.api=HfApi(token=token());self.cache=Path(cache);self.revisions={}
43
+ def revision(self,repo):
44
+ if repo not in self.revisions:
45
+ local=local_directory(repo)
46
+ self.revisions[repo]=fingerprint(local) if local else self.api.model_info(repo).sha
47
+ return self.revisions[repo]
48
+ def file(self,repo,name):
49
+ local=local_directory(repo)
50
+ if local:
51
+ p=local/name
52
+ if not p.is_file():raise FileNotFoundError(p)
53
+ return p
54
+ from huggingface_hub import hf_hub_download
55
+ return Path(hf_hub_download(repo,name,revision=self.revision(repo),cache_dir=str(self.cache),token=token()))
56
+ def snapshot(self,repo,dest):
57
+ local=local_directory(repo)
58
+ if local:return local
59
+ from huggingface_hub import snapshot_download
60
+ snapshot_download(repo,revision=self.revision(repo),local_dir=str(dest),token=token(),max_workers=8)
61
+ return Path(dest)
62
+ def complete_snapshot(path):
63
+ try:
64
+ wm=json.loads((Path(path)/'model.safetensors.index.json').read_text())['weight_map']
65
+ return bool(wm) and all((Path(path)/f).is_file() for f in set(wm.values()))
66
+ except (FileNotFoundError,KeyError,ValueError):return False
67
+ def valid_capture(path):
68
+ import torch
69
+ try:
70
+ d=torch.load(path,weights_only=True,map_location='cpu');n=len(d['x'])
71
+ return n>=32 and d['x'].shape==(n,6144) and d['topk_ids'].shape==(n,8) and d['topk_weights'].shape==(n,8) and bool(torch.isfinite(d['x']).all()) and bool(torch.isfinite(d['topk_weights']).all()) and bool(((d['topk_ids']>=0)&(d['topk_ids']<256)).all())
72
+ except (OSError,KeyError,ValueError,RuntimeError,EOFError):return False
73
+ def build_corpus(args,hub):
74
+ """Download a bounded, deterministic text corpus when local token shards are absent."""
75
+ from datasets import load_dataset
76
+ from transformers import AutoTokenizer
77
+ target=Path(args.calib_tokens);target.mkdir(parents=True,exist_ok=True)
78
+ marker=target/'corpus.building.json'
79
+ if marker.exists():
80
+ for shard in target.glob('shard_*.npy'):shard.unlink()
81
+ write(marker,{'builder':'arvq88 fallback corpus','target_tokens':args.calib_token_count})
82
+ tok=AutoTokenizer.from_pretrained(args.src,revision=None if local_directory(args.src) else hub.revision(args.src),token=token(),trust_remote_code=True)
83
+ specs=[('m-a-p/CodeFeedback-Filtered-Instruction',.4),('tatsu-lab/alpaca',.4),('databricks/databricks-dolly-15k',.2)]
84
+ pending=[];number=0;actual={};revisions={}
85
+ for repo,fraction in specs:
86
+ budget=int(args.calib_token_count*fraction);count=0
87
+ revisions[repo]=hub.api.dataset_info(repo).sha
88
+ data=load_dataset(repo,revision=revisions[repo],split='train',streaming=True,token=token()).shuffle(seed=42,buffer_size=10000)
89
+ while count<budget:
90
+ before=count
91
+ for row in data:
92
+ text='\n'.join(str(v) for v in row.values() if isinstance(v,(str,list)))
93
+ ids=tok.encode(text,add_special_tokens=False)+[tok.eos_token_id]
94
+ ids=ids[:budget-count];pending.extend(ids);count+=len(ids)
95
+ while len(pending)>=1_000_000:
96
+ np.save(target/f'shard_{number:05d}.npy',np.asarray(pending[:1_000_000],dtype=np.uint32));pending=pending[1_000_000:];number+=1
97
+ if count>=budget:break
98
+ if count==before:raise RuntimeError(f'Empty calibration dataset {repo}')
99
+ actual[repo]=count
100
+ if pending:np.save(target/f'shard_{number:05d}.npy',np.asarray(pending,dtype=np.uint32))
101
+ write(target/'corpus.json',{'builder':'fallback text calibration, not legacy calib-v3','sources_token_counts':actual,'dataset_revisions':revisions,'seed':42,'reasoning_rollouts':False,'tokenizer_repo':args.src,'tokenizer_revision':hub.revision(args.src)})
102
+ marker.unlink()
103
+ def ensure_captures(args,hub,work,need_stats=False):
104
+ import torch
105
+ shards=list(Path(args.calib_tokens).glob('shard_*.npy'))
106
+ if not shards or (Path(args.calib_tokens)/'corpus.building.json').exists():build_corpus(args,hub)
107
+ else:
108
+ for p in shards:
109
+ a=np.load(p,mmap_mode='r')
110
+ if a.ndim!=1 or not np.issubdtype(a.dtype,np.integer) or not a.size:raise ValueError(f'Invalid token shard {p}')
111
+ ready=all(valid_capture(Path(args.calib)/f'acts_layer{L}.pt') for L in LAYERS)
112
+ stats=Path(args.calib).parent/'expert_stats53.npz'
113
+ if ready and (not need_stats or stats.exists()):return stats
114
+ if not complete_snapshot(args.donor_cache):hub.snapshot(args.nvfp4_donor,args.donor_cache)
115
+ # Capture to a fresh directory so partial old ranks cannot enter the merge.
116
+ dest=work/'capture';dest.mkdir(exist_ok=True)
117
+ env={**os.environ,'GLM53_DONOR':args.donor_cache,'GLM53_CALIB':args.calib_tokens,'GLM53_CAPTURE_OUT':str(dest)}
118
+ script=ROOT.parent/'tools/capture53/stream_capture53.py';workers=[]
119
+ for rank,gpu in enumerate(args.gpus):
120
+ log=open(work/f'capture-rank{rank}.log','a')
121
+ workers.append(subprocess.Popen([sys.executable,'-u',str(script),'stats'],env={**env,'CUDA_VISIBLE_DEVICES':str(gpu),'RANK':str(rank),'WORLD':str(len(args.gpus))},stdout=log,stderr=subprocess.STDOUT,start_new_session=True));log.close()
122
+ codes=[p.wait() for p in workers]
123
+ if any(codes):raise RuntimeError('Capture workers failed; see capture-rank logs')
124
+ subprocess.run([sys.executable,str(script),'merge'],env={**env,'WORLD':str(len(args.gpus))},check=True)
125
+ args.calib=str(dest/'acts')
126
+ if not all(valid_capture(Path(args.calib)/f'acts_layer{L}.pt') for L in LAYERS):raise RuntimeError('Capture incomplete')
127
+ return dest/'expert_stats53.npz'
128
+ def assignment(args,hub,work):
129
+ if getattr(args,'reap_metric','aqlm')=='arvq':
130
+ from .arvq_reap import recompute
131
+ obj=recompute(args,hub,work)
132
+ write(work/'assignment.json',obj);args.assign=str(work/'assignment.json');return obj
133
+ cached=Path(args.assign)
134
+ if cached.exists():
135
+ obj=validate_assignment(json.loads(cached.read_text()),args.hot_count)
136
+ obj['provenance']={'mode':'reused allocation','path':str(cached),'sha256':sha256(cached)}
137
+ else:
138
+ from .reap import recompute
139
+ obj=recompute(args,hub,work)
140
+ write(work/'assignment.json',obj);args.assign=str(work/'assignment.json');return obj
141
+ def ensure_source(args,hub,work,allocation):
142
+ import torch
143
+ from safetensors import safe_open
144
+ from safetensors.torch import save_file
145
+ from ingest import probe_source,validate_meta_against_header
146
+ from remote_st import RemoteShard
147
+ import remote_st
148
+ cfg=json.loads(hub.file(args.src,'config.json').read_text());text=cfg.get('text_config',cfg)
149
+ if text.get('hidden_size')!=6144 or text.get('moe_intermediate_size')!=2048 or text.get('n_routed_experts',text.get('num_local_experts'))!=256:
150
+ raise ValueError('Input is not the supported GLM-5.3 expert geometry')
151
+ meta=probe_source(text);native={**vars(meta),'block':list(meta.block) if meta.block else None}
152
+ revision=hub.revision(args.src);cache=Path(args.src_cache);cache.mkdir(parents=True,exist_ok=True)
153
+ had_existing=any(cache.glob('layer_*'))
154
+ marker=cache/'source_identity.json'
155
+ if marker.exists():
156
+ identity=json.loads(marker.read_text())
157
+ if identity!={'repo':args.src,'revision':revision}:raise ValueError('Source cache identity mismatch; use another --src-cache')
158
+ elif args.src!='zai-org/GLM-5.3' and any(cache.glob('layer_*')):raise ValueError('Unidentified cache cannot be reused for a different source repo')
159
+ local=local_directory(args.src)
160
+ index=model_index(local) if local else json.loads(hub.file(args.src,'model.safetensors.index.json').read_text())['weight_map']
161
+ remote_st.resolve_url=lambda repo,name:f'https://huggingface.co/{repo}/resolve/{revision}/{name}'
162
+ headers={}
163
+ def shard(name):
164
+ fn=index[name]
165
+ if fn not in headers:headers[fn]=LocalShard(local/fn) if local else RemoteShard(args.src,fn)
166
+ return headers[fn]
167
+ dtype={'F8_E4M3':torch.float8_e4m3fn,'BF16':torch.bfloat16,'F16':torch.float16,'F32':torch.float32,'U8':torch.uint8}
168
+ def fetch(pair):
169
+ L,e=pair;path=cache/f'layer_{L}'/f'expert_{e}.safetensors'
170
+ if path.exists():
171
+ try:
172
+ with safe_open(str(path),framework='pt') as f:
173
+ metadata=f.metadata() or {}
174
+ if metadata.get('source',args.src)!=args.src:raise ValueError('Expert source mismatch')
175
+ for proj,shape in [('gate_proj',(2048,6144)),('up_proj',(2048,6144)),('down_proj',(6144,2048))]:
176
+ t=f.get_slice(proj+'.weight')
177
+ if tuple(t.get_shape())!=shape:raise ValueError('Invalid cached expert shape')
178
+ validate_meta_against_header(meta,t.get_dtype())
179
+ if meta.kind=='block_fp8' and tuple(f.get_slice(proj+'.weight_scale_inv').get_shape())!=(shape[0]//meta.block[0],shape[1]//meta.block[1]):raise ValueError('Invalid cached scales')
180
+ return
181
+ except Exception as exc:raise ValueError(f'Invalid cached expert {path}; remove or repair it before retry') from exc
182
+ tensors={}
183
+ for proj in ['gate_proj','up_proj','down_proj']:
184
+ for suffix in ['weight']+(['weight_scale_inv'] if meta.kind=='block_fp8' else []):
185
+ name=f'model.layers.{L}.mlp.experts.{e}.{proj}.{suffix}';sh=shard(name);m=sh.tensor_meta(name)
186
+ raw=bytearray(sh.read_tensor_bytes(name));expected=m['data_offsets'][1]-m['data_offsets'][0]
187
+ if len(raw)!=expected:raise IOError('Short source tensor read')
188
+ t=torch.frombuffer(raw,dtype=dtype[m['dtype']]).reshape(m['shape']).clone()
189
+ if suffix=='weight' and meta.kind=='block_fp8' and t.dtype==torch.uint8:t=t.view(torch.float8_e4m3fn)
190
+ tensors[proj+'.'+suffix]=t
191
+ path.parent.mkdir(parents=True,exist_ok=True);tmp=path.with_suffix('.part');save_file(tensors,str(tmp),metadata={'source':args.src,'revision':revision});tmp.replace(path)
192
+ jobs=[(L,e) for L in LAYERS for e in allocation['layers'][str(L)]['cold_5750']]
193
+ with ThreadPoolExecutor(max_workers=args.download_workers) as pool:list(pool.map(fetch,jobs))
194
+ # Existing canonical donor files predate revision markers; record their status honestly.
195
+ if local and fingerprint(local)!=revision:raise ValueError('Local source changed during extraction')
196
+ write(work/'source.json',{'repo':args.src,'revision':revision,'meta':native,'cache':str(cache),'legacy_cache_revision_unverified':not marker.exists() and had_existing})
197
+ if not had_existing:write(marker,{'repo':args.src,'revision':revision})
198
+ return meta
reproduce/source/btx53/arvq88/jobs.py ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Bounded subprocess jobs on explicitly reserved GPUs."""
2
+ import subprocess
3
+ import time
4
+ from pathlib import Path
5
+
6
+
7
+ def run_jobs(jobs, gpus, logs):
8
+ pending=list(jobs);running={};logs=Path(logs);logs.mkdir(parents=True,exist_ok=True)
9
+ try:
10
+ while pending or running:
11
+ for gpu in gpus:
12
+ if gpu in running or not pending:continue
13
+ name,command=pending.pop(0)
14
+ with (logs/(name+'.log')).open('a') as log:
15
+ p=subprocess.Popen(command(gpu),stdin=subprocess.DEVNULL,stdout=log,
16
+ stderr=subprocess.STDOUT,start_new_session=True)
17
+ running[gpu]=(name,p);print('START',name,'GPU',gpu,'PID',p.pid,flush=True)
18
+ time.sleep(5)
19
+ for gpu,(name,p) in list(running.items()):
20
+ rc=p.poll()
21
+ if rc is None:continue
22
+ if rc:raise RuntimeError(f'{name} failed ({rc}); see {logs/name}.log')
23
+ del running[gpu];print('DONE',name,flush=True)
24
+ except BaseException:
25
+ for name,p in running.values():
26
+ if p.poll() is None:p.terminate()
27
+ for name,p in running.values():p.wait()
28
+ raise
reproduce/source/btx53/arvq88/local_source.py ADDED
@@ -0,0 +1,43 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Read standard local safetensors snapshots without Hub requests."""
2
+ import hashlib,json,struct
3
+ from pathlib import Path
4
+
5
+ def local_directory(source):
6
+ p=Path(source).expanduser()
7
+ if p.exists():
8
+ if not p.is_dir():raise ValueError(f'Model input must be a directory: {p}')
9
+ return p.resolve()
10
+ if str(source).startswith(('/', './', '../', '~')):
11
+ raise FileNotFoundError(f'Local model directory does not exist: {p}')
12
+ return None
13
+
14
+ def model_index(root):
15
+ p=root/'model.safetensors.index.json'
16
+ if p.is_file():return json.loads(p.read_text())['weight_map']
17
+ p=root/'model.safetensors'
18
+ if not p.is_file():raise FileNotFoundError(f'Missing safetensors weights/index in {root}')
19
+ return {k:p.name for k in LocalShard(p).header if k!='__metadata__'}
20
+
21
+ def fingerprint(root):
22
+ # Metadata identity, not a full content hash of a multi-hundred-GB donor.
23
+ h=hashlib.sha256((root/'config.json').read_bytes())
24
+ index=model_index(root);h.update(json.dumps(index,sort_keys=True).encode())
25
+ for name in sorted(set(index.values())):
26
+ p=root/name;st=p.stat()
27
+ h.update(json.dumps([name,st.st_size,st.st_mtime_ns,st.st_ctime_ns]).encode())
28
+ return 'local-stat-sha256:'+h.hexdigest()
29
+
30
+ class LocalShard:
31
+ def __init__(self,path):
32
+ self.path=Path(path)
33
+ with self.path.open('rb') as f:
34
+ raw=f.read(8)
35
+ if len(raw)!=8:raise ValueError('Truncated safetensors header')
36
+ n=struct.unpack('<Q',raw)[0]
37
+ if n>100_000_000:raise ValueError('Invalid safetensors header length')
38
+ self.header=json.loads(f.read(n));self.offset=8+n
39
+ def tensor_meta(self,name):return self.header[name]
40
+ def read_tensor_bytes(self,name):
41
+ a,b=self.tensor_meta(name)['data_offsets']
42
+ with self.path.open('rb') as f:
43
+ f.seek(self.offset+a);return f.read(b-a)
reproduce/source/btx53/arvq88/pack.py ADDED
@@ -0,0 +1,142 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """ARVQ 8+8 test format: natural indices -> MMA fragments, 64 uint32 words/tile."""
2
+ import torch,json,hashlib
3
+ from pathlib import Path
4
+ LEVELS=[0.,.5,1.,1.5,2.,3.,4.,6.,-0.,-.5,-1.,-1.5,-2.,-3.,-4.,-6.]
5
+ def maps(device):
6
+ p=torch.arange(128,device=device);j=p//32;lane=p%32
7
+ return lane//4+8*(j%2),(j//2)*4+lane%4
8
+ def pack(a,b,N,K):
9
+ r,k=maps(a.device);rows=torch.arange(N//16,device=a.device)[:,None,None]*16+r[None,None,:];cols=torch.arange(K//64,device=a.device)[None,:,None]*8+k[None,None,:]
10
+ v=(a.long()|(b.long()<<8))[rows,cols]
11
+ return (v[...,::2]|(v[...,1::2]<<16)).to(torch.uint32)
12
+ def unpack(w,N,K):
13
+ w=w.long();v=torch.stack([w&65535,w>>16],-1).reshape(N//16,K//64,128)
14
+ r,k=maps(w.device);rows=(torch.arange(N//16,device=w.device)[:,None,None]*16+r[None,None,:]).expand_as(v);cols=(torch.arange(K//64,device=w.device)[None,:,None]*8+k[None,None,:]).expand_as(v)
15
+ a=torch.empty(N,K//8,dtype=torch.uint8,device=w.device);b=torch.empty_like(a)
16
+ a[rows,cols]=(v&255).byte();b[rows,cols]=(v>>8).byte();return a,b
17
+ def pack_codebooks(c0,c1):
18
+ cb=torch.cat([c0,c1],dim=-2)
19
+ if cb.shape[-2:]!=(512,8) or cb.ndim not in (2,3):raise ValueError('Expected shared or per-expert 8+8 books')
20
+ levels=torch.tensor(LEVELS,device=cb.device);n=(cb[...,None]-levels).abs().argmin(-1)
21
+ if not torch.equal(levels[n],cb):raise ValueError('Codebooks must be exactly on the FP4 grid')
22
+ return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
23
+ def pack_codebooks_mb(cb,factors):
24
+ """mcbook16: cb [E,256+B*256,8] in EFFECTIVE units. Rows are stored as plain-grid
25
+ nibbles; decode multiplies residual book m by factors[m] (exact powers of two)."""
26
+ B=len(factors)
27
+ if cb.ndim!=3 or cb.shape[1]!=256+B*256 or cb.shape[2]!=8:raise ValueError('Invalid mcbook codebook shape')
28
+ f=torch.cat([torch.ones(256),torch.tensor(factors).float().repeat_interleave(256)])[None,:,None]
29
+ grid=cb/f
30
+ levels=torch.tensor(LEVELS,device=cb.device);n=(grid[...,None]-levels).abs().argmin(-1)
31
+ if not torch.equal(levels[n]*f,cb):raise ValueError('Codebooks must be exactly on their per-book grids')
32
+ return (n.long()<<(torch.arange(8,device=cb.device)*4)).sum(-1).to(torch.uint32)
33
+ def decode_mb(t,layer,proj,e,N,K):
34
+ pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
35
+ a,b=unpack(t[pre+'packed'][e],N,K)
36
+ cb=t[pre+'codebooks'][e].long();factors=t[pre+'book_factors'].float();B=len(factors)
37
+ if cb.shape!=(256+B*256,):raise ValueError('Invalid mcbook packed codebook shape')
38
+ values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
39
+ values=torch.cat([values[:256],values[256:]*factors.repeat_interleave(256)[:,None]])
40
+ sel=t[pre+'selectors'][e].long()
41
+ m=sel.repeat_interleave(16,0).repeat_interleave(8,1)
42
+ stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
43
+ if stored.dtype==torch.float16:scales=stored.float()
44
+ elif stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
45
+ else:raise ValueError('Unsupported packed block scale dtype')
46
+ return ((values[a.long()]+values[256+m*256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
47
+ def export_layer_mb(source,dest,L):
48
+ """mcbook16 export (format version 5): 16 residual books, per-tile 4-bit selectors,
49
+ fp16 block scales; packed index stream and scales unchanged from v4."""
50
+ from safetensors.torch import save_file,load_file
51
+ source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
52
+ original=json.loads((source/'arvq-manifest.json').read_text());files={};first_factors=None
53
+ for proj,tag in [('w13','gateup'),('w2','down')]:
54
+ d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
55
+ if d.get('scale_dtype')!='fp16':raise ValueError('mcbook export requires fp16 block scales')
56
+ if not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
57
+ factors=[float(f) for f in d['book_factor']]
58
+ if first_factors is None:first_factors=factors
59
+ elif factors!=first_factors:raise ValueError('Mixed projection book factors')
60
+ sel=d['selector']
61
+ if sel.shape!=(E,N//16,K//64) or int(sel.max())>=len(factors):raise ValueError('Invalid selector tensor')
62
+ t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
63
+ pre+'codebooks':pack_codebooks_mb(d['cb'],factors),
64
+ pre+'selectors':sel.to(torch.uint8).contiguous(),
65
+ pre+'book_factors':torch.tensor(factors,dtype=torch.float32),
66
+ pre+'scales':d['s'].half().reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
67
+ pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
68
+ for e in range(E):
69
+ a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
70
+ for e in sorted({0,E//2,E-1}):
71
+ m=d['selector'][e].long().repeat_interleave(16,0).repeat_interleave(8,1)
72
+ cb=d['cb'][e]
73
+ ref=((cb[d['a'][e].long()]+cb[256+m*256+d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
74
+ assert torch.equal(ref,decode_mb(t,L,proj,e,N,K))
75
+ name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
76
+ save_file(t,str(tmp),metadata={'format':'pt','arvq_format':'rvq256_mb16_256x8_expert_fp16block','codebook_scope':'expert','residual_books':str(len(factors)),'book_factors':json.dumps(factors),'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires mcbook16 v5 loader/kernel (vllm-glm52-sm120 branch experiment/arvq-mcbook16)'})
77
+ tmp.replace(dest/name)
78
+ loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64) and loaded[pre+'selectors'].shape==(E,N//16,K//64);del loaded
79
+ files[proj]={'file':name,'sha256':sha(dest/name)};del t,d
80
+ m={**original,'format':'rvq256_mb16_256x8_expert_fp16block','version':5,'residual_scale_shift':0,'block_scale_dtype':'float16','codebook_scope':'expert','bits':2.125+4/1024,'codebook_sizes':[256,16*256],'index_bits':[8,8],'selector_bits_per_tile':4,'words_per_tile':64,'book_factors':first_factors,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True,'requires_mcbook16_loader':True}
81
+ (dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
82
+ def decode(t,layer,proj,e,N=None,K=None,residual_scale_shift=0):
83
+ pre=f'model.layers.{layer}.mlp.experts.arvq_{proj}_'
84
+ default=(4096,6144) if proj=='w13' else (6144,2048)
85
+ N,K=default if N is None else (N,K)
86
+ a,b=unpack(t[pre+'packed'][e],N,K);cb=t[pre+'codebooks'].long()
87
+ if cb.ndim==2:cb=cb[e]
88
+ if cb.shape!=(512,):raise ValueError('Invalid packed codebook shape')
89
+ values=torch.tensor(LEVELS,device=cb.device)[(cb[:,None]>>(torch.arange(8,device=cb.device)*4))&15]
90
+ stored=t[pre+'scales'][e].permute(0,2,1).contiguous().reshape(N,K//128)
91
+ if stored.dtype==torch.uint8:scales=stored.view(torch.float8_e4m3fn).float()
92
+ elif stored.dtype==torch.float16:scales=stored.float()
93
+ else:raise ValueError('Unsupported packed block scale dtype')
94
+ f=2.0**(-int(residual_scale_shift))
95
+ return ((values[a.long()]+f*values[256+b.long()])*scales.repeat_interleave(16,1)[...,None]*t[pre+'global']).reshape(N,K)
96
+ def sha(p):
97
+ h=hashlib.sha256()
98
+ with open(p,'rb') as f:
99
+ for chunk in iter(lambda:f.read(8<<20),b''):h.update(chunk)
100
+ return h.hexdigest()
101
+ def export_layer(source,dest,L,residual_scale_shift=0):
102
+ from safetensors.torch import save_file,load_file
103
+ source=Path(source);dest=Path(dest);dest.mkdir(parents=True,exist_ok=True)
104
+ probe=torch.load(source/'w13.pt',map_location='cpu',weights_only=True,mmap=True)['w13']
105
+ if 'selector' in probe:
106
+ del probe
107
+ if int(residual_scale_shift):raise ValueError('mcbook exports are shift-0 (per-book factors)')
108
+ return export_layer_mb(source,dest,L)
109
+ del probe
110
+ rss=int(residual_scale_shift);f=2.0**(-rss)
111
+ original=json.loads((source/'arvq-manifest.json').read_text());files={}
112
+ for proj,tag in [('w13','gateup'),('w2','down')]:
113
+ d=torch.load(source/f'{proj}.pt',map_location='cpu',weights_only=True)[proj];N,K=d['N'],d['K'];E=len(d['a']);pre=f'model.layers.{L}.mlp.experts.arvq_{proj}_'
114
+ expert=d['c0'].ndim==3
115
+ if expert and d['c0'].shape!=(E,256,8):raise ValueError('Codebook expert count mismatch')
116
+ if proj=='w13':scope='expert' if expert else 'layer'
117
+ elif scope!=('expert' if expert else 'layer'):raise ValueError('Mixed projection codebook scopes')
118
+ fp16=d.get('scale_dtype','fp8_e4m3')=='fp16'
119
+ if proj=='w13':fp16_blocks=fp16
120
+ elif fp16_blocks!=fp16:raise ValueError('Mixed projection scale precision')
121
+ if fp16 and not torch.equal(d['s'],d['s'].half().float()):raise ValueError('Scales are not FP16-representable')
122
+ stored_scales=d['s'].half() if fp16 else d['s'].to(torch.float8_e4m3fn).view(torch.uint8)
123
+ t={pre+'packed':torch.stack([pack(a,b,N,K) for a,b in zip(d['a'],d['b'])]),
124
+ pre+'codebooks':pack_codebooks(d['c0'],d['c1']),
125
+ pre+'scales':stored_scales.reshape(E,N//16,16,K//128).permute(0,1,3,2).contiguous(),
126
+ pre+'global':torch.tensor([d['glob']],dtype=torch.float32)}
127
+ for e in range(E):
128
+ a,b=unpack(t[pre+'packed'][e],N,K);assert torch.equal(a,d['a'][e]) and torch.equal(b,d['b'][e])
129
+ for e in sorted({0,E//2,E-1}):
130
+ c0,c1=(d['c0'][e],d['c1'][e]) if expert else (d['c0'],d['c1'])
131
+ ref=((c0[d['a'][e].long()]+f*c1[d['b'][e].long()])*d['s'][e].repeat_interleave(16,1)[...,None]*d['glob']).reshape(N,K)
132
+ assert torch.equal(ref,decode(t,L,proj,e,N,K,rss))
133
+ base_fmt='rvq256_256x8_expert_fp16block' if fp16 else ('rvq256_256x8_expert' if expert else 'rvq256_256x8')
134
+ arvq_fmt=base_fmt+(f'_rs{int(2**rss)}' if rss else '')
135
+ name=f'arvq-layer-{L:03d}-{tag}.safetensors';tmp=dest/(name+'.tmp')
136
+ save_file(t,str(tmp),metadata={'format':'pt','arvq_format':arvq_fmt,'residual_scale_shift':str(rss),'codebook_scope':scope,'fit':str(original.get('fit','initial Hessian fit, no PV')),'serving_compatibility':'requires FP16 block-scale v4 loader/kernel' if fp16 else ('requires per-expert v3 loader/kernel' if expert else 'requires 8+8 loader/kernel')})
137
+ tmp.replace(dest/name);del t,d
138
+ loaded=load_file(str(dest/name));assert loaded[pre+'packed'].shape==(E,N//16,K//64,64);del loaded
139
+ files[proj]={'file':name,'sha256':sha(dest/name)}
140
+ base_mfmt='rvq256_256x8_expert_fp16block' if fp16_blocks else ('rvq256_256x8_expert' if scope=='expert' else 'rvq256_256x8')
141
+ m={**original,'format':base_mfmt+(f'_rs{int(2**rss)}' if rss else ''),'version':4 if fp16_blocks else (3 if scope=='expert' else 2),'residual_scale_shift':rss,'block_scale_dtype':'float16' if fp16_blocks else 'float8_e4m3fn','codebook_scope':scope,'bits':2.125 if fp16_blocks else 2.0625,'codebook_sizes':[256,256],'index_bits':[8,8],'words_per_tile':64,'files':files,'fit':str(original.get('fit','initial Hessian fit, no PV')),'production_ready':False,'roundtrip':'all expert indices exact; 3 decoded weights/projection exact','requires_8x8_loader_kernel':True}
142
+ (dest/f'layer-{L:03d}-manifest.json').write_text(json.dumps(m,indent=2));return m
reproduce/source/btx53/arvq88/perf/FP16_BLOCK_SCALE_HANDOFF.md ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ARVQ FP16 block-scale format (v4)
2
+
3
+ This replaces FP8 E4M3 block scales with FP16 block scales. There are no added row scales. Books remain per-expert packed FP4, indices remain two 8-bit indices per group of eight weights, and global scale remains FP32 per projection. Initial FP8 values are exactly representable in FP16.
4
+
5
+ Checkpoint tensors use the existing `arvq_{w13,w2}_scales` names and logical tiled shape `[cold_experts, N/16, K/128, 16]`, now dtype F16 instead of U8 (FP8 bytes). Manifest version4, format `rvq256_256x8_expert_fp16block`, `block_scale_dtype: float16`. Packed codebooks and indices retain v3 layout. Index+block-scale storage is2.125 bits/weight, excluding books/global metadata. Each scale belongs to one expert, one output row, and128 input weights. Scales cannot be folded into one whole-row epilogue multiplier.
6
+
7
+ Decode for a group of8 weights:
8
+ `W[e,r,8*g:8*g+8] = global * scale16[e,r,g//16] * (book0[e,index0] + book1[e,index1])`.
9
+
10
+ The existing native MMA instruction accepts UE4M3 weight scales, not FP16. A compatible native path must preserve FP16 scale precision, e.g. use unit weight scales in FP4 MMA, accumulate the128-weight block (two K64 tiles, both codebooks and activation planes), multiply its FP32 contribution by the FP16 block scale converted to FP32, and accumulate blocks. Do not silently cast the new scales back to E4M3. Activation quantization and hot NVFP4 paths remain unchanged. Benchmark native latency and validate arithmetic before publication.
11
+
12
+ Training: `/tmp/glm53-fp16-sequential`, sequential layers3–77 from original pre-PV initialization, full18,001,846-token corpus, batch262144/micro65536, Adam and periodic output-gradient index reassignment. Trainable latent log-scales are FP32; every forward/export projects scales to FP16. Global scale is frozen. Max69 updates; three validation checks >0.1% worse trigger a4x LR drop, otherwise drop after45; stop after15 lower-LR updates without a new best. Validation, development-audit and export-replay gates remain enforced.
13
+
14
+ Current fitting/export modules:
15
+ - `sequential_pv_full_corpus.py`: config `block_scale_dtype: fp16` selects precision.
16
+ - `arvq88/pv.py`: encoded `scale_dtype` travels with the representation and is honored during replay/propagation.
17
+ - `arvq88/pack.py`: F16 tiled scales, v4 metadata and exact decoded-weight roundtrip.
18
+
19
+ Local fitting and propagation can proceed. Automatic HF publication is disabled until the loader/kernel supports this format. Do not publish v4 tensors under v3 metadata.
reproduce/source/btx53/arvq88/perf/FULL_CORPUS_PIPELINE.md ADDED
@@ -0,0 +1,80 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Full-corpus sequential PV
2
+
3
+ Run `full_pipeline.py --work WORK --first 3 --last N` with the repository Python.
4
+ The work directory needs config.json, the initial fixed capture3, and token file.
5
+ The driver runs eight GPU ranks per stage and never uploads. Set explicit last
6
+ layer: a bounded pilot is the default operational choice until it qualifies.
7
+
8
+ The existing layer-3 fit uses 18,001,846 training tokens, one shuffled sequence
9
+ pass, 262144 effective batch, 65536 microbatch, Adam codebook LR .048 and scale
10
+ LR .032. Gradient accumulation normalizes every microbatch by the full batch
11
+ size; the regularizer, clipping and Adam step execute once per effective batch.
12
+ Validation every five updates retains the best state. Audit is evaluated on the
13
+ retained state; the driver refuses propagation if validation or audit regresses.
14
+
15
+ Capture stores student normalized inputs/routing, frozen branch output, the
16
+ reference-minus-frozen target, and the BF16 reference output. Both full reference
17
+ and full student trajectories advance one block at a time. Layer 3 needs a
18
+ one-time recapture because the earlier training capture omitted frozen/reference
19
+ outputs, which cannot be reconstructed from their difference alone. Dense prefix
20
+ weights now remain resident during that recapture. Layers 4 onward read cached
21
+ boundaries and never execute the earlier prefix.
22
+
23
+ Propagation uses the retained serialized FP4 books, FP8 scales and indices with
24
+ the same activation emulation as training. It caches decoded expert weights per
25
+ GPU across all chunks. A fixed-evaluation replay gate allows relative difference
26
+ at most 1e-4 after BF16 output rounding, accounting for different FP32 reduction
27
+ order; it does not claim bitwise parity or native SM120 execution. Partial final
28
+ sequences are zero-padded at boundaries and masked from capture/training losses.
29
+
30
+ Rank zero alone assembles and prefetches each CPU training batch, then broadcasts
31
+ identical tensors to the expert ranks over NCCL; capture and propagation shard writes
32
+ run on one background thread with bounded outstanding data. Metadata is committed
33
+ only after the data file is atomically renamed. A partial smoke run cannot publish
34
+ a rank-complete marker. `timings.json` records rank-0 data wait, transfer, optimizer,
35
+ reassignment and validation/checkpoint time; these are not all-rank max timings.
36
+
37
+ The earlier monolithic 256K run exhausted GPU memory on fresh data. Its first five
38
+ updates took 158.9 seconds. The 64K accumulated version with prefetch took 65.0
39
+ seconds with validation 0.0047233582 versus 0.0047233591 (relative difference
40
+ ~2e-7). This is an observed first-five-update comparison, not a full-run speedup.
41
+
42
+ Pending qualification: the scheduled 262144-token integration smoke exercises
43
+ capture3 -> serialized propagation3 -> capture4. Full-layer transition throughput
44
+ must be measured before extrapolating all 75 layers. CUDA graphs are not enabled:
45
+ variable routing and fresh sequence batches require a separately validated static
46
+ buffer/shape design. Native serving-kernel numerical qualification remains separate.
47
+
48
+ The driver keeps two layers of regenerable full-corpus intermediates and all fit
49
+ outputs/checkpoints. It prunes only directories with its own completed stage
50
+ receipt after a later layer passes held-out checks; it never prunes the original
51
+ legacy training_capture directory or fixed validation/audit captures.
52
+
53
+ New captures store normalized inputs as BF16 only after an exact FP32 roundtrip
54
+ assertion for every chunk; arithmetic restores FP32 before use. This halves input
55
+ storage and CPU assembly traffic without rounding any captured input values.
56
+
57
+ ## Capture optimizations added during the live campaign
58
+
59
+ Layer-14 measurements: capture setup ~90 s, capture computation ~154 s,
60
+ CPU copying/statistics ~99 s, input transfers ~31 s. Separate fixed-evaluation
61
+ capture costs another ~100 s. Propagation CPU boundary packing costs ~27 s.
62
+
63
+ New capture uses two reusable pinned host banks and asynchronous CUDA copies,
64
+ with an event that the shard writer waits on before serialization. The writer's
65
+ single-outstanding-job contract prevents overwriting a bank still being saved.
66
+ Each rank compares its first staged production chunk exactly with the original
67
+ blocking CPU transfer. A real CUDA test also exercises six writes across bank
68
+ reuse and verifies every saved value. Propagation now copies the full output in
69
+ one transfer and reshapes complete sequences without CPU copies; partial tails
70
+ retain explicit zero padding and have exact parity tests.
71
+
72
+ Full training capture also produces a candidate fixed-evaluation capture using
73
+ resident weights. The existing standalone capture still runs for the first
74
+ qualification layer, compares every tensor on all eight ranks, and publishes
75
+ fused_eval_approved.json only on exact equality. Subsequent fixed capture stages
76
+ reuse that candidate instead of reloading the reference/donor weights. If parity
77
+ fails, standalone capture continues to supply the production tensors.
78
+
79
+ Production speedups are not yet measured. Inspect training_capture*/rank*_timings.json,
80
+ fused_eval_qualification, fused_eval_approved.json and pipeline stage receipts.
reproduce/source/btx53/arvq88/perf/README.md ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PV performance investigation (18 September 2026, Melbourne)
2
+
3
+ The active sequential quality pilot was not modified. These files are isolated
4
+ prototypes and a separately invocable trainer fork; no changes were published.
5
+
6
+ ## Measured single-expert forward + backward
7
+
8
+ Actual layer-3 expert 3, 64 routed training rows, original fitted per-expert
9
+ books, float32 GEMMs and the same FP4/FP16 activation emulation. GPU 0 was shared
10
+ with the running pilot. Four timing repeats per path after warmup, no optimizer
11
+ step/communication/validation included. Results repeated in two invocations.
12
+
13
+ | Path | Median, second invocation |
14
+ | --- | ---: |
15
+ | Original eager arithmetic, four 16-row microbatches | 386.47 ms |
16
+ | Original eager arithmetic, one 64-row batch | 132.02 ms |
17
+ | Graph-compatible arithmetic, graphed one-batch forward/backward | 119.08 ms |
18
+
19
+ Larger batch: 2.93x throughput for this expert calculation. Graph replay adds
20
+ about 10% lower latency relative to eager one-batch; combined factor 3.25x.
21
+ These are NOT full-layer or whole-campaign speedup measurements.
22
+
23
+ Graph versus eager one-batch output and parameter-gradient relative differences
24
+ were zero in this test. After in-place index and row-scale changes, replay still
25
+ matched eager output exactly. Four-microbatch versus one-batch maximum parameter
26
+ -gradient relative difference was 5.99e-5 (FP32 accumulation/GEMM differences),
27
+ with losses 0.1189905256 versus 0.1189905405. End-to-end optimization trajectories
28
+ therefore need comparison, not an assertion of bit-identical training.
29
+
30
+ Artifacts: `/tmp/glm53-pv-graph-benchmark.json` and
31
+ `/tmp/glm53-pv-graph-benchmark-check.json`. Benchmark peak allocated memory:
32
+ 0.514 GiB, isolated expert only; not a full-layer graph memory estimate.
33
+
34
+ ## Why larger batches help
35
+
36
+ The production pilot calls local_prediction separately for each of four
37
+ microbatches. Every call reconstructs full expert matrices and backpropagates
38
+ through their codebook gathers. We repeat this large decode/scatter workload
39
+ four times despite unchanged parameters within the optimizer step. The GPUs
40
+ can report high utilization while doing this redundant work.
41
+
42
+ `sequential_pv_fast.py` is a separate fork with configurable `--microbatch`
43
+ (default 1024 versus the original 256). Global batch, objective, selected rows,
44
+ regularization and update count remain unchanged. Full-layer peak-memory and
45
+ held-out parity/timing tests are needed before deployment.
46
+
47
+ ## Communication optimization
48
+
49
+ The pilot broadcasts an 8192 x 6144 float32 delta (192 MiB) for each proposal
50
+ attempt, mostly zeros outside an expert's routed rows. `sparse_delta.py` sends
51
+ only row IDs and their nonzero payload, then reconstructs the same dense buffer
52
+ before the unchanged acceptance arithmetic. For eight-of-256 routing the average
53
+ payload is roughly 1/32 as large, though expert popularity varies.
54
+
55
+ Two-rank CPU distributed tests confirm bit-exact reconstruction versus dense
56
+ broadcast, for either owner and empty row sets. This optimization is included
57
+ in the separate fast trainer fork. Eight-GPU NCCL qualification remains pending.
58
+
59
+ ## CUDA Graph integration boundary
60
+
61
+ `graph_expert.py` and `bench_graph_expert.py` demonstrate captured forward and
62
+ backward with live index/parameter buffers. They do not graph the production
63
+ trainer. Original per-expert use.any(), nonzero(), variable routed batch lengths,
64
+ and finite-value host checks prevent simply wrapping the loop in CUDAGraph.
65
+
66
+ Next integration: precompute routing outside capture, use fixed/padded row-count
67
+ buckets with correct zero masks, capture expert forward/backward, and keep the
68
+ coupled-output all-reduce and discrete acceptance outside the initial graph.
69
+ Stable parameter/code buffer addresses and graph-memory budgets are required.
70
+ Validate finiteness outside captured execution; do not silently remove checks.
71
+
72
+ ## Verification and next qualification
73
+
74
+ - Graph-compatible activation forward and STE derivative match the existing path.
75
+ - Sparse communication matches dense broadcast exactly in two-rank tests.
76
+ - Real GPU graph benchmark checks loss, gradients, and live code/scale updates.
77
+ - All tests in `perf/test_perf.py` pass; new modules compile.
78
+
79
+ After the quality pilot, compare baseline and fast fork on the same eight-GPU
80
+ layer and sample schedule, including an actual reassignment cycle. Measure
81
+ continuous steps, proposals, validation, peak memory and held-out error separately.
82
+ Then integrate graph buckets if their additional gain justifies the complexity.
83
+ No full-model runtime estimate should be reduced by the single-expert 3.25x factor.
84
+
85
+ ## Follow-up: specialized codebook backward
86
+
87
+ An actual CPU/CUDA operator profile (`/tmp/glm53-pv-operator-profile.txt`)
88
+ identified general advanced-index backward as 83% of self CUDA time in the
89
+ measured expert step; matrix multiplication accounted for <1%. This motivated
90
+ using `torch.nn.functional.embedding` for each book lookup, instead of `cb[a]`.
91
+ It changes the lookup/reduction implementation, not the representation or loss.
92
+
93
+ Two separate real-expert benchmarks while the pilot was active:
94
+
95
+ | Layer/expert | Four microbatches, original | One batch, original | One batch, embedding | Embedding + graphs |
96
+ | --- | ---: | ---: | ---: | ---: |
97
+ | 3 / 3 | 442.67 ms | 115.18 ms | 35.39 ms | 25.37 ms |
98
+ | 4 / 0 | 373.97 ms | 133.65 ms | 35.19 ms | 34.75 ms |
99
+
100
+ These include contention and are single-expert forward/backward measurements.
101
+ The combined ratio is about 11–17x for this path, NOT the whole PV campaign.
102
+ Graph benefit is variable here; specialized embedding backward and batching are
103
+ more consistently material than launch replay alone.
104
+
105
+ Outputs and losses match exactly between original one-batch, embedding and
106
+ graphed embedding. GPU parameter-gradient relative differences are ~1.5–2.2e-6,
107
+ consistent with a changed FP32 reduction order. CPU float64 tests independently
108
+ verify matching parameter gradients, including after changing index buffers.
109
+ Graph replay after changing indices and row corrections still matches eager.
110
+
111
+ Artifacts: `/tmp/glm53-pv-embedding-benchmark.json` and
112
+ `/tmp/glm53-pv-embedding-layer4.json`. `EmbeddingRowProjection` implements the
113
+ isolated prototype; the specialized embedding lookup is now also present in
114
+ `sequential_pv_fast.py`. All three performance tests pass. The active pilot still
115
+ uses its original code. Full-layer NCCL, memory, quality and timing qualification
116
+ is required before switching a production run to the faster fork.
117
+
118
+ ## Qualified full sequential campaign (18 September 2026)
119
+
120
+ The completed faster rerun matched both validation and development-audit error
121
+ exactly for reference layers 3 and 4. All propagated BF16 values matched exactly.
122
+ Measured training loops (including reassignment/validation) improved from
123
+ 1525.16 to 336.60 seconds for layer 3 and 1226.31 to 303.00 seconds for layer 4.
124
+ This qualifies the larger microbatch, embedding backward and sparse broadcasts;
125
+ CUDA graphs remain a separate prototype and are not in the full trainer.
126
+
127
+ `full_reference.py --work /tmp/glm53-sequential-reference-full --pilot
128
+ /tmp/glm53-sequential-pilot-fast` reuses those two fitted layers and chains capture
129
+ and 200-step reference tuning through layer 77. `sequential_capture.py` now reads
130
+ reference and retained student outputs from the immediately preceding layer.
131
+ It retains the original corpus, allocation, seeds and tuning hyperparameters.
132
+ An additional initial development-audit evaluation permits an audit regression
133
+ gate before propagation/publication. Validation and audit gates allow only
134
+ 1e-6 relative numerical slack; failure stops for investigation. The audit is
135
+ therefore development data, not an untouched final generalization estimate.
136
+
137
+ `publish_reference.py --work /tmp/glm53-sequential-reference-full` runs separately
138
+ on CPU/network. It reconstructs cold-slot order across eight expert shards,
139
+ checks slot coverage/allocation/global metadata, invokes the existing exact
140
+ index roundtrip and sampled decoded-weight parity checks, then atomically
141
+ replaces both projection files and reports on HF with parent-commit protection.
142
+ Remote hashes are verified before recording completion. Progress explicitly
143
+ identifies remaining legacy-PV and initial-fit layers. The old campaign and
144
+ publisher remain stopped. Full-model quality/native-SM120 qualification remains
145
+ pending; neither pilot proves that outcome.