Matryoshka NLA (earlier tests)
Collection
v2 item-truncation matryoshka NLA (Qwen2.5-7B L20): warm-start AV verbalizer + AR critic checkpoints, warm-start data, base NLA. โข 10 items โข Updated
Experimental reference checkpoint. The v2 warm-start AV
(syvb/nla-qwen2.5-7b-L20-av-matryoshka-sonnet46-v2) after 60 steps of v2 RL on a
4xH100 (actor2/critic1/rollout1 โ wiring-validation sharding, not the canonical critic2).
Config: GRPO; UNIFORM item-truncation (taper=1.0, max_items=10); KL=0.02 (2x v1); per-item length penalty coef=0.002, TARGET=10 tokens/item (token-efficiency pressure); rollout 16x8=128, lr 1e-5.
Result (100 held-out docs) โ items got terser without quality loss:
Matched readout / critic: use syvb/nla-qwen2.5-7b-L20-ar-matryoshka-sonnet46-v2
(this run's co-trained AR DCP was left incomplete by a post-save teardown crash, and is
only 60 steps off the warm-start AR โ not exported).