Matryoshka NLA (Qwen2.5-7B L20)
Collection
An NLA trained with random-length truncation of the verbalizer's output, so the most important information comes first. Checkpoints, data, baseline. โข 8 items โข Updated
Actor (AV) of the NLA Qwen2.5-7B
(layer-20) pair, v3 warm-start: SFT from the base
kitft/nla-qwen2.5-7b-L20-av
on Claude Sonnet-4.6 explanations in the v3 format.
v3 changes vs v2 (see the repo's experiment README):
nla_meta.yaml
sidecar carries the real trained template (v1/v2 shipped a stale tagged one).Warm-start data: v3 splits (v3/ folder).
Hyperparameters: global batch 256, lr 2e-5โ2e-6 cosine, warmup 50, 1 epoch, injection_scale 150.