nla-qwen2.5-7b-L20 AV โ€” v2 RL, per-item length penalty target=10 (smoke, 60 steps)

Experimental reference checkpoint. The v2 warm-start AV (syvb/nla-qwen2.5-7b-L20-av-matryoshka-sonnet46-v2) after 60 steps of v2 RL on a 4xH100 (actor2/critic1/rollout1 โ€” wiring-validation sharding, not the canonical critic2).

Config: GRPO; UNIFORM item-truncation (taper=1.0, max_items=10); KL=0.02 (2x v1); per-item length penalty coef=0.002, TARGET=10 tokens/item (token-efficiency pressure); rollout 16x8=128, lr 1e-5.

Result (100 held-out docs) โ€” items got terser without quality loss:

  • per-item tokens: mean 14.0 -> 9.7, median 12 -> 9, p99 33 -> 21, max 222 -> 36
  • round-trip reward (-MSE) improved -0.46 -> -0.32; fve_nrm ~0.55-0.62; 0 CJK.

Matched readout / critic: use syvb/nla-qwen2.5-7b-L20-ar-matryoshka-sonnet46-v2 (this run's co-trained AR DCP was left incomplete by a post-save teardown crash, and is only 60 steps off the warm-start AR โ€” not exported).

Downloads last month
4
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for syvb/nla-qwen2.5-7b-L20-av-matryoshka-sonnet46-v2-rl-target10

Collection including syvb/nla-qwen2.5-7b-L20-av-matryoshka-sonnet46-v2-rl-target10