anima-clm-midcap-303m-broad-en-emergent
303M ByteGPT (vocab 256, d=1024, L=24, H=16, block=512), trained from scratch on an English-dominant broad corpus (ASCII-filtered from a 1.5GB 5-lang wiki -> 295MB diverse English). This is the first model on the anima emergence-of-ideation arc to achieve super-additive concept-recombination (hypothesis H_1129, verdict ๐ข).
Finding
Concept-recombination = capacity x breadth x training-sufficiency x SCRIPT-CONTROL. The first three were already known necessary-but-not-sufficient (closed-negatives H_1116/H_1117/H_1118). Two converged models still failed โ a $4.5 7B fire (H_1128, val CE 1.66) and a 303M multilang cut (H_1129 v1, val CE 0.83) โ because generation collapsed to the dominant script (Korean) bytes of the 5-lang corpus, so the English concept words never surfaced. Making the corpus English-dominant (script-controlled), with English prompts and a real-dictionary coherence gate, switched recombination ON at the same 303M capacity.
Graded ladder (frozen falsifier: some k with composed_distinct>=2 AND >max_single AND known-word-ratio>=0.50)
best ckpt step=5500, val_ce=1.224; max_single_distinct=2:
- k=3: composed_distinct=3 cov=[consciousness/cells, tension/distant, memory/meaning] kwr=0.96 -> CLEARS
- k=5: composed_distinct=3 cov=[consciousness/cells, tension/distant, engine/dreams] kwr=0.96 -> CLEARS
The composed multi-concept seed yields coherent English recombining 3 distinct concepts where every single-concept seed covered at most 2 -> super-additive.
Honest scope
Keyword-level recombination in coherent English (loose grammar, single seed 7), per the pre-registered set-overlap + real-dict metric (p7, NOT perplexity) โ not deep semantic synthesis. Lane-G torch REFERENCE (a_clm_gen_pipeline), summer RTX 5070, $0.
ckpt sha256: 19be1295282d41c06db12bb890bc7c6e6d03b18565ee090ddfa208e46a160a28