MC7ever commited on
Commit
6baebab
·
verified ·
1 Parent(s): 5f5a478

card: certification status, voided runs, layer-sweep negative result

Browse files
Files changed (1) hide show
  1. README.md +13 -0
README.md CHANGED
@@ -56,6 +56,19 @@ are width-1024 — cross-width tensors cannot be weight-averaged. L/14
56
  (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video
57
  share width/depth/patch, so true weight fusion is possible.
58
 
 
 
 
 
 
 
 
 
 
 
 
 
 
59
  ## How it was built
60
 
61
  1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
 
56
  (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video
57
  share width/depth/patch, so true weight fusion is possible.
58
 
59
+ ## Status (2026-10-03): this tag is the grid-search winner, still current best
60
+
61
+ - Evolution rerun in progress after two voided attempts (documented below) —
62
+ if it certifies a better genome, that becomes the next revision.
63
+ - Layer-depth probe: naive mid-network readout does NOT beat the joint output
64
+ (t2i R@1 ≤ 0.06 vs 0.79) — mid-network features need trained per-layer
65
+ pooling (fine-tune phase), not a free readout change.
66
+ - Voided runs: (1) blend-of-blend contamination via shared-storage state_dict;
67
+ (2) silent no-op `load_state_dict` on this model's dual-prefix key layout;
68
+ (3) joint-path audio mismatch that faked ESC collapse at high audio weight.
69
+ All fixed (pristine clones, `copy_` by param name, standalone-tower-only
70
+ audio mapping) before trusting any evolution number.
71
+
72
  ## How it was built
73
 
74
  1. Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio