# Claim audit and provenance ledger This document records the manuscript's principal empirical claims, the public source page from which each value was taken, and the wording boundary preserved in the paper. It is meant to prevent accidental strengthening during later edits, announcements, or peer-review revisions. ## Library scope **Public index:** https://porkicoder.com/research/ The index states that the library contains ten published studies: nine tab-title studies and one matched-control study. It also explicitly requires the SO-board and Terminal-board to remain separate. The paper therefore synthesizes ten study pages. The user's count of eleven is reconciled by counting the index page as the eleventh public research page, not by inventing an extra study. ## Major claims | Topic | Public source | Value retained | Boundary that must remain | |---|---|---|---| | Original seed-lottery observation | https://porkicoder.com/research/the-sniff-test.html | Repeated titles fell from 145/907 (16.0%) to 92/907 (10.1%) in one fixed pair; mean teacher similarity changed by -0.00066. | This is behavior of two realized checkpoints, not evidence that literary sequence order caused the effect. | | Matched seed control | https://porkicoder.com/research/the-sniff-test.html | Thirty-one runs completed. Ordered Verne averaged 14.75% repetition versus 13.79% for shuffled Verne and was worse in 6/9 ordered-versus-shuffled cells. | A mechanism screen over three body and three donor seeds, not a population estimate across architectures or corpora. | | Early SO-board comparison | https://porkicoder.com/research/tab-titles-flan-centroid.html | 12.7M student 3.01; centroid 4.21; early 35M+centroid 4.72; raw FLAN 5.22; FLAN+centroid 5.43 on the same 1,000-row packet. | Same-packet SO-board result only. The usefulness threshold in this early study is not merged with later `>=6` rates. | | Mid-GSG improvement | https://porkicoder.com/research/tab-namer-mid-gsg.html | B-9500 with old centroid 5.65 versus Hybrid A 5.10 and raw FLAN 5.64; later B-9500+v2 5.96 versus FLAN+v2 6.11. | Development comparison, later superseded by beam search; no terminal-domain claim. | | Beam-4 sealed relative result | https://porkicoder.com/research/tab-namer-four-beams.html | On sealed SO 1,000, 35M beam-4+v2 scored 6.20 versus FLAN+v2 5.45; paired +0.752, 95% bootstrap interval [0.611, 0.891], W/T/L 603/67/330. | Always say SO-board. This is a relative win under the shared packet, not a general claim that the 35M model beats FLAN. | | Beam-4 absolute gate | https://porkicoder.com/research/tab-namer-four-beams.html | Identity-blind Codex audit: 19/40 useful, below the preregistered 24/40 gate; later dev audit 80/200 useful. | Codex is an auxiliary LLM auditor, not a human evaluator. The conjunctive final gate failed. The sealed set was consumed. | | Decoder versus weights | https://porkicoder.com/research/tab-namer-four-beams.html | Beam-4 beat corrected greedy by roughly +0.45 to +0.48; every scored continuation under beam-4 lost to frozen B-9500. | Attribute the gain to decoding. Do not describe it as a new checkpoint or parameter improvement. | | Domain transfer | https://porkicoder.com/research/session4_results.html | Canonical SO gap: FLAN+v2 6.11 versus pin 5.96 (-0.15). Canonical Terminal gap: 7.11 versus 5.94 (-1.17). | Absolute means come from different boards and packets. Compare only within-board gaps; never average or rank across boards. | | Selective terminal SFT | https://porkicoder.com/research/session4_results.html | r3 5.67 versus pin 5.56, paired +0.12 [0.04, 0.20]; earlier independent packet +0.07 [0.01, 0.14]. | A small replicated improvement over the local pin, not a close race with FLAN on Terminal-board. | | Six-teacher judged SFT | https://porkicoder.com/research/tab-namer-techniques.html | 6t improved over the pin by +0.19 [0.14, 0.25] on a 1,000-row packet. | Keep the negative registry: self-distillation, hinge, unconstrained decode, terminal IDF, and tested rankers did not add a reliable gain. | | Large weak-label negative result | https://porkicoder.com/research/2026-08-16_g300k-flashlite-test.html | 219,800 raw rows; 175,642 legal titles; best CE 0.91; best judged arm 6.05 versus 6t 6.16, paired -0.112. | The finding is that this labeling/filtering/CE recipe failed. It does not prove that large synthetic datasets are generally harmful. | | Pointer model | https://porkicoder.com/research/session9_results.html | P1+Hybrid A 6.43 versus 6t 6.30 (+0.13 [0.04, 0.22]); FLAN 7.02 (-0.60 [0.49, 0.70]); 24.5M live parameters. | Sealed-v2 Terminal-board result. Quality and size do not by themselves establish the final product choice. | | Common Session-10 board | https://porkicoder.com/research/session10_results.html | Raw title-SFT FLAN 7.28; P1+HA 6.79; 6t+HA 6.65; stock FLAN 3.12 on reusable-dev 500. | Same-packet quality comparison. Do not compare those means directly with Session-11 holdouts. | | Quality-latency trade-off | https://porkicoder.com/research/session10_results.html | P1 3.88 ms, 6t 10.1 ms, CPU picker 11.72 ms, neural picker 17.5 ms, title-SFT FLAN 29.3 ms, stock FLAN 25.7 ms on Apple M4 Max CPU, one thread. | The published 6t quality and latency correspond to different precision artifacts; the SFT INT8 payload is estimated. Session 11 inherited, rather than reran, the timing. | | Competence-dependent glue | https://porkicoder.com/research/session10_results.html | On reusable-dev, SFT+HA 7.41 versus raw 7.28; on fresh confirm-v1, raw 6.55 versus HA 5.78. | Heuristic glue is conditional. Do not promote the reused-packet gain over the fresh confirmation reversal. | | Final Session-11 result | https://porkicoder.com/research/session11_results.html | Holdout 1000: 7.59 versus 7.38 (+0.21), W/T/L 296/551/153. Holdout 1000b: 7.55 versus 7.34 (+0.20), 281/570/149. | These two fresh holdouts are the paper's main claim. Keep them separate from Session-10 reusable-dev means and from the closed terminal holdout 200. | | Final architecture and product state | https://porkicoder.com/research/session11_results.html | Same untied 77M graph; four epochs, lr 1e-4, seed 1, flat weights; no Hybrid A; inherited 29.3 ms and ~193 MB INT8 estimate. Public 2.17.42 still served 6t at publication. | “Selected for shipment” or “new ship” is accurate for the research decision. Do not state that the public installer had already shipped the new worker. | | Session-11 data ledger | https://porkicoder.com/research/session11_results.html | Published note states 7,623 starting rows, 1,148 task-hash leaks dropped, and 6,419 retained rows (372 gold, 6,047 teacher), split 6,219/200. | These counts do not arithmetically reconcile. The manuscript explicitly reproduces them as reported and does not infer an undocumented filtering step. | ## Evaluation boundaries preserved in the manuscript 1. **Board separation:** SO-board and Terminal-board are independent leaderboards. 2. **Packet separation:** judge means are compared only when systems appeared in the same blinded packet. 3. **Judge identity:** Gemini Flash-Lite is the primary LLM judge; Codex is an auxiliary LLM auditor. Neither is described as human evaluation. 4. **Threshold separation:** early `useful >=4`, later `solid >=6`, and Codex 0-to-4 usefulness are different instruments. 5. **Fresh evidence priority:** the headline rests on two Session-11 1,000-row holdouts, not the repeatedly reused development packets. 6. **Closed data:** terminal holdout 200, sealed SO data after consumption, raw private prompts, and employer-identifying examples are not released. 7. **Latency precision caveat:** 6t quality and latency are not measured on the same precision artifact; Session-11 timing is inherited. 8. **Model-loading caveat:** the final FLAN graph requires untied output weights (`tie_word_embeddings=false`). 9. **Scope:** conclusions apply to short English coding-session titles under the stated contract, not general summarization or open-ended generation. 10. **No universal synthetic-data claim:** the 176k result rejects one weakly filtered teacher-CE recipe, not synthetic supervision as a class. ## High-risk phrases to avoid - “We beat FLAN” without naming **SO-board** and the exact shared stack. - “Human audit” for the Codex audit. - “The literary initialization works” or “Verne reduced repetition” as a mechanism claim. - “More data hurts” as a universal conclusion. - “Session 11 is already in the public installer.” - “29.3 ms was remeasured in Session 11.” - “All 11 studies” unless the index page is explicitly included in that count. - Any average that combines SO-board and Terminal-board scores.