porkicoder.com · research index · 19 August 2026 · synthesis of the public library, not a new experiment

Local Tab-Title Generation under Product Constraints

A synthesis of eleven public PorkiCoder research pages. MD Ishtiaque Hossain, PorkiCoder Research, Vancouver. This document is a reading of the published notes; it does not add unpublished numbers.

Abstract. Naming a coding-agent terminal tab is extreme summarization under a product contract: one to three words, at most 32 characters, locally generated, useful enough to rediscover a session. The public PorkiCoder library contains ten studies plus the research index. This paper reads them as one campaign, not one preregistered experiment. A 12.7M student looked like a literary-init win until a 31-run matched control killed the mechanism. On a Stack Overflow proxy, constrained four-beam search on an unchanged 35M model beat title-tuned FLAN by a paired +0.752, then failed a 60% absolute usefulness gate and did not transfer to real terminal tasks. Selective judged labels moved the local pin a little; 175,642 weak teacher labels lowered CE and not utility. A 24.5M pointer system won the latency frontier (3.88 ms) but not quality. A 77M Title-SFT FLAN model became the quality ship. Continuing it on a leak-checked Aug-12 mix improved two never-seen 1,000-row holdouts by +0.21 and +0.20 (7.59 / 85.7% and 7.55 / 84.5%). Decoding, dialect, board, glue, and leakage control moved the product more than raw label count. SO-board and Terminal-board stay separate.

1. What this is

The product is PorkiCoder, a local-first coding environment. Many concurrent agent sessions need a tab title such as Fix SSE Disconnect, not Terminal 7. The incumbent public worker at ship time was still the 35M 6t ONNX path; the research decision in Session 11 selected Title-SFT FLAN raw, arm flat1e4 (76,961,152 parameters, untied lm_head).

Eleven public pages: the research index, the sniff / matched-control paper, FLAN+centroid, mid-GSG, four beams, Session 4 transfer, the techniques log (Sessions 5–8), the 176k Flash-Lite test, Session 9 pointer, Session 10 quality vs ms, and Session 11 ship. Sessions 5–8 live in one techniques page. That is the “eleven” count. There is no twelfth unpublished study in this synthesis.

Evidence levels used below: exploratory diagnosis; same-packet paired comparison; sealed or first-look packet; replicated fresh holdouts. Absolute judge means are not compared across packets or boards.

2. Task, boards, judge

Contract. 1–3 words, ≤32 characters, [A-Za-z0-9 ], Title Case. Prefer two words and on-task content words. Input is always title: agent=<a> [cwd=<c>] task=<z> with the first paragraph then 400 characters. Exact match has zero product weight.

Two boards. SO-board: Stack Overflow page bodies; clean-dev 1,000 and one consumed sealed 1,000. Canonical early gap: FLAN+centroid v2 6.11 vs 35M pin 5.96 (−0.15). Terminal-board: redacted coding-agent tasks. Canonical early gap: FLAN+centroid v2 7.11 vs pin 5.94 (−1.17). Never average the boards. An exact-three-word rule once gained 0.39 on SO and lost 0.36 on Terminal.

Judge. Gemini 3.5 Flash-Lite, identity-blinded, 1.00–10.00. Later product reports treat ≥6 as solid. Codex was an auxiliary auditor on the four-beam absolute gate, not a human study. Session 11 clocks (29.3 ms) were inherited from Session 10: same 77M graph, not remeasured.

3. Session ledger

PageWhat was testedWhat to keepEvidence
Sniff / matched control 12.7M student; Jules Verne init vs shuffle One pair cut repeats 145/907 → 92/907; mean teacher S almost unchanged (−0.00066). 31-run screen: ordered Verne 14.75% repeats vs shuffled 13.79%, worse in 6/9 cells. Mechanism claim dies. Matched control
FLAN + centroid 12.7M vs centroid vs early 35M vs FLAN Same 1,000-row SO packet: student 3.01; centroid 4.21; 35M+centroid 4.72; raw FLAN 5.22; FLAN+centroid 5.43. Extractive highlighter beat the tiny student. Same packet
Mid-GSG B-9500 35M + old vs v2 centroid B-9500 + old centroid 5.65 vs Hybrid A 5.10 vs raw FLAN 5.64. Later B-9500+v2 5.96 vs FLAN+v2 6.11. Development only; later superseded by beam search. Dev comparison
Four beams Decode-only change on frozen B-9500 Sealed SO 1,000: 35M beam-4+v2 6.20 vs FLAN+v2 5.45, paired +0.752 [0.611, 0.891], W/T/L 603/67/330. Beam-4 vs corrected greedy +0.45 to +0.48. Codex audit 19/40 useful, below 24/40 gate. Sealed set consumed. Sealed relative; failed absolute
Session 4 SO→Terminal transfer; judged SFT SO gap −0.15; Terminal gap −1.17. Selective SFT r3 vs pin +0.12 [0.04, 0.20] (earlier packet +0.07). Small win vs pin, not vs FLAN. Two boards; paired
Techniques (5–8) 6t judged SFT, glue, rankers 6t vs pin +0.19 [0.14, 0.25] on a 1,000-row packet. Self-distill, hinge, unconstrained 35M decode, terminal-IDF glue, tested rankers: no reliable add. Same packet + negatives
176k Flash-Lite Weak large teacher CE 219,800 raw / 175,642 legal titles. Best CE 0.91. Best judged arm 6.05 vs 6t 6.16, paired −0.112. This recipe failed. Not a proof that all synthetic data fails. Negative scale test
Session 9 Dense P1 pointer + Hybrid A Sealed-v2 Terminal: P1+HA 6.43 vs 6t 6.30 (+0.13); FLAN 7.02 (−0.60). 24.5M live params, 3.88 ms. Fast, not the quality ship. Sealed-v2
Session 10 Common packet, quality vs ms Reusable-dev 500: raw Title-SFT FLAN 7.28; P1+HA 6.79; 6t+HA 6.65; stock FLAN 3.12. Confirm-v1 335: SFT raw 6.55; SFT+HA 5.78 (glue reverses). Clocks: P1 3.88, 6t 10.1, SFT 29.3 ms (M4 Max, 1 thread). Do not mix these means with Session 11 holdouts. Same packet + fresh confirm
Session 11 Continue SFT on leak-checked Aug-12 mix Holdout 1000: 7.59 / 85.7% vs S10 7.38 / 83.5%, +0.21, 296/551/153. Holdout 1000b: 7.55 / 84.5% vs 7.34 / 82.5%, +0.20. Arm flat1e4, 4 ep, 1e-4, seed 1, flat weights, no Hybrid A. SHA 0dbd3a33…. Two fresh holdouts

Published Session 11 ledger as written: 7,623 start, 1,148 hash leaks dropped, 6,419 keep (372 gold + 6,047 teacher) → 6,219/200. Those integers do not fully reconcile; this synthesis does not invent a missing filter.

4. What actually moved quality

Decode before weights. Four-beam complete-sequence search on frozen B-9500 flipped the SO ranking without a new checkpoint. Later continuation training under that decoder lost to the frozen pin. Attribute that +0.45–0.48 to search.

Glue is competence-dependent. Hybrid A (keep ≤2 namer words, fill from centroid v2) helped weak extractive stacks and hurt Title-SFT FLAN on confirm-v1 (7.28 raw vs 5.78 glued on a different packet in Session 10). Session 11 ships raw greedy, no Hybrid A.

Dialect over volume. 6.2k leak-checked gold+Grok/Opus/Codex titles beat 175k weakly filtered Flash-Lite CE. Session-9 schema-argmax continue-SFT lost because it was the wrong dialect.

Capacity became worth the payload only after the dialect was learned. P1 is 3.88 ms and 24.5M. SFT FLAN is ~193 MB INT8 estimate and 29.3 ms. Session 10 accepted that cost for quality. Session 11 did not change the graph.

Boards do not transfer. The SO sealed win is not a Terminal win. The paper’s ship claim is Terminal-board Session 11 only.

5. The ship model

Title-SFT FLAN-T5-small descendant, arm flat1e4. Init: google/flan-t5-small → Session-10 ship b798d78f… → this file. Load with tie_word_embeddings=false. Share encoder/decoder input embeddings; refuse if lm_head is tied (decode becomes reheat / blackjack).

Holdoutflat1e4 mean / ≥6Session-10 shipΔW/T/L
holdout10007.59 / 85.7%7.38 / 83.5%+0.21296 / 551 / 153
holdout1000b7.55 / 84.5%7.34 / 82.5%+0.20281 / 570 / 149

Hugging Face: porkr/porkicoder-tab-namer-77m. Product: porkicoder.com. Public installer 2.17.42 still served 6t when Session 11 was written.

6. What this is not

7. Sources

All numbers above are from the linked public notes and the Session 11 aug12_scores.json snapshot. Companion claim ledger: claim_audit.md in the Hugging Face paper folder. Two independently typeset manuscripts in this repository (PorkiCoder_Tab_Title_Paper.pdf, porkicoder_tab_namer_paper.pdf) cover the same campaign at greater length; this page is the short synthesis.