# Accuracy and inference profile Frozen selection: **gx10-4b-expanded**, checkpoint step **0**. Selection used validation only; the following evaluation does not change the winner. The selected branch's step 0 retains the exact warm-start parent weights from step 1500; it is not an untrained base model. The saved CPU proof verifies equality of all trainable tensors. Validation-score tie: `gx10-4b-expanded`, `spark-b-4b-refinement`. The coordinator selected the first name in deterministic alphabetical order. The branch's latest checkpoint, step 159, was evaluated but not promoted by the fixed validation rule. Its expansion-diagnostic results are post-selection comparisons. ## Complete final evaluation | Split | Decisions | Calibrated selected accuracy | Calibrated base accuracy | Selected NLL | Base NLL | | --- | ---: | ---: | ---: | ---: | ---: | | test | 2042 | 92.90% | 84.48% | 0.2051 | 0.4527 | | holdout | 768 | 72.92% | 70.31% | 0.6783 | 0.7425 | Paired 95% source-group bootstrap intervals are **calibrated selected minus calibrated base**. Positive accuracy differences favor the selected model; negative NLL/Brier differences favor it. - test: accuracy +8.42 pp [+6.85, +9.89]; NLL -0.2477 [-0.2769, -0.2188]; Brier -0.1450 [-0.1637, -0.1258]. - holdout: accuracy +2.60 pp [+0.13, +5.34]; NLL -0.0642 [-0.1144, -0.0160]; Brier -0.0509 [-0.0755, -0.0264]. ## Matched profiling accuracy 320 held-out decisions and 383 expansion diagnostics are identical across methods. Each family row reports its exact sample count. Differences below are descriptive point estimates. | Sample / family | N | Selected scorer | Base yes/no verifier | Base one-token label | Expanded scorer | | --- | ---: | ---: | ---: | ---: | ---: | | heldout / overall | 320 | 89.06% | 80.94% | 86.25% | 88.75% | | heldout / arc | 64 | 96.88% | 90.62% | 93.75% | 96.88% | | heldout / banking | 64 | 96.88% | 84.38% | 96.88% | 96.88% | | heldout / boolq | 64 | 95.31% | 90.62% | 89.06% | 95.31% | | heldout / snli | 64 | 82.81% | 70.31% | 75.00% | 81.25% | | heldout / social | 64 | 73.44% | 68.75% | 76.56% | 73.44% | | diagnostics / overall | 383 | 79.11% | 77.81% | 79.11% | 81.46% | | diagnostics / commonsenseqa | 128 | 76.56% | 70.31% | 70.31% | 75.00% | | diagnostics / hellaswag | 128 | 75.78% | 78.91% | 85.16% | 82.81% | | diagnostics / piqa | 127 | 85.04% | 84.25% | 81.89% | 86.61% | Expanded scorer checkpoint step: **159**. Paired expanded-minus-selected accuracy: - heldout/overall: -0.31 pp; expanded alone correct 2, selected alone correct 3 (N=320). - heldout/arc: +0.00 pp; expanded alone correct 0, selected alone correct 0 (N=64). - heldout/banking: +0.00 pp; expanded alone correct 0, selected alone correct 0 (N=64). - heldout/boolq: +0.00 pp; expanded alone correct 0, selected alone correct 0 (N=64). - heldout/snli: -1.56 pp; expanded alone correct 1, selected alone correct 2 (N=64). - heldout/social: +0.00 pp; expanded alone correct 1, selected alone correct 1 (N=64). - diagnostics/overall: +2.35 pp; expanded alone correct 14, selected alone correct 5 (N=383). - diagnostics/commonsenseqa: -1.56 pp; expanded alone correct 0, selected alone correct 2 (N=128). - diagnostics/hellaswag: +7.03 pp; expanded alone correct 11, selected alone correct 2 (N=128). - diagnostics/piqa: +1.57 pp; expanded alone correct 3, selected alone correct 1 (N=127). The base label method computes one constrained next-token label from a prompt containing all options. The two verifier methods score each candidate independently. No free-text reasoning or JSON generation is timed. ## Warm local speed Each cell uses 2 warmups and 10 measured repeats. Times are median / exploratory p95 in seconds. Ratios are **base median ÷ selected median**: **above 1 means the selected scorer is faster; below 1 means the baseline is faster**. `summary.json` and `speed.csv` also report requests, questions and choice probabilities per second, derived from serial warm median latency; these do not measure concurrent serving. | State tokens | Questions × choices | Selected median / p95 | Base verifier median / p95 | Base label median / p95 | Verifier / selected | Label / selected | | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | 128 | 1 × 2 | 0.4507 / 0.4536 | 0.4048 / 0.4083 | 0.2056 / 0.2076 | 0.90× | 0.46× | | 128 | 1 × 4 | 0.9023 / 0.9065 | 0.8059 / 0.8119 | 0.2128 / 0.2154 | 0.89× | 0.24× | | 128 | 1 × 16 | 3.6218 / 3.7210 | 3.2337 / 3.2489 | 0.3059 / 0.3089 | 0.89× | 0.08× | | 128 | 4 × 2 | 1.8013 / 1.8038 | 1.6120 / 1.6169 | 0.8250 / 0.8318 | 0.89× | 0.46× | | 128 | 4 × 4 | 3.6008 / 3.6126 | 3.2206 / 3.2300 | 0.8501 / 0.8513 | 0.89× | 0.24× | | 128 | 16 × 2 | 7.2279 / 7.2993 | 6.4655 / 6.4843 | 3.3011 / 3.3170 | 0.89× | 0.46× | | 768 | 1 × 2 | 1.8398 / 1.8417 | 1.5885 / 1.5957 | 0.8080 / 0.8120 | 0.86× | 0.44× | | 768 | 1 × 4 | 3.7095 / 3.7701 | 3.1766 / 3.1828 | 0.8184 / 0.8239 | 0.86× | 0.22× | | 768 | 1 × 16 | 14.6790 / 14.6895 | 12.7095 / 12.7259 | 0.9420 / 0.9437 | 0.87× | 0.06× | | 768 | 4 × 2 | 7.3425 / 7.3540 | 6.3598 / 6.3755 | 3.2318 / 3.2416 | 0.87× | 0.44× | | 768 | 4 × 4 | 14.6808 / 14.6930 | 12.6692 / 12.7184 | 3.2784 / 3.2847 | 0.86× | 0.22× | | 768 | 16 × 2 | 29.3552 / 29.3689 | 25.4125 / 25.4531 | 12.9224 / 12.9458 | 0.87× | 0.44× | ![Warm latency comparison](latency.png) ![Matched sample accuracy](accuracy.png) ## Scope and evidence - Profile accuracy uses fixed matched samples; point differences have no significance claim. - Base labels jointly condition on all options; verifier paths score each option independently. - Profiles use raw probabilities without applying an artifact temperature. - p95 is an exploratory nearest-rank statistic from the recorded small repeat count. - Warm local timings exclude model loading and network latency; no shared-prefix optimization is used. - Throughput is derived from serial warm median latency; it makes no concurrent-serving capacity claim. - Expanded comparisons are post-selection diagnostics and cannot change the frozen winner. `summary.json` contains metrics, intervals and provenance. `final_evaluation.csv`, `accuracy.csv` and `speed.csv` contain table/chart data. Frozen protocol SHA-256: `b6dde5e425239455e9e31a44bc22b151b0ee66ec827c538aaca3308e29441c15`.