# Official typed-decisions TEST The fixed final Lex checkpoint and the official `laya-typed-decisions` specialist were each evaluated once on the original 2,000 TEST decisions, from 400 state components. Both models support all 2,000 complete inputs within 1,024 tokens, their native windows and their untruncated official-default support. No TRAIN state component or full input overlaps this TEST support. The test contains four English workflows and 20 question schemas. This is not evidence of multilingual or unseen-schema generalization. | Metric | Lex | Official specialist (shipped temperature) | |---|---:|---:| | Original-label hard accuracy | **78.15%** (1,563/2,000) | 76.60% (1,532/2,000) | | Soft NLL | 0.961902 | **0.884390** | | Mean soft Brier | 0.030858 | **0.018452** | | Score RPS | 0.022994 | **0.015602** | | Score expected-value MAE | 0.251408 | **0.242408** | Lex's accuracy gain is 1.55 percentage points (31 decisions). The paired 95% interval is **[-0.20, +3.15125] percentage points**, which includes zero. The interval uses 2,000 state-component bootstrap samples, seed 20260921, stratified by each component's workflow set; complete components are retained. This supports a higher observed hard-accuracy point, not statistically significant or uniformly better performance. Probability metrics are worse for Lex in this comparison. No calibration was fitted. | Type | Decisions | Lex accuracy | Official accuracy | |---|---:|---:|---:| | Choice | 600 | 74.00% | 73.33% | | Noul | 600 | 84.67% | 85.67% | | Score | 800 | 76.375% | 72.25% | | Workflow | Decisions | Lex accuracy | Official accuracy | |---|---:|---:|---:| | Agent trace observability | 500 | 72.4% | 73.0% | | Customer service | 500 | 79.6% | 76.4% | | Invoice processing | 500 | 84.4% | 80.4% | | Security incidents | 500 | 76.2% | 76.6% | Hard accuracy uses the source's original semantic hard label. Choice and Score use the first maximum in supplied order; Noul uses `p_yes >= 0.5`. Soft NLL uses the explicitly normalized original soft label distribution and natural logarithms (probability floor 1e-12). Brier averages squared probability error over the valid candidates; ordinal Score RPS averages cumulative-distribution squared error over K−1 boundaries. Score MAE compares the predicted expected value against the original source score. The official raw-temperature accuracy is also 76.60%; raw soft NLL is 0.869598. The frozen TEST parquet SHA256 is `4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c`. The actual saved-summary SHA256 is `1a21955b24be16cefa2963779e71edf4d6b5adcd14679fc438e6a87e3e604b88`. No TEST text, labels or predictions are redistributed. The annotations are synthetic/teacher-derived; accuracy does not certify operational safety or calibrated business utility.