--- license: apache-2.0 base_model: jhu-clsp/ettin-encoder-400m library_name: transformers pipeline_tag: text-classification language: - en tags: - oscar-1 - decision-model - system-one - calibrated-decisions - rlcd - laya - ettin - classification - routing - scoring - guardrails - moderation - reinforcement-learning - commercial-use datasets: - LocalLLaMA/typed-decisions model-index: - name: oscar-1-400m results: - task: type: text-classification name: decision making dataset: type: LocalLLaMA/typed-decisions name: typed-decisions (official test split, measured by us) metrics: - type: accuracy value: 0.7770 - type: brier_score value: 0.0616 - type: expected_calibration_error value: 0.1375 - type: mean_absolute_error value: 0.2269 - type: accuracy name: sysone sealed overall (9-suite harness, 1,240 decisions, seed 42) value: 0.7000 --- # Oscar-1 400M The 400M top of the v3 ladder: the recipe that closed the family's transfer gap. The number that matters: **0.7770** typed-decisions test accuracy — over the 421M laya specialist's published 0.766 and beyond its own previous generation by +21.4 pp on the sealed harness. **What it is.** A decision model: one document (the state) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation — nothing to parse, nothing to hallucinate. It is a decision head on a JHU CLSP Ettin encoder (`ettin-encoder-400m`, fully fine-tuned; ~403M parameters with the head), serving the same `answers` contract the hosted [Jev](https://jev.ai) API serves, and loads only with the `laya` runtime — `pip install laya`, then `laya.load("oscar-1-400m")`. The shipped temperatures (choice 1.044, score 1.036, noul 1.159) were fitted post-hoc in-distribution on a held-out 400-item split. **Part of the [Oscar-1 collection](https://huggingface.co/collections/mgoeckel/oscar-1-6abb984de72aa5d5f6978fcc)** (all checkpoints + live demo Space). ## This version (2026-09-30) Paired against the previous generation (sweep-v2 full-encoder generation (2026-09-28, unpublished)): +21.4 pp on the sealed harness (0.4863 → 0.7000); see [Paired reads](#paired-reads). Read this before relying on these numbers: - **typed-decisions is in-distribution.** The training mix contains the same public sources the benchmark's five public suites come from (agnews, mnli, emotion, sst5, banking77) plus the five decision workflows; the test split was held out of training, but the *capabilities* it measures were trained. Read 0.7770 as what this size learned from the mix, not a general read. - **The sealed 9-suite harness is the transfer check**: byte-identical cases at seed 42, labels are the datasets' own (not adjudicated for this harness), 1,240 decisions per participant. With one read and case-clustered resampling, CIs here are wide; treat sub-point gaps as noise. - **As served, confident errors are 26.6% out of domain** (decisions at reported confidence ≥ 0.9 that are wrong; the laya anchor: 5.3%). The shipped temperatures were fitted in-distribution; refit on your own decisions before gating anything on confidence. - **No pre-registered release rule for this generation.** The family's release rule is registered from the next corpus delta onward (`docs/release-rule.md`); this release was decided by reading the results and shipping, which is exactly what the rule is meant to prevent. - mnli 0.533 / sst5 0.525 as served — the v3 recipe's soft-ordinal golds are what keeps these cells alive. - guardrails as served: 0.781 — better than a small generalist has any right to be, but route prompt-injection and policy screens to a dedicated stack. ## Results (as served: each checkpoint at its own fitted temperature) | | Oscar-1 400M (this checkpoint) | previous version (sweep-v2, unpublished) | laya anchor | |---|---|---|---| | typed-decisions test **accuracy** (400 cases / 2,000 decisions) | **0.7770** | — | 0.7685 | | typed-decisions Brier / ECE (as served) | 0.0616 / 0.1375 | — | 0.0657 / 0.2156 | | typed-decisions raw accuracy (T = 1.0) | 0.7770 | — | 0.7685 | | typed raw Brier / ECE | 0.0638 / 0.1243 | — | 0.0537 / 0.1301 | | typed score MAE / within-1 | 0.2269 / 0.990 | — | 0.2425 / 0.995 | | typed latency p50 | 10.0 ms | — | 10.0 ms | | sysone sealed overall (952 cases / 1,240 decisions) | **0.7000** | 0.4863 | 0.7258 | | sysone agnews | 0.863 | 0.562 | 0.844 | | sysone banking77 (12) | 0.854 | 0.531 | 0.812 | | sysone emotion | 0.578 | 0.385 | 0.693 | | sysone guardrails screens | 0.781 | 0.292 | 0.750 | | sysone mnli | 0.533 | 0.308 | 0.575 | | sysone moderation | 0.882 | 0.729 | 0.847 | | sysone multilingual intent | 0.500 | 0.475 | 0.567 | | sysone sst5 (5-way sentiment) | 0.525 | 0.133 | 0.375 | | sysone support triage | 0.771 | 0.755 | 0.927 | | confident errors out of domain (p ≥ 0.9 and wrong, as served) | 26.6% | — | 5.3% | | coverage at ≤ 5% error (share of decisions automatable) | 0.002 | 0.007 | 0.135 | Latency: typed-decisions readout on an RTX 5060 Ti (bf16); the sealed harness runs the same checkpoint as served. The laya anchor is our own re-measure on this host, using the same evaluator that produced the Oscar rows. Its published figures (accuracy 0.766, Tesla T4) differ from the re-measure within temperature-fitting rounding; its shipped temperatures include one invalid entry (`choice:11+`, clamped at load) — accuracy is unaffected. ## Paired reads Paired with the previous generation: sealed overall +21.4 pp [+18.0, +24.6]; 119 decisions right only in the previous generation, 384 only in this checkpoint (exact McNemar p ≈ 0). The delta is clear. Paired with `oscar-1-150m`: sealed overall +2.3 pp [+0.4, +4.3]; 61 decisions right only in `oscar-1-150m`, 90 only in this checkpoint (exact McNemar p = 0.0224). The delta is clear. ## Calibration (raw vs served) Temperature fitting moved typed-decisions accuracy 0.7770 → 0.7770 and Brier 0.0638 → 0.0616 (raw = all three temperatures at 1.0, set via `agent.cfg['temperature'] = [1, 1, 1]`). The sealed-harness runs above are as served only: the harness answers every case through the stock laya contract with the shipped temperatures. ## The family ladder | member | typed-decisions test acc | sysone sealed overall | confident errors (p ≥ 0.9) | coverage at ≤ 5% error | |---|---:|---:|---:|---:| | [`oscar-1-17m`](https://huggingface.co/mgoeckel/oscar-1-17m) | 0.6775 | 0.5702 | 19.9% | 0.006 | | [`oscar-1-32m`](https://huggingface.co/mgoeckel/oscar-1-32m) | 0.7005 | 0.5847 | 20.3% | 0.002 | | [`oscar-1-68m`](https://huggingface.co/mgoeckel/oscar-1-68m) | 0.7455 | 0.6629 | 25.0% | 0.060 | | [`oscar-1-150m`](https://huggingface.co/mgoeckel/oscar-1-150m) | 0.7750 | 0.6766 | 25.4% | 0.019 | | **Oscar-1 400M (this checkpoint)** | 0.7770 | 0.7000 | 26.6% | 0.002 | | [`laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) (anchor, measured by us) | 0.7685 | 0.7258 | 5.3% | 0.135 | All rows measured by us on this host: typed-decisions official test split; sealed 9-suite harness (seed 42). Method and per-decision provenance: see [Reproduce](#reproduce). ## How it was built - **Backbone:** [JHU CLSP Ettin](https://huggingface.co/jhu-clsp) `ettin-encoder-400m` (bidirectional encoder, fully fine-tuned) + a decision head trained from scratch: 2 transformer layers, an option-marker scorer, an act/escalate head. ~403M parameters together, 842.6 MB shipped on disk. - **Recipe (RLCD, the Laya method):** the policy reports a full distribution per question; exploration adds zero-mean Gaussian noise to the logits (sigma annealed 0.4 → 0.1); the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal `score` questions), so expected reward is maximised only by honest probabilities. Updates are REINFORCE with a group-mean baseline (GRPO-style), alongside soft cross-entropy against teacher targets (w_ce = 1.0). 4 epochs, 5,300 updates, AdEMAMix, bf16, 2-GPU DDP (5.21 GPU-h per run). - **Lineage:** trained from the released base encoder (no warm start). - **Corpus:** expansion chunk v3 — typed-decisions train mix (customer service, invoice, agent-trace, security, moderation) + public corpora (agnews, mnli, emotion, sst5, banking77) + minted cross-primitive variants + counterfactual repairs + soft-ordinal one-hots + translated multilingual = 84,854 train items (400 held-out calib). - **Calibration:** one temperature per primitive fitted post-hoc on a held-out 400-item split (no id overlap with training): choice 1.044, score 1.036, noul 1.159; set `agent.cfg['temperature'] = [1, 1, 1]` for the raw readout. - **Weights:** `model.safetensors` sha-256 `fde674ab599199c9…` (full digest in results provenance below). ## Known limits - **Routing.** The family top (842 MB): typed 0.7770 over its own specialist anchor and the best transfer (0.700). Use it when states will be strange; use the 150M when they will not. - **Out-of-domain calibration is the open problem.** As served, decisions at reported confidence ≥ 0.9 are wrong 26.6% of the time on the sealed harness (laya: 5.3%); coverage at a 5% error budget is 0.002 (laya: 0.135). Do not treat confidence as an abstention signal out of distribution; refit or threshold your own data first. - guardrails 0.781 as served; useful as a coarse screen, never as the policy gate. - **Ordinal `score` is the weakest primitive** (within-1 0.990, argmax accuracy falls). Keep rubrics to 3-5 levels and read the expected score, not the argmax. - **Long states**: trained and benchmarked at `max_len = 512` (the measured p50 reflects the 512 budget); the ModernBERT encoder itself reads up to ~8,000 tokens — raise `agent.cfg["max_len"]` for long states and verify on your own data (the same caveat the laya cards document for their 8k mode). Keep choice sets under ~20 options unless `head_max_len` is raised too. - **English-only.** The multilingual intent cell is 0.500 as served; for non-English states use [`laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual). - **A decision model, not an assistant.** It cannot answer free-form questions, generate text, or abstain with an "I don't know" option unless you add one to the criteria. ## Use ```bash pip install laya ``` Python 3.10 or newer; CPU inference works, `device="cuda"` with any modern GPU. ```python import laya agent = laya.load("mgoeckel/oscar-1-400m") # downloads and builds state = {"text": "I was charged twice, please return the money."} questions = { "intent": {"type": "choice", "instructions": "What does the customer want?", "criteria": ["refund", "cancel", "information", "other"]}, "urgency": {"type": "score", "instructions": "Rate urgency 0 to 4.", "criteria": ["routine", "low", "moderate", "high", "critical"]}, "sensitive": {"type": "noul", "instructions": "Is this a fraud/security issue?"}, } r = agent.predict(state, questions) print(r["answers"]["intent"]["choice"]) # -> refund print(r["answers"]["urgency"]["score"]) # -> 0.898 (expected level, 0..1) print(r["answers"]["sensitive"]["noul"]) # -> probability the answer is true ``` All questions in a call are answered in **one forward pass**; latency follows the state's length, not the question count. Answer shape per primitive: | type | `criteria` | answer fields | |---|---|---| | `choice` | 2-16 option IDs to descriptions | `choice`, `probabilities`, `answer_confidence` | | `score` | ordered rubric levels of 2-16 | `score` (expected level, 0-1), `probabilities`, `answer_confidence` | | `noul` | optional true/false descriptions | `noul` (probability true), `confidence` = max(p, 1-p) | The raw JSON is the same `answers` contract the hosted Jev API and the [laya-demo](https://huggingface.co/spaces/convaiinnovations/laya-demo) Space serve — a client written for either reads Oscar's answers unchanged. `act_probability` is present on every answer (act/escalate head, escalate cost 0.5, wrong-act cost 3.0) but carries no calibrated escalation signal yet — gate decisions on `answer_confidence`. ## Reproduce - Sealed harness: sysone-bench v2 orchestrator, dataset version 2.0.0 checksum-verified, seed 42; per-decision predictions, checksums and manifests are archived in `/tmp/rlcd-research/sysone-bench/runs-v2/` (regenerable with `scripts/run_sysone_gpu.py `). - Typed-decisions readout: `encoder_rlcd/eval.py` single mode over the official test split (400 cases / 2,000 decisions; dataset revision as served through 2026-09). - Result files: `results/oscar-1-400m-sysone.json` (sealed), `results/ettin-mixed-400m-v3-typeddec.json` (typed; raw and as-served blocks), `results/oscar-1-paired-sysone.json` (paired reads), `results/oscar-1-family-sysone.json` (as-served confidence metrics), `reports/oscar-1-400m-gpu-sweep.md` and `reports/oscar-1-68m-150m-gpu-sweep.md` (sweep narratives). Pairing method: exact McNemar on discordant decisions + case-clustered bootstrap (10,000 resamples, seed 42). Oscar accuracy is read post-temperature with the per-primitive temperatures in `rl_agent_config.json`; set all temperatures to 1.0 for the raw readout. ## Links - **Collection (all checkpoints + demo):** https://huggingface.co/collections/mgoeckel/oscar-1-6abb984de72aa5d5f6978fcc - **Live demo Space:** https://huggingface.co/spaces/mgoeckel/oscar-1-demo - **Sibling checkpoints:** [`oscar-1-17m`](https://huggingface.co/mgoeckel/oscar-1-17m) · [`oscar-1-32m`](https://huggingface.co/mgoeckel/oscar-1-32m) · [`oscar-1-68m`](https://huggingface.co/mgoeckel/oscar-1-68m) · [`oscar-1-150m`](https://huggingface.co/mgoeckel/oscar-1-150m) - **Method & runtime:** [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya) · `pip install laya` - **Benchmark dataset:** https://huggingface.co/datasets/LocalLLaMA/typed-decisions - **Companion models:** [`convaiinnovations/laya-typed-decisions`](https://huggingface.co/convaiinnovations/laya-typed-decisions) · [`SupersonicLabs/Julia-1`](https://huggingface.co/SupersonicLabs/Julia-1) · [`openjev/openjev`](https://huggingface.co/openjev/openjev) Apache 2.0 · Oscar-1 (mgoeckel) · Ettin encoders MIT (JHU CLSP) · RLCD method & Laya runtime Apache-2.0 (Convai Innovations)