Oscar-1 400M

The 400M top of the v3 ladder: the recipe that closed the family's transfer gap. The number that matters: 0.7770 typed-decisions test accuracy — over the 421M laya specialist's published 0.766 and beyond its own previous generation by +21.4 pp on the sealed harness.

What it is. A decision model: one document (the state) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation — nothing to parse, nothing to hallucinate. It is a decision head on a JHU CLSP Ettin encoder (ettin-encoder-400m, fully fine-tuned; ~403M parameters with the head), serving the same answers contract the hosted Jev API serves, and loads only with the laya runtime — pip install laya, then laya.load("oscar-1-400m"). The shipped temperatures (choice 1.044, score 1.036, noul 1.159) were fitted post-hoc in-distribution on a held-out 400-item split. Part of the Oscar-1 collection (all checkpoints + live demo Space).

This version (2026-09-30)

Paired against the previous generation (sweep-v2 full-encoder generation (2026-09-28, unpublished)): +21.4 pp on the sealed harness (0.4863 → 0.7000); see Paired reads.

Read this before relying on these numbers:

  • typed-decisions is in-distribution. The training mix contains the same public sources the benchmark's five public suites come from (agnews, mnli, emotion, sst5, banking77) plus the five decision workflows; the test split was held out of training, but the capabilities it measures were trained. Read 0.7770 as what this size learned from the mix, not a general read.
  • The sealed 9-suite harness is the transfer check: byte-identical cases at seed 42, labels are the datasets' own (not adjudicated for this harness), 1,240 decisions per participant. With one read and case-clustered resampling, CIs here are wide; treat sub-point gaps as noise.
  • As served, confident errors are 26.6% out of domain (decisions at reported confidence ≥ 0.9 that are wrong; the laya anchor: 5.3%). The shipped temperatures were fitted in-distribution; refit on your own decisions before gating anything on confidence.
  • No pre-registered release rule for this generation. The family's release rule is registered from the next corpus delta onward (docs/release-rule.md); this release was decided by reading the results and shipping, which is exactly what the rule is meant to prevent.
  • mnli 0.533 / sst5 0.525 as served — the v3 recipe's soft-ordinal golds are what keeps these cells alive.
  • guardrails as served: 0.781 — better than a small generalist has any right to be, but route prompt-injection and policy screens to a dedicated stack.

Results (as served: each checkpoint at its own fitted temperature)

Oscar-1 400M (this checkpoint) previous version (sweep-v2, unpublished) laya anchor
typed-decisions test accuracy (400 cases / 2,000 decisions) 0.7770 — 0.7685
typed-decisions Brier / ECE (as served) 0.0616 / 0.1375 — 0.0657 / 0.2156
typed-decisions raw accuracy (T = 1.0) 0.7770 — 0.7685
typed raw Brier / ECE 0.0638 / 0.1243 — 0.0537 / 0.1301
typed score MAE / within-1 0.2269 / 0.990 — 0.2425 / 0.995
typed latency p50 10.0 ms — 10.0 ms
sysone sealed overall (952 cases / 1,240 decisions) 0.7000 0.4863 0.7258
sysone agnews 0.863 0.562 0.844
sysone banking77 (12) 0.854 0.531 0.812
sysone emotion 0.578 0.385 0.693
sysone guardrails screens 0.781 0.292 0.750
sysone mnli 0.533 0.308 0.575
sysone moderation 0.882 0.729 0.847
sysone multilingual intent 0.500 0.475 0.567
sysone sst5 (5-way sentiment) 0.525 0.133 0.375
sysone support triage 0.771 0.755 0.927
confident errors out of domain (p ≥ 0.9 and wrong, as served) 26.6% — 5.3%
coverage at ≤ 5% error (share of decisions automatable) 0.002 0.007 0.135

Latency: typed-decisions readout on an RTX 5060 Ti (bf16); the sealed harness runs the same checkpoint as served. The laya anchor is our own re-measure on this host, using the same evaluator that produced the Oscar rows. Its published figures (accuracy 0.766, Tesla T4) differ from the re-measure within temperature-fitting rounding; its shipped temperatures include one invalid entry (choice:11+, clamped at load) — accuracy is unaffected.

Paired reads

Paired with the previous generation: sealed overall +21.4 pp [+18.0, +24.6]; 119 decisions right only in the previous generation, 384 only in this checkpoint (exact McNemar p ≈ 0). The delta is clear.

Paired with oscar-1-150m: sealed overall +2.3 pp [+0.4, +4.3]; 61 decisions right only in oscar-1-150m, 90 only in this checkpoint (exact McNemar p = 0.0224). The delta is clear.

Calibration (raw vs served)

Temperature fitting moved typed-decisions accuracy 0.7770 → 0.7770 and Brier 0.0638 → 0.0616 (raw = all three temperatures at 1.0, set via agent.cfg['temperature'] = [1, 1, 1]). The sealed-harness runs above are as served only: the harness answers every case through the stock laya contract with the shipped temperatures.

The family ladder

member typed-decisions test acc sysone sealed overall confident errors (p ≥ 0.9) coverage at ≤ 5% error
oscar-1-17m 0.6775 0.5702 19.9% 0.006
oscar-1-32m 0.7005 0.5847 20.3% 0.002
oscar-1-68m 0.7455 0.6629 25.0% 0.060
oscar-1-150m 0.7750 0.6766 25.4% 0.019
Oscar-1 400M (this checkpoint) 0.7770 0.7000 26.6% 0.002
laya-typed-decisions (anchor, measured by us) 0.7685 0.7258 5.3% 0.135

All rows measured by us on this host: typed-decisions official test split; sealed 9-suite harness (seed 42). Method and per-decision provenance: see Reproduce.

How it was built

  • Backbone: JHU CLSP Ettin ettin-encoder-400m (bidirectional encoder, fully fine-tuned) + a decision head trained from scratch: 2 transformer layers, an option-marker scorer, an act/escalate head. ~403M parameters together, 842.6 MB shipped on disk.
  • Recipe (RLCD, the Laya method): the policy reports a full distribution per question; exploration adds zero-mean Gaussian noise to the logits (sigma annealed 0.4 → 0.1); the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal score questions), so expected reward is maximised only by honest probabilities. Updates are REINFORCE with a group-mean baseline (GRPO-style), alongside soft cross-entropy against teacher targets (w_ce = 1.0). 4 epochs, 5,300 updates, AdEMAMix, bf16, 2-GPU DDP (5.21 GPU-h per run).
  • Lineage: trained from the released base encoder (no warm start).
  • Corpus: expansion chunk v3 — typed-decisions train mix (customer service, invoice, agent-trace, security, moderation) + public corpora (agnews, mnli, emotion, sst5, banking77) + minted cross-primitive variants + counterfactual repairs + soft-ordinal one-hots + translated multilingual = 84,854 train items (400 held-out calib).
  • Calibration: one temperature per primitive fitted post-hoc on a held-out 400-item split (no id overlap with training): choice 1.044, score 1.036, noul 1.159; set agent.cfg['temperature'] = [1, 1, 1] for the raw readout.
  • Weights: model.safetensors sha-256 fde674ab599199c9… (full digest in results provenance below).

Known limits

  • Routing. The family top (842 MB): typed 0.7770 over its own specialist anchor and the best transfer (0.700). Use it when states will be strange; use the 150M when they will not.
  • Out-of-domain calibration is the open problem. As served, decisions at reported confidence ≥ 0.9 are wrong 26.6% of the time on the sealed harness (laya: 5.3%); coverage at a 5% error budget is 0.002 (laya: 0.135). Do not treat confidence as an abstention signal out of distribution; refit or threshold your own data first.
  • guardrails 0.781 as served; useful as a coarse screen, never as the policy gate.
  • Ordinal score is the weakest primitive (within-1 0.990, argmax accuracy falls). Keep rubrics to 3-5 levels and read the expected score, not the argmax.
  • Long states: trained and benchmarked at max_len = 512 (the measured p50 reflects the 512 budget); the ModernBERT encoder itself reads up to ~8,000 tokens — raise agent.cfg["max_len"] for long states and verify on your own data (the same caveat the laya cards document for their 8k mode). Keep choice sets under ~20 options unless head_max_len is raised too.
  • English-only. The multilingual intent cell is 0.500 as served; for non-English states use laya-multilingual.
  • A decision model, not an assistant. It cannot answer free-form questions, generate text, or abstain with an "I don't know" option unless you add one to the criteria.

Use

pip install laya

Python 3.10 or newer; CPU inference works, device="cuda" with any modern GPU.

import laya

agent = laya.load("mgoeckel/oscar-1-400m")   # downloads and builds

state = {"text": "I was charged twice, please return the money."}
questions = {
    "intent": {"type": "choice", "instructions": "What does the customer want?",
               "criteria": ["refund", "cancel", "information", "other"]},
    "urgency": {"type": "score", "instructions": "Rate urgency 0 to 4.",
                "criteria": ["routine", "low", "moderate", "high", "critical"]},
    "sensitive": {"type": "noul", "instructions": "Is this a fraud/security issue?"},
}
r = agent.predict(state, questions)
print(r["answers"]["intent"]["choice"])          # -> refund
print(r["answers"]["urgency"]["score"])           # -> 0.898 (expected level, 0..1)
print(r["answers"]["sensitive"]["noul"])          # -> probability the answer is true

All questions in a call are answered in one forward pass; latency follows the state's length, not the question count. Answer shape per primitive:

type criteria answer fields
choice 2-16 option IDs to descriptions choice, probabilities, answer_confidence
score ordered rubric levels of 2-16 score (expected level, 0-1), probabilities, answer_confidence
noul optional true/false descriptions noul (probability true), confidence = max(p, 1-p)

The raw JSON is the same answers contract the hosted Jev API and the laya-demo Space serve — a client written for either reads Oscar's answers unchanged.

act_probability is present on every answer (act/escalate head, escalate cost 0.5, wrong-act cost 3.0) but carries no calibrated escalation signal yet — gate decisions on answer_confidence.

Reproduce

  • Sealed harness: sysone-bench v2 orchestrator, dataset version 2.0.0 checksum-verified, seed 42; per-decision predictions, checksums and manifests are archived in /tmp/rlcd-research/sysone-bench/runs-v2/ (regenerable with scripts/run_sysone_gpu.py <name> <checkpoint> <run_id>).
  • Typed-decisions readout: encoder_rlcd/eval.py single mode over the official test split (400 cases / 2,000 decisions; dataset revision as served through 2026-09).
  • Result files: results/oscar-1-400m-sysone.json (sealed), results/ettin-mixed-400m-v3-typeddec.json (typed; raw and as-served blocks), results/oscar-1-paired-sysone.json (paired reads), results/oscar-1-family-sysone.json (as-served confidence metrics), reports/oscar-1-400m-gpu-sweep.md and reports/oscar-1-68m-150m-gpu-sweep.md (sweep narratives). Pairing method: exact McNemar on discordant decisions + case-clustered bootstrap (10,000 resamples, seed 42). Oscar accuracy is read post-temperature with the per-primitive temperatures in rl_agent_config.json; set all temperatures to 1.0 for the raw readout.

Links

Apache 2.0 · Oscar-1 (mgoeckel) · Ettin encoders MIT (JHU CLSP) · RLCD method & Laya runtime Apache-2.0 (Convai Innovations)

Downloads last month
14
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mgoeckel/oscar-1-400m

Finetuned
(12)
this model

Dataset used to train mgoeckel/oscar-1-400m

Collection including mgoeckel/oscar-1-400m

Evaluation results

  • accuracy on typed-decisions (official test split, measured by us)
    self-reported
    0.777
  • brier_score on typed-decisions (official test split, measured by us)
    self-reported
    0.062
  • expected_calibration_error on typed-decisions (official test split, measured by us)
    self-reported
    0.138
  • mean_absolute_error on typed-decisions (official test split, measured by us)
    self-reported
    0.227
  • sysone sealed overall (9-suite harness, 1,240 decisions, seed 42) on typed-decisions (official test split, measured by us)
    self-reported
    0.700