Instructions to use mgoeckel/oscar-1-400m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mgoeckel/oscar-1-400m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mgoeckel/oscar-1-400m")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mgoeckel/oscar-1-400m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Oscar-1 400M
The 400M top of the v3 ladder: the recipe that closed the family's transfer gap. The number that matters: 0.7770 typed-decisions test accuracy — over the 421M laya specialist's published 0.766 and beyond its own previous generation by +21.4 pp on the sealed harness.
What it is. A decision model: one document (the state) and a set of typed questions in, a
probability distribution per question out, in one forward pass. No text generation — nothing
to parse, nothing to hallucinate. It is a decision head on a JHU CLSP Ettin encoder
(ettin-encoder-400m, fully fine-tuned; ~403M parameters with the head), serving the same
answers contract the hosted Jev API serves, and loads only with the
laya runtime — pip install laya, then laya.load("oscar-1-400m"). The shipped
temperatures (choice 1.044, score 1.036, noul 1.159) were fitted
post-hoc in-distribution on a held-out 400-item split.
Part of the Oscar-1 collection
(all checkpoints + live demo Space).
This version (2026-09-30)
Paired against the previous generation (sweep-v2 full-encoder generation (2026-09-28, unpublished)): +21.4 pp on the sealed harness (0.4863 → 0.7000); see Paired reads.
Read this before relying on these numbers:
- typed-decisions is in-distribution. The training mix contains the same public sources the benchmark's five public suites come from (agnews, mnli, emotion, sst5, banking77) plus the five decision workflows; the test split was held out of training, but the capabilities it measures were trained. Read 0.7770 as what this size learned from the mix, not a general read.
- The sealed 9-suite harness is the transfer check: byte-identical cases at seed 42, labels are the datasets' own (not adjudicated for this harness), 1,240 decisions per participant. With one read and case-clustered resampling, CIs here are wide; treat sub-point gaps as noise.
- As served, confident errors are 26.6% out of domain (decisions at reported confidence ≥ 0.9 that are wrong; the laya anchor: 5.3%). The shipped temperatures were fitted in-distribution; refit on your own decisions before gating anything on confidence.
- No pre-registered release rule for this generation. The family's release rule is registered from the next corpus delta onward (
docs/release-rule.md); this release was decided by reading the results and shipping, which is exactly what the rule is meant to prevent. - mnli 0.533 / sst5 0.525 as served — the v3 recipe's soft-ordinal golds are what keeps these cells alive.
- guardrails as served: 0.781 — better than a small generalist has any right to be, but route prompt-injection and policy screens to a dedicated stack.
Results (as served: each checkpoint at its own fitted temperature)
| Oscar-1 400M (this checkpoint) | previous version (sweep-v2, unpublished) | laya anchor | |
|---|---|---|---|
| typed-decisions test accuracy (400 cases / 2,000 decisions) | 0.7770 | — | 0.7685 |
| typed-decisions Brier / ECE (as served) | 0.0616 / 0.1375 | — | 0.0657 / 0.2156 |
| typed-decisions raw accuracy (T = 1.0) | 0.7770 | — | 0.7685 |
| typed raw Brier / ECE | 0.0638 / 0.1243 | — | 0.0537 / 0.1301 |
| typed score MAE / within-1 | 0.2269 / 0.990 | — | 0.2425 / 0.995 |
| typed latency p50 | 10.0 ms | — | 10.0 ms |
| sysone sealed overall (952 cases / 1,240 decisions) | 0.7000 | 0.4863 | 0.7258 |
| sysone agnews | 0.863 | 0.562 | 0.844 |
| sysone banking77 (12) | 0.854 | 0.531 | 0.812 |
| sysone emotion | 0.578 | 0.385 | 0.693 |
| sysone guardrails screens | 0.781 | 0.292 | 0.750 |
| sysone mnli | 0.533 | 0.308 | 0.575 |
| sysone moderation | 0.882 | 0.729 | 0.847 |
| sysone multilingual intent | 0.500 | 0.475 | 0.567 |
| sysone sst5 (5-way sentiment) | 0.525 | 0.133 | 0.375 |
| sysone support triage | 0.771 | 0.755 | 0.927 |
| confident errors out of domain (p ≥ 0.9 and wrong, as served) | 26.6% | — | 5.3% |
| coverage at ≤ 5% error (share of decisions automatable) | 0.002 | 0.007 | 0.135 |
Latency: typed-decisions readout on an RTX 5060 Ti (bf16); the sealed harness runs the same
checkpoint as served. The laya anchor is our own re-measure on this host, using the same
evaluator that produced the Oscar rows. Its published figures (accuracy 0.766, Tesla T4)
differ from the re-measure within temperature-fitting rounding; its shipped temperatures
include one invalid entry (choice:11+, clamped at load) — accuracy is unaffected.
Paired reads
Paired with the previous generation: sealed overall +21.4 pp [+18.0, +24.6]; 119 decisions right only in the previous generation, 384 only in this checkpoint (exact McNemar p ≈ 0). The delta is clear.
Paired with oscar-1-150m: sealed overall +2.3 pp [+0.4, +4.3]; 61 decisions right only in oscar-1-150m, 90 only in this checkpoint (exact McNemar p = 0.0224). The delta is clear.
Calibration (raw vs served)
Temperature fitting moved typed-decisions accuracy 0.7770 →
0.7770 and Brier 0.0638 → 0.0616 (raw = all three
temperatures at 1.0, set via agent.cfg['temperature'] = [1, 1, 1]). The sealed-harness
runs above are as served only: the harness answers every case through the stock laya
contract with the shipped temperatures.
The family ladder
| member | typed-decisions test acc | sysone sealed overall | confident errors (p ≥ 0.9) | coverage at ≤ 5% error |
|---|---|---|---|---|
oscar-1-17m |
0.6775 | 0.5702 | 19.9% | 0.006 |
oscar-1-32m |
0.7005 | 0.5847 | 20.3% | 0.002 |
oscar-1-68m |
0.7455 | 0.6629 | 25.0% | 0.060 |
oscar-1-150m |
0.7750 | 0.6766 | 25.4% | 0.019 |
| Oscar-1 400M (this checkpoint) | 0.7770 | 0.7000 | 26.6% | 0.002 |
laya-typed-decisions (anchor, measured by us) |
0.7685 | 0.7258 | 5.3% | 0.135 |
All rows measured by us on this host: typed-decisions official test split; sealed 9-suite harness (seed 42). Method and per-decision provenance: see Reproduce.
How it was built
- Backbone: JHU CLSP Ettin
ettin-encoder-400m(bidirectional encoder, fully fine-tuned) + a decision head trained from scratch: 2 transformer layers, an option-marker scorer, an act/escalate head. ~403M parameters together, 842.6 MB shipped on disk. - Recipe (RLCD, the Laya method): the policy reports a full distribution per question; exploration adds zero-mean Gaussian noise to the logits (sigma annealed 0.4 → 0.1); the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal
scorequestions), so expected reward is maximised only by honest probabilities. Updates are REINFORCE with a group-mean baseline (GRPO-style), alongside soft cross-entropy against teacher targets (w_ce = 1.0). 4 epochs, 5,300 updates, AdEMAMix, bf16, 2-GPU DDP (5.21 GPU-h per run). - Lineage: trained from the released base encoder (no warm start).
- Corpus: expansion chunk v3 — typed-decisions train mix (customer service, invoice, agent-trace, security, moderation) + public corpora (agnews, mnli, emotion, sst5, banking77) + minted cross-primitive variants + counterfactual repairs + soft-ordinal one-hots + translated multilingual = 84,854 train items (400 held-out calib).
- Calibration: one temperature per primitive fitted post-hoc on a held-out 400-item split (no id overlap with training): choice 1.044, score 1.036, noul 1.159; set
agent.cfg['temperature'] = [1, 1, 1]for the raw readout. - Weights:
model.safetensorssha-256fde674ab599199c9…(full digest in results provenance below).
Known limits
- Routing. The family top (842 MB): typed 0.7770 over its own specialist anchor and the best transfer (0.700). Use it when states will be strange; use the 150M when they will not.
- Out-of-domain calibration is the open problem. As served, decisions at reported confidence ≥ 0.9 are wrong 26.6% of the time on the sealed harness (laya: 5.3%); coverage at a 5% error budget is 0.002 (laya: 0.135). Do not treat confidence as an abstention signal out of distribution; refit or threshold your own data first.
- guardrails 0.781 as served; useful as a coarse screen, never as the policy gate.
- Ordinal
scoreis the weakest primitive (within-1 0.990, argmax accuracy falls). Keep rubrics to 3-5 levels and read the expected score, not the argmax. - Long states: trained and benchmarked at
max_len = 512(the measured p50 reflects the 512 budget); the ModernBERT encoder itself reads up to ~8,000 tokens — raiseagent.cfg["max_len"]for long states and verify on your own data (the same caveat the laya cards document for their 8k mode). Keep choice sets under ~20 options unlesshead_max_lenis raised too. - English-only. The multilingual intent cell is 0.500 as
served; for non-English states use
laya-multilingual. - A decision model, not an assistant. It cannot answer free-form questions, generate text, or abstain with an "I don't know" option unless you add one to the criteria.
Use
pip install laya
Python 3.10 or newer; CPU inference works, device="cuda" with any modern GPU.
import laya
agent = laya.load("mgoeckel/oscar-1-400m") # downloads and builds
state = {"text": "I was charged twice, please return the money."}
questions = {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": ["refund", "cancel", "information", "other"]},
"urgency": {"type": "score", "instructions": "Rate urgency 0 to 4.",
"criteria": ["routine", "low", "moderate", "high", "critical"]},
"sensitive": {"type": "noul", "instructions": "Is this a fraud/security issue?"},
}
r = agent.predict(state, questions)
print(r["answers"]["intent"]["choice"]) # -> refund
print(r["answers"]["urgency"]["score"]) # -> 0.898 (expected level, 0..1)
print(r["answers"]["sensitive"]["noul"]) # -> probability the answer is true
All questions in a call are answered in one forward pass; latency follows the state's length, not the question count. Answer shape per primitive:
| type | criteria |
answer fields |
|---|---|---|
choice |
2-16 option IDs to descriptions | choice, probabilities, answer_confidence |
score |
ordered rubric levels of 2-16 | score (expected level, 0-1), probabilities, answer_confidence |
noul |
optional true/false descriptions | noul (probability true), confidence = max(p, 1-p) |
The raw JSON is the same answers contract the hosted Jev API and the
laya-demo Space serve — a client
written for either reads Oscar's answers unchanged.
act_probability is present on every answer (act/escalate head, escalate cost 0.5,
wrong-act cost 3.0) but carries no calibrated escalation signal yet — gate decisions on
answer_confidence.
Reproduce
- Sealed harness: sysone-bench v2 orchestrator, dataset version 2.0.0 checksum-verified, seed 42; per-decision predictions, checksums and manifests are archived in
/tmp/rlcd-research/sysone-bench/runs-v2/(regenerable withscripts/run_sysone_gpu.py <name> <checkpoint> <run_id>). - Typed-decisions readout:
encoder_rlcd/eval.pysingle mode over the official test split (400 cases / 2,000 decisions; dataset revision as served through 2026-09). - Result files:
results/oscar-1-400m-sysone.json(sealed),results/ettin-mixed-400m-v3-typeddec.json(typed; raw and as-served blocks),results/oscar-1-paired-sysone.json(paired reads),results/oscar-1-family-sysone.json(as-served confidence metrics),reports/oscar-1-400m-gpu-sweep.mdandreports/oscar-1-68m-150m-gpu-sweep.md(sweep narratives). Pairing method: exact McNemar on discordant decisions + case-clustered bootstrap (10,000 resamples, seed 42). Oscar accuracy is read post-temperature with the per-primitive temperatures inrl_agent_config.json; set all temperatures to 1.0 for the raw readout.
Links
- Collection (all checkpoints + demo): https://huggingface.co/collections/mgoeckel/oscar-1-6abb984de72aa5d5f6978fcc
- Live demo Space: https://huggingface.co/spaces/mgoeckel/oscar-1-demo
- Sibling checkpoints:
oscar-1-17m·oscar-1-32m·oscar-1-68m·oscar-1-150m - Method & runtime: NandhaKishorM/laya ·
pip install laya - Benchmark dataset: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
- Companion models:
convaiinnovations/laya-typed-decisions·SupersonicLabs/Julia-1·openjev/openjev
Apache 2.0 · Oscar-1 (mgoeckel) · Ettin encoders MIT (JHU CLSP) · RLCD method & Laya runtime Apache-2.0 (Convai Innovations)
- Downloads last month
- 14
Model tree for mgoeckel/oscar-1-400m
Base model
jhu-clsp/ettin-encoder-400mDataset used to train mgoeckel/oscar-1-400m
Collection including mgoeckel/oscar-1-400m
Evaluation results
- accuracy on typed-decisions (official test split, measured by us)self-reported0.777
- brier_score on typed-decisions (official test split, measured by us)self-reported0.062
- expected_calibration_error on typed-decisions (official test split, measured by us)self-reported0.138
- mean_absolute_error on typed-decisions (official test split, measured by us)self-reported0.227
- sysone sealed overall (9-suite harness, 1,240 decisions, seed 42) on typed-decisions (official test split, measured by us)self-reported0.700