| --- |
| license: apache-2.0 |
| language: |
| - en |
| - ko |
| - multilingual |
| base_model: |
| - jhu-clsp/mmBERT-base |
| library_name: laya |
| tags: |
| - decision-model |
| - system-1 |
| - non-autoregressive |
| - typed-decisions |
| - calibration |
| - multilingual |
| pipeline_tag: text-classification |
| --- |
| |
| # XERON-0.1 ๐ฏ |
|
|
| **XERON-0.1** is a fine-tuned **typed-decision (System 1) model** built on [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)'s |
| multilingual checkpoint (backbone: `jhu-clsp/mmBERT-base`, 322M params). |
|
|
| It answers **typed questions over a state** โ `choice` / `score` / `noul` (boolean) โ in a **single forward pass**, |
| returning calibrated probabilities. It never generates text, so it cannot hallucinate and cannot emit malformed schemas. |
|
|
| Trained by [PIXELZX](https://huggingface.co/PIXELZX) as the first release of the XERON family |
| (training plan & data pipeline: https://github.com/PIXELZX0/XERON). |
|
|
| ## โจ What makes XERON-0.1 different |
|
|
| | | | |
| |---|---| |
| | **Base** | `jhu-clsp/mmBERT-base` (Laya `multilingual` subfolder) | |
| | **Fine-tuned on** | 148,281 sequences โ English typed-decisions + Korean (KLUE ynat/nli/sts) + browser/web decisions (BBC + AG News, SMS spam, phishing URLs) + web-agent actions (Mind2Web) + long-context (SCOTUS, 20 Newsgroups) | |
| | **Training** | RLCD โ strictly proper scoring rule + GRPO-style baseline; post-hoc temperature calibration. 4 epochs ยท 15,987 updates ยท ~5h on 1รGPU (bf16) | |
| | **Context** | 4,096 tokens trained (`CTX_CAP` up to 32,768 via RoPE extension) | |
| | **Languages** | 100+ via mmBERT (English + Korean trained explicitly) | |
| | **License** | Apache-2.0 | |
|
|
| ## ๐ Head-to-head: Laya vs Jev vs XERON-0.1 |
|
|
| ### A. ์ค์ธก ๋น๊ต (๋์ผ ํ๊ฐ์
ยท ๋์ผ ์คํฌ๋ฆฝํธ) |
|
|
| ํ๊ฐ์
`LocalLLaMA/typed-decisions` **test** split โ 400 cases = 600 `choice` + 400 `score` + 400 `noul` (**1,400 decisions**). |
| ๋์ผํ `scripts/evaluate.py` / `scripts/recompute_metrics.py`, CPU ๋จ์ผ ํ๋ก์ธ์ค, 2026-09-22. ์๋ณธ: `eval_comparison.md`, `eval_baseline_*.json`. |
|
|
| | ๋ชจ๋ธ | ๋ฐฑ๋ณธ / ํ๋ผ๋ฏธํฐ | choice acc | soft acc | Brier | ECE โ | score MAE โ | noul acc | |
| |---|---|---|---|---|---|---|---| |
| | **XERON-0.1** (ours) | mmBERT-base ยท 322M ยท fine-tuned | 0.7000 | **0.5171** | **0.4493** | **0.2143** | 0.4213 | 0.7817 | |
| | `laya-typed-decisions` (Convai) | ModernBERT-large ยท 421M ยท ๋ฒค๋ ํ๋ | **0.7333** | 0.4460 | 0.4669 | 0.2380 | **0.2963** | **0.8583** | |
| | `laya-multilingual` (Convai) | mmBERT-base ยท 322M ยท base(๋ฏธํ๋) | 0.2900 | 0.2758 | 0.9952 | 0.3282 | 1.0625 | 0.4983 | |
| | Jev (TypeSafe AI) | ๋น๊ณต๊ฐ ยท ํธ์คํ
API | ์ธก์ ๋ถ๊ฐ (์จ์ดํธ ๋น๊ณต๊ฐ) | โ | โ | โ | โ | โ | |
|
|
| - ๋ฏธํ๋ Laya base(0.290) โ XERON-0.1(0.700) = **+41.0%p**. ํ์ธํ๋์ด ๊ฒฐ์ ์ (Laya base๋ ํ์คํฌ๋ณ FT ์ ์ ๋ผ๋ ๊ณต๊ฐ ์ค๋ช
๊ณผ ์ผ์น). |
| - ๋ฒค๋๊ฐ ์ด ๋ฒค์น๋งํฌ์ ์ง์ ํ๋ํ 421M ์์ด ๋ชจ๋ธ ๋๋น ํ๋ ๋ผ๋ฒจ **-3.3%p**, ๊ทธ๋ฌ๋ **soft accuracy +7.1%p / ECE โ0.024 ๋ XERON ์ฐ์** โ ํ๋ฅ (์ ๋ขฐ๋) ํ์ง์ด ๋ ์ข์. |
|
|
| ### B. ์คํ ยท ๊ณต๊ฐ ์งํ ๋น๊ต |
|
|
| Laya/Jev ์์น๋ ๋ฒค๋ยท์ 3์ ๊ณต๊ฐ ์๋ฃ์์ ์ธ์ฉ (2026-09-22 ์กฐํ). XERON ์์น๋ ์ค์ธก. |
|
|
| | ํญ๋ชฉ | **XERON-0.1** | **Laya** (Convai) | **Jev** (TypeSafe AI) | |
| |---|---|---|---| |
| | ๋ฐฐํฌ ํํ | ์คํ ์จ์ดํธ (self-host) | ์คํ ์จ์ดํธ (self-host) | ํธ์คํ
API (closed) | |
| | ๋ผ์ด์ ์ค | Apache-2.0 | Apache-2.0 | ์์ฉ API | |
| | ๋ฐฑ๋ณธ / ํ๋ผ๋ฏธํฐ | mmBERT-base ยท 322M | EN 421M / multi 322M | ๋น๊ณต๊ฐ | |
| | ์ปจํ
์คํธ | 4,096 (cap 32,768) | 512 (EN) / 1,024 (multi) | 64k | |
| | ์ธ์ด | 100+ (mmBERT), ENยทKR ํ์ต | 100+ (multi ์ฒดํฌํฌ์ธํธ) | ๋ค๊ตญ์ด (๋น๊ณต๊ฐ) | |
| | ๊ฒฐ์ ํ๋ฆฌ๋ฏธํฐ๋ธ | choice / score / noul | choice / score / noul | choice / score / noul | |
| | ํ์ธํ๋ ํ์ | โ (์ด๋ฏธ ํ๋๋จ) | โ
(base๋ FT ์ ์ ) | โ (์ ๋ก์ท) | |
| | ์ ํ๋ (๊ณต๊ฐ) | โ | in-task 0.753 ยท zero-shot 0.651 ยท typed-decisions 0.766 | JevBench v1.3.0 composite **74.4 (#1)** | |
| | ๋์ด๋๋ณ (Easy/Std/Judge/Hard) | โ | 94.4 / 72.9 / 69.2 / 34.1 % | 100 / 99 / 94.5 / **74.1 %** | |
| | ์บ๋ฆฌ๋ธ๋ ์ด์
(ECE) | 0.2143 (๋์ผ ํ๊ฐ์
) | in-task 0.030 ยท zero-shot 0.204 ยท multi 0.081 (๋ฒค๋) | ๋ฏธ๊ณต๊ฐ | |
| | Brier | 0.4493 (๋์ผ ํ๊ฐ์
) | 0.308 in-task / 0.532 zero-shot (๋ฒค๋) | ๋ฏธ๊ณต๊ฐ | |
| | ์ง์ฐ | 2.2 s/case (CPU, ๋ก์ปฌ ์ธก์ ) | p50 38.4 ms (1๋ฌธํญ, ๋ฒค๋) ยท 32.8 ms (T4) | ~150 ms (์ 3์ 236โ276 ms) | |
| | ๋น์ฉ | self-host (GPU ๋น์ฉ๋ง) | self-host (GPU ๋น์ฉ๋ง) | $0.042 / 1M input tokens | |
| | JevBench v1.3.0 ์์ | ๋ฏธ๋ฑ์ฌ | #33 (54.4) | **#1 (74.4)** | |
|
|
| > โ ๏ธ A์ B๋ **๋ค๋ฅธ ํ๊ฐ์
ยท๋ค๋ฅธ ํ๋์จ์ด**์
๋๋ค. A๋ ๋์ผ ์กฐ๊ฑด ์ค์ธก(๋ชจ๋ธ ๊ฐ ์ง์ ๋น๊ต ๊ฐ๋ฅ), B๋ ๊ณต๊ฐ ์๋ฃ ์ธ์ฉ(์ฐธ๊ณ ์ฉ). |
| > Jev๋ ์จ์ดํธ๊ฐ ๊ณต๊ฐ๋์ง ์์ A์ ๋ฃ์ ์ ์์ต๋๋ค. |
|
|
| ### C. JevBench v1.3 ๋ฐฉ์ ํ๊ฐ (๊ณต๊ฐ ์์ดํ
231/534) |
|
|
| JevBench v1.3.0(Benchmark Heaven) ํ๋ค์คยท์ด๋ํฐยท์ฑ์ ์ฝ๋๋ฅผ ๊ทธ๋๋ก ์ฌ์ฉํด ์ธก์ ํ์ต๋๋ค. |
| **๊ณต๊ฐ ์์ดํ
์ 534๊ฐ ์ค 231๊ฐ๋ฟ**(judge tier๋ ์ ๋ถ ๋น๊ณต๊ฐ)์ด๋ฏ๋ก **๊ณต์ ๋ณด๋ ์์์ ์ง์ ๋น๊ตํ ์ ์์ต๋๋ค.** |
| ์ ์ฒด ๋ฐฉ๋ฒยทfamily๋ณ ๋ถ์ยท์ฌํ ์ฝ๋: https://github.com/PIXELZX0/XERON/tree/main/results/jevbench-public |
|
|
| ๋์ผํ 231๊ฐ ๊ณต๊ฐ ์์ดํ
์์ 4๊ฐ ์์คํ
์ ๋น๊ต (Jev/Laya๋ ๊ณต์ per-task ์ํฐํฉํธ์ ๊ณต๊ฐ ์์ดํ
๊ฒฐ๊ณผ): |
|
|
| | ์์คํ
| easy (48) | standard (72) | hard (111) | ์ ์ฒด (231) | Intelligence | |
| |---|---|---|---|---|---| |
| | Jev 1.13.0 (TypeSafe, API) | 1.000 | 0.986 | 0.730 | **0.866** | 82.2 | |
| | laya-typed-decisions (Convai, 421M, ๋ฒค๋ ํ๋) | 0.979 | 0.653 | 0.270 | **0.537** | 38.0 | |
| | **XERON-0.1** (ours) | 0.875 | 0.444 | **0.306** | 0.468 | 23.3 | |
| | laya-multilingual (๋ฏธํ๋ base) | 0.896 | 0.403 | 0.324 | 0.468 | 21.5 | |
|
|
| **์์งํ ๊ฒฐ๋ก ** |
| - XERON-0.1์ **JevBench๋ฅ ํ์คํฌ(์ ์ฑ
๋ฌธ์ + ๋ฃจ๋ธ๋ฆญ ์ ํ)์์ ๋ฒค๋ Laya ํ๋ํ๋ณด๋ค ์ฝํฉ๋๋ค.** ํ์ต ๋ฐ์ดํฐ(KLUEยท๋ธ๋ผ์ฐ์ ยทMind2WebยทSCOTUS)์ ๋๋ฉ์ธ์ด ๊ฒน์น์ง ์๊ธฐ ๋๋ฌธ์
๋๋ค. |
| - ๋จ, **๋์ผ ๋ฐฑ๋ณธ ๋ฏธํ๋ base ๋๋น ํ์ธํ๋ ํจ๊ณผ๋ ๊ฒ์ฆ**๋์ต๋๋ค (ECE hard 0.289 โ 0.211, ํ๋ฅ TVD 0.573 โ 0.389). |
| - **hard tier์์๋ XERON์ด ๋ฒค๋ ํ๋ํ์ ์ด๊น๋๋ค** (0.306 vs 0.270) โ adversarial / trap / judge_hard ๊ณ์ด. |
| - ๋ถ๋ถ JevBench Score(judge tier ์์, ์ฌ์ ๊ทํ): laya-typed-decisions 37.5 ยท XERON-0.1 11.5 ยท laya-multilingual 8.8. Intelligence<50 ํจ๋ํฐ `(I/50)ยฒ`๊ฐ XERON ์ ์๋ฅผ ํฌ๊ฒ ๊น์ต๋๋ค. |
| - ๋ค์ ๋จ๊ณ: JevBench ์คํ์ผ ๋ฐ์ดํฐ(๊ณต๊ฐ 231๊ฑด ๋๋ HF `Praveenrajus/jev-bench` 166k rows)๋ก ์ถ๊ฐ ํ์ธํ๋ ํ ์ฌํ๊ฐ. |
| |
| **์ฝ๋ ๋ฒ** |
| - Jev = ์ ๋ก์ท ๊ฐ์ (ํ๋ ๋์ด๋ 74.1%). |
| - Laya / XERON = ์คํ ์จ์ดํธ, **ํ์ธํ๋ ํ ์๊ธฐ ๋๋ฉ์ธ์์ ๊ฐํด์ง๋** ์ง์. |
| - XERON-0.1์ Laya multilingual๊ณผ ๊ฐ์ ๋ฐฑ๋ณธ + ํ๊ตญ์ดยท๋ธ๋ผ์ฐ์ ยท์น ์์ด์ ํธยท์ฅ๋ฌธ ๋ฐ์ดํฐ ์ถ๊ฐ ํ์ตํ. |
| |
| ## ๐ ์ฌ์ฉ๋ฒ |
| |
| ```bash |
| pip install laya |
| ``` |
| |
| ```python |
| import laya |
| |
| agent = laya.load("PIXELZX/XERON-0.1") |
| |
| state = "Customer email: 'I was charged twice, please refund immediately.' tier=premium, sla=4h" |
| questions = { |
| "intent": {"type": "choice", "options": ["billing", "technical", "cancellation", "other"]}, |
| "urgency": {"type": "score", "levels": ["0 โ no time pressure", "1 โ routine", "2 โ elevated", "3 โ critical"]}, |
| "needs_refund": {"type": "noul"}, |
| } |
|
|
| res = agent.predict(state, questions) |
| print(res["answers"]) |
| ``` |
| |
| Returned per question: selected key / expected level + full probability distribution + calibrated confidence. |
| |
| ## ๐ ํ๊ฐ ์ฌํ |
| |
| ```bash |
| python scripts/evaluate.py \ |
| --model PIXELZX/XERON-0.1 \ |
| --dataset LocalLLaMA/typed-decisions --split test \ |
| --device cpu --output eval_results.json |
| python scripts/recompute_metrics.py eval_results.json # score MAE / ํ์
๋ณ ์ ํ๋ |
| ``` |
| |
| `eval_results.json` (per-decision predictions ํฌํจ)์ด ์ด ์ ์ฅ์์ ํฌํจ๋์ด ์์ต๋๋ค. |
|
|
| ## โ ๏ธ ํ๊ณ |
|
|
| - 4,096 ํ ํฐ ํ์ต โ ์ด์ฅ๋ฌธ์ `CTX_CAP` ํ์ฅ ํ ์ฌํ์ต ํ์. |
| - ํ๊ฐ๋ ์์ด typed-decisions split ์ค์ฌ. ํ๊ตญ์ด/๋ธ๋ผ์ฐ์ /์น ์์ด์ ํธ ํธ๋์ ๋ณ๋ ๋ฒค์น๋งํฌ ๋ฏธ๊ณต๊ฐ. |
| - `score` ํ์
์ ์์ํ(ordinal) rubric ์ ์ฉ. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{xeron01, |
| title = {XERON-0.1: a fine-tuned multilingual typed-decision model}, |
| author = {PIXELZX}, |
| year = {2026}, |
| url = {https://huggingface.co/PIXELZX/XERON-0.1} |
| } |
| ``` |
|
|
| Built on [Laya](https://huggingface.co/convaiinnovations/laya) by Convai Innovations (Apache-2.0) and `jhu-clsp/mmBERT-base`. |
|
|