Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -77,34 +77,50 @@ split (`training/fit_calibration.py`). Configs and seed live under `training/con
|
|
| 77 |
Reported by `training/evaluate.py` and rendered by `benchmarks/report.py` (accuracy, ECE, p50/p95
|
| 78 |
latency per primitive and language).
|
| 79 |
|
| 80 |
-
**Full-scale run (single RTX 3060 12GB):**
|
| 81 |
-
deterministic synthetic records (fully localized per language, a learnable `other`
|
| 82 |
-
descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English,
|
| 83 |
-
plus a dedicated low-rank `choice` head (
|
| 84 |
-
|
|
|
|
|
|
|
|
|
|
| 85 |
|
| 86 |
| Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
|
| 87 |
| --- | --- | --- | --- |
|
| 88 |
-
| English (ModernBERT-large + LoRA r=16 + choice
|
| 89 |
| Multilingual (mmBERT-base + LoRA r=64 + choice head) | **0.743** | 0.089 | 16.0 |
|
| 90 |
|
| 91 |
-
Per primitive (English): `choice`
|
| 92 |
0.468, `noul` 0.892, `score` 0.870. **Label audit:** every `noul` row is judged against its own
|
| 93 |
text β **0 contradictory** in both eval sets (positive rates 0.482 / 0.486), per language in
|
| 94 |
`benchmarks/report.md`.
|
| 95 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 96 |
`choice` is the weak primitive on the multilingual side (0.468) and it is a *training* trade, not
|
| 97 |
a label problem: `choice` labels never changed, and retraining on the corrected labels makes
|
| 98 |
`noul` (1.000) and `score` (0.998) trivial β the same tone detector β while the shared trunk
|
| 99 |
starves `choice` (an identical-recipe English control landed at 0.841, the five-domain retrain's
|
| 100 |
-
`choice` collapsed to 0.303).
|
| 101 |
-
|
|
|
|
|
|
|
|
|
|
| 102 |
`fr` 0.101, `de` 0.118, `nl` 0.192 and `it` 0.197 remain above target (NFR-C06), and multilingual
|
| 103 |
`choice`/`noul` ECE (0.099 / 0.108) is declared with them. The CUDA-graph fast path
|
| 104 |
(`TACHYONE_FAST=1`) gives a 2.65Γ p50 speedup (8.97 β 3.39 ms) with 0 top-label flips.
|
| 105 |
|
| 106 |
**Robustness (B-4).** On a noisy view (one surface edit β typo/accents/casing β applied to 15% of
|
| 107 |
-
states) English drops only 0.
|
| 108 |
adapters are robust to this noise model.
|
| 109 |
|
| 110 |
Full tables and environment are in
|
|
|
|
| 77 |
Reported by `training/evaluate.py` and rendered by `benchmarks/report.py` (accuracy, ECE, p50/p95
|
| 78 |
latency per primitive and language).
|
| 79 |
|
| 80 |
+
**Full-scale run (single RTX 3060 12GB):** 21,000 five-domain English / 18,000 multilingual train
|
| 81 |
+
/ 1,500 eval deterministic synthetic records (fully localized per language, a learnable `other`
|
| 82 |
+
option with rich descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English,
|
| 83 |
+
r=64 multilingual) plus a dedicated low-rank `choice` head (near-identity init); **8 epochs for
|
| 84 |
+
multilingual**, batch 16, bf16 + gradient checkpointing. The **English artifact** is the B-5 bank
|
| 85 |
+
(ADR-0016): run 5's six-epoch trunk kept **frozen** while the `choice` head was re-fitted as a
|
| 86 |
+
bank β shared + one head per domain, rank 128, 8 epochs at lr 1e-4 β by
|
| 87 |
+
`training/fit_choice_bank.py`, whose recipe ships next to the weights as `choice_bank_fit.json`.
|
| 88 |
|
| 89 |
| Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
|
| 90 |
| --- | --- | --- | --- |
|
| 91 |
+
| English (ModernBERT-large + five-domain LoRA r=16 + choice-head bank) | **0.964** | 0.023 | 54.4 |
|
| 92 |
| Multilingual (mmBERT-base + LoRA r=64 + choice head) | **0.743** | 0.089 | 16.0 |
|
| 93 |
|
| 94 |
+
Per primitive (English): `choice` **1.000**, `noul` 0.946, `score` 0.946; (multilingual): `choice`
|
| 95 |
0.468, `noul` 0.892, `score` 0.870. **Label audit:** every `noul` row is judged against its own
|
| 96 |
text β **0 contradictory** in both eval sets (positive rates 0.482 / 0.486), per language in
|
| 97 |
`benchmarks/report.md`.
|
| 98 |
|
| 99 |
+
**What this English row trades (published in full, not summarized away).** On the support-only
|
| 100 |
+
split the previous artifact scored 0.972 β `noul` 0.992, `score` 0.978, `choice` 0.946. The
|
| 101 |
+
five-domain bank scores 0.964 there: `choice` becomes **1.000** while `noul`/`score` give up 4.6
|
| 102 |
+
and 3.2 points, and in exchange the adapter covers four domains it could not answer at all
|
| 103 |
+
before β on the five-domain split the previous artifact scores **0.511** overall (worst domain
|
| 104 |
+
0.328) against this one's **0.964** (worst domain 0.963). The gate that routes each `choice`
|
| 105 |
+
question to its domain head scores **strict 1.000** (0 to shared, 0 wrong domain) over 2,500 rows,
|
| 106 |
+
and accuracy on the 473 rows whose text never occurs in training is **0.998**. Full tables:
|
| 107 |
+
[`docs/benchmarks.md`](https://github.com/munod/tachyone/blob/main/docs/benchmarks.md).
|
| 108 |
+
|
| 109 |
`choice` is the weak primitive on the multilingual side (0.468) and it is a *training* trade, not
|
| 110 |
a label problem: `choice` labels never changed, and retraining on the corrected labels makes
|
| 111 |
`noul` (1.000) and `score` (0.998) trivial β the same tone detector β while the shared trunk
|
| 112 |
starves `choice` (an identical-recipe English control landed at 0.841, the five-domain retrain's
|
| 113 |
+
`choice` collapsed to 0.303). The English side did **not** retrain the trunk to escape that: it
|
| 114 |
+
re-fitted the `choice` head on a **frozen** trunk (B-5 / ADR-0016), which lifts `choice` to 1.000
|
| 115 |
+
and leaves `noul`/`score` where the trunk already had them β the cost and the gain in the table
|
| 116 |
+
above are one artifact, not two runs (provenance: `choice_bank_fit.json` next to the weights).
|
| 117 |
+
One of six languages meets ECE β€ 0.05 (`es` 0.045); `pt` 0.062,
|
| 118 |
`fr` 0.101, `de` 0.118, `nl` 0.192 and `it` 0.197 remain above target (NFR-C06), and multilingual
|
| 119 |
`choice`/`noul` ECE (0.099 / 0.108) is declared with them. The CUDA-graph fast path
|
| 120 |
(`TACHYONE_FAST=1`) gives a 2.65Γ p50 speedup (8.97 β 3.39 ms) with 0 top-label flips.
|
| 121 |
|
| 122 |
**Robustness (B-4).** On a noisy view (one surface edit β typo/accents/casing β applied to 15% of
|
| 123 |
+
states) English drops only 0.964 β 0.963 and multilingual 0.743 β 0.744, so the released
|
| 124 |
adapters are robust to this noise model.
|
| 125 |
|
| 126 |
Full tables and environment are in
|