munod commited on
Commit
d3f64cc
Β·
verified Β·
1 Parent(s): 5f68805

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +26 -10
README.md CHANGED
@@ -77,34 +77,50 @@ split (`training/fit_calibration.py`). Configs and seed live under `training/con
77
  Reported by `training/evaluate.py` and rendered by `benchmarks/report.py` (accuracy, ECE, p50/p95
78
  latency per primitive and language).
79
 
80
- **Full-scale run (single RTX 3060 12GB):** 9,000 English / 18,000 multilingual train / 1,500 eval
81
- deterministic synthetic records (fully localized per language, a learnable `other` team with rich
82
- descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64 multilingual)
83
- plus a dedicated low-rank `choice` head (r=32, near-identity init); 4 epochs for English and
84
- **8 for multilingual**, batch 16, bf16 + gradient checkpointing.
 
 
 
85
 
86
  | Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
87
  | --- | --- | --- | --- |
88
- | English (ModernBERT-large + LoRA r=16 + choice head) | **0.972** | 0.020 | 23.1 |
89
  | Multilingual (mmBERT-base + LoRA r=64 + choice head) | **0.743** | 0.089 | 16.0 |
90
 
91
- Per primitive (English): `choice` 0.946, `noul` 0.992, `score` 0.978; (multilingual): `choice`
92
  0.468, `noul` 0.892, `score` 0.870. **Label audit:** every `noul` row is judged against its own
93
  text β€” **0 contradictory** in both eval sets (positive rates 0.482 / 0.486), per language in
94
  `benchmarks/report.md`.
95
 
 
 
 
 
 
 
 
 
 
 
96
  `choice` is the weak primitive on the multilingual side (0.468) and it is a *training* trade, not
97
  a label problem: `choice` labels never changed, and retraining on the corrected labels makes
98
  `noul` (1.000) and `score` (0.998) trivial β€” the same tone detector β€” while the shared trunk
99
  starves `choice` (an identical-recipe English control landed at 0.841, the five-domain retrain's
100
- `choice` collapsed to 0.303). That is why the published English adapter keeps its B-11 weights
101
- (provenance stated in the header). One of six languages meets ECE ≀ 0.05 (`es` 0.045); `pt` 0.062,
 
 
 
102
  `fr` 0.101, `de` 0.118, `nl` 0.192 and `it` 0.197 remain above target (NFR-C06), and multilingual
103
  `choice`/`noul` ECE (0.099 / 0.108) is declared with them. The CUDA-graph fast path
104
  (`TACHYONE_FAST=1`) gives a 2.65Γ— p50 speedup (8.97 β†’ 3.39 ms) with 0 top-label flips.
105
 
106
  **Robustness (B-4).** On a noisy view (one surface edit β€” typo/accents/casing β€” applied to 15% of
107
- states) English drops only 0.972 β†’ 0.969 and multilingual 0.743 β†’ 0.744, so the released
108
  adapters are robust to this noise model.
109
 
110
  Full tables and environment are in
 
77
  Reported by `training/evaluate.py` and rendered by `benchmarks/report.py` (accuracy, ECE, p50/p95
78
  latency per primitive and language).
79
 
80
+ **Full-scale run (single RTX 3060 12GB):** 21,000 five-domain English / 18,000 multilingual train
81
+ / 1,500 eval deterministic synthetic records (fully localized per language, a learnable `other`
82
+ option with rich descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English,
83
+ r=64 multilingual) plus a dedicated low-rank `choice` head (near-identity init); **8 epochs for
84
+ multilingual**, batch 16, bf16 + gradient checkpointing. The **English artifact** is the B-5 bank
85
+ (ADR-0016): run 5's six-epoch trunk kept **frozen** while the `choice` head was re-fitted as a
86
+ bank β€” shared + one head per domain, rank 128, 8 epochs at lr 1e-4 β€” by
87
+ `training/fit_choice_bank.py`, whose recipe ships next to the weights as `choice_bank_fit.json`.
88
 
89
  | Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
90
  | --- | --- | --- | --- |
91
+ | English (ModernBERT-large + five-domain LoRA r=16 + choice-head bank) | **0.964** | 0.023 | 54.4 |
92
  | Multilingual (mmBERT-base + LoRA r=64 + choice head) | **0.743** | 0.089 | 16.0 |
93
 
94
+ Per primitive (English): `choice` **1.000**, `noul` 0.946, `score` 0.946; (multilingual): `choice`
95
  0.468, `noul` 0.892, `score` 0.870. **Label audit:** every `noul` row is judged against its own
96
  text β€” **0 contradictory** in both eval sets (positive rates 0.482 / 0.486), per language in
97
  `benchmarks/report.md`.
98
 
99
+ **What this English row trades (published in full, not summarized away).** On the support-only
100
+ split the previous artifact scored 0.972 β€” `noul` 0.992, `score` 0.978, `choice` 0.946. The
101
+ five-domain bank scores 0.964 there: `choice` becomes **1.000** while `noul`/`score` give up 4.6
102
+ and 3.2 points, and in exchange the adapter covers four domains it could not answer at all
103
+ before β€” on the five-domain split the previous artifact scores **0.511** overall (worst domain
104
+ 0.328) against this one's **0.964** (worst domain 0.963). The gate that routes each `choice`
105
+ question to its domain head scores **strict 1.000** (0 to shared, 0 wrong domain) over 2,500 rows,
106
+ and accuracy on the 473 rows whose text never occurs in training is **0.998**. Full tables:
107
+ [`docs/benchmarks.md`](https://github.com/munod/tachyone/blob/main/docs/benchmarks.md).
108
+
109
  `choice` is the weak primitive on the multilingual side (0.468) and it is a *training* trade, not
110
  a label problem: `choice` labels never changed, and retraining on the corrected labels makes
111
  `noul` (1.000) and `score` (0.998) trivial β€” the same tone detector β€” while the shared trunk
112
  starves `choice` (an identical-recipe English control landed at 0.841, the five-domain retrain's
113
+ `choice` collapsed to 0.303). The English side did **not** retrain the trunk to escape that: it
114
+ re-fitted the `choice` head on a **frozen** trunk (B-5 / ADR-0016), which lifts `choice` to 1.000
115
+ and leaves `noul`/`score` where the trunk already had them β€” the cost and the gain in the table
116
+ above are one artifact, not two runs (provenance: `choice_bank_fit.json` next to the weights).
117
+ One of six languages meets ECE ≀ 0.05 (`es` 0.045); `pt` 0.062,
118
  `fr` 0.101, `de` 0.118, `nl` 0.192 and `it` 0.197 remain above target (NFR-C06), and multilingual
119
  `choice`/`noul` ECE (0.099 / 0.108) is declared with them. The CUDA-graph fast path
120
  (`TACHYONE_FAST=1`) gives a 2.65Γ— p50 speedup (8.97 β†’ 3.39 ms) with 0 top-label flips.
121
 
122
  **Robustness (B-4).** On a noisy view (one surface edit β€” typo/accents/casing β€” applied to 15% of
123
+ states) English drops only 0.964 β†’ 0.963 and multilingual 0.743 β†’ 0.744, so the released
124
  adapters are robust to this noise model.
125
 
126
  Full tables and environment are in