oddadmix commited on
Commit
4f789c2
·
verified ·
1 Parent(s): e7d1ec5

Add out-of-domain English->Egyptian leaderboard eval

Browse files
Files changed (1) hide show
  1. README.md +10 -8
README.md CHANGED
@@ -106,13 +106,13 @@ def translate(text, system=SYSTEM):
106
  - **Eval split:** 3,000 deterministic held-out pairs (`seed=42`), scored both directions.
107
 
108
  <!-- BEGIN: leaderboard-en2egy-eval -->
109
- ## Out-of-domain evaluation — Egyptian leaderboard set
110
 
111
  The results above are **in-domain**: a held-out split of the same corpus this model was
112
- trained on. The numbers below are **out-of-domain** — the same model scored on
113
- [`oddadmix/egyptian_translation_eval_nourhan`](https://huggingface.co/datasets/oddadmix/egyptian_translation_eval_nourhan) (319 English→Egyptian pairs,
114
- a different annotator with different orthographic conventions), the eval set used by the
115
- Egyptian translation leaderboard.
116
 
117
  Expect these to be substantially lower than the in-domain scores. That gap is the
118
  generalization penalty, not a regression — both numbers are real, they measure different things.
@@ -149,9 +149,11 @@ translation that picks a different valid word is penalized — e.g. `فريش` v
149
  good Egyptian; only one matches the reference. **chrF and METEOR track perceived quality more
150
  closely here.**
151
 
152
- These are small models — 5M to 50M parameters, orders of magnitude below the large systems that
153
- top this leaderboard. The result of interest is the **scaling curve and per-parameter
154
- efficiency**, not absolute rank against models 100–1000× the size.
 
 
155
  <!-- END: leaderboard-en2egy-eval -->
156
 
157
  ## Limitations
 
106
  - **Eval split:** 3,000 deterministic held-out pairs (`seed=42`), scored both directions.
107
 
108
  <!-- BEGIN: leaderboard-en2egy-eval -->
109
+ ## Out-of-domain evaluation — Egyptian Arabic Translation Benchmark
110
 
111
  The results above are **in-domain**: a held-out split of the same corpus this model was
112
+ trained on. The numbers below are **out-of-domain** — the same model scored on the
113
+ [Egyptian Arabic Translation Benchmark](https://huggingface.co/datasets/oddadmix/egyptian-arabic-translation-benchmark)
114
+ (`oddadmix/egyptian-arabic-translation-benchmark`, 319 English→Egyptian pairs written by a different annotator with different
115
+ orthographic conventions).
116
 
117
  Expect these to be substantially lower than the in-domain scores. That gap is the
118
  generalization penalty, not a regression — both numbers are real, they measure different things.
 
149
  good Egyptian; only one matches the reference. **chrF and METEOR track perceived quality more
150
  closely here.**
151
 
152
+ At 319 rows, differences of roughly 1–2 BLEU between adjacent rungs are within noise.
153
+
154
+ These are small models — 5M to 50M parameters, orders of magnitude below the large systems
155
+ typically evaluated on this benchmark. The result of interest is the **scaling curve and
156
+ per-parameter efficiency**, not absolute rank against models 100–1000× the size.
157
  <!-- END: leaderboard-en2egy-eval -->
158
 
159
  ## Limitations