File size: 11,536 Bytes
4a8bc16 92a5cbe 4a8bc16 deb9835 4a8bc16 f16b495 82aa0f0 4a8bc16 deb9835 4a8bc16 13de6cf 4a8bc16 13de6cf 4a8bc16 13de6cf 4a8bc16 13de6cf 599fb9a 4a8bc16 13de6cf 4a8bc16 13de6cf 4a8bc16 599fb9a 4a8bc16 deb9835 4a8bc16 13de6cf 4a8bc16 deb9835 4a8bc16 93c90c6 deb9835 4a8bc16 93c90c6 4a8bc16 deb9835 4a8bc16 13de6cf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 | ---
license: cc-by-nc-sa-4.0
language:
- he
- en
library_name: laya
pipeline_tag: zero-shot-classification
base_model: dicta-il/neodictabert-bilingual
tags:
- laya
- hebrew
- decision-model
- calibrated
- rlcd
datasets:
- LocalLLaMA/typed-decisions
- clinc/clinc_oos
- legacy-datasets/banking77
- fancyzhx/ag_news
- fancyzhx/dbpedia_14
- community-datasets/yahoo_answers_topics
- facebook/anli
- Yelp/yelp_review_full
- HebArabNlpProject/HebNLI
- Etelis/HeQ_v1
- HebArabNlpProject/HebrewSentiment
- Tobi-Bueck/customer-support-tickets
- nyu-mll/glue
- google/civil_comments
- google-research-datasets/go_emotions
- ucberkeley-dlab/measuring-hate-speech
- tasksource/tasksource-instruct-v0
- wikimedia/wikipedia
- HuggingFaceFW/fineweb-2
- google/boolq
- ehovy/race
---
# Laya-Hebrew: a HebrewโEnglish decision model
A [Laya](https://github.com/NandhaKishorM/laya)-style decision model for Hebrew and English. You give it a **state** (the
text or fields to judge) and **questions**: a choice between options, a score on a scale, or a yes/no claim. In one
forward pass it returns a **calibrated probability for every answer**. It is a fast classifier that you configure at
call time. It is not a chatbot and it does not generate text.
- **Encoder:** [`dicta-il/neodictabert-bilingual`](https://huggingface.co/dicta-il/neodictabert-bilingual) (NeoBERT, 28
layers, Hebrew + English)
- **Head:** Laya's `DecisionModel` architecture (2 transformer layers and a scorer over the option markers), trained
from scratch. No weights come from Laya's published checkpoints; the only pretrained weights are the encoder's.
- **Size:** 378M parameters (encoder 363M, head 15M), stored in fp16 (755 MB)
- **Speed on an Apple M1 CPU (4 threads):** about 0.1 s per question for a short message, and about 1 s at the full
input length
- **Input:** 1,024 tokens in total, per question
- The instructions and options share 256 of those tokens, and each option is cut at 48 tokens.
- The state gets the rest. A longer state is cut from the end without a warning, so put what matters first.
- **Training:** Laya's RLCD objective (proper-scoring-rule rewards plus soft cross-entropy), with a temperature for each
question type fitted on held-out human- or rule-labeled items
- **Use:** non-commercial only (see [License](#license))
## Usage
This checkpoint needs Laya 0.3.7 (commit `010bace`) with `neobert.patch`. The patch loads NeoBERT's remote code,
recomputes its rotary tables (without that, every output is NaN under transformers 5) and keeps the encoder in fp32.
```bash
git clone https://github.com/NandhaKishorM/laya && cd laya && git checkout 010bace
git apply /path/to/neobert.patch # from this repository
pip install -e . # torch, transformers 5.x
```
```python
import laya
agent = laya.load("RoeiG/laya-hebrew", device="cpu") # or "cuda"
out = agent.predict(
{"message": "ืืืคืืืงืฆืื ืงืืจืกืช ืืฉืื ื ืคืืชื ืืช ืืืฆืืื"},
{
"team": {"type": "choice", "instructions": "ืืืื ืฆืืืช ืฆืจืื ืืืคื ืืืืืขื?",
"criteria": {"billing": "ืชืฉืืืืื ืืืืืจืื", "tech": "ืืืืื ืืงืจืืกืืช", "shipping": "ืืฉืืืืื"}},
"upset": {"type": "noul", "instructions": "ืืืงืื ืืืขืก."},
},
)
# out["answers"]["team"] -> {"choice": "tech", "probabilities": {"billing": 0.0014, "tech": 0.9973, "shipping": 0.0013}, ...}
# out["answers"]["upset"] -> {"noul": 0.0551, ...} (P(true))
```
Question types:
- **`choice`:** `criteria` maps each option to a description.
- **`score`:** `criteria` is a list of levels from lowest to highest. The answer includes the expected level.
- **`noul`:** yes/no. The answer is P(true) for the statement in `instructions`.
## How to get good answers
1. **Compute numbers, dates, units and relations in code.** Pass the result as a field, such as
`"age_ok": "ืืืื ืขืืื ืืชื ืื"` or `"relation": "The sender is the receiver's direct manager"`. The model does not do
arithmetic reliably (see Limitations).
2. **Prefer the claim form for yes/no.** "ืืืงืื ืืืขืก." discriminates better than "ืืื ืืืงืื ืืืขืก?" (gap 0.67 against
0.38 on he_bench). The question form works, but it is weaker.
3. **Describe every option in a line.** Bare labels or codes route much worse than labels with a one-line description.
Name what each option owns, not only its keywords.
4. **Read the probabilities, not only the top answer.** A top answer below about 0.6 means the model is unsure.
5. **Ignore `act_probability`.** It comes from a head that no Hebrew checkpoint trained.
## Evaluation
Every evaluation set below was held out of training. he_bench v1 has 8 Hebrew tasks, each asked in 3 phrasings; the
set is frozen by sha256. Brier and ECE are better when lower. "laya-multilingual" is Laya's published multilingual
checkpoint. "Previous" is this project's previous checkpoint, trained without the reading and teacher data. **Bold**
marks the best value in each row: the highest, or the lowest for Brier and ECE. Differences smaller than the run-to-run
noise below are not meaningful.
| Metric | laya-multilingual | Previous | Laya-Hebrew |
|---|---|---|---|
| MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
| MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
| MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
| he_bench accuracy (4,290) | 0.439 | 0.527 | **0.595** |
| he_bench Brier (lower is better) | 0.701 | 0.538 | **0.499** |
| he_bench ECE (lower is better) | 0.266 | **0.119** | 0.125 |
| Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
| SIB-200-he topic | | **0.808** | 0.801 |
| Yes/no as a question: P(yes \| true) โ P(yes \| false) | | 0.047 | **0.379** |
| Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
| Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
| Rule-direction probe (208) | | 0.798 | **0.832** |
| Held-out soft-label set (900), Brier (lower is better) | | **0.230** | 0.231 |
he_bench accuracy by task:
| Task | Accuracy | Chance |
|---|---|---|
| relevance (question form) | 0.785 | 0.50 |
| qa_verify (question form) | 0.704 | 0.50 |
| copa | 0.687 | 0.50 |
| sentiment | 0.633 | 0.33 |
| winograd | 0.607 | 0.50 |
| hellaswag | 0.447 | 0.25 |
| tone arousal | 0.389 | 0.21 |
| tone valence | 0.328 | 0.23 |
Chance is the accuracy of a uniformly random pick. The tone tasks are 5-level scales, and their chance is slightly above
0.20 because some items tie between two levels.
**Run-to-run noise.** A second training seed, with the same data and settings, differed by these amounts:
- he_bench overall: 0.2 points
- Belebele and MASSIVE: 0.6โ1.0 points
- single he_bench tasks: 1โ3 points
- each half of the rule probe: 5โ7 points
Differences smaller than these are noise. This checkpoint is seed 1, which was fixed as the release before training.
## Limitations
- **Reasoning is the weak spot.**
- Hellaswag, winograd and copa are well above chance but far from solved.
- Multi-step inferences (e.g. "A is taller than B, B is taller than C: who is shortest?") often fail.
- **Numeric, date and unit rules are unreliable, and often confidently wrong.**
- Examples: 2.5 hours against a 2-hour limit, a purchase 19 days ago against a 14-day window, or age 17 against an
English "18 and up" rule. These got P(true) of 0.93โ0.98.
- Compute them in code.
- **Irony and sarcasm are read literally.** "ืืืื, ืฉืืจืืช ืืืืืโฆ ื ืืชืงื ืื ๐" is scored as positive, at 0.97.
- **Routing leans on keywords.**
- In a small hand-written check of tech tickets, anything that mentioned "ืืืคืืื" (deploy) was pulled towards DevOps.
That happened even when the cause was a code bug, a network path, an expired certificate or a locked account.
- Strong keywords in the text can outweigh the option descriptions.
- **Claim-form relevance is slightly below the previous checkpoint:** he_bench 0.78 against 0.80, and BEIR-he 0.78 against 0.82.
- **Scales:** sentiment and tone are about 0.45 accuracy, and the tone probabilities are overconfident (valence Brier
0.53).
- **Calculated fields steer less than in the previous checkpoint.** On 8 test emails, a "direct manager" sender field
raised importance by 0.11 of a level, against 0.22 before, and a "mailing list" field lowered it in only 2 of 8.
- **Calibration does not catch everything.** The failures above are often high-confidence, so a confidence threshold
will not filter them out.
## Training data
The training start was an earlier checkpoint of this project, with the same encoder, trained on part of the public
data below. The run was one epoch over 396,508 items (6,196 updates, 1.4 A100 hours). The learning rates were 5e-6 for the
encoder and 1e-4 for the head. Every case was converted to Laya's format, with a random subset of options, paraphrased
instructions and varied field names.
- **Public labeled data**, English unless marked:
- typed-decisions
- CLINC, Banking77, AG News, DBpedia, Yahoo Answers
- ANLI, Yelp, GLUE STS-B, Civil Comments, customer-support tickets
- Hebrew NLI (HebNLI), Hebrew QA (HeQ) and Hebrew sentiment
- **Soft labels from annotator disagreement:**
- GoEmotions and Measuring Hate Speech
- SNLI and MultiNLI votes, used as P(true)
- plus breadth from tasksource-instruct
- **Native Hebrew:** Hebrew Wikipedia topics and facts, labeled from Wikidata.
- **Relevance and spam:** BEIR-he relevance (biunlp's Hebrew translation of BEIR) and UCI SMS spam.
- **Machine-translated:** part of the English data was translated to Hebrew with NLLB-200-distilled-600M: 60% of the
soft-label cases, and half of the NLI votes and the spam. Only the state text was translated; the questions and
options stayed as they were.
- **Generated with code-computed labels:** rule and unit checks.
- **Added in the final training run:**
- RACE (30,000 questions) and BoolQ, in English with human labels.
- About 58,000 Hebrew items written and labeled by
[DictaLM-3.0-24B](https://huggingface.co/dicta-il) (Apache-2.0), on passages from Hebrew Wikipedia and FineWeb-2
Hebrew:
- yes/no questions about a passage, and answer checking
- queryโpassage relevance
- 4-option reading comprehension
- routing of invented business messages
- BoolQ translated to Hebrew, which keeps its human answers
- An item was kept only when the teacher's label agreed with the answer it was written for.
- Subjective scales were never teacher-labeled.
- The teacher-generated files are not published.
Decontamination: no teacher text shares an 8-word run with any evaluation set. The evaluation instructions and the
yes/no wordings used by he_bench were kept out of training. MASSIVE was never trained on.
No private data and no personal data were used.
## License
**CC-BY-NC-SA-4.0: non-commercial use only.** Several training sources are non-commercial or research-only: ANLI, Yelp,
AG News, Yahoo Answers, RACE, and the NLLB translation model (CC-BY-NC-4.0). Others are ShareAlike: Wikipedia, BEIR-he,
SNLI and BoolQ. The encoder is by Dicta (`neodictabert-bilingual`, CC-BY-4.0). The Laya architecture, training method
and runtime are by Laya's authors (Apache-2.0).
## Acknowledgements
- Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
- Laya's authors, for the architecture, the RLCD training method and the runtime
- The creators of every dataset listed above
|