File size: 11,536 Bytes
4a8bc16
 
 
 
 
 
 
 
 
 
 
 
 
92a5cbe
4a8bc16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
deb9835
4a8bc16
 
 
 
 
 
 
 
f16b495
 
82aa0f0
 
 
 
 
 
4a8bc16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
deb9835
4a8bc16
 
 
 
 
 
 
 
 
13de6cf
 
 
4a8bc16
13de6cf
4a8bc16
13de6cf
4a8bc16
13de6cf
599fb9a
 
 
4a8bc16
13de6cf
4a8bc16
13de6cf
4a8bc16
 
599fb9a
4a8bc16
deb9835
4a8bc16
 
 
 
 
 
 
 
 
13de6cf
 
 
 
 
4a8bc16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
deb9835
4a8bc16
93c90c6
deb9835
 
4a8bc16
 
 
 
 
93c90c6
 
4a8bc16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
deb9835
4a8bc16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13de6cf
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
---
license: cc-by-nc-sa-4.0
language:
- he
- en
library_name: laya
pipeline_tag: zero-shot-classification
base_model: dicta-il/neodictabert-bilingual
tags:
- laya
- hebrew
- decision-model
- calibrated
- rlcd
datasets:
- LocalLLaMA/typed-decisions
- clinc/clinc_oos
- legacy-datasets/banking77
- fancyzhx/ag_news
- fancyzhx/dbpedia_14
- community-datasets/yahoo_answers_topics
- facebook/anli
- Yelp/yelp_review_full
- HebArabNlpProject/HebNLI
- Etelis/HeQ_v1
- HebArabNlpProject/HebrewSentiment
- Tobi-Bueck/customer-support-tickets
- nyu-mll/glue
- google/civil_comments
- google-research-datasets/go_emotions
- ucberkeley-dlab/measuring-hate-speech
- tasksource/tasksource-instruct-v0
- wikimedia/wikipedia
- HuggingFaceFW/fineweb-2
- google/boolq
- ehovy/race
---

# Laya-Hebrew: a Hebrewโ€“English decision model

A [Laya](https://github.com/NandhaKishorM/laya)-style decision model for Hebrew and English. You give it a **state** (the
text or fields to judge) and **questions**: a choice between options, a score on a scale, or a yes/no claim. In one
forward pass it returns a **calibrated probability for every answer**. It is a fast classifier that you configure at
call time. It is not a chatbot and it does not generate text.

- **Encoder:** [`dicta-il/neodictabert-bilingual`](https://huggingface.co/dicta-il/neodictabert-bilingual) (NeoBERT, 28
  layers, Hebrew + English)
- **Head:** Laya's `DecisionModel` architecture (2 transformer layers and a scorer over the option markers), trained
  from scratch. No weights come from Laya's published checkpoints; the only pretrained weights are the encoder's.
- **Size:** 378M parameters (encoder 363M, head 15M), stored in fp16 (755 MB)
- **Speed on an Apple M1 CPU (4 threads):** about 0.1 s per question for a short message, and about 1 s at the full
  input length
- **Input:** 1,024 tokens in total, per question
  - The instructions and options share 256 of those tokens, and each option is cut at 48 tokens.
  - The state gets the rest. A longer state is cut from the end without a warning, so put what matters first.
- **Training:** Laya's RLCD objective (proper-scoring-rule rewards plus soft cross-entropy), with a temperature for each
  question type fitted on held-out human- or rule-labeled items
- **Use:** non-commercial only (see [License](#license))

## Usage

This checkpoint needs Laya 0.3.7 (commit `010bace`) with `neobert.patch`. The patch loads NeoBERT's remote code,
recomputes its rotary tables (without that, every output is NaN under transformers 5) and keeps the encoder in fp32.

```bash
git clone https://github.com/NandhaKishorM/laya && cd laya && git checkout 010bace
git apply /path/to/neobert.patch          # from this repository
pip install -e .                          # torch, transformers 5.x
```

```python
import laya

agent = laya.load("RoeiG/laya-hebrew", device="cpu")   # or "cuda"
out = agent.predict(
    {"message": "ื”ืืคืœื™ืงืฆื™ื” ืงื•ืจืกืช ื›ืฉืื ื™ ืคื•ืชื— ืืช ื”ืžืฆืœืžื”"},
    {
        "team": {"type": "choice", "instructions": "ืื™ื–ื” ืฆื•ื•ืช ืฆืจื™ืš ืœื˜ืคืœ ื‘ื”ื•ื“ืขื”?",
                 "criteria": {"billing": "ืชืฉืœื•ืžื™ื ื•ื”ื—ื–ืจื™ื", "tech": "ื‘ืื’ื™ื ื•ืงืจื™ืกื•ืช", "shipping": "ืžืฉืœื•ื—ื™ื"}},
        "upset": {"type": "noul", "instructions": "ื”ืœืงื•ื— ื›ื•ืขืก."},
    },
)
# out["answers"]["team"]  -> {"choice": "tech", "probabilities": {"billing": 0.0014, "tech": 0.9973, "shipping": 0.0013}, ...}
# out["answers"]["upset"] -> {"noul": 0.0551, ...}   (P(true))
```

Question types:
- **`choice`:** `criteria` maps each option to a description.
- **`score`:** `criteria` is a list of levels from lowest to highest. The answer includes the expected level.
- **`noul`:** yes/no. The answer is P(true) for the statement in `instructions`.

## How to get good answers

1. **Compute numbers, dates, units and relations in code.** Pass the result as a field, such as
   `"age_ok": "ื”ื’ื™ืœ ืขื•ืžื“ ื‘ืชื ืื™"` or `"relation": "The sender is the receiver's direct manager"`. The model does not do
   arithmetic reliably (see Limitations).
2. **Prefer the claim form for yes/no.** "ื”ืœืงื•ื— ื›ื•ืขืก." discriminates better than "ื”ืื ื”ืœืงื•ื— ื›ื•ืขืก?" (gap 0.67 against
   0.38 on he_bench). The question form works, but it is weaker.
3. **Describe every option in a line.** Bare labels or codes route much worse than labels with a one-line description.
   Name what each option owns, not only its keywords.
4. **Read the probabilities, not only the top answer.** A top answer below about 0.6 means the model is unsure.
5. **Ignore `act_probability`.** It comes from a head that no Hebrew checkpoint trained.

## Evaluation

Every evaluation set below was held out of training. he_bench v1 has 8 Hebrew tasks, each asked in 3 phrasings; the
set is frozen by sha256. Brier and ECE are better when lower. "laya-multilingual" is Laya's published multilingual
checkpoint. "Previous" is this project's previous checkpoint, trained without the reading and teacher data. **Bold**
marks the best value in each row: the highest, or the lowest for Brier and ECE. Differences smaller than the run-to-run
noise below are not meaningful.

| Metric | laya-multilingual | Previous | Laya-Hebrew |
|---|---|---|---|
| MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
| MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
| MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
| he_bench accuracy (4,290) | 0.439 | 0.527 | **0.595** |
| he_bench Brier (lower is better) | 0.701 | 0.538 | **0.499** |
| he_bench ECE (lower is better) | 0.266 | **0.119** | 0.125 |
| Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
| SIB-200-he topic | | **0.808** | 0.801 |
| Yes/no as a question: P(yes \| true) โˆ’ P(yes \| false) | | 0.047 | **0.379** |
| Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
| Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
| Rule-direction probe (208) | | 0.798 | **0.832** |
| Held-out soft-label set (900), Brier (lower is better) | | **0.230** | 0.231 |

he_bench accuracy by task:

| Task | Accuracy | Chance |
|---|---|---|
| relevance (question form) | 0.785 | 0.50 |
| qa_verify (question form) | 0.704 | 0.50 |
| copa | 0.687 | 0.50 |
| sentiment | 0.633 | 0.33 |
| winograd | 0.607 | 0.50 |
| hellaswag | 0.447 | 0.25 |
| tone arousal | 0.389 | 0.21 |
| tone valence | 0.328 | 0.23 |

Chance is the accuracy of a uniformly random pick. The tone tasks are 5-level scales, and their chance is slightly above
0.20 because some items tie between two levels.

**Run-to-run noise.** A second training seed, with the same data and settings, differed by these amounts:
- he_bench overall: 0.2 points
- Belebele and MASSIVE: 0.6โ€“1.0 points
- single he_bench tasks: 1โ€“3 points
- each half of the rule probe: 5โ€“7 points

Differences smaller than these are noise. This checkpoint is seed 1, which was fixed as the release before training.

## Limitations

- **Reasoning is the weak spot.**
  - Hellaswag, winograd and copa are well above chance but far from solved.
  - Multi-step inferences (e.g. "A is taller than B, B is taller than C: who is shortest?") often fail.
- **Numeric, date and unit rules are unreliable, and often confidently wrong.**
  - Examples: 2.5 hours against a 2-hour limit, a purchase 19 days ago against a 14-day window, or age 17 against an
    English "18 and up" rule. These got P(true) of 0.93โ€“0.98.
  - Compute them in code.
- **Irony and sarcasm are read literally.** "ื•ื•ืื•, ืฉื™ืจื•ืช ืžื“ื”ื™ืโ€ฆ ื ื™ืชืงื• ืœื™ ๐Ÿ‘" is scored as positive, at 0.97.
- **Routing leans on keywords.**
  - In a small hand-written check of tech tickets, anything that mentioned "ื“ื™ืคืœื•ื™" (deploy) was pulled towards DevOps.
    That happened even when the cause was a code bug, a network path, an expired certificate or a locked account.
  - Strong keywords in the text can outweigh the option descriptions.
- **Claim-form relevance is slightly below the previous checkpoint:** he_bench 0.78 against 0.80, and BEIR-he 0.78 against 0.82.
- **Scales:** sentiment and tone are about 0.45 accuracy, and the tone probabilities are overconfident (valence Brier
  0.53).
- **Calculated fields steer less than in the previous checkpoint.** On 8 test emails, a "direct manager" sender field
  raised importance by 0.11 of a level, against 0.22 before, and a "mailing list" field lowered it in only 2 of 8.
- **Calibration does not catch everything.** The failures above are often high-confidence, so a confidence threshold
  will not filter them out.

## Training data

The training start was an earlier checkpoint of this project, with the same encoder, trained on part of the public
data below. The run was one epoch over 396,508 items (6,196 updates, 1.4 A100 hours). The learning rates were 5e-6 for the
encoder and 1e-4 for the head. Every case was converted to Laya's format, with a random subset of options, paraphrased
instructions and varied field names.

- **Public labeled data**, English unless marked:
  - typed-decisions
  - CLINC, Banking77, AG News, DBpedia, Yahoo Answers
  - ANLI, Yelp, GLUE STS-B, Civil Comments, customer-support tickets
  - Hebrew NLI (HebNLI), Hebrew QA (HeQ) and Hebrew sentiment
- **Soft labels from annotator disagreement:**
  - GoEmotions and Measuring Hate Speech
  - SNLI and MultiNLI votes, used as P(true)
  - plus breadth from tasksource-instruct
- **Native Hebrew:** Hebrew Wikipedia topics and facts, labeled from Wikidata.
- **Relevance and spam:** BEIR-he relevance (biunlp's Hebrew translation of BEIR) and UCI SMS spam.
- **Machine-translated:** part of the English data was translated to Hebrew with NLLB-200-distilled-600M: 60% of the
  soft-label cases, and half of the NLI votes and the spam. Only the state text was translated; the questions and
  options stayed as they were.
- **Generated with code-computed labels:** rule and unit checks.
- **Added in the final training run:**
  - RACE (30,000 questions) and BoolQ, in English with human labels.
  - About 58,000 Hebrew items written and labeled by
    [DictaLM-3.0-24B](https://huggingface.co/dicta-il) (Apache-2.0), on passages from Hebrew Wikipedia and FineWeb-2
    Hebrew:
    - yes/no questions about a passage, and answer checking
    - queryโ€“passage relevance
    - 4-option reading comprehension
    - routing of invented business messages
    - BoolQ translated to Hebrew, which keeps its human answers
  - An item was kept only when the teacher's label agreed with the answer it was written for.
  - Subjective scales were never teacher-labeled.
  - The teacher-generated files are not published.

Decontamination: no teacher text shares an 8-word run with any evaluation set. The evaluation instructions and the
yes/no wordings used by he_bench were kept out of training. MASSIVE was never trained on.

No private data and no personal data were used.

## License

**CC-BY-NC-SA-4.0: non-commercial use only.** Several training sources are non-commercial or research-only: ANLI, Yelp,
AG News, Yahoo Answers, RACE, and the NLLB translation model (CC-BY-NC-4.0). Others are ShareAlike: Wikipedia, BEIR-he,
SNLI and BoolQ. The encoder is by Dicta (`neodictabert-bilingual`, CC-BY-4.0). The Laya architecture, training method
and runtime are by Laya's authors (Apache-2.0).

## Acknowledgements

- Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
- Laya's authors, for the architecture, the RLCD training method and the runtime
- The creators of every dataset listed above