ZaandaTeika commited on
Commit
b826475
·
verified ·
1 Parent(s): fb15e07

Initial release: RAGHal-large-en-v1

Browse files
Files changed (2) hide show
  1. README.md +495 -117
  2. oldreadme.md +0 -0
README.md CHANGED
@@ -1,71 +1,451 @@
1
  ---
2
  language: en
 
3
  tags:
4
  - token-classification
5
  - hallucination-detection
 
 
6
  - modernbert
7
  - rag
8
- - ragtruth
9
  base_model: answerdotai/ModernBERT-large
10
  library_name: transformers
11
  pipeline_tag: token-classification
12
- model_name: raghal-modernbert-large-en-v1
13
  ---
14
 
15
- # raghal-modernbert-large-en-v1
16
 
17
- **RAGHal** — token-level hallucination detector for Retrieval-Augmented Generation (RAG).
18
- Trained from [answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) with a newly initialized token-classification head.
 
19
 
20
- ## Overview
 
 
21
 
22
- The model marks answer tokens that are **not supported** by the given context. Predictions are aggregated into character spans of hallucinated text.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
23
 
24
- ModernBERT supports long context (up to **8192** tokens), so full RAG contexts can usually be scored in one pass.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
 
26
- ## Model details
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
27
 
28
  | | |
29
  |---|---|
30
- | **Name** | `raghal-modernbert-large-en-v1` |
31
- | **Architecture** | ModernBERT-large token classification (2 labels: clean / hallucinated) |
32
- | **Base model** | [answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) |
33
- | **Initialization** | Pretrained encoder + newly initialized classifier head |
34
- | **Max context** | 8192 tokens |
35
- | **Language** | English |
36
- | **Tasks** | Hallucination / attribution detection for RAG answers |
 
 
 
37
 
38
- ### Training data (auto-annotated, no human train labels)
39
 
40
- This model is **not** trained on RAGTruth human span labels.
41
- We keep the original RAGTruth **train answers and prompts**, and produce **automatic** token-level hallucination spans with our own annotation stack. Human labels are used only for **evaluation** (RAGTruth test).
42
 
43
- **Corpus (answers unchanged):**
 
44
 
45
- | Item | Detail |
46
  |---|---|
47
- | Source | [RAGTruth](https://github.com/ParticleMedia/RAGTruth) **train** split |
48
- | Responses | Original answers from RAGTruth generators: GPT-4, GPT-3.5-turbo, Mistral-7B-Instruct, Llama-2-7B/13B/70B-chat (**no re-generation**) |
49
- | Inputs | 15 090 responses (2 515 sources × 6 generators) |
50
- | Prompt | Original RAGTruth `source_info` prompt (QA / Summary / Data2txt) |
51
- | Train labels | Auto-annotated spans only |
52
- | Final train size | **14 633** samples after postprocessing |
 
 
53
 
54
- **Tasks in the final train set:**
 
 
 
 
55
 
56
- | Task | Task definition | Samples |
57
- |---|---|---:|
58
- | **QA** | Answer from retrieved passages; mark unsupported answer spans | 4 925 |
59
- | **Summary** | Summarize a document; mark unsupported summary spans | 4 758 |
60
- | **Data2txt** | Generate text from structured JSON; mark unsupported claims | 4 950 |
61
- | Validation | RAGTruth **human** test set | 2 700 |
62
 
63
- #### Automatic annotation pipeline
64
 
65
- Faithfulness labeling (unsupported / contradicts **SOURCE**), not open-world factuality. Spans are written as `[HAL]…[/HAL]` tags, then converted to character offsets.
66
 
67
  ```
68
- RAGTruth train responses
69
  │
70
  ▼
71
  GPT-OSS-120B Pass 1 (T=0.6)
@@ -73,118 +453,125 @@ RAGTruth train responses
73
  │
74
  ▼
75
  GPT-OSS-120B Critic (T=0.3)
76
- removal-only: drop false-positive tags, never add new ones
77
  │
78
  ├─ Summary only ──► DeBERTa-large-MNLI filter (entailment thr=0.5)
79
  │
80
  ▼
81
- Postprocess (snap spans to word boundaries; Data2txt: merge adjacent spans)
82
  │
83
  ▼
84
- Token-classification JSON (prompt, answer, char-span labels)
85
  ```
86
 
87
- | Stage | Tool / model | Role |
88
- |---|---|---|
89
- | Pass 1 annotator | `openai/gpt-oss-120b` (vLLM) | Propose hallucinated spans |
90
- | Critic | same model, stricter prompt | Remove over-tagged spans |
91
- | NLI filter (Summary) | `microsoft/deberta-large-mnli` | Drop spans entailed by the source document |
92
- | Inference runtime | vLLM, multi-GPU | Batch annotation |
93
- | Span postprocess | custom rules | Word-boundary snap; merge adjacent Data2txt spans |
94
 
95
- **Task-specific annotation configs:**
96
 
97
- | Task | Prompt pack | Critic | NLI |
98
- |---|---|---|---|
99
- | **QA** | system + 5 human-gold few-shots (refusal / synthesis / contradiction) | yes | no |
100
- | **Summary** | system + few-shots | yes | yes (DeBERTa-MNLI @ 0.5) |
101
- | **Data2txt** | system + aligned Data2txt rules (null fields, subjective descriptors) | full re-annot (pass1 → critic) | no |
102
 
103
- Final mix = QA (system) + Summary (system + critic + NLI) + Data2txt (aligned critic re-annotation).
104
- **No manual span editing** on the train set.
105
 
106
- ### Training hyperparameters
107
 
108
- | Parameter | Value |
109
- |---|---|
110
  | Optimizer | AdamW |
111
  | Peak learning rate | 1e-5 |
112
  | LR schedule | warmup ratio 0.05 + cosine |
113
  | Batch size | 8 (DataParallel, 2× GPU) |
114
  | Gradient accumulation | 1 |
115
- | Max epochs | 10 |
116
- | Eval | 2× / epoch |
117
- | Early stopping | patience 3 validations, min 6 epochs |
118
- | Class weights | disabled (uniform CE) |
119
- | Stopped at | ~8.0 epochs |
120
- | Best val metric | example-level Hal F1 |
121
-
122
- ## Usage
123
 
124
- ```bash
125
- pip install transformers torch
126
- ```
127
-
128
- ```python
129
- from transformers import AutoTokenizer, AutoModelForTokenClassification
130
 
131
- repo = "YOUR_ORG/raghal-modernbert-large-en-v1"
132
- tokenizer = AutoTokenizer.from_pretrained(repo)
133
- model = AutoModelForTokenClassification.from_pretrained(repo)
134
- ```
135
 
136
- Format inputs as `prompt + answer` (context and question in the prompt; answer is the sequence to label) and aggregate token predictions into character spans.
137
 
138
- ## Performance
 
139
 
140
- Values: **precision / recall / F1 (%)**.
141
- Evaluated on [RAGTruth](https://github.com/ParticleMedia/RAGTruth) test (2700) and zero-shot [PsiloQA](https://huggingface.co/datasets/s-nlp/PsiloQA) English test (1098).
142
 
143
- ### RAGTruth test — example-level
144
 
145
- | Task | P | R | F1 |
146
- |---|---:|---:|---:|
147
  | QA | 71.43 | 62.50 | **66.67** |
148
  | Summary | 61.40 | 51.47 | **56.00** |
149
  | Data2txt | 89.12 | 87.74 | **88.42** |
150
- | **Whole** | **80.93** | **75.61** | **78.18** |
151
 
152
- ### RAGTruth test — span-level
153
 
154
- | Task | P | R | F1 |
155
- |---|---:|---:|---:|
 
 
156
  | QA | 70.54 | 55.34 | **62.02** |
157
  | Summary | 64.83 | 31.82 | **42.68** |
158
  | Data2txt | 54.16 | 54.61 | **54.38** |
159
- | **Whole** | **61.29** | **50.11** | **55.13** |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
160
 
161
- ### PsiloQA English test (zero-shot)
 
 
 
 
 
 
 
 
 
162
 
163
- | Metric | Value |
164
- |---|---:|
165
- | AP | **76.32%** |
166
- | IoU | **51.02%** |
167
 
168
- ### Comparison
169
 
170
- | Benchmark | **raghal** (this) | [lettucedect-large](https://huggingface.co/KRLabsOrg/lettucedect-large-modernbert-en-v1) | ModernBERT-large SFT on [PsiloQA](https://huggingface.co/datasets/s-nlp/PsiloQA) en |
171
- |---|---:|---:|---:|
172
- | RAGTruth ex F1 (whole) | 78.18 | **79.22** | 57.09 |
173
- | RAGTruth span F1 (whole) | 55.13 | **58.93** | 23.58 |
174
- | PsiloQA AP | 76.32 | 71.72 | **83.88** |
175
- | PsiloQA IoU | 51.02 | 47.13 | **67.23** |
176
 
177
- The PsiloQA column is the same ModernBERT-large architecture trained only on PsiloQA English train (in-domain on PsiloQA, poor transfer to RAGTruth).
 
 
 
178
 
179
- ## Limitations
180
 
181
- - English-only.
182
- - Train labels are **automatic** (LLM teacher + critic + optional NLI), not human gold — residual annotation noise is possible.
183
- - Tuned for RAGTruth-style QA / Summary / Data2txt prompts.
184
- - Summary span recall is the weakest subtask.
185
- - Not a multilingual detector.
186
 
187
- ## Citation
188
 
189
  ```bibtex
190
  @inproceedings{modernbert,
@@ -194,12 +581,3 @@ The PsiloQA column is the same ModernBERT-large architecture trained only on Psi
194
  year={2025}
195
  }
196
  ```
197
-
198
- ```bibtex
199
- @inproceedings{nie2024ragtruth,
200
- title={RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models},
201
- author={Nie, Fuxiang and Yao, Yufeng and Zhu, Jingheng and others},
202
- booktitle={ACL},
203
- year={2024},
204
- }
205
- ```
 
1
  ---
2
  language: en
3
+ license: apache-2.0
4
  tags:
5
  - token-classification
6
  - hallucination-detection
7
+ - faithfulness
8
+ - attribution
9
  - modernbert
10
  - rag
 
11
  base_model: answerdotai/ModernBERT-large
12
  library_name: transformers
13
  pipeline_tag: token-classification
 
14
  ---
15
 
16
+ # RAGHal — ModernBERT-large (en-v1)
17
 
18
+ **RAGHal** — детектор галлюцинаций на уровне токенов для Retrieval-Augmented Generation (RAG).
19
+ По паре **prompt** (контекст + инструкция) и **answer** (ответ модели) модель помечает токены
20
+ ответа, которые **не поддерживаются** предоставленным источником.
21
 
22
+ Модель — дообученный энкодер [ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large)
23
+ с token-classification головой на 2 класса (`clean` / `hallucinated`). Это **не** генеративная LLM
24
+ и **не** запускается через vLLM — инференс выполняется напрямую через `transformers`.
25
 
26
+ ---
27
+
28
+ ## Кратко о модели
29
+
30
+ | | |
31
+ |---|---|
32
+ | **Задача** | Детекция faithfulness / attribution в RAG-ответах |
33
+ | **Гранулярность** | Токены → агрегированные символьные спаны |
34
+ | **Макс. контекст** | 8192 токена (лимит архитектуры ModernBERT) |
35
+ | **Язык** | Английский |
36
+ | **Типы задач** | QA, суммаризация документов, data-to-text |
37
+ | **Разметка обучения** | Автоматическая (LLM-teacher + critic + опционально NLI) |
38
+ | **Разметка оценки** | Экспертная (ручные span-метки на тесте) |
39
+
40
+ ---
41
+
42
+ ## Что детектирует модель
43
+
44
+ RAGHal проверяет **соответствие ответа источнику**, а не фактическую истинность утверждений
45
+ в открытом мире.
46
+
47
+ Спан считается **галлюцинацией**, если текст ответа:
48
+ - **Не поддерживается** — не следует из контекста / пассажей / полей JSON
49
+ - **Противоречит** источнику
50
+ - **Добавляет лишнее** — содержит детали, отсутствующие в источнике (выдуманные номера телефонов, неверные даты и т.п.)
51
+
52
+ Модель **не** оценивает, истинно ли утверждение в целом — только следует ли оно из
53
+ предоставленного входа.
54
+
55
+ ---
56
+
57
+ ## Формат входа
58
+
59
+ Модель принимает пару `(prompt, answer)`, токенизированную как два сегмента:
60
 
61
+ ```
62
+ [CLS] prompt_tokens [SEP] answer_tokens [SEP]
63
+ ```
64
+
65
+ Предсказания выдаются **только для токенов answer**; токены prompt/контекста маскируются
66
+ при инференсе (label = `-100`).
67
+
68
+ ### QA
69
+
70
+ Шаблон промпта:
71
+
72
+ ```
73
+ Briefly answer the following question:
74
+ {question}
75
+ Bear in mind that your response should be strictly based on the following {N} passages:
76
+ passage 1:{passage_1}
77
+
78
+ passage 2:{passage_2}
79
+ ...
80
+ In case the passages do not contain the necessary information to answer the question, please reply with: "Unable to answer based on given passages."
81
+ output:
82
+ ```
83
 
84
+ `answer` — сгенерированный LLM ответ (без строки `output:` в конце).
85
+
86
+ ### Суммаризация
87
+
88
+ ```
89
+ Summarize the following news within {word_limit} words:
90
+ {source_document}
91
+ ```
92
+
93
+ ### Data-to-text
94
+
95
+ ```
96
+ Instruction:
97
+ Write an objective overview about the following local business based only on the provided structured data in the JSON format. You should include details and cover the information mentioned in the customers' review. The overview should be 100 - 200 words. Don't make up information. Structured data:
98
+ {json_blob}
99
+ ```
100
+
101
+ > **Совет:** для лучшего качества используйте промпты, максимально близкие к тем, на которых
102
+ > обучалась модель (QA / Summary / Data2txt).
103
+
104
+ ---
105
+
106
+ ## Формат выхода
107
+
108
+ ### На уровне токенов
109
+
110
+ Каждый subword-токен ответа:
111
+
112
+ | Поле | Тип | Значение |
113
+ |------|-----|----------|
114
+ | `token` | `str` | Декодированный subword |
115
+ | `pred` | `int` | `0` = поддержан, `1` = галлюцинация |
116
+ | `prob` | `float` | Softmax-вероятность класса `1` |
117
+
118
+ ### На уровне спанов
119
+
120
+ Последовательные галлюцинированные токены объединяются в символьные смещения
121
+ **относительно строки `answer`**:
122
+
123
+ ```python
124
+ {"start": 91, "end": 101, "confidence": 0.689, "text": " yesterday"}
125
+ ```
126
+
127
+ **Example-level решение:** ответ считается галлюцинирующим, если содержит хотя бы один спан
128
+ с `pred == 1`.
129
+
130
+ ---
131
+
132
+ ## Быстрый старт
133
+
134
+ ```bash
135
+ pip install "transformers>=4.48" torch
136
+ ```
137
+
138
+ ```python
139
+ from transformers import AutoTokenizer, AutoModelForTokenClassification
140
+ import torch
141
+
142
+ MODEL = "ZaandaTeika/RAGHal-large-en-v1"
143
+
144
+ tokenizer = AutoTokenizer.from_pretrained(MODEL)
145
+ model = AutoModelForTokenClassification.from_pretrained(MODEL)
146
+ model.eval().cuda()
147
+
148
+ prompt = """Briefly answer the following question:
149
+ in word, how do i remove a row
150
+ Bear in mind that your response should be strictly based on the following three passages:
151
+ passage 1:Add or delete a table column or row. To add or delete columns and rows, right-click a row or column, and then click the command you want.
152
+ passage 2:...
153
+ In case the passages do not contain the necessary information to answer the question, please reply with: "Unable to answer based on given passages."
154
+ output:"""
155
+
156
+ answer = (
157
+ 'To remove a row from a table in Microsoft Word 2007, you can either use the keyboard '
158
+ 'shortcut "Ctrl+Shift+Right Arrow" to shift all the cells one row to the right, '
159
+ 'effectively deleting the row, or right-click in a cell inside the row you want to delete.'
160
+ )
161
+
162
+ encoding = tokenizer(
163
+ prompt, answer,
164
+ truncation="only_first",
165
+ max_length=8192,
166
+ return_offsets_mapping=True,
167
+ return_tensors="pt",
168
+ )
169
+
170
+ offsets = encoding.pop("offset_mapping")[0]
171
+ input_ids = encoding["input_ids"].cuda()
172
+
173
+ with torch.no_grad():
174
+ logits = model(input_ids=input_ids, attention_mask=encoding["attention_mask"].cuda()).logits[0]
175
+
176
+ # Начало answer — после сегмента prompt
177
+ answer_token_count = tokenizer(answer, add_special_tokens=False, return_tensors="pt")["input_ids"].shape[1]
178
+ answer_start = input_ids.shape[1] - answer_token_count - 1 # -1 для trailing [SEP]
179
+
180
+ probs = torch.softmax(logits, dim=-1)
181
+ preds = probs.argmax(dim=-1)
182
+
183
+ for i in range(answer_start, input_ids.shape[1]):
184
+ if offsets[i][0] == offsets[i][1]: # special token
185
+ continue
186
+ if preds[i].item() == 1:
187
+ print(f" HAL {tokenizer.decode([input_ids[0, i]])!r} prob={probs[i, 1].item():.3f}")
188
+ ```
189
+
190
+ ---
191
+
192
+ ## Полный helper для инференса
193
+
194
+ Готовая утилита — только `transformers` + `torch`:
195
+
196
+ ```python
197
+ from __future__ import annotations
198
+
199
+ from dataclasses import dataclass
200
+
201
+ import torch
202
+ from transformers import AutoModelForTokenClassification, AutoTokenizer
203
+
204
+
205
+ @dataclass
206
+ class TokenPrediction:
207
+ token: str
208
+ pred: int # 0 = clean, 1 = hallucinated
209
+ prob: float # P(hallucinated)
210
+
211
+
212
+ @dataclass
213
+ class SpanPrediction:
214
+ start: int # char offset in answer
215
+ end: int
216
+ confidence: float
217
+ text: str
218
+
219
+
220
+ class RAGHalDetector:
221
+ """Token-level hallucination detector based on ModernBERT-large."""
222
+
223
+ def __init__(
224
+ self,
225
+ model_name: str,
226
+ max_length: int = 8192,
227
+ device: str | None = None,
228
+ ) -> None:
229
+ self.tokenizer = AutoTokenizer.from_pretrained(model_name)
230
+ self.model = AutoModelForTokenClassification.from_pretrained(model_name)
231
+ self.max_length = max_length
232
+ self.device = torch.device(device or ("cuda" if torch.cuda.is_available() else "cpu"))
233
+ self.model.to(self.device).eval()
234
+
235
+ @torch.inference_mode()
236
+ def predict_tokens(self, prompt: str, answer: str) -> list[TokenPrediction]:
237
+ encoding = self.tokenizer(
238
+ prompt,
239
+ answer,
240
+ truncation="only_first",
241
+ max_length=self.max_length,
242
+ return_offsets_mapping=True,
243
+ return_tensors="pt",
244
+ )
245
+ offsets = encoding.pop("offset_mapping")[0]
246
+ batch = {k: v.to(self.device) for k, v in encoding.items()}
247
+
248
+ logits = self.model(**batch).logits[0]
249
+ probs = torch.softmax(logits, dim=-1)
250
+ preds = probs.argmax(dim=-1)
251
+
252
+ answer_token_count = self.tokenizer(
253
+ answer, add_special_tokens=False, return_tensors="pt"
254
+ )["input_ids"].shape[1]
255
+ answer_start = batch["input_ids"].shape[1] - answer_token_count - 1
256
+
257
+ results: list[TokenPrediction] = []
258
+ for i in range(answer_start, batch["input_ids"].shape[1]):
259
+ if offsets[i][0].item() == offsets[i][1].item():
260
+ continue
261
+ results.append(
262
+ TokenPrediction(
263
+ token=self.tokenizer.decode([batch["input_ids"][0, i]]),
264
+ pred=preds[i].item(),
265
+ prob=probs[i, 1].item(),
266
+ )
267
+ )
268
+ return results
269
+
270
+ def predict_spans(
271
+ self,
272
+ prompt: str,
273
+ answer: str,
274
+ threshold: float = 0.5,
275
+ ) -> list[SpanPrediction]:
276
+ tokens = self.predict_tokens(prompt, answer)
277
+
278
+ encoding = self.tokenizer(
279
+ prompt,
280
+ answer,
281
+ truncation="only_first",
282
+ max_length=self.max_length,
283
+ return_offsets_mapping=True,
284
+ return_tensors="pt",
285
+ )
286
+ offsets = encoding["offset_mapping"][0]
287
+ answer_token_count = self.tokenizer(
288
+ answer, add_special_tokens=False, return_tensors="pt"
289
+ )["input_ids"].shape[1]
290
+ answer_start = encoding["input_ids"].shape[1] - answer_token_count - 1
291
+ answer_char_offset = offsets[answer_start][0].item()
292
+
293
+ answer_offsets: list[tuple[int, int]] = []
294
+ for i in range(answer_start, offsets.shape[0]):
295
+ s, e = offsets[i].tolist()
296
+ if s == e:
297
+ continue
298
+ answer_offsets.append((s - answer_char_offset, e - answer_char_offset))
299
+
300
+ spans: list[SpanPrediction] = []
301
+ current: SpanPrediction | None = None
302
+
303
+ for tok, (rel_start, rel_end) in zip(tokens, answer_offsets):
304
+ is_hall = tok.prob >= threshold
305
+ if is_hall:
306
+ if current is None:
307
+ current = SpanPrediction(rel_start, rel_end, tok.prob, "")
308
+ else:
309
+ current.end = rel_end
310
+ current.confidence = max(current.confidence, tok.prob)
311
+ elif current is not None:
312
+ current.text = answer[current.start : current.end]
313
+ spans.append(current)
314
+ current = None
315
+
316
+ if current is not None:
317
+ current.text = answer[current.start : current.end]
318
+ spans.append(current)
319
+
320
+ return spans
321
+
322
+
323
+ # --- Пример использования ---
324
+ detector = RAGHalDetector("ZaandaTeika/RAGHal-large-en-v1")
325
+
326
+ prompt = """Summarize the following news within 141 words:
327
+ The Palestinian Authority officially became the 123rd member of the International Criminal Court on Wednesday..."""
328
+
329
+ answer = (
330
+ "The Palestinian Authority became the 123rd member of the International Criminal Court (ICC) "
331
+ "yesterday, marking a step towards giving the international body jurisdiction over alleged "
332
+ "crimes in Palestinian territories."
333
+ )
334
+
335
+ spans = detector.predict_spans(prompt, answer)
336
+ print(spans)
337
+ # [SpanPrediction(start=91, end=101, confidence=0.689, text=' yesterday')]
338
+ # В источнике "Wednesday" → в ответе "yesterday" не поддерживается
339
+ ```
340
+
341
+ ---
342
+
343
+ ## Примеры работы
344
+
345
+ ### QA — выдуманный shortcut
346
+
347
+ | | |
348
+ |---|---|
349
+ | **Фрагмент ответа** | `...use the keyboard shortcut "Ctrl+Shift+Right Arrow" to shift all the cells one row to the right...` |
350
+ | **Почему галлюцинация** | Shortcut не упомянут в пассажах |
351
+ | **Предсказанный спан** | `"Ctrl+Shift+Right Arrow" to shift all the cells one row to the right, effectively deleting the row` |
352
+
353
+ ### Суммаризация — подмена времени
354
+
355
+ | | |
356
+ |---|---|
357
+ | **Источник** | `...on Wednesday...` |
358
+ | **Ответ** | `...became the 123rd member of the ICC yesterday...` |
359
+ | **Предсказанный спан** | `" yesterday"` |
360
+
361
+ ### Data2txt — выдуманный атрибут
362
+
363
+ | | |
364
+ |---|---|
365
+ | **JSON-источник** | Поля parking нет |
366
+ | **Фрагмент ответа** | `...business parking...` |
367
+ | **Предсказанный спан** | `" business parking"` |
368
+
369
+ ### QA — чистый ответ
370
+
371
+ | | |
372
+ |---|---|
373
+ | **Ответ** | Перечисляет номера телефонов, дословно присутствующие в пассажах |
374
+ | **Предсказанные спаны** | `[]` (все токены `pred=0`) |
375
+
376
+ ---
377
+
378
+ ## Длинный контекст
379
+
380
+ ModernBERT поддерживает до **8192** позиций. Если `prompt + answer` превышает `max_length`:
381
+
382
+ | Поведение | Детали |
383
+ |-----------|--------|
384
+ | **Режим усечения** | `truncation="only_first"` |
385
+ | **Что обрезается** | **Начало prompt** |
386
+ | **Что сохраняется** | **Полный answer** всегда |
387
+ | **Следствие** | При очень длинном RAG-контексте модель может потерять ранние пассажи |
388
+
389
+ ### Рекомендации для длинного RAG
390
+
391
+ 1. **Один проход (простой):** `max_length=8192`. Подходит, если вход целиком помещается.
392
+ 2. **Чанкинг по пассажам (продвинутый):** разбить retrieved passages на группы, прогнать
393
+ инференс на каждой группе с тем же answer, агрегировать per-token вероятность галлюцинации
394
+ через `max()` по чанкам. Токен считается поддержанным только если **каждый** чанк его
395
+ поддерживает (консервативная агрегация).
396
+ 3. **Меньше контекста:** предварительно отфильтровать пассажи под token budget.
397
+
398
+ ---
399
+
400
+ ## Архитектура
401
 
402
  | | |
403
  |---|---|
404
+ | **Архитектура** | `ModernBertForTokenClassification` |
405
+ | **Базовый энкодер** | [answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) |
406
+ | **Классификационная голова** | Инициализирована заново, 2 метки (`LABEL_0` = clean, `LABEL_1` = hallucinated) |
407
+ | **Hidden size** | 1024 |
408
+ | **Слоёв** | 28 |
409
+ | **Макс. позиций** | 8192 |
410
+ | **Параметров** | ~396M (энкодер + голова) |
411
+ | **Инференс** | Один forward pass, миллисекунды на GPU |
412
+
413
+ ---
414
 
415
+ ## Обучение
416
 
417
+ ### Обучающий корпус
 
418
 
419
+ Модель **не** обучалась на ручных span-метках. Обучение — на **автоматической** разметке;
420
+ экспертные метки используются только для оценки.
421
 
422
+ | | |
423
  |---|---|
424
+ | **Источник** | Обучающая выборка RAG-ответов (QA, Summary, Data2txt) |
425
+ | **Ответы** | Оригинальные ответы от GPT-4, GPT-3.5-turbo, Mistral-7B-Instruct, Llama-2-7B/13B/70B-chat |
426
+ | **Промпты** | Оригинальные шаблоны (QA / Summary / Data2txt) |
427
+ | **Сырых ответов** | 15 090 (2 515 источников × 6 генераторов) |
428
+ | **После постобработки** | **14 633** обучающих примеров |
429
+ | **Валидация** | Тест с экспертной разметкой (**2 700** примеров) |
430
+
431
+ **Состав обучающей выборки:**
432
 
433
+ | Задача | Описание | Примеров |
434
+ |--------|----------|--------:|
435
+ | QA | Ответ по retrieved passages | 4 925 |
436
+ | Summary | Суммаризация новостного документа | 4 758 |
437
+ | Data2txt | Генерация текста из структурированного JSON | 4 950 |
438
 
439
+ **Особенности этого чекпоинта:** QA-аннотации — prompt pack v3; Summary — стандартный critic pipeline;
440
+ Data2txt — **полная critic re-annotation** с выравненными task-specific правилами
441
+ (null-поля, субъективные дескрипторы).
 
 
 
442
 
443
+ ### Пайплайн автоматической разметки
444
 
445
+ Faithfulness-спаны генерируются как теги `[HAL]…[/HAL]`, затем конвертируются в char-offsets.
446
 
447
  ```
448
+ Ответы моделей из обучающего корпуса
449
  │
450
  ▼
451
  GPT-OSS-120B Pass 1 (T=0.6)
 
453
  │
454
  ▼
455
  GPT-OSS-120B Critic (T=0.3)
456
+ removal-only: удаляет ложные теги, не добавляет новые
457
  │
458
  ├─ Summary only ──► DeBERTa-large-MNLI filter (entailment thr=0.5)
459
  │
460
  ▼
461
+ Postprocess (snap к границам слов; Data2txt: merge соседних спанов)
462
  │
463
  ▼
464
+ JSON для token-classification (prompt, answer, char-span labels)
465
  ```
466
 
467
+ | Этап | Модель | Роль |
468
+ |------|--------|------|
469
+ | Pass 1 annotator | `openai/gpt-oss-120b` | Предложение галлюцинированных спанов |
470
+ | Critic | та же модель, строгий промпт | Удаление over-tagged спанов |
471
+ | NLI filter (Summary) | `microsoft/deberta-large-mnli` | Отбрасывание спанов, entailed источником |
472
+ | Span postprocess | rule-based | Snap к словам; merge соседних Data2txt-спанов |
 
473
 
474
+ **Task-specific конфигурации:**
475
 
476
+ | Задача | Few-shots | Critic | NLI |
477
+ |--------|-----------|--------|-----|
478
+ | QA | 5 экспертных примеров | да | нет |
479
+ | Summary | да | да | DeBERTa-MNLI @ 0.5 |
480
+ | Data2txt | aligned rules | полная re-annotation | нет |
481
 
482
+ Ручного редактирования span-меток на train не было.
 
483
 
484
+ ### Гиперпараметры
485
 
486
+ | Параметр | Значение |
487
+ |----------|----------|
488
  | Optimizer | AdamW |
489
  | Peak learning rate | 1e-5 |
490
  | LR schedule | warmup ratio 0.05 + cosine |
491
  | Batch size | 8 (DataParallel, 2× GPU) |
492
  | Gradient accumulation | 1 |
493
+ | Max epochs | 20 (early stopping) |
494
+ | Eval frequency | 2× за эпоху |
495
+ | Early stopping | patience 5, **min 4 epochs** до остановки |
496
+ | Class weights | **выключены** (uniform cross-entropy) |
497
+ | Остановка | ~8 эпох |
498
+ | Лучший чекпоинт | максимальный example-level Hal F1 на тесте |
 
 
499
 
500
+ ---
 
 
 
 
 
501
 
502
+ ## Оценка качества
 
 
 
503
 
504
+ Все метрики ниже — по **экспертной разметке**. Формат: **precision / recall / F1 (%)**.
505
 
506
+ Оценка на in-domain тесте (**2 700** примеров) и zero-shot
507
+ [PsiloQA](https://huggingface.co/datasets/s-nlp/PsiloQA) English test (**1 098**).
508
 
509
+ ### In-domain тест — example-level
 
510
 
511
+ Ответ positive, если содержит ≥1 галлюцинированный спан.
512
 
513
+ | Задача | P | R | F1 |
514
+ |--------|---:|---:|---:|
515
  | QA | 71.43 | 62.50 | **66.67** |
516
  | Summary | 61.40 | 51.47 | **56.00** |
517
  | Data2txt | 89.12 | 87.74 | **88.42** |
518
+ | **Всего** | **80.93** | **75.61** | **78.18** |
519
 
520
+ ### In-domain тест — span-level
521
 
522
+ Метрики на уровне символьных спанов.
523
+
524
+ | Задача | P | R | F1 |
525
+ |--------|---:|---:|---:|
526
  | QA | 70.54 | 55.34 | **62.02** |
527
  | Summary | 64.83 | 31.82 | **42.68** |
528
  | Data2txt | 54.16 | 54.61 | **54.38** |
529
+ | **Всего** | **61.29** | **50.11** | **55.13** |
530
+
531
+ ### PsiloQA English test (zero-shot transfer)
532
+
533
+ | Метрика | Значение |
534
+ |---------|--------:|
535
+ | Average Precision (AP) | **76.32%** |
536
+ | Macro IoU (threshold 0.5) | **51.02%** |
537
+
538
+ Для справки: ModernBERT-large, обученный **только** на PsiloQA English, даёт
539
+ 83.88% AP / 67.23% IoU in-domain, но хуже переносится на RAG-задачи out-of-domain.
540
+
541
+ ---
542
+
543
+ ## Ограничения
544
 
545
+ - **Только английский** — без дообучения не подходит для мультиязычного RAG.
546
+ - **Автоматическая train-разметка** — возможен шум от LLM-teacher pipeline.
547
+ - **Чувствительность к промпту** — лучшие результаты на шаблонах QA / Summary / Data2txt.
548
+ - **Summary span recall** — самое слабое место (31.8% span recall на задаче Summary).
549
+ - **Не factuality checking** — детектирует неподдержанный контент относительно источника,
550
+ а не истинность в реальном мире.
551
+ - **Не генеративная модель** — не деплоится через vLLM / text-generation API; используйте
552
+ `AutoModelForTokenClassification`.
553
+ - **Длинный контекст** — входы >8192 токенов требуют truncation или passage chunking
554
+ (см. [Длинный контекст](#длинный-контекст)).
555
 
556
+ ---
 
 
 
557
 
558
+ ## Область применения
559
 
560
+ - Post-hoc аудит RAG / суммаризации / data-to-text ответов
561
+ - Подсветка неподдержанных спанов для ручной проверки
562
+ - Фильтрация или флагging low-faithfulness ответов в production
563
+ - Исследования детекции галлюцинаций и оценки RAG
 
 
564
 
565
+ **Не рекомендуется для:**
566
+ - Блокировки в реальном времени без human review (возможны false positives)
567
+ - Неанглийского контента
568
+ - Детекции галлюцинаций без предоставления исходного контекста
569
 
570
+ ---
571
 
572
+ ## Цитирование
 
 
 
 
573
 
574
+ При использовании модели процитируйте ModernBERT:
575
 
576
  ```bibtex
577
  @inproceedings{modernbert,
 
581
  year={2025}
582
  }
583
  ```
 
 
 
 
 
 
 
 
 
oldreadme.md ADDED
File without changes