benediktnlp commited on
Commit
5164527
·
verified ·
1 Parent(s): ad070cd

Initial release

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ images/sui_icon.jpeg filter=lfs diff=lfs merge=lfs -text
37
+ tekken.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,572 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
4
+ tags:
5
+ - summarization
6
+ - long-context
7
+ - grounded-generation
8
+ - citation
9
+ - xml-tagging
10
+ language:
11
+ - en
12
+ - de
13
+ - es
14
+ - fr
15
+ - it
16
+ - pt
17
+ - nl
18
+ - pl
19
+ - zh
20
+ - ja
21
+ - ko
22
+ - ru
23
+ - ar
24
+ - tr
25
+ - vi
26
+ - id
27
+ - hi
28
+ - sv
29
+ - uk
30
+ - ro
31
+ - cs
32
+ - el
33
+ - hu
34
+ - th
35
+ pipeline_tag: text-generation
36
+ library_name: transformers
37
+ ---
38
+
39
+ <div align="center">
40
+ <img src="images/sui_icon.jpeg" alt="SUI Icon" width="200">
41
+ <h3>sui-1</h3>
42
+ </div>
43
+
44
+ **sui-1** (Summarization with Unique Identifiers) is a specialized model for high-quality summarization of very long texts with built-in source grounding. Every claim in the summary can be traced back to its source sentence, enabling verification and reducing hallucination risk.
45
+
46
+ ## Key Features
47
+
48
+ - **Very Long Document Processing**: Handles up to 128k tokens natively, with a two-step iterative approach for documents up to 2 million tokens
49
+ - **Single GPU Deployment**: The FP8 variant runs on a single A100 40GB or A6000 48GB GPU; the iterative approach enables deployment on even more modest hardware
50
+ - **Competitive Performance**: Significantly outperforms all tested open-weight baselines, including models with 3x more parameters
51
+ - **Multilingual Support**: Fine-tuned for English, German, Spanish, French, and Italian; inherits 20+ additional languages from Mistral Small 3.2
52
+ - **High-Quality Training Data**: Built using a sophisticated data generation pipeline that produced 22,000+ training examples from parliamentary documents, web sources, and Wikipedia using chain-of-thought reasoning with multi-stage verification
53
+ - **Verifiable Outputs**: Built-in citation mechanism links each claim to its source sentence for full traceability
54
+
55
+ ## Quick Start
56
+
57
+ Run the end-to-end [example.py](example.py) script (requires [uv](https://docs.astral.sh/uv/)):
58
+
59
+ ```bash
60
+ # Summarize a document
61
+ uv run example.py document.txt
62
+
63
+ # Or with inline text
64
+ uv run example.py --text "Your long text here..." --words 300 --tags 8
65
+ ```
66
+
67
+ The script handles everything: sentence tagging, model inference, and formatted output with source citations.
68
+
69
+ ## Evaluation
70
+
71
+ We evaluate sui-1-24b using an **LLM-as-a-Judge** methodology, where a strong judge model evaluates summary quality across multiple criteria. This approach captures nuanced quality aspects that traditional metrics like ROUGE cannot measure.
72
+
73
+ ### Overall Performance
74
+
75
+ ![Overall Performance](images/overall_performance.png)
76
+
77
+ The chart shows the overall success rate across all evaluation criteria. sui-1-24b significantly outperforms its base model (Mistral-Small-3.2-24B) on the summarization task.
78
+
79
+ ### Performance by Criteria
80
+
81
+ We evaluate summaries on five key dimensions:
82
+
83
+ | Criterion | Description |
84
+ |-----------|-------------|
85
+ | **Factual Accuracy** | Does the summary avoid introducing new facts, entities, numbers, or claims not supported by the source content? |
86
+ | **Coverage & Completeness**¹ | Does the summary cover the document's main points and key takeaways at appropriate granularity? |
87
+ | **Specificity & Informativeness** | Are claims specific and informative rather than generic filler (e.g., "there are several points")? |
88
+ | **Format Compliance** | Is the output compliant with formatting instructions including language consistency, semantic-aware planning, and paragraph structure? |
89
+ | **Custom Instruction**² | If a custom instruction is provided, is it followed appropriately? |
90
+
91
+ ![Criteria Breakdown](images/criteria_breakdown.png)
92
+
93
+ The evaluation was conducted on 100 diverse test samples covering multiple languages (English, German, Spanish, French, Italian) and document types. Scoring uses binary pass/fail per criterion, aggregated to success rates.
94
+
95
+ *¹ Coverage scores are lower when samples require constrained formats (bullet points, short summaries) that inherently limit content coverage.*
96
+ *² Tests whether the model deviates from its default prose style when users request specific formats.*
97
+
98
+ ### Grounding Metrics
99
+
100
+ In addition to LLM-as-a-Judge evaluation, we validate grounding quality using structural checks:
101
+
102
+ 1. **Tag Uniqueness**: All referenced tags in `xml_tags` must be unique
103
+ 2. **Tag Validity**: All referenced tags must exist in the input text
104
+ 3. **Tag Usage**: All tags in `xml_tags` must appear in the summary
105
+
106
+ ### Elluminate
107
+
108
+ The evaluation was performed using [Elluminate](https://elluminate.de), a collaborative evaluation platform for enterprise AI. Elluminate provides structured LLM-as-a-Judge workflows that enable teams to standardize quality metrics and systematically measure AI performance across defined criteria.
109
+
110
+ This model is a contribution by [ellamind](https://ellamind.ai) to the open-source community.
111
+
112
+ ---
113
+
114
+ ## Model Weights
115
+
116
+ We provide two variants:
117
+
118
+ | Variant | Description | Link |
119
+ |---------|-------------|------|
120
+ | **bfloat16** | Full precision (~48GB weights) | [ellamind/sui-1-24b](https://huggingface.co/ellamind/sui-1-24b) |
121
+ | **FP8** | Quantized (~24GB weights), lower VRAM | [ellamind/sui-1-24b-fp8](https://huggingface.co/ellamind/sui-1-24b-fp8) |
122
+
123
+ The FP8 version preserves high quality, scoring 81.05% overall on our benchmark—nearly identical to bfloat16.
124
+
125
+ ---
126
+
127
+ ## Hardware Requirements
128
+
129
+ We tested various GPU configurations using vLLM. The tables below show minimum requirements for different context lengths.
130
+
131
+ The bfloat16 variant requires **~55GB VRAM** for 8k context, scaling to **~76GB** for 128k. The FP8 variant requires **~38GB** for 8k and **~50GB** for 128k.
132
+
133
+ ### bfloat16 (Full Precision)
134
+
135
+ | Setup | 8k | 32k | 64k | 128k |
136
+ |-------|:--:|:---:|:---:|:----:|
137
+ | 1× A100 80GB / H100 96GB| ✓ | ✓ | ✓ | ✓ |
138
+ | 2× RTX 5090 (32GB) | ✓ | ✓ | ✗ | ✗ |
139
+ | 2× A100 40GB / A6000 | ✓ | ✓ | ✓ | ✓ |
140
+ | 4× RTX 4090 (24GB) | ✓ | ✓ | ✓ | ✓ |
141
+
142
+ ### FP8 Quantized (Recommended for Consumer GPUs)
143
+
144
+ | Setup | 8k | 32k | 64k | 128k |
145
+ |-------|:--:|:---:|:---:|:----:|
146
+ | 1× A100 40GB | ✓ | ✓ | ✗ | ✗ |
147
+ | 1× A6000 (48GB) | ✓ | ✓ | ✓ | ✗ |
148
+ | 1× A100 80GB / H100 96GB| ✓ | ✓ | ✓ | ✓ |
149
+ | 2× RTX 4090 (24GB) | ✓ | ✓ | ✓ | ✗ |
150
+ | 2× RTX 5090 (32GB) | ✓ | ✓ | ✓ | ✓ |
151
+ | 4× RTX 4090 (24GB) | ✓ | ✓ | ✓ | ✓ |
152
+
153
+ > **Tip**: The model supports both one-shot summarization (full document in context) and an iterative two-step approach for very long documents (see [Handling Very Long Contexts](#handling-very-long-contexts)). The 8k context configuration is sufficient to produce high-quality summaries using the iterative approach, making the model accessible on more modest hardware.
154
+
155
+ ---
156
+
157
+ ## How It Works
158
+
159
+ The model follows a three-phase approach:
160
+
161
+ 1. **Planning Phase**: Analyzes the input and plans the summary structure
162
+ 2. **Reference Selection**: Identifies the most important sentences to cite
163
+ 3. **Grounded Generation**: Produces a summary with inline citations to source sentences
164
+
165
+ Citations use XML tags assigned during preprocessing, enabling deterministic verification of each claim.
166
+
167
+ ---
168
+
169
+ ## Input Format
170
+
171
+ The input text must be preprocessed with XML sentence tags. Each sentence is wrapped in a unique 8-character hexadecimal tag:
172
+
173
+ ```
174
+ <a1b2c3d4>First sentence of the document.</a1b2c3d4><e5f6g7h8>Second sentence continues here.</e5f6g7h8>...
175
+ ```
176
+
177
+ ### Tag Format Requirements
178
+
179
+ - Tags must be **8 lowercase hexadecimal characters** (e.g., `a1b2c3d4`)
180
+ - Each tag must be **unique** within the document
181
+ - Tags wrap individual sentences: `<tag>sentence text</tag>`
182
+ - Tags should be contiguous (no whitespace between closing and opening tags)
183
+
184
+ ### Preprocessing with spaCy (Recommended)
185
+
186
+ ```python
187
+ import hashlib
188
+ import spacy
189
+
190
+ def generate_tag(index: int, sentence: str) -> str:
191
+ """Generate unique 8-char hex tag from sentence."""
192
+ return hashlib.md5(f"{index}_{sentence[:50]}".encode()).hexdigest()[:8]
193
+
194
+ def tag_text(text: str, language: str = "en") -> tuple[str, dict]:
195
+ """
196
+ Tag text with XML sentence markers.
197
+
198
+ Args:
199
+ text: Input text to tag
200
+ language: Language code (en, de, es, fr, it)
201
+
202
+ Returns:
203
+ tuple: (tagged_text, tag_to_sentence_mapping)
204
+ """
205
+ # Load appropriate spaCy model
206
+ models = {"en": "en_core_web_sm", "de": "de_core_news_sm",
207
+ "es": "es_core_news_sm", "fr": "fr_core_news_sm", "it": "it_core_news_sm"}
208
+ nlp = spacy.load(models.get(language, "en_core_web_sm"))
209
+
210
+ doc = nlp(text)
211
+ tagged_text = ""
212
+ tag_mapping = {}
213
+
214
+ for i, sent in enumerate(doc.sents):
215
+ sentence = sent.text.strip()
216
+ if sentence:
217
+ tag = generate_tag(i, sentence)
218
+ tag_mapping[tag] = sentence
219
+ tagged_text += f"<{tag}>{sentence}</{tag}>"
220
+
221
+ return tagged_text, tag_mapping
222
+
223
+ # Example usage
224
+ text = "This is the first sentence. Here is the second one. And a third."
225
+ tagged, mapping = tag_text(text)
226
+ print(tagged)
227
+ # Output: <a1b2c3d4>This is the first sentence.</a1b2c3d4><e5f67890>Here is the second one.</e5f67890>...
228
+ ```
229
+
230
+ ### Installation for Preprocessing
231
+
232
+ ```bash
233
+ pip install spacy langdetect
234
+ python -m spacy download en_core_web_sm # English
235
+ python -m spacy download de_core_news_sm # German (optional)
236
+ ```
237
+
238
+ To automatically detect the input language and select the appropriate spaCy model, you can use `langdetect`:
239
+
240
+ ```python
241
+ from langdetect import detect
242
+
243
+ def detect_language(text: str) -> str:
244
+ lang_code = detect(text[:1000]) # Sample first 1000 chars
245
+ lang_map = {"de": "German", "en": "English", "es": "Spanish",
246
+ "fr": "French", "it": "Italian"}
247
+ return lang_map.get(lang_code, "English") # Default to English
248
+ ```
249
+
250
+ For non-enhanced languages, English-style sentence segmentation is used as fallback, which may be suboptimal for languages with different punctuation conventions (e.g., Chinese, Japanese).
251
+
252
+ ---
253
+
254
+ ## Output Format
255
+
256
+ The model outputs a JSON object with three keys:
257
+
258
+ ```json
259
+ {
260
+ "structure": "Planning text describing how the summary will be organized...",
261
+ "xml_tags": ["<a1b2c3d4>", "<e5f67890>", "<12345678>"],
262
+ "summary": "The document discusses... [<a1b2c3d4>]. Furthermore... [<e5f67890>]."
263
+ }
264
+ ```
265
+
266
+ ### Output Keys
267
+
268
+ | Key | Type | Description |
269
+ |-----|------|-------------|
270
+ | `structure` | string | Internal reasoning about the content and planned summary structure. Shows which topics will be covered and in what order. |
271
+ | `xml_tags` | array | List of XML tags that will be cited in the summary. Each tag corresponds to a source sentence. Tags are listed in the order they appear in the source. |
272
+ | `summary` | string | The actual summary with inline XML tag references in square brackets (e.g., `[<a1b2c3d4>]`). Each referenced tag must appear in the `xml_tags` list and exist in the input. |
273
+
274
+ ### Reference Format in Summary
275
+
276
+ Citations appear inline immediately after the statement they support:
277
+
278
+ ```
279
+ The report found significant budget overruns [<a1b2c3d4>]. This was attributed to
280
+ supply chain disruptions [<e5f67890>] and labor shortages [<12345678>].
281
+ ```
282
+
283
+ <details>
284
+ <summary><b>Example Output (click to expand)</b></summary>
285
+
286
+ ```json
287
+ {
288
+ "structure": "This government report examines infrastructure spending over the past fiscal year. I will structure the summary as follows: 1) Executive overview with key findings, 2) Budget analysis section covering the main expenditures, 3) Regional breakdown of infrastructure projects, 4) Recommendations from the oversight committee. I will select 8 XML tags distributed across these sections to provide comprehensive source grounding.",
289
+ "xml_tags": [
290
+ "<9f2a1b3c>",
291
+ "<d4e5f678>",
292
+ "<90ab12cd>",
293
+ "<ef345678>",
294
+ "<9abc0def>",
295
+ "<12345678>",
296
+ "<abcdef90>",
297
+ "<78901234>"
298
+ ],
299
+ "summary": "The Congressional Budget Office's annual infrastructure report reveals that federal infrastructure spending reached $284 billion in fiscal year 2024, representing a 12% increase from the previous year [<9f2a1b3c>]. The majority of funds were allocated to transportation projects, with highway maintenance and bridge repair accounting for 45% of total expenditures [<d4e5f678>].\n\nThe report identifies significant regional disparities in infrastructure investment. Northeastern states received disproportionately higher funding due to aging infrastructure requiring immediate attention [<90ab12cd>], while southern and western regions saw increased allocations for new construction projects [<ef345678>]. Rural communities continued to face challenges in securing adequate funding for basic infrastructure needs [<9abc0def>].\n\nThe oversight committee raised concerns about cost overruns in several major projects, with the average project exceeding initial estimates by 23% [<12345678>]. The committee recommends implementing stricter oversight mechanisms and requiring more detailed feasibility studies before project approval [<abcdef90>]. Additionally, the report suggests exploring public-private partnerships as a means to supplement federal funding and improve project efficiency [<78901234>]."
300
+ }
301
+ ```
302
+
303
+ </details>
304
+
305
+ ---
306
+
307
+ ## Handling Very Long Contexts
308
+
309
+ The model supports a 128k token context window natively. For longer documents (tested up to 2 million tokens), use the iterative approach.
310
+
311
+ ### Approach 1: Oneshot (Up to 128k tokens)
312
+
313
+ For documents within the context limit, use the standard prompt with `PROMPT_SUMMARY`:
314
+
315
+ ```python
316
+ prompt = f"""You are a professional summarizer, following all given instructions with the utmost care.
317
+
318
+ <text>
319
+ {tagged_text}
320
+ </text>
321
+
322
+ # Output Format
323
+ The output must be in JSON format with the following structure:
324
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
325
+ 2. An "xml_tags" list containing objects with:
326
+ - "xml_tag": The XML tag identifier from the tagged text (e.g., "<a1b2c3d4>")
327
+ 3. A "summary" string containing the actual summary with inline XML tag references
328
+
329
+ # Instructions
330
+ ...
331
+
332
+ Parameters:
333
+ - Word count (excl. XML tags): {word_count}
334
+ - Number of XML tags: {number_of_xml_tags}
335
+ - Language: {language}
336
+ """
337
+ ```
338
+
339
+ **Output:** JSON with `structure`, `xml_tags`, and `summary`
340
+
341
+ ---
342
+
343
+ ### Approach 2: Iterative (128k+ tokens)
344
+
345
+ For documents exceeding the context limit, use a two-step iterative approach that preserves grounding quality:
346
+
347
+ #### Step 1: Partial Summaries (`PROMPT_SUMMARY_PARTIAL`)
348
+
349
+ Split the document into chunks and summarize each independently:
350
+
351
+ ```python
352
+ prompt_partial = f"""You are a professional summarizer, following all given instructions with the utmost care.
353
+
354
+ This is a section of a larger document. Create a partial summary that will later be combined with other sections.
355
+
356
+ <text>
357
+ {chunk_tagged_text}
358
+ </text>
359
+
360
+ # Output Format
361
+ The output must be in JSON format with the following structure:
362
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
363
+ 2. An "xml_tags" list containing objects with:
364
+ - "xml_tag": The XML tag identifier from the tagged text (e.g., "<a1b2c3d4>")
365
+ 3. A "summary" string containing the actual summary with inline XML tag references
366
+
367
+ # Instructions
368
+ 1. Select {number_of_xml_tags} XML tags that capture the most significant data and facts.
369
+ 2. Begin with a brief introduction of the section's main topics (no executive summary for partial summaries).
370
+ 3. Structure the summary in coherent paragraphs with at least one XML tag reference each.
371
+ 4. The summary should be 300-600 words long (without the XML tags).
372
+ 5. Only include title/author if explicitly mentioned in this section.
373
+ ...
374
+ """
375
+ ```
376
+
377
+ **Output per chunk:** JSON with `structure`, `xml_tags`, and `summary` (300-600 words each)
378
+
379
+ #### Step 2: Final Merge (`PROMPT_SUMMARY_PARTIAL_LAST`)
380
+
381
+ Combine all partial summaries into a coherent final summary:
382
+
383
+ ```python
384
+ # Concatenate all partial summary outputs
385
+ partial_summaries_text = "\n\n".join([
386
+ f"--- Section {i+1} ---\n{partial_output}"
387
+ for i, partial_output in enumerate(partial_outputs)
388
+ ])
389
+
390
+ prompt_final = f"""You are a professional summarizer, following all given instructions with the utmost care.
391
+
392
+ You are given partial summaries from a larger document. Combine them into a coherent final summary.
393
+
394
+ <partial_summaries>
395
+ {partial_summaries_text}
396
+ </partial_summaries>
397
+
398
+ # Output Format
399
+ The output must be in JSON format with the following structure:
400
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
401
+ 2. An "xml_tags" list containing objects with:
402
+ - "xml_tag": The XML tag identifier from the tagged text (e.g., "<a1b2c3d4>")
403
+ 3. A "summary" string containing the actual summary with inline XML tag references
404
+
405
+ # Instructions
406
+ 1. Select the {number_of_xml_tags} most significant XML tags from the partial summaries.
407
+ Copy the XML tags verbatim, ensuring they represent key points from different sections.
408
+ 2. Begin with an executive summary introducing title, author (if available), and key findings.
409
+ 3. Structure the summary in coherent paragraphs following a coherent thread.
410
+ 4. Each XML tag must appear exactly once. Use only XML tags from the partial summaries.
411
+ 5. Don't repeat content that is very similar or identical in multiple partial summaries.
412
+ ...
413
+ """
414
+ ```
415
+
416
+ **Final Output:** JSON with `structure`, `xml_tags`, and `summary`
417
+
418
+ #### How Grounding Quality is Maintained
419
+
420
+ The iterative approach preserves source grounding through careful XML tag propagation:
421
+
422
+ 1. **Tag Extraction**: Each partial summary extracts XML tags from its chunk, linking claims to source sentences
423
+ 2. **Tag Preservation**: The final merge prompt explicitly instructs to "copy XML tags verbatim" from partials
424
+ 3. **No Hallucinated Tags**: The final summary can only reference tags that were already validated in partial summaries
425
+ 4. **Distributed Coverage**: By selecting tags "from different sections," the final summary maintains broad source coverage
426
+
427
+ This ensures that even for 2M+ token documents, every claim in the final summary traces back to a specific source sentence.
428
+
429
+ ### Recommended Parameters
430
+
431
+ | Summary Length | Word Count | XML Tags |
432
+ |---------------|------------|----------|
433
+ | Short | ~100 words | 3 tags |
434
+ | Medium | ~250 words | 6 tags |
435
+ | Long | ~500 words | 12 tags |
436
+
437
+ ---
438
+
439
+ ## Usage
440
+
441
+ For production use, we provide ready-to-use prompt templates in [`prompts.py`](prompts.py). This file contains:
442
+
443
+ - `PROMPT_SUMMARY`: Standard single-pass summarization prompt
444
+ - `PROMPT_SUMMARY_PARTIAL`: Prompt for creating partial summaries of document chunks
445
+ - `PROMPT_SUMMARY_PARTIAL_LAST`: Prompt for merging partial summaries into a final summary
446
+
447
+ **Resource-constrained environments:** The iterative two-step approach is not only useful for very long documents—it also enables deployment on hardware with limited VRAM. The model was trained on a broad range of chunk sizes, so partial summaries work reliably even with smaller context windows (e.g., 5k token chunks). This flexibility allows you to adjust chunk sizes to match your available GPU memory.
448
+
449
+ ### With vLLM (Recommended for Production)
450
+
451
+ ```python
452
+ from vllm import LLM, SamplingParams
453
+
454
+ # Load model
455
+ llm = LLM(
456
+ model="ellamind/sui-1-24b-fp8",
457
+ tensor_parallel_size=4, # Adjust based on available GPUs
458
+ dtype="bfloat16",
459
+ tokenizer_mode="mistral",
460
+ max_model_len=128000,
461
+ trust_remote_code=True,
462
+ )
463
+
464
+ # Prepare prompt
465
+ prompt = f"""You are a professional summarizer, following all given instructions with the utmost care.
466
+
467
+ <text>
468
+ {tagged_text}
469
+ </text>
470
+
471
+ # Output Format
472
+ The output must be in JSON format with the following structure:
473
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
474
+ 2. An "xml_tags" list containing the XML tag identifiers from the tagged text (e.g., "<a1b2c3d4>")
475
+ 3. A "summary" string containing the actual summary with inline XML tag references
476
+
477
+ # Instructions
478
+ 1. Start by thinking about and explaining the structure and content of your summary. Select {num_tags} XML tags from the tagged text that capture the most significant data and facts.
479
+ 2. Begin with an executive summary introducing the title, author (if available), and key findings.
480
+ 3. Structure the summary in coherent paragraphs. Every paragraph should contain at least one XML tag reference.
481
+ 4. Reference XML tags inline in square brackets (e.g., [<a1b2c3d4>]) immediately after the statement they support.
482
+ 5. Each XML tag must appear exactly once in the summary.
483
+ 6. Avoid a concluding paragraph that merely restates points.
484
+ 7. Do not use bullet points or headings unless explicitly requested.
485
+
486
+ # Custom Instruction
487
+ {custom_instruction}
488
+
489
+ Parameters:
490
+ - Word count (excl. XML tags): {word_count}
491
+ - Number of XML tags: {num_tags}
492
+ - Language: {language}
493
+ """
494
+
495
+ # Generate
496
+ sampling_params = SamplingParams(max_tokens=8192, temperature=0.0)
497
+ outputs = llm.chat([[{"role": "user", "content": prompt}]], sampling_params)
498
+ result = outputs[0].outputs[0].text
499
+ ```
500
+
501
+ ### With Transformers
502
+
503
+ ```python
504
+ from transformers import AutoModelForCausalLM, AutoTokenizer
505
+ import torch
506
+
507
+ model = AutoModelForCausalLM.from_pretrained(
508
+ "ellamind/sui-1-24b",
509
+ torch_dtype=torch.bfloat16,
510
+ device_map="auto",
511
+ trust_remote_code=True,
512
+ )
513
+ tokenizer = AutoTokenizer.from_pretrained("ellamind/sui-1-24b")
514
+
515
+ messages = [{"role": "user", "content": prompt}]
516
+ inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
517
+ outputs = model.generate(inputs, max_new_tokens=8192, temperature=0.0, do_sample=False)
518
+ result = tokenizer.decode(outputs[0], skip_special_tokens=True)
519
+ ```
520
+
521
+ ---
522
+
523
+ ## Language Support
524
+
525
+ ### Enhanced Languages (Fine-tuned)
526
+
527
+ The model was fine-tuned with training data in these languages, providing optimal summarization quality:
528
+
529
+ | Language | Code | Tagging Support |
530
+ |----------|------|-----------------|
531
+ | English | `en` | `en_core_web_sm` |
532
+ | German | `de` | `de_core_news_sm` |
533
+ | Spanish | `es` | `es_core_news_sm` |
534
+ | French | `fr` | `fr_core_news_sm` |
535
+ | Italian | `it` | `it_core_news_sm` |
536
+
537
+ ### Inherited Languages (Base Model)
538
+
539
+ The following languages are supported through the Mistral Small 3.2 base model. Summarization works but may have reduced quality compared to enhanced languages:
540
+
541
+ | Category | Languages |
542
+ |----------|-----------|
543
+ | European | Portuguese, Dutch, Polish, Russian, Swedish, Ukrainian, Romanian, Czech, Greek, Hungarian |
544
+ | Asian | Chinese, Japanese, Korean, Vietnamese, Indonesian, Thai, Hindi |
545
+ | Middle Eastern | Arabic, Turkish, Persian |
546
+
547
+ ---
548
+
549
+ ## Limitations
550
+
551
+ - Requires preprocessing of input text with XML tags
552
+ - Maximum single-pass context of 128k tokens
553
+ - JSON output parsing may occasionally fail; implement retry logic for production use
554
+
555
+ ---
556
+
557
+ ## Citation
558
+
559
+ ```bibtex
560
+ @article{droste2025sui1,
561
+ title={sui-1: Grounded and Verifiable Long-Form Summarization},
562
+ author={Droste, Benedikt and Harries, Jan Philipp and Idahl, Maximilian and Pl{\"u}ster, Bj{\"o}rn},
563
+ journal={arXiv preprint arXiv:2601.08472},
564
+ year={2025}
565
+ }
566
+ ```
567
+
568
+ ---
569
+
570
+ ## License
571
+
572
+ This model is released under the Apache 2.0 license, consistent with the base Mistral model.
config.json ADDED
@@ -0,0 +1,96 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Mistral3ForConditionalGeneration"
4
+ ],
5
+ "dtype": "bfloat16",
6
+ "image_token_index": 10,
7
+ "model_type": "mistral3",
8
+ "multimodal_projector_bias": false,
9
+ "projector_hidden_act": "gelu",
10
+ "quantization_config": {
11
+ "config_groups": {
12
+ "group_0": {
13
+ "format": "float-quantized",
14
+ "input_activations": {
15
+ "actorder": null,
16
+ "block_structure": null,
17
+ "dynamic": true,
18
+ "group_size": null,
19
+ "num_bits": 8,
20
+ "observer": null,
21
+ "observer_kwargs": {},
22
+ "strategy": "token",
23
+ "symmetric": true,
24
+ "type": "float"
25
+ },
26
+ "output_activations": null,
27
+ "targets": [
28
+ "Linear"
29
+ ],
30
+ "weights": {
31
+ "actorder": null,
32
+ "block_structure": null,
33
+ "dynamic": false,
34
+ "group_size": null,
35
+ "num_bits": 8,
36
+ "observer": "minmax",
37
+ "observer_kwargs": {},
38
+ "strategy": "channel",
39
+ "symmetric": true,
40
+ "type": "float"
41
+ }
42
+ }
43
+ },
44
+ "format": "float-quantized",
45
+ "global_compression_ratio": null,
46
+ "ignore": [
47
+ "lm_head"
48
+ ],
49
+ "kv_cache_scheme": null,
50
+ "quant_method": "compressed-tensors",
51
+ "quantization_status": "compressed",
52
+ "sparsity_config": {},
53
+ "transform_config": {},
54
+ "version": "0.11.0"
55
+ },
56
+ "spatial_merge_size": 2,
57
+ "text_config": {
58
+ "attention_dropout": 0.0,
59
+ "dtype": "bfloat16",
60
+ "head_dim": 128,
61
+ "hidden_act": "silu",
62
+ "hidden_size": 5120,
63
+ "initializer_range": 0.02,
64
+ "intermediate_size": 32768,
65
+ "max_position_embeddings": 131072,
66
+ "model_type": "mistral",
67
+ "num_attention_heads": 32,
68
+ "num_hidden_layers": 40,
69
+ "num_key_value_heads": 8,
70
+ "rms_norm_eps": 1e-05,
71
+ "rope_theta": 1000000000.0,
72
+ "sliding_window": null,
73
+ "use_cache": true,
74
+ "vocab_size": 131072
75
+ },
76
+ "tie_word_embeddings": false,
77
+ "torch_dtype": "bfloat16",
78
+ "transformers_version": "4.55.2",
79
+ "vision_config": {
80
+ "attention_dropout": 0.0,
81
+ "dtype": "bfloat16",
82
+ "head_dim": 64,
83
+ "hidden_act": "silu",
84
+ "hidden_size": 1024,
85
+ "image_size": 1540,
86
+ "initializer_range": 0.02,
87
+ "intermediate_size": 4096,
88
+ "model_type": "pixtral",
89
+ "num_attention_heads": 16,
90
+ "num_channels": 3,
91
+ "num_hidden_layers": 24,
92
+ "patch_size": 14,
93
+ "rope_theta": 10000.0
94
+ },
95
+ "vision_feature_layer": -1
96
+ }
example.py ADDED
@@ -0,0 +1,203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env -S uv run --script
2
+ # /// script
3
+ # requires-python = ">=3.12"
4
+ # dependencies = [
5
+ # "vllm>=0.11.0",
6
+ # "spacy>=3.7.0",
7
+ # "mistral_common>=1.5.0",
8
+ # "en-core-web-sm @ https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl",
9
+ # ]
10
+ #
11
+ # [tool.uv]
12
+ # no-build = true
13
+ # index-strategy = "unsafe-best-match"
14
+ # extra-index-url = ["https://download.pytorch.org/whl/cu128"]
15
+ # ///
16
+ """
17
+ Minimal end-to-end example for sui-1-24b summarization.
18
+
19
+ Usage:
20
+ # Summarize a file
21
+ uv run example.py document.txt
22
+
23
+ # Summarize inline text
24
+ uv run example.py --text "Your long text here..."
25
+
26
+ # With custom parameters
27
+ uv run example.py document.txt --words 300 --tags 8 --language en
28
+ """
29
+
30
+ import argparse
31
+ import hashlib
32
+ import json
33
+ import re
34
+ import sys
35
+ from pathlib import Path
36
+
37
+ # Lazy imports for faster --help
38
+ def main():
39
+ parser = argparse.ArgumentParser(
40
+ description="Summarize text using sui-1-24b with source grounding",
41
+ formatter_class=argparse.RawDescriptionHelpFormatter,
42
+ epilog=__doc__,
43
+ )
44
+ parser.add_argument("input", nargs="?", help="Input file path (or use --text)")
45
+ parser.add_argument("--text", "-t", help="Input text directly")
46
+ parser.add_argument("--words", "-w", type=int, default=250, help="Target word count (default: 400)")
47
+ parser.add_argument("--tags", "-n", type=int, default=4, help="Number of XML tags to cite (default: 10)")
48
+ parser.add_argument("--language", "-l", default="en", choices=["en", "de", "es", "fr", "it"], help="Language (default: en)")
49
+ parser.add_argument("--model", "-m", default="ellamind/sui-1-24b", help="Model path or HF repo")
50
+ parser.add_argument("--tensor-parallel", "-tp", type=int, default=1, help="Tensor parallel size (default: 1)")
51
+ parser.add_argument("--raw", action="store_true", help="Print raw JSON output instead of formatted")
52
+ args = parser.parse_args()
53
+
54
+ # Get input text
55
+ if args.text:
56
+ text = args.text
57
+ elif args.input:
58
+ text = Path(args.input).read_text()
59
+ else:
60
+ parser.error("Provide input file or --text")
61
+
62
+ # Import heavy dependencies only when needed
63
+ import spacy
64
+ from vllm import LLM, SamplingParams
65
+
66
+ # Load spaCy model for sentence segmentation
67
+ # Note: Only English is bundled. For other languages, install the model first:
68
+ # pip install https://github.com/explosion/spacy-models/releases/download/de_core_news_sm-3.8.0/de_core_news_sm-3.8.0-py3-none-any.whl
69
+ spacy_models = {
70
+ "en": "en_core_web_sm",
71
+ "de": "de_core_news_sm",
72
+ "es": "es_core_news_sm",
73
+ "fr": "fr_core_news_sm",
74
+ "it": "it_core_news_sm",
75
+ }
76
+ try:
77
+ nlp = spacy.load(spacy_models[args.language])
78
+ except OSError:
79
+ print(f"Error: spaCy model '{spacy_models[args.language]}' not found.")
80
+ print(f"For English, this should be bundled automatically.")
81
+ print(f"For other languages, install the model first:")
82
+ print(f" pip install https://github.com/explosion/spacy-models/releases/download/{spacy_models[args.language]}-3.8.0/{spacy_models[args.language]}-3.8.0-py3-none-any.whl")
83
+ sys.exit(1)
84
+
85
+ # Tag sentences with unique XML identifiers
86
+ print("Tagging sentences...")
87
+ doc = nlp(text)
88
+ tagged_text = ""
89
+ tag_mapping = {}
90
+
91
+ for i, sent in enumerate(doc.sents):
92
+ sentence = sent.text.strip()
93
+ if sentence:
94
+ tag = hashlib.md5(f"{i}_{sentence[:50]}".encode()).hexdigest()[:8]
95
+ tag_mapping[tag] = sentence
96
+ tagged_text += f"<{tag}>{sentence}</{tag}>"
97
+
98
+ print(f"Tagged {len(tag_mapping)} sentences")
99
+
100
+ # Build prompt
101
+ language_names = {"en": "English", "de": "German", "es": "Spanish", "fr": "French", "it": "Italian"}
102
+ prompt = f"""You are a professional summarizer, following all given instructions with the utmost care.
103
+
104
+ <text>
105
+ {tagged_text}
106
+ </text>
107
+
108
+ # Output Format
109
+ The output must be in JSON format with the following structure:
110
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
111
+ 2. An "xml_tags" list containing the XML tag identifiers from the tagged text (e.g., "<a1b2c3d4>")
112
+ 3. A "summary" string containing the actual summary with inline XML tag references
113
+
114
+ # Instructions
115
+ 1. Start by thinking about and explaining the structure and content of your summary. Select {args.tags} XML tags from the tagged text that capture the most significant data and facts.
116
+ 2. Begin with an executive summary introducing the title, author (if available), and key findings.
117
+ 3. Structure the summary in coherent paragraphs. Every paragraph should contain at least one XML tag reference.
118
+ 4. Reference XML tags inline in square brackets (e.g., [<a1b2c3d4>]) immediately after the statement they support.
119
+ 5. Each XML tag must appear exactly once in the summary.
120
+ 6. Avoid a concluding paragraph that merely restates points.
121
+ 7. Do not use bullet points or headings unless explicitly requested.
122
+
123
+ Parameters:
124
+ - Word count (excl. XML tags): {args.words}
125
+ - Number of XML tags: {args.tags}
126
+ - Language: {language_names[args.language]}
127
+ """
128
+
129
+ # Load model and generate
130
+ print(f"Loading model: {args.model}")
131
+ llm = LLM(
132
+ model=args.model,
133
+ tensor_parallel_size=args.tensor_parallel,
134
+ dtype="bfloat16",
135
+ tokenizer_mode="mistral",
136
+ trust_remote_code=True,
137
+ limit_mm_per_prompt={"image": 0}, # Disable vision encoder for text-only
138
+ )
139
+
140
+ print("Generating summary...")
141
+ sampling_params = SamplingParams(max_tokens=4096, temperature=0.0)
142
+ outputs = llm.chat([[{"role": "user", "content": prompt}]], sampling_params)
143
+ result = outputs[0].outputs[0].text
144
+
145
+ # Parse and display output
146
+ if args.raw:
147
+ print(result)
148
+ return
149
+
150
+ try:
151
+ # Extract JSON from response
152
+ json_match = re.search(r'\{[\s\S]*\}', result)
153
+ if json_match:
154
+ data = json.loads(json_match.group())
155
+
156
+ print("\n" + "=" * 60)
157
+ print("SUMMARY")
158
+ print("=" * 60 + "\n")
159
+
160
+ summary = data.get("summary", "")
161
+
162
+ # Replace XML tags with highlighted source references
163
+ def replace_tag(match):
164
+ tag = match.group(1)
165
+ source = tag_mapping.get(tag, "???")
166
+ # Truncate long sources
167
+ if len(source) > 80:
168
+ source = source[:77] + "..."
169
+ return f"[{tag}]"
170
+
171
+ clean_summary = re.sub(r'\[<([a-f0-9]{8})>\]', replace_tag, summary)
172
+ print(clean_summary)
173
+
174
+ print("\n" + "-" * 60)
175
+ print("SOURCES")
176
+ print("-" * 60)
177
+
178
+ # Show referenced sources
179
+ # Handle both formats: ["<tag>"] or [{"xml_tag": "<tag>"}]
180
+ xml_tags = data.get("xml_tags", [])
181
+ for tag in xml_tags:
182
+ if isinstance(tag, str):
183
+ clean_tag = tag.strip("<>")
184
+ elif isinstance(tag, dict) and "xml_tag" in tag:
185
+ clean_tag = tag["xml_tag"].strip("<>")
186
+ else:
187
+ continue
188
+ source = tag_mapping.get(clean_tag, "Not found")
189
+ if len(source) > 100:
190
+ source = source[:97] + "..."
191
+ print(f"[{clean_tag}] {source}")
192
+
193
+ else:
194
+ print("Could not parse JSON response:")
195
+ print(result)
196
+
197
+ except json.JSONDecodeError as e:
198
+ print(f"JSON parse error: {e}")
199
+ print(result)
200
+
201
+
202
+ if __name__ == "__main__":
203
+ main()
generation_config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 1,
4
+ "do_sample": true,
5
+ "eos_token_id": 2,
6
+ "temperature": 0.15,
7
+ "transformers_version": "4.55.2"
8
+ }
images/criteria_breakdown.png ADDED
images/overall_performance.png ADDED
images/sui_icon.jpeg ADDED

Git LFS Details

  • SHA256: 94fe6b6803e458880fe305c9aac977fd139c29c7c1f80b6f24cc8cc4e80c67b6
  • Pointer size: 132 Bytes
  • Size of remote file: 1.55 MB
model-00001-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83bb1e8affb27de2c7a2fd5455352a6f53567ea2442dfca313921eb11dc9aedc
3
+ size 4950292144
model-00002-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b096aa10c2af04426e7a0f3d7b469cbecf26c005d1e15497cf356cf7dc53102d
3
+ size 4835547160
model-00003-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e6095515abb71b4429064b21dfeca791f2324bdda4eebf70ccaec57b35b8b19f
3
+ size 4835547224
model-00004-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:075d83d6e06fada90d24004ae33b3f0fdf8225560e077d86692adc7a2b8b1ed6
3
+ size 4982403168
model-00005-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:34e61b252e1262ef6d220d34481cddc12c6a3cbdcaa52eddb2b1dec7e805af7f
3
+ size 4415993472
model-00006-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5b72834082387ce74af5414bf353165eab35661f3261f7c9c9f16da505b64e93
3
+ size 1342177424
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
prompts.py ADDED
@@ -0,0 +1,253 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ Ready-to-use prompt templates for sui-1-24b summarization model.
3
+
4
+ Usage:
5
+ from prompts import format_prompt, format_partial_prompt, format_merge_prompt
6
+
7
+ # Single-pass summarization
8
+ prompt = format_prompt(
9
+ text=tagged_text,
10
+ word_count=500,
11
+ number_of_xml_tags=10,
12
+ language="English"
13
+ )
14
+
15
+ # Iterative approach for long documents
16
+ partial_prompt = format_partial_prompt(text=chunk, ...)
17
+ merge_prompt = format_merge_prompt(text=partial_summaries, ...)
18
+ """
19
+
20
+ # =============================================================================
21
+ # Single-pass summarization prompt
22
+ # =============================================================================
23
+
24
+ PROMPT_SUMMARY = """You are a professional summarizer, following all given instructions with the utmost care.
25
+
26
+ <text>
27
+ {text}
28
+ </text>
29
+
30
+ # Output Format
31
+ The output must be in JSON format with the following structure:
32
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
33
+ 2. An "xml_tags" list containing the XML tag identifiers from the tagged text (e.g., "<a1b2c3d4>")
34
+ 3. A "summary" string containing the actual summary with inline XML tag references
35
+
36
+ # Instructions
37
+ 1. Start by thinking about and explaining the structure and content of your summary. Select {number_of_xml_tags} XML tags from the tagged text that capture the most significant data and facts. Ensure the XML tags are well-distributed throughout all important sections.
38
+ 2. Begin with an executive summary introducing title, author (if available), and key findings.
39
+ 3. Structure the summary in coherent paragraphs. Every paragraph should contain at least one XML tag reference.
40
+ 4. Reference XML tags inline in square brackets (e.g., [<a1b2c3d4>]) immediately after the statement they support.
41
+ 5. Each XML tag must appear exactly once in the summary.
42
+ 6. Avoid a concluding paragraph that merely restates points. Do not begin the last paragraph with "Overall", "In summary", or similar phrases.
43
+ 7. Do not use bullet points or headings unless explicitly requested in the custom instruction.
44
+ 8. If the text lacks meaningful content, return a refusal message.
45
+ {custom_instruction_section}
46
+ Parameters:
47
+ - Word count (excl. XML tags): {word_count}
48
+ - Number of XML tags: {number_of_xml_tags}
49
+ - Language: {language}
50
+ """
51
+
52
+ # =============================================================================
53
+ # Partial summarization prompt (for chunks of long documents)
54
+ # =============================================================================
55
+
56
+ PROMPT_SUMMARY_PARTIAL = """You are a professional summarizer, following all given instructions with the utmost care.
57
+
58
+ This is a section of a larger document. Create a partial summary that will later be combined with other sections.
59
+
60
+ <text>
61
+ {text}
62
+ </text>
63
+
64
+ # Output Format
65
+ The output must be in JSON format with the following structure:
66
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
67
+ 2. An "xml_tags" list containing the XML tag identifiers from the tagged text (e.g., "<a1b2c3d4>")
68
+ 3. A "summary" string containing the actual summary with inline XML tag references
69
+
70
+ # Instructions
71
+ 1. Start by thinking about and explaining the structure and content of your summary. Select {number_of_xml_tags} XML tags from the tagged text that capture the most significant data and facts. Ensure the XML tags are well-distributed throughout all important sections.
72
+ 2. Begin with a brief introduction of the section's main topics (no executive summary for partial summaries).
73
+ 3. Structure the summary in coherent paragraphs. Every paragraph should contain at least one XML tag reference.
74
+ 4. Reference XML tags inline in square brackets (e.g., [<a1b2c3d4>]) immediately after the statement they support.
75
+ 5. Each XML tag must appear exactly once in the summary.
76
+ 6. Avoid a concluding paragraph that merely restates points.
77
+ 7. The summary should be 300-600 words long (without the XML tags).
78
+ 8. Only include title/author if explicitly mentioned in this section.
79
+ {custom_instruction_section}
80
+ Parameters:
81
+ - Word count (excl. XML tags): {word_count}
82
+ - Number of XML tags: {number_of_xml_tags}
83
+ - Language: {language}
84
+ """
85
+
86
+ # =============================================================================
87
+ # Merge prompt (for combining partial summaries)
88
+ # =============================================================================
89
+
90
+ PROMPT_SUMMARY_PARTIAL_LAST = """You are a professional summarizer, following all given instructions with the utmost care.
91
+
92
+ You are given partial summaries from a larger document. Combine them into a coherent final summary.
93
+
94
+ <partial_summaries>
95
+ {text}
96
+ </partial_summaries>
97
+
98
+ # Output Format
99
+ The output must be in JSON format with the following structure:
100
+ 1. A "structure" string containing your thoughts about the content and structure of the summary
101
+ 2. An "xml_tags" list containing the XML tag identifiers from the tagged text (e.g., "<a1b2c3d4>")
102
+ 3. A "summary" string containing the actual summary with inline XML tag references
103
+
104
+ # Instructions
105
+ 1. Start by thinking about and explaining the structure and content of your summary. Select the {number_of_xml_tags} most significant XML tags from the partial summaries. Copy the XML tags verbatim, ensuring they represent key points from different sections.
106
+ 2. Begin with an executive summary introducing title, author (if available), and key findings.
107
+ 3. Structure the summary in coherent paragraphs following a coherent thread. Every paragraph should contain at least one XML tag reference.
108
+ 4. Reference XML tags inline in square brackets (e.g., [<a1b2c3d4>]) immediately after the statement they support.
109
+ 5. Each XML tag must appear exactly once in the summary. Use only XML tags from the partial summaries.
110
+ 6. Avoid a concluding paragraph that merely restates points. Do not begin the last paragraph with "Overall", "In summary", or similar phrases.
111
+ 7. Don't repeat content that is very similar or identical in multiple partial summaries.
112
+ 8. Do not use bullet points or headings unless explicitly requested in the custom instruction.
113
+ {custom_instruction_section}
114
+ Parameters:
115
+ - Word count (excl. XML tags): {word_count}
116
+ - Number of XML tags: {number_of_xml_tags}
117
+ - Language: {language}
118
+ """
119
+
120
+ # =============================================================================
121
+ # Custom instruction template (inserted when custom_instruction is provided)
122
+ # =============================================================================
123
+
124
+ CUSTOM_INSTRUCTION_SECTION = """
125
+ # Custom Instruction
126
+ The user has provided a custom instruction below. It takes priority over default formatting or tone rules.
127
+ However, if the custom instruction is unrelated to summarization (e.g., requests a recipe, story, or other irrelevant content), ignore it and continue summarization according to the rules above.
128
+
129
+ <custom_instruction>{custom_instruction}</custom_instruction>
130
+ """
131
+
132
+
133
+ # =============================================================================
134
+ # Helper functions
135
+ # =============================================================================
136
+
137
+ def format_prompt(
138
+ text: str,
139
+ word_count: int,
140
+ number_of_xml_tags: int,
141
+ language: str = "English",
142
+ custom_instruction: str = ""
143
+ ) -> str:
144
+ """
145
+ Format the single-pass summarization prompt.
146
+
147
+ Args:
148
+ text: XML-tagged input text
149
+ word_count: Target word count for the summary (excluding XML tags)
150
+ number_of_xml_tags: Number of source sentences to cite
151
+ language: Output language (e.g., "English", "German")
152
+ custom_instruction: Optional custom formatting or content instructions
153
+
154
+ Returns:
155
+ Formatted prompt string ready for model input
156
+ """
157
+ custom_section = ""
158
+ if custom_instruction.strip():
159
+ custom_section = CUSTOM_INSTRUCTION_SECTION.format(
160
+ custom_instruction=custom_instruction
161
+ )
162
+
163
+ return PROMPT_SUMMARY.format(
164
+ text=text,
165
+ word_count=word_count,
166
+ number_of_xml_tags=number_of_xml_tags,
167
+ language=language,
168
+ custom_instruction_section=custom_section
169
+ )
170
+
171
+
172
+ def format_partial_prompt(
173
+ text: str,
174
+ word_count: int = 450,
175
+ number_of_xml_tags: int = 8,
176
+ language: str = "English",
177
+ custom_instruction: str = ""
178
+ ) -> str:
179
+ """
180
+ Format the partial summarization prompt for document chunks.
181
+
182
+ Args:
183
+ text: XML-tagged chunk of the document
184
+ word_count: Target word count (default 450, recommended 300-600)
185
+ number_of_xml_tags: Number of source sentences to cite per chunk
186
+ language: Output language
187
+ custom_instruction: Optional custom instructions (format constraints are
188
+ automatically relaxed for partial summaries)
189
+
190
+ Returns:
191
+ Formatted prompt string ready for model input
192
+ """
193
+ custom_section = ""
194
+ if custom_instruction.strip():
195
+ custom_section = CUSTOM_INSTRUCTION_SECTION.format(
196
+ custom_instruction=custom_instruction
197
+ )
198
+
199
+ return PROMPT_SUMMARY_PARTIAL.format(
200
+ text=text,
201
+ word_count=word_count,
202
+ number_of_xml_tags=number_of_xml_tags,
203
+ language=language,
204
+ custom_instruction_section=custom_section
205
+ )
206
+
207
+
208
+ def format_merge_prompt(
209
+ text: str,
210
+ word_count: int,
211
+ number_of_xml_tags: int,
212
+ language: str = "English",
213
+ custom_instruction: str = ""
214
+ ) -> str:
215
+ """
216
+ Format the merge prompt for combining partial summaries.
217
+
218
+ Args:
219
+ text: Concatenated partial summaries (JSON outputs from partial prompts)
220
+ word_count: Target word count for the final summary
221
+ number_of_xml_tags: Number of XML tags to retain in final summary
222
+ language: Output language
223
+ custom_instruction: Optional custom instructions
224
+
225
+ Returns:
226
+ Formatted prompt string ready for model input
227
+
228
+ Example:
229
+ # Combine partial outputs
230
+ partial_text = "\\n\\n".join([
231
+ f"--- Section {i+1} ---\\n{output}"
232
+ for i, output in enumerate(partial_outputs)
233
+ ])
234
+ prompt = format_merge_prompt(
235
+ text=partial_text,
236
+ word_count=800,
237
+ number_of_xml_tags=15,
238
+ language="English"
239
+ )
240
+ """
241
+ custom_section = ""
242
+ if custom_instruction.strip():
243
+ custom_section = CUSTOM_INSTRUCTION_SECTION.format(
244
+ custom_instruction=custom_instruction
245
+ )
246
+
247
+ return PROMPT_SUMMARY_PARTIAL_LAST.format(
248
+ text=text,
249
+ word_count=word_count,
250
+ number_of_xml_tags=number_of_xml_tags,
251
+ language=language,
252
+ custom_instruction_section=custom_section
253
+ )
recipe.yaml ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ QuantizationModifier:
4
+ targets: [Linear]
5
+ ignore: [lm_head, 're:vision_tower.*', 're:multi_modal_projector.*']
6
+ scheme: FP8_DYNAMIC
tekken.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6e2501687ccd0e1f30f36319eaf2b46958b897811e246cd8eb5d385b9e3de7d1
3
+ size 19399895