ai-forever Cursor commited on
Commit
bb7d347
Β·
1 Parent(s): 6d9a5f4

Mark LIBRA Mini datasets with asterisk on leaderboard; refresh docs and LiquidAI LongContext scores

Browse files
datasets_config.json CHANGED
@@ -22,7 +22,7 @@
22
  ]
23
  },
24
  "matreshka_names": {
25
- "name": "MatreshkaNames",
26
  "lengths": [
27
  "4k",
28
  "8k",
@@ -53,7 +53,7 @@
53
  ]
54
  },
55
  "ru_sci_passage_count": {
56
- "name": "ruSciPassageCount",
57
  "lengths": [
58
  "4k",
59
  "8k",
@@ -64,7 +64,7 @@
64
  ]
65
  },
66
  "ru_2wikimultihopqa": {
67
- "name": "ru2WikiMultihopQA",
68
  "lengths": [
69
  "8k",
70
  "16k",
@@ -72,14 +72,13 @@
72
  ]
73
  },
74
  "long_context_multiq": {
75
- "name": "LongContextMultiQ",
76
  "lengths": [
77
  "4k",
78
  "8k",
79
  "16k",
80
  "32k",
81
- "64k",
82
- "128k"
83
  ]
84
  },
85
  "ru_sci_abstract_retrieval": {
@@ -101,7 +100,7 @@
101
  ]
102
  },
103
  "librusec_mhqa": {
104
- "name": "LibrusecMHQA",
105
  "lengths": [
106
  "8k"
107
  ]
@@ -129,7 +128,7 @@
129
  ]
130
  },
131
  "ru_babilong_qa3": {
132
- "name": "ruBABILongQA3",
133
  "lengths": [
134
  "4k",
135
  "8k",
 
22
  ]
23
  },
24
  "matreshka_names": {
25
+ "name": "MatreshkaNames *",
26
  "lengths": [
27
  "4k",
28
  "8k",
 
53
  ]
54
  },
55
  "ru_sci_passage_count": {
56
+ "name": "ruSciPassageCount *",
57
  "lengths": [
58
  "4k",
59
  "8k",
 
64
  ]
65
  },
66
  "ru_2wikimultihopqa": {
67
+ "name": "ru2WikiMultihopQA *",
68
  "lengths": [
69
  "8k",
70
  "16k",
 
72
  ]
73
  },
74
  "long_context_multiq": {
75
+ "name": "LongContextMultiQ *",
76
  "lengths": [
77
  "4k",
78
  "8k",
79
  "16k",
80
  "32k",
81
+ "64k"
 
82
  ]
83
  },
84
  "ru_sci_abstract_retrieval": {
 
100
  ]
101
  },
102
  "librusec_mhqa": {
103
+ "name": "LibrusecMHQA *",
104
  "lengths": [
105
  "8k"
106
  ]
 
128
  ]
129
  },
130
  "ru_babilong_qa3": {
131
+ "name": "ruBABILongQA3 *",
132
  "lengths": [
133
  "4k",
134
  "8k",
docs/description.md CHANGED
@@ -1,71 +1,128 @@
1
  # LIBRA: Long Input Benchmark for Russian Analysis
 
 
 
2
 
3
- <img src="https://i.imgur.com/BNleRrG.png" width="800" />
4
 
5
- ## Dataset Summary
6
 
7
- LIBRA (Long Input Benchmark for Russian Analysis) is designed to evaluate the capabilities of large language models (LLMs) in understanding and processing long texts in Russian. This benchmark includes 21 datasets adapted for different tasks and complexities. The tasks are divided into four complexity groups and allow evaluation across various context lengths ranging from 4k up to 128k tokens.
8
 
9
- ## Tasks and Complexity Groups
10
 
11
- ### Group I: Simple Information Retrieval
12
- - **Passkey**: Extract a relevant piece of code number from a long text fragment. Based on the original [PassKey test](https://github.com/CStanKonrad/long_llama/blob/main/examples/passkey.py) from the m LongLLaMA’s GitHub repo.
13
- - **PasskeyWithLibrusec**: Similar to Passkey but with added noise from Librusec texts.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
15
  ### Group II: Question Answering and Multiple Choice
 
16
  - **MatreshkaNames**: Identify the person in dialogues based on the discussed topic. We used [Matreshka](https://huggingface.co/datasets/zjkarina/matreshka) dataset and [Russian Names](https://www.kaggle.com/datasets/rai220/russian-cyrillic-names-and-sex/data) dataset to create this and the next task.
17
  - **MatreshkaYesNo**: Indicate whether a specific topic was mentioned in the dialog.
18
- - **LibrusecHistory**: Answer questions based on historical texts. Ideologically similiar to the [PassageRetrieval dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/passage_retrieval_en) from LongBench.
19
- - **ruTREC**: Few-shot in-context learning for topic classification. Created by translating the [TREC dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/trec_e) from LongBench.
20
  - **ruSciFi**: Answer true/false based on context and general world knowledge. Translation of [SciFi dataset](https://huggingface.co/datasets/L4NLP/LEval/viewer/sci_f) from L-Eval which originally was based on [SF-Gram](https://github.com/nschaetti/SFGram-dataset).
21
  - **ruSciAbstractRetrieval**: Retrieve relevant paragraphs from scientific abstracts.
22
  - **ruTPO**: Multiple-choice questions similar to TOEFL exams. Translation of the [TPO dataset](https://huggingface.co/datasets/L4NLP/LEval/viewer/tpo) from L-Eval.
23
  - **ruQuALITY**: Multiple-choice QA tasks based on detailed texts. Created by translating the [QuALITY dataset](https://huggingface.co/datasets/L4NLP/LEval/viewer/quality) from L-Eval.
24
-
25
  ### Group III: Multi-hop Question Answering
26
- - **ruBABILongQA**: 5 long-context reasoning tasks for QA using facts hidden among irrelevant information.
 
27
  - **LongContextMultiQ**: Multi-hop QA based on Wikidata and Wikipedia.
28
  - **LibrusecMHQA**: Multi-hop QA requiring information distributed across several text parts.
29
  - **ru2WikiMultihopQA**: Translation of the [2WikiMultihopQA dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/2wikimqa_e) from LongBench.
30
-
31
  ### Group IV: Complex Reasoning and Mathematical Problems
 
32
  - **ruSciPassageCount**: Count unique paragraphs in a long text. Uses the basic idea of the original [PassageCount dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/passage_count) from LongBench.
33
- - **ruQasper**: Question Answering over academic research papers. Created by translating the [Qasper dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/qasper_e) from LongBench.
34
- - **ruGSM100**: Solve math problems using Chain-of-Thought reasoning. Created by translating the [GSM100](https://huggingface.co/datasets/L4NLP/LEval/viewer/gsm100) dataset from L-Eval.
35
-
36
- ## Dataset Structure
37
-
38
- The datasets are divided into subsets based on context lengths: 4k, 8k, 16k, 32k, 64k, and 128k tokens. Each subset contains a different number of samples depending on the task complexity.
39
-
40
- ## Add your model
41
-
42
- For placing your model to leaderboard you need to:
43
-
44
- 1. Score the model on our repository.
45
- 2. Add the result in the format **"\<model_name\>.json"** in the "results" folder.
46
- 3. Create a Pull Request.
47
-
48
- ## *GPT-4o
49
-
50
- Because of limited resources, we assessed GPT-4o on just 10% of each dataset in our benchmark, including each context length. Consequently, the results might not be exact.
51
-
 
 
 
 
 
 
 
 
 
52
  ## Citation
53
-
54
- _TODO_
55
-
56
  ```
57
- @article{LIBRA2024,
58
- title={Long Input Benchmark for Russian Analysis},
59
- author={Anonymous},
60
- journal={ACL},
61
- year={2024}
 
 
 
62
  }
63
- ```
64
-
65
- ## License
66
-
67
- The datasets are published under the MIT license.
68
-
69
- ## Acknowledgments
70
-
71
- For more details and code, please visit our [GitHub repository](https://github.com/ai-forever/LIBRA/).
 
1
  # LIBRA: Long Input Benchmark for Russian Analysis
2
+ <p align="center">
3
+ <img src="logo.png" width="500" />
4
+ </p>
5
 
6
+ ## LIBRA
7
 
8
+ LIBRA (Long Input Benchmark for Russian Analysis) is designed to evaluate the capabilities of large language models (LLMs) in understanding and processing long texts in Russian. This benchmark includes 18 datasets adapted for different tasks and complexities. The tasks are divided into four complexity groups and allow evaluation across various context lengths ranging from 4k up to 512k tokens.
9
 
10
+ For model comparison and results, see the [LIBRA Leaderboard](https://huggingface.co/spaces/ai-forever/LIBRA-Leaderboard). The benchmark is described in detail in our [paper](https://arxiv.org/abs/2408.02439).
11
 
12
+ **NOTE:** This is a new benchmark version released in May 2026. The original version (described in the paper) can be found [here](https://huggingface.co/datasets/ai-forever/LIBRA_old). We strongly encourage using the current version as it contains cleaned and extended datasets with ensured data quality.
13
 
14
+ ## LIBRA Mini
15
+
16
+ Running a full LIBRA evaluation can be prohibitively expensive and time-consuming due to the large number of datasets and long context lengths involved. Moreover, some of the included tasks have become less informative as benchmarks, with modern models achieving near-saturated scores on them.
17
+
18
+ To address this, we introduce **LIBRA Mini** β€” a compact, curated subset of 6 datasets selected from the full benchmark. These datasets represent the most challenging and diagnostically informative tasks in LIBRA, covering diverse task types and complexity levels (see [Task Description](#task-description) for more information):
19
+
20
+ - **ruBABILongQA3** * β€” multi-fact reasoning over long contexts
21
+ - **ruSciPassageCount** * β€” counting unique paragraphs in extended scientific texts
22
+ - **LibrusecMHQA** * β€” multi-hop QA with information spread across multiple text parts
23
+ - **LongContextMultiQ** * β€” multi-hop QA based on Wikidata and Wikipedia
24
+ - **ru2WikiMultihopQA** * β€” multi-hop reasoning across multiple Wikipedia articles
25
+ - **MatreshkaNames** * β€” identifying persons in dialogues based on discussed topics
26
+
27
+ **Note:** The same asterisk (*) is used on the leaderboard to label LIBRA Mini datasets in the results tables.
28
+
29
+ LIBRA Mini uses the same **Exact Match (EM)** metric and evaluation methodology as the full benchmark. Results for LIBRA Mini are reported in a dedicated section on the [leaderboard](https://huggingface.co/spaces/ai-forever/LIBRA-Leaderboard).
30
+
31
+ We recommend **using LIBRA Mini as the primary evaluation suite for model comparisons**, while the full LIBRA benchmark remains available for comprehensive analysis.
32
+
33
+ ## Dataset Structure
34
 
35
+ <p align="center">
36
+ <img src="libra_structure.svg" width="600" />
37
+ </p>
38
+
39
+ The datasets are divided into subsets based on context lengths. The table below shows the number of examples per context length for each dataset. Datasets included in **LIBRA Mini** are highlighted in bold. Note that not all datasets cover the full range of context lengths β€” some are designed for specific length ranges that best suit their task type.
40
+
41
+ | **Task** | **4k** | **8k** | **16k** | **32k** | **64k** | **128k** | **256k** | **512k** | | **Total** |
42
+ |:---|---:|---:|---:|---:|---:|---:|---:|---:|---|---:|
43
+ | *β€” Group I β€”* | | | | | | | | | | |
44
+ | passkey | 200 | 200 | 200 | 200 | 200 | 200 | 200 | 200 | | **1600** |
45
+ | passkey_with_librusec | 200 | 200 | 200 | 200 | 200 | 200 | 200 | 200 | | **1600** |
46
+ | *β€” Group II β€”* | | | | | | | | | | |
47
+ | librusec_history | - | 32 | 32 | 32 | 32 | - | - | - | | **128** |
48
+ | **matreshka_names** | 150 | 150 | 145 | 150 | 45 | - | - | - | | **640** |
49
+ | matreshka_yes_no | 300 | 300 | 300 | 300 | 290 | 280 | - | - | | **1770** |
50
+ | ru_quality | - | 18 | 184 | - | - | - | - | - | | **202** |
51
+ | ru_sci_abstract_retrieval | 209 | 210 | 210 | 206 | 185 | 200 | 200 | - | | **1420** |
52
+ | ru_sci_fi | - | - | - | 216 | 213 | - | - | - | | **429** |
53
+ | ru_tpo | - | 900 | - | - | - | - | - | - | | **900** |
54
+ | *β€” Group III β€”* | | | | | | | | | | |
55
+ | **librusec_mhqa** | - | 384 | - | - | - | - | - | - | | **384** |
56
+ | **long_context_multiq** | 158 | 121 | 83 | 109 | 41 | - | - | - | | **512** |
57
+ | **ru_2wikimultihopqa** | - | 147 | 384 | 369 | - | - | - | - | | **900** |
58
+ | ru_babilong_qa1 | 99 | 99 | 99 | 94 | 91 | 99 | 200 | 98 | | **879** |
59
+ | ru_babilong_qa2 | 85 | 76 | 77 | 73 | 65 | 99 | 200 | 98 | | **773** |
60
+ | **ru_babilong_qa3** | 60 | 68 | 69 | 65 | 65 | 100 | 198 | 98 | | **723** |
61
+ | ru_babilong_qa4 | 78 | 85 | 83 | 87 | 75 | 99 | 200 | 98 | | **805** |
62
+ | ru_babilong_qa5 | 99 | 99 | 98 | 96 | 96 | 99 | 200 | 98 | | **885** |
63
+ | *β€” Group IV β€”* | | | | | | | | | | |
64
+ | **ru_sci_passage_count** | 104 | 99 | 100 | 87 | 91 | 90 | 43 | 60 | | **674** |
65
+ ## Task Description
66
+ The benchmark tasks are organized into four complexity groups, ranging from simple retrieval to complex reasoning. The grouping reflects both the cognitive difficulty of the task and the degree to which models must integrate information across the full context window. Group I serves as a basic sanity check, verifying that a model can process long inputs at all. Groups II and III progressively require deeper language understanding, multi-step reasoning, and the ability to locate and combine information from distant parts of the context. Group IV represents the most demanding tasks, requiring complex reasoning that goes beyond standard question answering formats. The total score on the leaderboard is computed across all four groups.
67
+ ### Group I: Simple Information Retrieval (sanity check)
68
+ This group includes the most simple tasks which serve as a sanity check for models to work with such amount of tokens.
69
+ - **Passkey**: Extract a relevant piece of code number from a long text fragment. Based on the original [PassKey test](https://github.com/CStanKonrad/long_llama/blob/main/examples/passkey.py) from the LongLLaMA's GitHub repo.
70
+ - **PasskeyWithLibrusec**: Similar to Passkey but with added noise from Librusec texts.
71
  ### Group II: Question Answering and Multiple Choice
72
+ This group consists of standard QA and multiple choice tasks adapted for the long-context setting.
73
  - **MatreshkaNames**: Identify the person in dialogues based on the discussed topic. We used [Matreshka](https://huggingface.co/datasets/zjkarina/matreshka) dataset and [Russian Names](https://www.kaggle.com/datasets/rai220/russian-cyrillic-names-and-sex/data) dataset to create this and the next task.
74
  - **MatreshkaYesNo**: Indicate whether a specific topic was mentioned in the dialog.
75
+ - **LibrusecHistory**: Answer questions based on historical texts. Ideologically similar to the [PassageRetrieval dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/passage_retrieval_en) from LongBench.
 
76
  - **ruSciFi**: Answer true/false based on context and general world knowledge. Translation of [SciFi dataset](https://huggingface.co/datasets/L4NLP/LEval/viewer/sci_f) from L-Eval which originally was based on [SF-Gram](https://github.com/nschaetti/SFGram-dataset).
77
  - **ruSciAbstractRetrieval**: Retrieve relevant paragraphs from scientific abstracts.
78
  - **ruTPO**: Multiple-choice questions similar to TOEFL exams. Translation of the [TPO dataset](https://huggingface.co/datasets/L4NLP/LEval/viewer/tpo) from L-Eval.
79
  - **ruQuALITY**: Multiple-choice QA tasks based on detailed texts. Created by translating the [QuALITY dataset](https://huggingface.co/datasets/L4NLP/LEval/viewer/quality) from L-Eval.
 
80
  ### Group III: Multi-hop Question Answering
81
+ This group includes long-context multi-hop QA problems where the answer requires combining multiple pieces of information distributed across the context.
82
+ - **ruBABILongQA (1-5)**: 5 long-context reasoning tasks for QA using facts hidden among irrelevant information.
83
  - **LongContextMultiQ**: Multi-hop QA based on Wikidata and Wikipedia.
84
  - **LibrusecMHQA**: Multi-hop QA requiring information distributed across several text parts.
85
  - **ru2WikiMultihopQA**: Translation of the [2WikiMultihopQA dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/2wikimqa_e) from LongBench.
 
86
  ### Group IV: Complex Reasoning and Mathematical Problems
87
+ This group includes the most complex long-context tasks, which span beyond multiple-choice and multi-hop QA. At this point Group IV comprises only one task and we invite the community to contribute to it.
88
  - **ruSciPassageCount**: Count unique paragraphs in a long text. Uses the basic idea of the original [PassageCount dataset](https://huggingface.co/datasets/THUDM/LongBench/viewer/passage_count) from LongBench.
89
+ ## Metrics
90
+ We use **Exact Match (EM)** as a primary metric for all tasks. **EM** is used to evaluate the accuracy of the model's responses by comparing the predicted answers to the ground truth.
91
+ ## Changes from the Original Version
92
+ This version of LIBRA includes both automatic and manual improvements over the original release. All datasets underwent automatic quality filtering to ensure consistency and reliability of annotations. In addition, several datasets were manually revised and extended with the help of human annotators: **LibrusecMHQA**, **LongContextMultiQ**, **MatreshkaNames**, **MatreshkaYesNo**, **ru2WikiMultihopQA**, **ruSciFi**, and **ruTPO** received targeted corrections and additional examples. The datasets **ruGSM100** and **ruQasper** were removed from the benchmark as they did not meet the updated quality criteria.
93
+ The maximum supported context length has been extended from 128k to 512k tokens. The following datasets now include examples at longer context lengths not present in the original version: **ruBABILongQA (1–5)**, **ruSciAbstractRetrieval**, **ruSciPassageCount**, **LongContextMultiQ**, **MatreshkaYesNo**, **Passkey**, and **PasskeyWithLibrusec**.
94
+ ## Evaluation
95
+ Starting from this version, LIBRA supports evaluation via [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) β€” a widely adopted framework for standardized LLM evaluation. Both full **LIBRA** and compact **LIBRA Mini** evaluations are supported.
96
+ To get started:
97
+ ```bash
98
+ pip install lm-eval[vllm]
99
+ ```
100
+ Run evaluation on the full LIBRA benchmark:
101
+ ```bash
102
+ lm_eval --model vllm \
103
+ --model_args pretrained=Qwen/Qwen3-30B-A3B,max_model_len=262144 \
104
+ --tasks libra \
105
+ --apply_chat_template \
106
+ --device cuda:0
107
+ ```
108
+ Run evaluation on LIBRA Mini only:
109
+ ```bash
110
+ lm_eval --model vllm \
111
+ --model_args pretrained=Qwen/Qwen3-30B-A3B,max_model_len=262144 \
112
+ --tasks libra_mini \
113
+ --apply_chat_template \
114
+ --device cuda:0
115
+ ```
116
+ For the full list of configuration options and instructions on adding new models, please refer to the [lm-evaluation-harness documentation](https://github.com/EleutherAI/lm-evaluation-harness).
117
  ## Citation
 
 
 
118
  ```
119
+ @misc{churin2024longinputbenchmarkrussian,
120
+ title={Long Input Benchmark for Russian Analysis},
121
+ author={Igor Churin and Murat Apishev and Maria Tikhonova and Denis Shevelev and Aydar Bulatov and Yuri Kuratov and Sergei Averkiev and Alena Fenogenova},
122
+ year={2024},
123
+ eprint={2408.02439},
124
+ archivePrefix={arXiv},
125
+ primaryClass={cs.CL},
126
+ url={https://arxiv.org/abs/2408.02439},
127
  }
128
+ ```
 
 
 
 
 
 
 
 
results/LiquidAI_LFM2.5-1.2B-Instruct.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "total_score": 0.2860481481481481,
3
  "passkey": {
4
  "dataset_total_score": 0.6666666666666666,
5
  "4k": 1.0,
@@ -58,13 +58,12 @@
58
  "32k": 0.287
59
  },
60
  "long_context_multiq": {
61
- "dataset_total_score": 0.17666666666666667,
62
  "4k": 0.418,
63
  "8k": 0.364,
64
  "16k": 0.241,
65
  "32k": 0.037,
66
- "64k": 0.0,
67
- "128k": 0.0
68
  },
69
  "ru_sci_abstract_retrieval": {
70
  "dataset_total_score": 0.2318333333333333,
 
1
  {
2
+ "total_score": 0.2880111111111111,
3
  "passkey": {
4
  "dataset_total_score": 0.6666666666666666,
5
  "4k": 1.0,
 
58
  "32k": 0.287
59
  },
60
  "long_context_multiq": {
61
+ "dataset_total_score": 0.21200000000000002,
62
  "4k": 0.418,
63
  "8k": 0.364,
64
  "16k": 0.241,
65
  "32k": 0.037,
66
+ "64k": 0.0
 
67
  },
68
  "ru_sci_abstract_retrieval": {
69
  "dataset_total_score": 0.2318333333333333,