Update citation generation model card with latest results and reference to granite-4.0-micro model card (#7)
Browse files- Update citation generation model card with latest results and reference to granite-4.0-micro model card (1fd194df60b990691b66a30fd3b21bf5a1d2175e)
Co-authored-by: Yannis Katsis <yannisk2@users.noreply.huggingface.co>
- citations/README.md +154 -147
citations/README.md
CHANGED
|
@@ -7,208 +7,215 @@ library_name: peft
|
|
| 7 |
library_name: transformers
|
| 8 |
---
|
| 9 |
|
| 10 |
-
#
|
| 11 |
|
| 12 |
## Model Summary
|
| 13 |
|
| 14 |
-
This is a RAG-specific
|
| 15 |
-
1. **Fine-grained citations:** The intrinsic generates citations for each sentence
|
| 16 |
-
2. **Post-hoc citation generation:** Since the intrinsic takes the assistant response as input, it can generate citations for responses generated by any LLM. Pick your favorite LLM and use the intrinsic to generate post-hoc citations!
|
| 17 |
|
| 18 |
-
We
|
| 19 |
-
|
| 20 |
-
</br>
|
| 21 |
|
| 22 |
- **Developer:** IBM Research
|
| 23 |
-
- **Model type:** LoRA adapter for [
|
| 24 |
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
|
| 25 |
|
| 26 |
## Intended use
|
| 27 |
-
This is a
|
| 28 |
|
| 29 |
> [!TIP]
|
| 30 |
-
> Note: While you can invoke
|
| 31 |
-
|
| 32 |
-
**Intrinsic input**: The
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
"
|
| 68 |
-
"
|
| 69 |
-
"
|
| 70 |
-
"
|
| 71 |
-
"
|
| 72 |
-
"
|
| 73 |
-
"
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
"
|
| 77 |
-
"
|
| 78 |
-
"
|
| 79 |
-
"
|
| 80 |
-
|
| 81 |
-
"
|
| 82 |
-
"
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
"
|
| 86 |
-
"
|
| 87 |
-
"
|
| 88 |
-
"
|
| 89 |
-
"
|
| 90 |
-
"
|
| 91 |
-
"
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
"
|
| 95 |
-
"
|
| 96 |
-
"
|
| 97 |
-
"
|
| 98 |
-
"
|
| 99 |
-
"
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
doc_id="2",
|
| 103 |
-
text="This document has nothing to do with Rupert Murdoch. This document is "
|
| 104 |
-
"two sentences long.",
|
| 105 |
-
),
|
| 106 |
]
|
|
|
|
| 107 |
|
|
|
|
| 108 |
|
| 109 |
-
|
| 110 |
-
print(f"Result of citations intrinsic:\n{json.dumps(result, indent=2)}")
|
| 111 |
-
```
|
| 112 |
|
| 113 |
## Training Details
|
| 114 |
|
| 115 |
-
The citation generation
|
| 116 |
-
-
|
| 117 |
-
-
|
| 118 |
|
| 119 |
-
The resulting data instances were used to train the citation generation
|
| 120 |
|
| 121 |
### Training Data
|
| 122 |
|
| 123 |
The following public datasets were used as seed datasets for the multi-turn RAG conversation generation process:
|
| 124 |
-
- [CoQA](https://stanfordnlp.github.io/coqa/) - Wikipedia passages
|
| 125 |
- [MultiDoc2Dial](https://huggingface.co/datasets/IBM/multidoc2dial)
|
| 126 |
- [QuAC](https://huggingface.co/datasets/allenai/quac)
|
| 127 |
|
| 128 |
-
|
| 129 |
## Evaluation
|
| 130 |
|
| 131 |
-
We
|
| 132 |
-
- [ALCE](https://aclanthology.org/2023.emnlp-main.398/): Evaluates the ability of models to produce document/passage-level citations (i.e., identify the documents/passages that support a statement in the response).
|
| 133 |
-
- [LongBench-Cite](https://arxiv.org/abs/2409.02897): Evaluates the ability of models to produce fine-grained span-level citations (i.e., identify the spans within the input documents/passages that support a statement in the response) with a focus on long contexts.
|
| 134 |
-
|
| 135 |
-
Since the intrinsics correspond to a post-hoc citation generation approach, their performance on the two benchmarks depends on the assistant responses for which they are asked to generate citations. To facilitate an apples-to-apples comparison, for each experiment, we keep the assistant responses the same and change the model that is used to generate the citations. In particular, we prompt an LLM to create an assistant response together with citations and evaluate the generated citations on the corresponding benchmark. Then, we compute and evaluate the citations generated for the same LLM response by each of the citation generation intrinsics. We provide results for the two intrinsics, implemented as LoRA adapters over Granite-3.3-2b-instruct and Granite-3.3-8b-instruct, respectively.
|
| 136 |
-
|
| 137 |
-
### Evaluation on ALCE
|
| 138 |
-
|
| 139 |
-
For the ALCE evaluation, we prompt Llama-3.1-70B-Instruct and Mixtral-8x22B-Instruct to generate both the assistant response and corresponding passage-level citations. We first calculate the performance of the citations generated by these models on ALCE. Subsequently, we feed the responses of these models (leaving out the citations) to the citation generation intrinsics and evaluate their generated citations. The results are shown in the table below:
|
| 140 |
-
|
| 141 |
-
Model used to generate response | Model used to generate citations | Recall | Precision | F1 |
|
| 142 |
-
|--------------| ----------------------------- | --------------- | ----------------- | --------- |
|
| 143 |
-
| Llama-3.1-70B-Instruct | Llama-3.1-70B-Instruct | 61.4 | 58.1 | 59.7 |
|
| 144 |
-
| Llama-3.1-70B-Instruct | Granite-3.3-2B LoRA citations | 51.5 | 64.2 | 57.2 |
|
| 145 |
-
| Llama-3.1-70B-Instruct | Granite-3.3-8B LoRA citations | 55.4 | 64.2 | 59.5 |
|
| 146 |
-
| Mixtral-8x22B-Instruct | Mixtral-8x22B-Instruct | 62.2 | 62.5 | 62.3 |
|
| 147 |
-
| Mixtral-8x22B-Instruct | Granite-3.3-2B LoRA citations | 51.4 | 67.3 | 58.3 |
|
| 148 |
-
| Mixtral-8x22B-Instruct | Granite-3.3-8B LoRA citations | 55.8 | 68.5 | 61.5 |
|
| 149 |
-
|
| 150 |
-
We observe that the LoRA adapter over Granite-3.3-8b-instruct performs on par with much bigger models when those are prompted to create passage-level citations (with the LoRA adapter over over Granite-3.3-2b-instruct being slightly worse). It is interesting to note that while the adapter's F1 performance is similar to the baselines, it exhibits a different precision-recall trade-off, trading lower recall for higher precision.
|
| 151 |
-
|
| 152 |
-
Notes:
|
| 153 |
-
- All results are reported on the ELI5 dataset using the ORACLE (5-psg) setting.
|
| 154 |
-
- To prompt Llama and Mixtral, we employ a setting similar to the one proposed in the ALCE paper; in particular we use a two-shot prompt comprised of two of the ICL examples from ALCE as well as a slightly modified version of the instruction from the paper.
|
| 155 |
-
- Sentence splitting of context/response is performed using NLTK.
|
| 156 |
-
- Finally, since ALCE expects passage-level citations, we elevate the finer-grained citations produced by the LoRA adapter to the passage level before running the ALCE evaluation.
|
| 157 |
-
|
| 158 |
|
| 159 |
-
|
|
|
|
|
|
|
| 160 |
|
| 161 |
-
|
| 162 |
|
| 163 |
<table>
|
| 164 |
<tr>
|
| 165 |
-
<th>Model
|
| 166 |
-
<th>Model used to generate citations</th>
|
| 167 |
<th colspan="3">Longbench-Chat (en)</th>
|
| 168 |
<th colspan="3">MultifieldQA (en)</th>
|
| 169 |
<th colspan="3">HotpotQA</th>
|
| 170 |
<th colspan="3">GovReport</th>
|
|
|
|
| 171 |
</tr>
|
| 172 |
<tr>
|
| 173 |
-
<th></th>
|
| 174 |
<th></th>
|
| 175 |
<th>R</th><th>P</th><th>F1</th>
|
| 176 |
<th>R</th><th>P</th><th>F1</th>
|
| 177 |
<th>R</th><th>P</th><th>F1</th>
|
| 178 |
<th>R</th><th>P</th><th>F1</th>
|
|
|
|
| 179 |
</tr>
|
| 180 |
<tr>
|
| 181 |
-
<
|
| 182 |
-
<td>Llama-3.1-70B-Instruct</td>
|
| 183 |
-
<td>27.0</td><td>34.4</td><td>26.1</td>
|
| 184 |
-
<td>46.1</td><td>63.3</td><td>49.7</td>
|
| 185 |
-
<td>34.0</td><td>39.4</td><td>30.2</td>
|
| 186 |
-
<td>55.0</td><td>77.5</td><td>62.0</td>
|
| 187 |
</tr>
|
| 188 |
<tr>
|
| 189 |
-
<td>
|
| 190 |
-
<td>
|
| 191 |
-
<td>
|
| 192 |
-
<td>
|
| 193 |
-
<td>
|
| 194 |
-
<td>
|
| 195 |
</tr>
|
| 196 |
<tr>
|
| 197 |
-
<td>
|
| 198 |
-
<td>
|
| 199 |
-
<td>
|
| 200 |
-
<td>
|
| 201 |
-
<td>
|
| 202 |
-
<td>
|
| 203 |
</tr>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 204 |
</table>
|
| 205 |
|
| 206 |
-
We observe that both
|
| 207 |
|
| 208 |
Notes:
|
| 209 |
- The evaluation results are reported on the English subset of LongBench-Cite (i.e., restricted to instances whose `language` field equals to `en`).
|
| 210 |
-
- To
|
| 211 |
-
-
|
| 212 |
|
| 213 |
## Model Card Authors
|
| 214 |
|
|
|
|
| 7 |
library_name: transformers
|
| 8 |
---
|
| 9 |
|
| 10 |
+
# Citation Generation Intrinsic
|
| 11 |
|
| 12 |
## Model Summary
|
| 13 |
|
| 14 |
+
This is a RAG-specific intrinsic fine-tuned for the citation generation task. Given a multi-turn conversation between a user and an AI assistant ending with an assistant response and a set of documents/passages on which the last assistant response is supposed to be based, the intrinsic generates citations for the last assistant response from the provided documents/passages. The citation generation intrinsic has the following features:
|
| 15 |
+
1. **Fine-grained citations:** The intrinsic generates citations for each sentence of the assistant response (when available). Each citation consists of a set of sentences from the documents/passages that support the corresponding sentence in the assistant response.
|
| 16 |
+
2. **Post-hoc citation generation:** Since the intrinsic takes the assistant response as input, it can generate citations for responses generated by any LLM. Pick your favorite LLM for response generation and use the citation generation intrinsic to generate post-hoc citations!
|
| 17 |
|
| 18 |
+
We have created two implementation of the intrinsic as LoRA adapters trained over granite-4.0-micro and gpt-oss-20b, respectively. This is the model card for the LoRA adapter trained over gpt-oss-20b. The model card for the LoRA adapter trained over granite-4.0-micro can be found [here](https://huggingface.co/ibm-granite/granitelib-rag-r1.0/blob/main/citations/README.md).
|
|
|
|
|
|
|
| 19 |
|
| 20 |
- **Developer:** IBM Research
|
| 21 |
+
- **Model type:** LoRA adapter for [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
|
| 22 |
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
|
| 23 |
|
| 24 |
## Intended use
|
| 25 |
+
This is a citation generation intrinsic that gives the ability to generate citations for the last assistant response in a multi-turn RAG conversation based on a set of provided documents/passages. It can be used to generate post-hoc citations for assistant responses generated by any LLM in a RAG setting.
|
| 26 |
|
| 27 |
> [!TIP]
|
| 28 |
+
> Note: While you can invoke the citation generation intrinsic directly, it is strongly recommended to call it through the [Mellea](https://mellea.ai) framework, which wraps the model with a tailored I/O processor, enabling a friendlier development interface. We next describe the input/output of the citation generation intrinsic when invoked through Mellea.
|
| 29 |
+
|
| 30 |
+
**Intrinsic input**: The citation generation intrinsic takes as input the following:
|
| 31 |
+
- _Conversation:_ A list of conversational turns ending with the last user question, encoded as a list of user/assistant messages.
|
| 32 |
+
- _Assistant response:_ The assistant response to the last user question, which is also the response for which citations will be generated, provided as a string.
|
| 33 |
+
- _Documents:_ A list of documents from which the citations should be drawn, encoded as a collection of Document objects.
|
| 34 |
+
|
| 35 |
+
**Intrinsic output**: The output of the citation generation intrinsic contains the citations for the last assistant response. The citations are provided in the form of a JSON array, whose items include the text and begin/end of a response span together with the text, document id and begin/end of a document span that serves as a citation for that response span. When there are more than one document spans that serve as citations for a single response span, they are represented as separate objects in the JSON array.
|
| 36 |
+
|
| 37 |
+
**Going from input to output**: When calling the intrinsic through Mellea, the framework internally performs multiple steps to transform the intrinsic input to the corresponding output. While you do not have to explicitly invoke these steps, we next provide a brief overview of this process. Given an input to the citation generation intrinsic, Mellea performs the following tasks:
|
| 38 |
+
- _Convert user input to the appropriate format expected by the underlying citation generation model._ This includes, among others, splitting the last assistant response and the documents into sentences and prepending them with sentence IDs as well as introducing an appropriate task-specific instruction.
|
| 39 |
+
- _Call underlying citation generation model for inference._ The model generates citations using a compact representation consisting of sentence IDs in the last assistant response and documents.
|
| 40 |
+
- _Convert model output to final output._ The low-level raw model output is converted to the final output by, among others, mapping the sentence IDs back to response and document spans. The result is an application-friendly JSON format ready for consumption by downstream applications.
|
| 41 |
+
|
| 42 |
+
## Example
|
| 43 |
+
|
| 44 |
+
You can find below an example of the input and corresponding output of the citations generation intrinsic:
|
| 45 |
+
|
| 46 |
+
### Input
|
| 47 |
+
|
| 48 |
+
**Conversation:**
|
| 49 |
+
|
| 50 |
+
User: What is the visibility level of Git Repos and Issue Tracking projects?
|
| 51 |
+
|
| 52 |
+
**Assistant response:**
|
| 53 |
+
|
| 54 |
+
Git Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. Private projects are visible only to project members, internal projects are visible to all users that are logged in to IBM Cloud, and public projects are visible to anyone. By default, new projects are set to private visibility level, which is the most secure for your data.
|
| 55 |
+
|
| 56 |
+
**Documents:**
|
| 57 |
+
|
| 58 |
+
Git Repos and Issue Tracking is an IBM-hosted component of the Continuous Delivery service. All of the data that you provide to Git Repos and Issue Tracking, including but not limited to source files, issues, pull requests, and project configuration properties, is managed securely within Continuous Delivery. However, Git Repos and Issue Tracking supports various mechanisms for exporting, sending, or otherwise sharing data to users and third parties. The ability of Git Repos and Issue Tracking to share information is typical of many social coding platforms. However, such sharing might conflict with regulatory controls that apply to your business. After you create a project in Git Repos and Issue Tracking, but before you entrust any files, issues, records, or other data with the project, review the project settings and change any settings that you deem necessary to protect your data. Settings to review include visibility levels, email notifications, integrations, web hooks, access tokens, deploy tokens, and deploy keys. Project visibility levels \n\nGit Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. * Private projects are visible only to project members. This setting is the default visibility level for new projects, and is the most secure visibility level for your data. * Internal projects are visible to all users that are logged in to IBM Cloud. * Public projects are visible to anyone. To limit project access to only project members, complete the following steps:\n\n\n\n1. From the project sidebar, click Settings > General. 2. On the General Settings page, click Visibility > project features > permissions. 3. Locate the Project visibility setting. 4. Select Private, if it is not already selected. 5. Click Save changes. Project membership \n\nGit Repos and Issue Tracking is a cloud hosted social coding environment that is available to all Continuous Delivery users. If you are a Git Repos and Issue Tracking project Maintainer or Owner, you can invite any user and group members to the project. IBM Cloud places no restrictions on who you can invite to a project.
|
| 59 |
+
|
| 60 |
+
### Output
|
| 61 |
+
|
| 62 |
+
```json
|
| 63 |
+
[
|
| 64 |
+
{
|
| 65 |
+
"response_begin": 0,
|
| 66 |
+
"response_end": 117,
|
| 67 |
+
"response_text": "Git Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. ",
|
| 68 |
+
"citation_doc_id": "1",
|
| 69 |
+
"citation_begin": 1034,
|
| 70 |
+
"citation_end": 1179,
|
| 71 |
+
"citation_text": "Project visibility levels \n\nGit Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. "
|
| 72 |
+
},
|
| 73 |
+
{
|
| 74 |
+
"response_begin": 117,
|
| 75 |
+
"response_end": 290,
|
| 76 |
+
"response_text": "Private projects are visible only to project members, internal projects are visible to all users that are logged in to IBM Cloud, and public projects are visible to anyone. ",
|
| 77 |
+
"citation_doc_id": "1",
|
| 78 |
+
"citation_begin": 1179,
|
| 79 |
+
"citation_end": 1235,
|
| 80 |
+
"citation_text": "* Private projects are visible only to project members. "
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"response_begin": 117,
|
| 84 |
+
"response_end": 290,
|
| 85 |
+
"response_text": "Private projects are visible only to project members, internal projects are visible to all users that are logged in to IBM Cloud, and public projects are visible to anyone. ",
|
| 86 |
+
"citation_doc_id": "1",
|
| 87 |
+
"citation_begin": 1353,
|
| 88 |
+
"citation_end": 1472,
|
| 89 |
+
"citation_text": "* Internal projects are visible to all users that are logged in to IBM Cloud. * Public projects are visible to anyone. "
|
| 90 |
+
},
|
| 91 |
+
{
|
| 92 |
+
"response_begin": 290,
|
| 93 |
+
"response_end": 391,
|
| 94 |
+
"response_text": "By default, new projects are set to private visibility level, which is the most secure for your data.",
|
| 95 |
+
"citation_doc_id": "1",
|
| 96 |
+
"citation_begin": 1235,
|
| 97 |
+
"citation_end": 1353,
|
| 98 |
+
"citation_text": "This setting is the default visibility level for new projects, and is the most secure visibility level for your data. "
|
| 99 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
| 100 |
]
|
| 101 |
+
```
|
| 102 |
|
| 103 |
+
## Quickstart
|
| 104 |
|
| 105 |
+
The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework. For code snippets demonstrating how to use this and other intrinsics, please refer to the [Mellea intrinsics examples](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
|
|
|
|
|
|
|
| 106 |
|
| 107 |
## Training Details
|
| 108 |
|
| 109 |
+
The citation generation intrinsic was trained on synthetically-generated citation datasets. The process of generating the training data consisted of two main steps:
|
| 110 |
+
- _Multi-turn RAG conversation generation:_ Starting from publicly available document corpora, we generated a set of multi-turn RAG data, consisting of multi-turn conversations grounded on passages retrieved from the corpora. For details on the RAG conversation generation process please refer to the [Granite Technical Report](https://github.com/ibm-granite/granite-3.0-language-models/blob/main/paper.pdf) and [Lee, Young-Suk, et al.](https://arxiv.org/pdf/2409.11500)
|
| 111 |
+
- _Citation generation:_ For each turn of the multi-turn RAG conversations from the previous step, we used a multi-step synthetic citation generation pipeline to generate citations for the assistant response.
|
| 112 |
|
| 113 |
+
The resulting data instances were used to train the citation generation intrinsic.
|
| 114 |
|
| 115 |
### Training Data
|
| 116 |
|
| 117 |
The following public datasets were used as seed datasets for the multi-turn RAG conversation generation process:
|
|
|
|
| 118 |
- [MultiDoc2Dial](https://huggingface.co/datasets/IBM/multidoc2dial)
|
| 119 |
- [QuAC](https://huggingface.co/datasets/allenai/quac)
|
| 120 |
|
|
|
|
| 121 |
## Evaluation
|
| 122 |
|
| 123 |
+
We evaluated the citation generation intrinsic on a revised version of the [LongBench-Cite](https://arxiv.org/abs/2409.02897) benchmark; a benchmark evaluating the ability of models to produce fine-grained span-level citations (i.e., identify the spans within the input documents/passages that support a statement in the response) with a focus on long contexts. Being originally designed to evaluate inline citation generation approaches (i.e., approaches generating the assistant response and the citations at the same time), we adapt the benchmark for the evaluation of post-hoc citation generation (where citations are generated for a given assistant response generated by an upstream model).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 124 |
|
| 125 |
+
For the following experiments, we prompted Llama-3.1-70B-Instruct to generate the assistant response for the LongBench-Cite tasks. Then, two types of models were asked to create citations for these responses:
|
| 126 |
+
- _Citation generation LoRA adapters:_ These are the two citation generation LoRA adapter implementations of the citation intrinsic, as described above.
|
| 127 |
+
- _Prompt-based baselines:_ These are out-of-the-box LLMs prompted to generate post-hoc citations for the given assistant responses. Prompting was performed through a version of the 1-shot prompt used in the original benchmark, adapted for post-hoc citation generation.
|
| 128 |
|
| 129 |
+
The evaluation results are shown in the table below:
|
| 130 |
|
| 131 |
<table>
|
| 132 |
<tr>
|
| 133 |
+
<th>Model</th>
|
|
|
|
| 134 |
<th colspan="3">Longbench-Chat (en)</th>
|
| 135 |
<th colspan="3">MultifieldQA (en)</th>
|
| 136 |
<th colspan="3">HotpotQA</th>
|
| 137 |
<th colspan="3">GovReport</th>
|
| 138 |
+
<th>AVG F1</th>
|
| 139 |
</tr>
|
| 140 |
<tr>
|
|
|
|
| 141 |
<th></th>
|
| 142 |
<th>R</th><th>P</th><th>F1</th>
|
| 143 |
<th>R</th><th>P</th><th>F1</th>
|
| 144 |
<th>R</th><th>P</th><th>F1</th>
|
| 145 |
<th>R</th><th>P</th><th>F1</th>
|
| 146 |
+
<th></th>
|
| 147 |
</tr>
|
| 148 |
<tr>
|
| 149 |
+
<th colspan="14" style="background-color: #f5f5f5;">Citation Generation LoRA Adapters</th>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 150 |
</tr>
|
| 151 |
<tr>
|
| 152 |
+
<td>granite-4.0-micro LoRA</td>
|
| 153 |
+
<td>42.7</td><td>46.5</td><td>41.4</td>
|
| 154 |
+
<td>68.5</td><td>81.1</td><td>72.0</td>
|
| 155 |
+
<td>62.9</td><td>67.9</td><td>61.0</td>
|
| 156 |
+
<td>70.2</td><td>79.3</td><td>74.1</td>
|
| 157 |
+
<td><b>62.1</b></td>
|
| 158 |
</tr>
|
| 159 |
<tr>
|
| 160 |
+
<td>gpt-oss-20b LoRA</td>
|
| 161 |
+
<td>56.1</td><td>61.4</td><td>55.3</td>
|
| 162 |
+
<td>71.6</td><td>87.1</td><td>76.8</td>
|
| 163 |
+
<td>69.9</td><td>71.5</td><td>66.1</td>
|
| 164 |
+
<td>73.8</td><td>84.8</td><td>78.2</td>
|
| 165 |
+
<td><b>69.1</b></td>
|
| 166 |
</tr>
|
| 167 |
+
<tr>
|
| 168 |
+
<th></th>
|
| 169 |
+
<th></th><th></th><th></th>
|
| 170 |
+
<th></th><th></th><th></th>
|
| 171 |
+
<th></th><th></th><th></th>
|
| 172 |
+
<th></th><th></th><th></th>
|
| 173 |
+
<th></th>
|
| 174 |
+
</tr>
|
| 175 |
+
<tr>
|
| 176 |
+
<th colspan="14" style="background-color: #f5f5f5;">Prompting-based Baselines</th>
|
| 177 |
+
</tr>
|
| 178 |
+
<tr>
|
| 179 |
+
<td>granite-4.0-micro Prompted</td>
|
| 180 |
+
<td>5.4</td><td>6.1</td><td>3.5</td>
|
| 181 |
+
<td>22.3</td><td>32.1</td><td>22.3</td>
|
| 182 |
+
<td>18.0</td><td>21.4</td><td>14.2</td>
|
| 183 |
+
<td>8.9</td><td>18.0</td><td>10.1</td>
|
| 184 |
+
<td><b>12.5</b></td>
|
| 185 |
+
</tr>
|
| 186 |
+
<tr>
|
| 187 |
+
<td>gpt-oss-20b Prompted</td>
|
| 188 |
+
<td>38.0</td><td>38.7</td><td>34.7</td>
|
| 189 |
+
<td>56.8</td><td>68.1</td><td>59.5</td>
|
| 190 |
+
<td>54.0</td><td>60.4</td><td>52.2</td>
|
| 191 |
+
<td>48.3</td><td>59.2</td><td>52.4</td>
|
| 192 |
+
<td><b>49.7</b></td>
|
| 193 |
+
</tr>
|
| 194 |
+
<tr>
|
| 195 |
+
<td>gpt-oss-120b Prompted</td>
|
| 196 |
+
<td>46.4</td><td>49.1</td><td>45.2</td>
|
| 197 |
+
<td>68.9</td><td>76.6</td><td>70.1</td>
|
| 198 |
+
<td>65.4</td><td>66.5</td><td>62.2</td>
|
| 199 |
+
<td>70.8</td><td>75.4</td><td>72.2</td>
|
| 200 |
+
<td><b>62.4</b></td>
|
| 201 |
+
</tr>
|
| 202 |
+
<tr>
|
| 203 |
+
<td>gpt-4o Prompted</td>
|
| 204 |
+
<td>56.9</td><td>60.1</td><td>56.2</td>
|
| 205 |
+
<td>68.7</td><td>79.6</td><td>71.4</td>
|
| 206 |
+
<td>65.3</td><td>73.6</td><td>65.3</td>
|
| 207 |
+
<td>74.3</td><td>81.5</td><td>76.9</td>
|
| 208 |
+
<td><b>67.5</b></td>
|
| 209 |
+
</tr>
|
| 210 |
+
|
| 211 |
</table>
|
| 212 |
|
| 213 |
+
We observe that both citation generation LoRA adapters perform better not only than the corresponding base models prompted out of the box but also better than bigger models. For instance, the granite-4.0-micro LoRA performs on par with prompting the significantly larger gpt-oss-120b. Similarly, the gpt-oss-20b LoRA outperforms prompting the much larger gpt-4o.
|
| 214 |
|
| 215 |
Notes:
|
| 216 |
- The evaluation results are reported on the English subset of LongBench-Cite (i.e., restricted to instances whose `language` field equals to `en`).
|
| 217 |
+
- To generate the assistant responses fed to all evaluated models, we prompted Llama-3.1-70B-Instruct using the one-shot prompt described in the LongBench-Cite paper (which asks the model to generate a grounded response with citations) and removed the citations in post-processing.
|
| 218 |
+
- The AVG F1 column contains the average of the four dataset-specific F1 scores.
|
| 219 |
|
| 220 |
## Model Card Authors
|
| 221 |
|