frreiss yannisk2 commited on
Commit
664dfb1
·
1 Parent(s): ed250f0

Update citation generation model card with latest results and reference to granite-4.0-micro model card (#7)

Browse files

- Update citation generation model card with latest results and reference to granite-4.0-micro model card (1fd194df60b990691b66a30fd3b21bf5a1d2175e)


Co-authored-by: Yannis Katsis <yannisk2@users.noreply.huggingface.co>

Files changed (1) hide show
  1. citations/README.md +154 -147
citations/README.md CHANGED
@@ -7,208 +7,215 @@ library_name: peft
7
  library_name: transformers
8
  ---
9
 
10
- # Intrinsics for Citation Generation
11
 
12
  ## Model Summary
13
 
14
- This is a RAG-specific family of intrinsics fine-tuned for the citation generation task. Given a multi-turn conversation between a user and an AI assistant ending with an assistant response and a set of documents/passages on which the last assistant response is supposed to be based, each intrinsic in the family generates citations for the last assistant response from the provided documents/passages. The intrinsic has the following features:
15
- 1. **Fine-grained citations:** The intrinsic generates citations for each sentence in the assistant response (when available). Moreover, each citation consists of a set of sentences from the documents/passages that support the corresponding sentence in the assistant response.
16
- 2. **Post-hoc citation generation:** Since the intrinsic takes the assistant response as input, it can generate citations for responses generated by any LLM. Pick your favorite LLM and use the intrinsic to generate post-hoc citations!
17
 
18
- We provide two intrinsics implemented as LoRA adapters trained over Granite-3.3-2b-instruct and Granite-3.3-8b-instruct, respectively.
19
-
20
- </br>
21
 
22
  - **Developer:** IBM Research
23
- - **Model type:** LoRA adapter for [ibm-granite/granite-3.3-2b-instruct](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct) and [ibm-granite/granite-3.3-8b-instruct](https://huggingface.co/ibm-granite/granite-3.3-8b-instruct)
24
  - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
25
 
26
  ## Intended use
27
- This is a family of citation generation intrinsics that give the ability to generate citations for the last assistant response in a multi-turn RAG conversation based on a set of provided documents/passages. They can be used to generate post-hoc citations for assistant responses generated by any LLM in a RAG setting.
28
 
29
  > [!TIP]
30
- > Note: While you can invoke a citation generation intrinsic directly, it is strongly recommended to call it through [granite-common](https://github.com/ibm-granite/granite-common), which wraps the model with a tailored I/O processor, enabling a friendlier development interface. The I/O processor takes care of several data transformation/validation tasks that would be otherwise required (incl. splitting the input documents and assistant response into sentences before calling the intrinsic as well as validating the intrinsic's output and transforming the returned sentence IDs into spans over the documents and the response). We next describe the input/output of the citation generation intrinsics when invoked through granite-common.
31
-
32
- **Intrinsic input**: The input to the citation generation intrinsic is an OpenAI-compatible chat completion request, containing a list of conversation turns ending with the assistant response for which the citations should be generated as well as the list of documents from which the citations should be drawn. Please see the code snippets in the Quickstart Example section below for examples on how to specify the chat completion request as a JSON object.
33
-
34
- **Intrinsic output**: The output of the citation generation intrinsic is formatted as the result of the original chat completion request containing the citations for the last assistant response. The citations are provided in the form of a JSON array, whose items include the text and begin/end of a response span together with the text, document id and begin/end of a document span that serves as a citation for the response span. When there are more than one document spans that serve as citations for a single response span, they are represented as separate objects in the JSON array.
35
-
36
- **Going from input to output**: When calling the intrinsic through granite-common one should follow the steps below to transform the intrinsic input to the corresponding output. These steps are also exemplified in the code snippets included in the Quickstart Example section below. Given an input chat completion request, the request should be passed to the corresponding input processor (also referred to as IntrinsicsRewriter) provided by granite-common. The input processor converts the request to the appropriate format expected by the underlying citation generation model. This includes, among others, splitting the last assistant response and the documents into sentences and prepending them with sentence IDs as well as introducing an appropriate task-specific instruction. The input processor's result should then be passed to the underlying citation generation model for inference. The model generates citations using a compact representation consisting of sentence IDs in the last assistant response and documents. This output should finally be passed to the appropriate output processor (also referred to as IntrinsicsResultProcessor) provided by granite-common. The output processor converts the low-level raw model output to the final output by, among others, mapping the sentence IDs back to response and document spans. The result is an application-friendly format ready for consumption by downstream applications.
37
-
38
- ## Quickstart Example
39
-
40
- The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework.
41
- Here is some example code for calling this intrinsic from Mellea:
42
- ```
43
- from mellea.backends.huggingface import LocalHFBackend
44
- from mellea.stdlib.base import ChatContext, Document
45
- from mellea.stdlib.chat import Message
46
- from mellea.stdlib.intrinsics import rag
47
- import json
48
-
49
-
50
- backend = LocalHFBackend(model_id="ibm-granite/granite-3.3-2b-instruct")
51
- context = ChatContext().add(
52
- Message(
53
- "user",
54
- "How does Murdoch's expansion in Australia compare to his expansion "
55
- "in New Zealand?",
56
- )
57
- )
58
- assistant_response = (
59
- "Murdoch expanded in Australia and New Zealand by acquiring and expanding local "
60
- "newspapers. I do not have information about his expansion in New Zealand after "
61
- "purchasing The Dominion."
62
- )
63
- documents = [
64
- Document(
65
- doc_id="1",
66
- text="Keith Rupert Murdoch was born on 11 March 1931 in Melbourne, Australia, "
67
- "the son of Sir Keith Murdoch (1885-1952) and Dame Elisabeth Murdoch (nee "
68
- "Greene; 1909-2012). He is of English, Irish, and Scottish ancestry. Murdoch's "
69
- "parents were also born in Melbourne. Keith Murdoch was a war correspondent "
70
- "and later a regional newspaper magnate owning two newspapers in Adelaide, "
71
- "South Australia, and a radio station in a faraway mining town. Following his "
72
- "father's death, when he was 21, Murdoch returned from Oxford to take charge "
73
- "of the family business News Limited, which had been established in 1923. "
74
- "Rupert Murdoch turned its Adelaide newspaper, The News, its main asset, into "
75
- "a major success. He began to direct his attention to acquisition and "
76
- "expansion, buying the troubled Sunday Times in Perth, Western Australia "
77
- "(1956) and over the next few years acquiring suburban and provincial "
78
- "newspapers in New South Wales, Queensland, Victoria and the Northern "
79
- "Territory, including the Sydney afternoon tabloid, The Daily Mirror (1960). "
80
- 'The Economist describes Murdoch as "inventing the modern tabloid", as he '
81
- "developed a pattern for his newspapers, increasing sports and scandal "
82
- "coverage and adopting eye-catching headlines. Murdoch's first foray outside "
83
- "Australia involved the purchase of a controlling interest in the New Zealand "
84
- "daily The Dominion. In January 1964, while touring New Zealand with friends "
85
- "in a rented Morris Minor after sailing across the Tasman, Murdoch read of a "
86
- "takeover bid for the Wellington paper by the British-based Canadian newspaper "
87
- "magnate, Lord Thomson of Fleet. On the spur of the moment, he launched a "
88
- "counter-bid. A four-way battle for control ensued in which the 32-year-old "
89
- "Murdoch was ultimately successful. Later in 1964, Murdoch launched The "
90
- "Australian, Australia's first national daily newspaper, which was based "
91
- "first in Canberra and later in Sydney. In 1972, Murdoch acquired the Sydney "
92
- "morning tabloid The Daily Telegraph from Australian media mogul Sir Frank "
93
- "Packer, who later regretted selling it to him. In 1984, Murdoch was appointed "
94
- "Companion of the Order of Australia (AC) for services to publishing. In 1999, "
95
- "Murdoch significantly expanded his music holdings in Australia by acquiring "
96
- "the controlling share in a leading Australian independent label, Michael "
97
- "Gudinski's Mushroom Records; he merged that with Festival Records, and the "
98
- "result was Festival Mushroom Records (FMR). Both Festival and FMR were "
99
- "managed by Murdoch's son James Murdoch for several years.",
100
- ),
101
- Document(
102
- doc_id="2",
103
- text="This document has nothing to do with Rupert Murdoch. This document is "
104
- "two sentences long.",
105
- ),
106
  ]
 
107
 
 
108
 
109
- result = rag.find_citations(assistant_response, documents, context, backend)
110
- print(f"Result of citations intrinsic:\n{json.dumps(result, indent=2)}")
111
- ```
112
 
113
  ## Training Details
114
 
115
- The citation generation intrinsics were trained on synthetically-generated citation datasets. The process of generating the training data consisted of two main steps:
116
- - **Multi-turn RAG conversation generation:** Starting from publicly available document corpora, we generated a set of multi-turn RAG data, consisting of multi-turn conversations grounded on passages retrieved from the corpora. For details on the RAG conversation generation process please refer to the [Granite Technical Report](https://github.com/ibm-granite/granite-3.0-language-models/blob/main/paper.pdf) and [Lee, Young-Suk, et al.](https://arxiv.org/pdf/2409.11500).
117
- - **Citation generation:** For each turn of the multi-turn RAG conversations from the previous step, we used a multi-step synthetic citation generation pipeline to generate citations for the assistant response.
118
 
119
- The resulting data instances were used to train the citation generation intrinsics.
120
 
121
  ### Training Data
122
 
123
  The following public datasets were used as seed datasets for the multi-turn RAG conversation generation process:
124
- - [CoQA](https://stanfordnlp.github.io/coqa/) - Wikipedia passages
125
  - [MultiDoc2Dial](https://huggingface.co/datasets/IBM/multidoc2dial)
126
  - [QuAC](https://huggingface.co/datasets/allenai/quac)
127
 
128
-
129
  ## Evaluation
130
 
131
- We evaluate the citation generation intrinsics on two citation benchmarks:
132
- - [ALCE](https://aclanthology.org/2023.emnlp-main.398/): Evaluates the ability of models to produce document/passage-level citations (i.e., identify the documents/passages that support a statement in the response).
133
- - [LongBench-Cite](https://arxiv.org/abs/2409.02897): Evaluates the ability of models to produce fine-grained span-level citations (i.e., identify the spans within the input documents/passages that support a statement in the response) with a focus on long contexts.
134
-
135
- Since the intrinsics correspond to a post-hoc citation generation approach, their performance on the two benchmarks depends on the assistant responses for which they are asked to generate citations. To facilitate an apples-to-apples comparison, for each experiment, we keep the assistant responses the same and change the model that is used to generate the citations. In particular, we prompt an LLM to create an assistant response together with citations and evaluate the generated citations on the corresponding benchmark. Then, we compute and evaluate the citations generated for the same LLM response by each of the citation generation intrinsics. We provide results for the two intrinsics, implemented as LoRA adapters over Granite-3.3-2b-instruct and Granite-3.3-8b-instruct, respectively.
136
-
137
- ### Evaluation on ALCE
138
-
139
- For the ALCE evaluation, we prompt Llama-3.1-70B-Instruct and Mixtral-8x22B-Instruct to generate both the assistant response and corresponding passage-level citations. We first calculate the performance of the citations generated by these models on ALCE. Subsequently, we feed the responses of these models (leaving out the citations) to the citation generation intrinsics and evaluate their generated citations. The results are shown in the table below:
140
-
141
- Model used to generate response | Model used to generate citations | Recall | Precision | F1 |
142
- |--------------| ----------------------------- | --------------- | ----------------- | --------- |
143
- | Llama-3.1-70B-Instruct | Llama-3.1-70B-Instruct | 61.4 | 58.1 | 59.7 |
144
- | Llama-3.1-70B-Instruct | Granite-3.3-2B LoRA citations | 51.5 | 64.2 | 57.2 |
145
- | Llama-3.1-70B-Instruct | Granite-3.3-8B LoRA citations | 55.4 | 64.2 | 59.5 |
146
- | Mixtral-8x22B-Instruct | Mixtral-8x22B-Instruct | 62.2 | 62.5 | 62.3 |
147
- | Mixtral-8x22B-Instruct | Granite-3.3-2B LoRA citations | 51.4 | 67.3 | 58.3 |
148
- | Mixtral-8x22B-Instruct | Granite-3.3-8B LoRA citations | 55.8 | 68.5 | 61.5 |
149
-
150
- We observe that the LoRA adapter over Granite-3.3-8b-instruct performs on par with much bigger models when those are prompted to create passage-level citations (with the LoRA adapter over over Granite-3.3-2b-instruct being slightly worse). It is interesting to note that while the adapter's F1 performance is similar to the baselines, it exhibits a different precision-recall trade-off, trading lower recall for higher precision.
151
-
152
- Notes:
153
- - All results are reported on the ELI5 dataset using the ORACLE (5-psg) setting.
154
- - To prompt Llama and Mixtral, we employ a setting similar to the one proposed in the ALCE paper; in particular we use a two-shot prompt comprised of two of the ICL examples from ALCE as well as a slightly modified version of the instruction from the paper.
155
- - Sentence splitting of context/response is performed using NLTK.
156
- - Finally, since ALCE expects passage-level citations, we elevate the finer-grained citations produced by the LoRA adapter to the passage level before running the ALCE evaluation.
157
-
158
 
159
- ### Evaluation on LongBench-Cite
 
 
160
 
161
- For the LonBench-Cite evaluation, we prompt Llama-3.1-70B-Instruct to generate both the assistant response and corresponding citations. Then we evaluate the citations generated by Llama as well as the post-hoc citations generated by the citation generation intrinsics when invoked on the Llama responses. The results are shown in the table below:
162
 
163
  <table>
164
  <tr>
165
- <th>Model used to generate response</th>
166
- <th>Model used to generate citations</th>
167
  <th colspan="3">Longbench-Chat (en)</th>
168
  <th colspan="3">MultifieldQA (en)</th>
169
  <th colspan="3">HotpotQA</th>
170
  <th colspan="3">GovReport</th>
 
171
  </tr>
172
  <tr>
173
- <th></th>
174
  <th></th>
175
  <th>R</th><th>P</th><th>F1</th>
176
  <th>R</th><th>P</th><th>F1</th>
177
  <th>R</th><th>P</th><th>F1</th>
178
  <th>R</th><th>P</th><th>F1</th>
 
179
  </tr>
180
  <tr>
181
- <td>Llama-3.1-70B-Instruct</td>
182
- <td>Llama-3.1-70B-Instruct</td>
183
- <td>27.0</td><td>34.4</td><td>26.1</td>
184
- <td>46.1</td><td>63.3</td><td>49.7</td>
185
- <td>34.0</td><td>39.4</td><td>30.2</td>
186
- <td>55.0</td><td>77.5</td><td>62.0</td>
187
  </tr>
188
  <tr>
189
- <td>Llama-3.1-70B-Instruct</td>
190
- <td>Granite-3.3-2B LoRA citations</td>
191
- <td>38.7</td><td>47.4</td><td>39.3</td>
192
- <td>66.4</td><td>81.8</td><td>70.4</td>
193
- <td>60.7</td><td>68.5</td><td>59.7</td>
194
- <td>60.1</td><td>72.4</td><td>64.7</td>
195
  </tr>
196
  <tr>
197
- <td>Llama-3.1-70B-Instruct</td>
198
- <td>Granite-3.3-8B LoRA citations</td>
199
- <td>54.5</td><td>59.9</td><td>55.6</td>
200
- <td>73.0</td><td>82.9</td><td>75.7</td>
201
- <td>68.5</td><td>73.8</td><td>66.4</td>
202
- <td>73.5</td><td>84.6</td><td>78.2</td>
203
  </tr>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
204
  </table>
205
 
206
- We observe that both variants of the LoRA adapter (even the one trained over Granite-3.3-2b-instruct) perform across the board significantly better than Llama-3.1-70B-Instruct when prompted to create span-level citations. This demonstrates the value of the adapter to create post-hoc citations even for assistant responses generated by much bigger LLMs.
207
 
208
  Notes:
209
  - The evaluation results are reported on the English subset of LongBench-Cite (i.e., restricted to instances whose `language` field equals to `en`).
210
- - To prompt Llama to generate a response with citations, we use the one-shot prompt described in the paper.
211
- - For the LoRA adapter, sentence splitting of the context is performed using NLTK. For the response, we reuse the splitting in Llama's output (since the LongBench-Cite prompt instructs the model to output a response split into sentences/statements).
212
 
213
  ## Model Card Authors
214
 
 
7
  library_name: transformers
8
  ---
9
 
10
+ # Citation Generation Intrinsic
11
 
12
  ## Model Summary
13
 
14
+ This is a RAG-specific intrinsic fine-tuned for the citation generation task. Given a multi-turn conversation between a user and an AI assistant ending with an assistant response and a set of documents/passages on which the last assistant response is supposed to be based, the intrinsic generates citations for the last assistant response from the provided documents/passages. The citation generation intrinsic has the following features:
15
+ 1. **Fine-grained citations:** The intrinsic generates citations for each sentence of the assistant response (when available). Each citation consists of a set of sentences from the documents/passages that support the corresponding sentence in the assistant response.
16
+ 2. **Post-hoc citation generation:** Since the intrinsic takes the assistant response as input, it can generate citations for responses generated by any LLM. Pick your favorite LLM for response generation and use the citation generation intrinsic to generate post-hoc citations!
17
 
18
+ We have created two implementation of the intrinsic as LoRA adapters trained over granite-4.0-micro and gpt-oss-20b, respectively. This is the model card for the LoRA adapter trained over gpt-oss-20b. The model card for the LoRA adapter trained over granite-4.0-micro can be found [here](https://huggingface.co/ibm-granite/granitelib-rag-r1.0/blob/main/citations/README.md).
 
 
19
 
20
  - **Developer:** IBM Research
21
+ - **Model type:** LoRA adapter for [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
22
  - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
23
 
24
  ## Intended use
25
+ This is a citation generation intrinsic that gives the ability to generate citations for the last assistant response in a multi-turn RAG conversation based on a set of provided documents/passages. It can be used to generate post-hoc citations for assistant responses generated by any LLM in a RAG setting.
26
 
27
  > [!TIP]
28
+ > Note: While you can invoke the citation generation intrinsic directly, it is strongly recommended to call it through the [Mellea](https://mellea.ai) framework, which wraps the model with a tailored I/O processor, enabling a friendlier development interface. We next describe the input/output of the citation generation intrinsic when invoked through Mellea.
29
+
30
+ **Intrinsic input**: The citation generation intrinsic takes as input the following:
31
+ - _Conversation:_ A list of conversational turns ending with the last user question, encoded as a list of user/assistant messages.
32
+ - _Assistant response:_ The assistant response to the last user question, which is also the response for which citations will be generated, provided as a string.
33
+ - _Documents:_ A list of documents from which the citations should be drawn, encoded as a collection of Document objects.
34
+
35
+ **Intrinsic output**: The output of the citation generation intrinsic contains the citations for the last assistant response. The citations are provided in the form of a JSON array, whose items include the text and begin/end of a response span together with the text, document id and begin/end of a document span that serves as a citation for that response span. When there are more than one document spans that serve as citations for a single response span, they are represented as separate objects in the JSON array.
36
+
37
+ **Going from input to output**: When calling the intrinsic through Mellea, the framework internally performs multiple steps to transform the intrinsic input to the corresponding output. While you do not have to explicitly invoke these steps, we next provide a brief overview of this process. Given an input to the citation generation intrinsic, Mellea performs the following tasks:
38
+ - _Convert user input to the appropriate format expected by the underlying citation generation model._ This includes, among others, splitting the last assistant response and the documents into sentences and prepending them with sentence IDs as well as introducing an appropriate task-specific instruction.
39
+ - _Call underlying citation generation model for inference._ The model generates citations using a compact representation consisting of sentence IDs in the last assistant response and documents.
40
+ - _Convert model output to final output._ The low-level raw model output is converted to the final output by, among others, mapping the sentence IDs back to response and document spans. The result is an application-friendly JSON format ready for consumption by downstream applications.
41
+
42
+ ## Example
43
+
44
+ You can find below an example of the input and corresponding output of the citations generation intrinsic:
45
+
46
+ ### Input
47
+
48
+ **Conversation:**
49
+
50
+ User: What is the visibility level of Git Repos and Issue Tracking projects?
51
+
52
+ **Assistant response:**
53
+
54
+ Git Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. Private projects are visible only to project members, internal projects are visible to all users that are logged in to IBM Cloud, and public projects are visible to anyone. By default, new projects are set to private visibility level, which is the most secure for your data.
55
+
56
+ **Documents:**
57
+
58
+ Git Repos and Issue Tracking is an IBM-hosted component of the Continuous Delivery service. All of the data that you provide to Git Repos and Issue Tracking, including but not limited to source files, issues, pull requests, and project configuration properties, is managed securely within Continuous Delivery. However, Git Repos and Issue Tracking supports various mechanisms for exporting, sending, or otherwise sharing data to users and third parties. The ability of Git Repos and Issue Tracking to share information is typical of many social coding platforms. However, such sharing might conflict with regulatory controls that apply to your business. After you create a project in Git Repos and Issue Tracking, but before you entrust any files, issues, records, or other data with the project, review the project settings and change any settings that you deem necessary to protect your data. Settings to review include visibility levels, email notifications, integrations, web hooks, access tokens, deploy tokens, and deploy keys. Project visibility levels \n\nGit Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. * Private projects are visible only to project members. This setting is the default visibility level for new projects, and is the most secure visibility level for your data. * Internal projects are visible to all users that are logged in to IBM Cloud. * Public projects are visible to anyone. To limit project access to only project members, complete the following steps:\n\n\n\n1. From the project sidebar, click Settings > General. 2. On the General Settings page, click Visibility > project features > permissions. 3. Locate the Project visibility setting. 4. Select Private, if it is not already selected. 5. Click Save changes. Project membership \n\nGit Repos and Issue Tracking is a cloud hosted social coding environment that is available to all Continuous Delivery users. If you are a Git Repos and Issue Tracking project Maintainer or Owner, you can invite any user and group members to the project. IBM Cloud places no restrictions on who you can invite to a project.
59
+
60
+ ### Output
61
+
62
+ ```json
63
+ [
64
+ {
65
+ "response_begin": 0,
66
+ "response_end": 117,
67
+ "response_text": "Git Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. ",
68
+ "citation_doc_id": "1",
69
+ "citation_begin": 1034,
70
+ "citation_end": 1179,
71
+ "citation_text": "Project visibility levels \n\nGit Repos and Issue Tracking projects can have one of the following visibility levels: private, internal, or public. "
72
+ },
73
+ {
74
+ "response_begin": 117,
75
+ "response_end": 290,
76
+ "response_text": "Private projects are visible only to project members, internal projects are visible to all users that are logged in to IBM Cloud, and public projects are visible to anyone. ",
77
+ "citation_doc_id": "1",
78
+ "citation_begin": 1179,
79
+ "citation_end": 1235,
80
+ "citation_text": "* Private projects are visible only to project members. "
81
+ },
82
+ {
83
+ "response_begin": 117,
84
+ "response_end": 290,
85
+ "response_text": "Private projects are visible only to project members, internal projects are visible to all users that are logged in to IBM Cloud, and public projects are visible to anyone. ",
86
+ "citation_doc_id": "1",
87
+ "citation_begin": 1353,
88
+ "citation_end": 1472,
89
+ "citation_text": "* Internal projects are visible to all users that are logged in to IBM Cloud. * Public projects are visible to anyone. "
90
+ },
91
+ {
92
+ "response_begin": 290,
93
+ "response_end": 391,
94
+ "response_text": "By default, new projects are set to private visibility level, which is the most secure for your data.",
95
+ "citation_doc_id": "1",
96
+ "citation_begin": 1235,
97
+ "citation_end": 1353,
98
+ "citation_text": "This setting is the default visibility level for new projects, and is the most secure visibility level for your data. "
99
+ }
 
 
 
 
100
  ]
101
+ ```
102
 
103
+ ## Quickstart
104
 
105
+ The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework. For code snippets demonstrating how to use this and other intrinsics, please refer to the [Mellea intrinsics examples](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
 
 
106
 
107
  ## Training Details
108
 
109
+ The citation generation intrinsic was trained on synthetically-generated citation datasets. The process of generating the training data consisted of two main steps:
110
+ - _Multi-turn RAG conversation generation:_ Starting from publicly available document corpora, we generated a set of multi-turn RAG data, consisting of multi-turn conversations grounded on passages retrieved from the corpora. For details on the RAG conversation generation process please refer to the [Granite Technical Report](https://github.com/ibm-granite/granite-3.0-language-models/blob/main/paper.pdf) and [Lee, Young-Suk, et al.](https://arxiv.org/pdf/2409.11500)
111
+ - _Citation generation:_ For each turn of the multi-turn RAG conversations from the previous step, we used a multi-step synthetic citation generation pipeline to generate citations for the assistant response.
112
 
113
+ The resulting data instances were used to train the citation generation intrinsic.
114
 
115
  ### Training Data
116
 
117
  The following public datasets were used as seed datasets for the multi-turn RAG conversation generation process:
 
118
  - [MultiDoc2Dial](https://huggingface.co/datasets/IBM/multidoc2dial)
119
  - [QuAC](https://huggingface.co/datasets/allenai/quac)
120
 
 
121
  ## Evaluation
122
 
123
+ We evaluated the citation generation intrinsic on a revised version of the [LongBench-Cite](https://arxiv.org/abs/2409.02897) benchmark; a benchmark evaluating the ability of models to produce fine-grained span-level citations (i.e., identify the spans within the input documents/passages that support a statement in the response) with a focus on long contexts. Being originally designed to evaluate inline citation generation approaches (i.e., approaches generating the assistant response and the citations at the same time), we adapt the benchmark for the evaluation of post-hoc citation generation (where citations are generated for a given assistant response generated by an upstream model).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
124
 
125
+ For the following experiments, we prompted Llama-3.1-70B-Instruct to generate the assistant response for the LongBench-Cite tasks. Then, two types of models were asked to create citations for these responses:
126
+ - _Citation generation LoRA adapters:_ These are the two citation generation LoRA adapter implementations of the citation intrinsic, as described above.
127
+ - _Prompt-based baselines:_ These are out-of-the-box LLMs prompted to generate post-hoc citations for the given assistant responses. Prompting was performed through a version of the 1-shot prompt used in the original benchmark, adapted for post-hoc citation generation.
128
 
129
+ The evaluation results are shown in the table below:
130
 
131
  <table>
132
  <tr>
133
+ <th>Model</th>
 
134
  <th colspan="3">Longbench-Chat (en)</th>
135
  <th colspan="3">MultifieldQA (en)</th>
136
  <th colspan="3">HotpotQA</th>
137
  <th colspan="3">GovReport</th>
138
+ <th>AVG F1</th>
139
  </tr>
140
  <tr>
 
141
  <th></th>
142
  <th>R</th><th>P</th><th>F1</th>
143
  <th>R</th><th>P</th><th>F1</th>
144
  <th>R</th><th>P</th><th>F1</th>
145
  <th>R</th><th>P</th><th>F1</th>
146
+ <th></th>
147
  </tr>
148
  <tr>
149
+ <th colspan="14" style="background-color: #f5f5f5;">Citation Generation LoRA Adapters</th>
 
 
 
 
 
150
  </tr>
151
  <tr>
152
+ <td>granite-4.0-micro LoRA</td>
153
+ <td>42.7</td><td>46.5</td><td>41.4</td>
154
+ <td>68.5</td><td>81.1</td><td>72.0</td>
155
+ <td>62.9</td><td>67.9</td><td>61.0</td>
156
+ <td>70.2</td><td>79.3</td><td>74.1</td>
157
+ <td><b>62.1</b></td>
158
  </tr>
159
  <tr>
160
+ <td>gpt-oss-20b LoRA</td>
161
+ <td>56.1</td><td>61.4</td><td>55.3</td>
162
+ <td>71.6</td><td>87.1</td><td>76.8</td>
163
+ <td>69.9</td><td>71.5</td><td>66.1</td>
164
+ <td>73.8</td><td>84.8</td><td>78.2</td>
165
+ <td><b>69.1</b></td>
166
  </tr>
167
+ <tr>
168
+ <th></th>
169
+ <th></th><th></th><th></th>
170
+ <th></th><th></th><th></th>
171
+ <th></th><th></th><th></th>
172
+ <th></th><th></th><th></th>
173
+ <th></th>
174
+ </tr>
175
+ <tr>
176
+ <th colspan="14" style="background-color: #f5f5f5;">Prompting-based Baselines</th>
177
+ </tr>
178
+ <tr>
179
+ <td>granite-4.0-micro Prompted</td>
180
+ <td>5.4</td><td>6.1</td><td>3.5</td>
181
+ <td>22.3</td><td>32.1</td><td>22.3</td>
182
+ <td>18.0</td><td>21.4</td><td>14.2</td>
183
+ <td>8.9</td><td>18.0</td><td>10.1</td>
184
+ <td><b>12.5</b></td>
185
+ </tr>
186
+ <tr>
187
+ <td>gpt-oss-20b Prompted</td>
188
+ <td>38.0</td><td>38.7</td><td>34.7</td>
189
+ <td>56.8</td><td>68.1</td><td>59.5</td>
190
+ <td>54.0</td><td>60.4</td><td>52.2</td>
191
+ <td>48.3</td><td>59.2</td><td>52.4</td>
192
+ <td><b>49.7</b></td>
193
+ </tr>
194
+ <tr>
195
+ <td>gpt-oss-120b Prompted</td>
196
+ <td>46.4</td><td>49.1</td><td>45.2</td>
197
+ <td>68.9</td><td>76.6</td><td>70.1</td>
198
+ <td>65.4</td><td>66.5</td><td>62.2</td>
199
+ <td>70.8</td><td>75.4</td><td>72.2</td>
200
+ <td><b>62.4</b></td>
201
+ </tr>
202
+ <tr>
203
+ <td>gpt-4o Prompted</td>
204
+ <td>56.9</td><td>60.1</td><td>56.2</td>
205
+ <td>68.7</td><td>79.6</td><td>71.4</td>
206
+ <td>65.3</td><td>73.6</td><td>65.3</td>
207
+ <td>74.3</td><td>81.5</td><td>76.9</td>
208
+ <td><b>67.5</b></td>
209
+ </tr>
210
+
211
  </table>
212
 
213
+ We observe that both citation generation LoRA adapters perform better not only than the corresponding base models prompted out of the box but also better than bigger models. For instance, the granite-4.0-micro LoRA performs on par with prompting the significantly larger gpt-oss-120b. Similarly, the gpt-oss-20b LoRA outperforms prompting the much larger gpt-4o.
214
 
215
  Notes:
216
  - The evaluation results are reported on the English subset of LongBench-Cite (i.e., restricted to instances whose `language` field equals to `en`).
217
+ - To generate the assistant responses fed to all evaluated models, we prompted Llama-3.1-70B-Instruct using the one-shot prompt described in the LongBench-Cite paper (which asks the model to generate a grounded response with citations) and removed the citations in post-processing.
218
+ - The AVG F1 column contains the average of the four dataset-specific F1 scores.
219
 
220
  ## Model Card Authors
221