Granite Library
Safetensors
GGUF
English
frreiss Hong23 commited on
Commit
d83fa4c
·
1 Parent(s): 17b2c9f

Update context relevance model card with granite 4 micro lora eval (#14)

Browse files

- Update context relevance model card with granite 4 micro lora eval (664b90c175b03e67608efd3c82ac0afa75d155f6)


Co-authored-by: Li <Hong23@users.noreply.huggingface.co>

Files changed (1) hide show
  1. context_relevance/README.md +32 -79
context_relevance/README.md CHANGED
@@ -11,20 +11,20 @@ library_name: transformers
11
 
12
  ## Model Summary
13
 
14
- This is a RAG-specific family of intrinsics fine-tuned for the context relevancy task:
15
 
16
  Given (1) a document and (2) a multi-turn conversation between a user and an AI assistant, identify whether the document is relevant (including partially relevant) and useful to answering the last user question.
17
  While this adapter is general purpose, it is especially effective in RAG settings right after the retrieval model's step where the adapter can be used to identify documents or passages that may mislead or harm the downstream generator model's response generation.
18
 
19
  - **Developer:** IBM Research
20
- - **Model type:** LoRA and aLoRA adapter for [ibm-granite/granite-3.3-8b-instruct](https://huggingface.co/ibm-granite/granite-3.3-8b-instruct)
21
  and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
22
  - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
23
 
24
  ## Intended use
25
  This is a family of intrinsincs that enables classification of relevant/irrelevant documents for the final user query in a multi-turn conversation. The model has been trained to determine whether a document is useful and relevant for answering the last user question in a multi-turn conversation.
26
 
27
- The classification output from the context relevancy model can be used in several downstream applications, including but not limited to:
28
 
29
  - Filter out irrelevant and misleading documents before sending them to a generator model in a RAG setting. By removing irrelevant documents upfront, the downsteam generator model reduces the likelihood of outputing misleading responses. Moreover, removing irrelevant documents or passages can help reduce eat up downstream context-length limits , thereby improving inference time.
30
 
@@ -35,7 +35,7 @@ The classification output from the context relevancy model can be used in severa
35
  1. A conversation formatted using the chat template
36
  2. The final user query extracted from the conversation
37
  3. A document to evaluate for relevance
38
- 4. A special context relevancy invocation prompt
39
 
40
  The model uses a specific format with separate roles for each component:
41
  - Conversation: Applied via `tokenizer.apply_chat_template()`
@@ -61,34 +61,8 @@ The model uses a specific format with separate roles for each component:
61
  }
62
  ```
63
 
64
- ## Quickstart Example
65
-
66
- The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework.
67
- Here is some example code for calling this intrinsic from Mellea:
68
- ```
69
- from mellea.backends.huggingface import LocalHFBackend
70
- from mellea.stdlib.base import ChatContext, Document
71
- from mellea.stdlib.chat import Message
72
- from mellea.stdlib.intrinsics import rag
73
-
74
-
75
- backend = LocalHFBackend(model_id="ibm-granite/granite-3.3-2b-instruct")
76
- context = ChatContext()
77
- question = "Who is the CEO of Microsoft?"
78
- document = Document(
79
- # Document text does not say who is the CEO.
80
- "Microsoft Corporation is an American multinational corporation and technology "
81
- "conglomerate headquartered in Redmond, Washington.[2] Founded in 1975, the "
82
- "company became influential in the rise of personal computers through software "
83
- "like Windows, and the company has since expanded to Internet services, cloud "
84
- "computing, video gaming and other fields. Microsoft is the largest software "
85
- "maker, one of the most valuable public U.S. companies,[a] and one of the most "
86
- "valuable brands globally."
87
- )
88
-
89
- result = rag.check_context_relevance(question, document, context, backend)
90
- print(f"Result of context relevance check: {result}")
91
- ```
92
 
93
 
94
  ## Training Details
@@ -106,7 +80,7 @@ The LoRA adapter was fine-tuned using PEFT under the following regime: rank = 32
106
 
107
  Our model was evaluated against several existing benchmarks to measure its ability to determine whether document is relevant or not to the last user question in a multi-turn conversation.
108
 
109
- ### Forming evaluation datasets for context relevancy
110
 
111
  For each of the datasets, we created pairs of document and question pairs and generated a label (based on the available labels in the dataset) as to whether the document is relevant or useful to answering the question.
112
 
@@ -117,62 +91,41 @@ Given the dataset groundtruth and the LoRA outputs, we then calculated the preci
117
 
118
  ### Evaluation Datasets
119
 
120
- We evaluated on the following datasets from which we form documents and question pairs for an evaluation dataset (we also denote how we derive negative labels per dataset):
121
 
122
- - Multi-turn conversational RAG benchmarks: MTRAG [MT-RAG](https://arxiv.org/pdf/2501.03468). The MT-RAG dataset comes with three settings of RAG retrieval: (1) Reference: perfect retriever (all passages are relevant to the question); (2) Reference + RAG: references passages + additional passages resulting in 5 total passages; (3) RAG: standard RAG setting (top 5 retrieved passages). We split the evaluation dataset according to the three settings and formed pairs of questions and passages.
123
- - Pre-existing negative labels.
 
 
 
124
 
125
- - SAGE RAG benchmarks used in evaluating Granite models: [FinanceBench](https://arxiv.org/abs/2311.11944), [BioASQ](https://huggingface.co/datasets/enelpol/rag-mini-bioasq), [Open Australian Legal Corpus](https://huggingface.co/datasets/isaacus/open-australian-legal-qa).
126
- - Negative labels created by random pairings.
127
 
128
- - [Granite Guardian](https://arxiv.org/abs/2412.07724) datasets for context relevancy: HotpotQA and SquadV2.
129
- - Pre-existing negative labels which were obtained via prompting a separate model over the original datasets.
130
-
131
- - Question Answering datasets, that [implicit-fact retrieval](https://arxiv.org/abs/2409.14924): [CLAPNQ](https://arxiv.org/abs/2404.02103) and [Drop](https://aclanthology.org/N19-1246/).
132
- - Pre-existing negative labels.
133
 
134
- - [BEIR](https://arxiv.org/abs/2104.08663) benchmark datasets for information retrieval: [BioASQ](https://huggingface.co/datasets/enelpol/rag-mini-bioasq)
135
- - Negative labels created by random pairings.
136
 
 
 
 
 
 
 
 
 
 
 
 
 
 
137
 
138
- ### Evaluation Results
 
139
 
140
- We evaluated the LoRA adapter against `meta-llama/Llama-3.3-70B-Instruct` and the base Granite model, `ibm-granite/granite-3.3-8b-instruct`. For Llama, we prompted the model using the prompt [proven to perform best for context relevancy](https://arxiv.org/abs/2502.13908).
141
-
142
- #### Precision for relevant labels
143
-
144
- | Benchmark Type | Dataset | meta-llama/Llama-3.3-70B-Instruct | ibm-granite/granite-guardian-3.1-5b | ibm-granite/granite-3.3-8b-instruct | Three-label LoRA | Delta from base |
145
- |-----------------------------|-----------------------------------------|-----------------------------------|-------------------------------------|-------------------------------------|----------------------------------------------------------|-----------------|
146
- | MTRAG | MTRAG (RAG) | 0.93 | 0.90 | 0.93 | 0.96 | 0.034 |
147
- | MTRAG | MTRAG Clean (RAG) (71 fixed) | 0.95 | 0.92 | 0.94 | 0.97 | 0.03 |
148
- | MTRAG | MTRAG (Reference + RAG) | 0.87 | 0.83 | 0.86 | 0.92 | 0.06 |
149
- | MTRAG | MTRAG (Reference) | 1 | 1 | 1 | 1 | 0 |
150
- | QA, Implicit Fact Retrieval | PrimeQA/clapnq | 0.97 | 0.31 | 0.69 | 0.99 | 0.30 |
151
- | QA, Implicit Fact Retrieval | ucinlp/drop | 0.86 | 0.35 | 0.73 | 0.94 | 0.21 |
152
- | Sage Benchmarks + BeIR | PatronusAI/financebench | 0.92 | 0.44 | 0.66 | 0.99 | 0.33 |
153
- | Sage Benchmarks + BeIR | enelpol/rag-mini-bioasq | 0.98 | 0.83 | 0.94 | 0.99 | 0.05 |
154
- | Sage Benchmarks + BeIR | isaacus/open-australian-legal-qa | 0.93 | 0.08 | 0.60 | 1 | 0.40 |
155
- | Granite Guardian | HotpotQA + SquadV2 | 0.54 | 0.13 | 0.54 | 0.65 | 0.11 |
156
-
157
-
158
- #### Recall for irrelevant labels:
159
-
160
- | Benchmark Type | Dataset | meta-llama/Llama-3.3-70B-Instruct | ibm-granite/granite-guardian-3.1-5b | ibm-granite/granite-3.3-8b-instruct | Three-label LoRA | Delta from base |
161
- |-----------------------------|-----------------------------------------|-----------------------------------|-------------------------------------|-------------------------------------|------------------|-----------------|
162
- | MTRAG | MTRAG (RAG) | 0.41 | 0.21 | 0.29 | 0.73 | 0.44 |
163
- | MTRAG | MTRAG Clean (RAG) (71 fixed) | 0.44 | 0.21 | 0.33 | 0.74 | 0.41 |
164
- | MTRAG | MTRAG (Reference + RAG) | 0.39 | 0.21 | 0.29 | 0.73 | 0.44 |
165
- | MTRAG | MTRAG (Reference) | 0 | 0 | 0 | 0 | 0 |
166
- | QA, Implicit Fact Retrieval | PrimeQA/clapnq | 0.97 | 0 | 0.65 | 0.99 | 0.34 |
167
- | QA, Implicit Fact Retrieval | ucinlp/drop | 0.84 | 0.02 | 0.64 | 0.94 | 0.30 |
168
- | Sage Benchmarks + BeIR | PatronusAI/financebench | 0.89 | 0 | 0.34 | 0.99 | 0.65 |
169
- | Sage Benchmarks + BeIR | enelpol/rag-mini-bioasq | 0.87 | 0 | 0.53 | 0.99 | 0.46 |
170
- | Sage Benchmarks + BeIR | isaacus/open-australian-legal-qa | 0.92 | 0 | 0.36 | 1.0 | 0.64 |
171
- | Granite Guardian | HotpotQA + SquadV2 | 0.16 | 0.08 | 0.15 | 0.35 | 0.20 |
172
 
173
  ## Contact
174
 
175
- [Maeda Hanafi](mailto:maeda.hanafi@ibm.com)
176
 
177
  ### Framework versions
178
 
 
11
 
12
  ## Model Summary
13
 
14
+ This is a RAG-specific family of intrinsics fine-tuned for the context relevance task:
15
 
16
  Given (1) a document and (2) a multi-turn conversation between a user and an AI assistant, identify whether the document is relevant (including partially relevant) and useful to answering the last user question.
17
  While this adapter is general purpose, it is especially effective in RAG settings right after the retrieval model's step where the adapter can be used to identify documents or passages that may mislead or harm the downstream generator model's response generation.
18
 
19
  - **Developer:** IBM Research
20
+ - **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granitelib-rag-r1.0/tree/main/context_relevance/granite-4.0-micro)
21
  and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
22
  - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
23
 
24
  ## Intended use
25
  This is a family of intrinsincs that enables classification of relevant/irrelevant documents for the final user query in a multi-turn conversation. The model has been trained to determine whether a document is useful and relevant for answering the last user question in a multi-turn conversation.
26
 
27
+ The classification output from the context relevance model can be used in several downstream applications, including but not limited to:
28
 
29
  - Filter out irrelevant and misleading documents before sending them to a generator model in a RAG setting. By removing irrelevant documents upfront, the downsteam generator model reduces the likelihood of outputing misleading responses. Moreover, removing irrelevant documents or passages can help reduce eat up downstream context-length limits , thereby improving inference time.
30
 
 
35
  1. A conversation formatted using the chat template
36
  2. The final user query extracted from the conversation
37
  3. A document to evaluate for relevance
38
+ 4. A special context relevance invocation prompt
39
 
40
  The model uses a specific format with separate roles for each component:
41
  - Conversation: Applied via `tokenizer.apply_chat_template()`
 
61
  }
62
  ```
63
 
64
+ **Usage Example**:
65
+ The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework, following the examples [here](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
66
 
67
 
68
  ## Training Details
 
80
 
81
  Our model was evaluated against several existing benchmarks to measure its ability to determine whether document is relevant or not to the last user question in a multi-turn conversation.
82
 
83
+ ### Forming evaluation datasets for context relevance
84
 
85
  For each of the datasets, we created pairs of document and question pairs and generated a label (based on the available labels in the dataset) as to whether the document is relevant or useful to answering the question.
86
 
 
91
 
92
  ### Evaluation Datasets
93
 
94
+ We evaluated on the following five datasets:
95
 
96
+ - [CLAPNQ](https://arxiv.org/abs/2404.02103)
97
+ - [RAG MINI BIOASQ](https://huggingface.co/datasets/enelpol/rag-mini-bioasq)
98
+ - [Open Australian Legal QA](https://huggingface.co/datasets/isaacus/open-australian-legal-qa)
99
+ - [FinanceBench](https://arxiv.org/abs/2311.11944)
100
+ - [UCINLP DROP](https://aclanthology.org/N19-1246/)
101
 
 
 
102
 
103
+ ### Evaluation Results
 
 
 
 
104
 
105
+ We evaluated the LoRA adapters against various frontier models on the five evaluation datasets. Here, we create strong prompt instruction (three labels) for frontier models, but no instruction for LoRA adapters.
106
+ The results show average F1 scores for relevant and irrelevant classifications, along with overall accuracy.
107
 
108
+ | Model | Avg. F1 (Relevant) | Avg. F1 (Irrelevant) | Accuracy |
109
+ |-------|-------------------|---------------------|----------|
110
+ | **Frontier Models (Out of Box)** | | | |
111
+ | Granite 3.3-8b Instruct | 94.9% | 92.3% | 94.7% |
112
+ | Granite 4 micro | 93.8% | 88.6% | 93.1% |
113
+ | GPT-OSS-20b | 90.1% | 85.7% | 90.0% |
114
+ | GPT-OSS-120b | 93.7% | 91.4% | 93.6% |
115
+ | GPT 4o mini | 93.7% | 90.8% | 93.6% |
116
+ | GPT 4o | 95.6% | 93.9% | 95.7% |
117
+ | GPT 5 mini | 95.8% | 93.9% | 95.6% |
118
+ | **Trained LoRAs** | | | |
119
+ | Granite 4 micro-context-relevance-LoRA | 93.2% | 87.1% | 92.1% |
120
+ | Granite 4 micro-context-relevance-aLoRA | 87.0% | 72.8% | 83.4% |
121
 
122
+ **Takeaways:**
123
+ - LoRA without instruction is comparable to frontier models with strong instruction
124
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
125
 
126
  ## Contact
127
 
128
+ [Lihong He](mailto:lihong.he@ibm.com)
129
 
130
  ### Framework versions
131