Instructions to use ibm-granite/granitelib-rag-r1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Granite Library
How to use ibm-granite/granitelib-rag-r1.0 with Granite Library:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Update context relevance model card with granite 4 micro lora eval (#14)
Browse files- Update context relevance model card with granite 4 micro lora eval (664b90c175b03e67608efd3c82ac0afa75d155f6)
Co-authored-by: Li <Hong23@users.noreply.huggingface.co>
- context_relevance/README.md +32 -79
context_relevance/README.md
CHANGED
|
@@ -11,20 +11,20 @@ library_name: transformers
|
|
| 11 |
|
| 12 |
## Model Summary
|
| 13 |
|
| 14 |
-
This is a RAG-specific family of intrinsics fine-tuned for the context
|
| 15 |
|
| 16 |
Given (1) a document and (2) a multi-turn conversation between a user and an AI assistant, identify whether the document is relevant (including partially relevant) and useful to answering the last user question.
|
| 17 |
While this adapter is general purpose, it is especially effective in RAG settings right after the retrieval model's step where the adapter can be used to identify documents or passages that may mislead or harm the downstream generator model's response generation.
|
| 18 |
|
| 19 |
- **Developer:** IBM Research
|
| 20 |
-
- **Model type:** LoRA
|
| 21 |
and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
|
| 22 |
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
|
| 23 |
|
| 24 |
## Intended use
|
| 25 |
This is a family of intrinsincs that enables classification of relevant/irrelevant documents for the final user query in a multi-turn conversation. The model has been trained to determine whether a document is useful and relevant for answering the last user question in a multi-turn conversation.
|
| 26 |
|
| 27 |
-
The classification output from the context
|
| 28 |
|
| 29 |
- Filter out irrelevant and misleading documents before sending them to a generator model in a RAG setting. By removing irrelevant documents upfront, the downsteam generator model reduces the likelihood of outputing misleading responses. Moreover, removing irrelevant documents or passages can help reduce eat up downstream context-length limits , thereby improving inference time.
|
| 30 |
|
|
@@ -35,7 +35,7 @@ The classification output from the context relevancy model can be used in severa
|
|
| 35 |
1. A conversation formatted using the chat template
|
| 36 |
2. The final user query extracted from the conversation
|
| 37 |
3. A document to evaluate for relevance
|
| 38 |
-
4. A special context
|
| 39 |
|
| 40 |
The model uses a specific format with separate roles for each component:
|
| 41 |
- Conversation: Applied via `tokenizer.apply_chat_template()`
|
|
@@ -61,34 +61,8 @@ The model uses a specific format with separate roles for each component:
|
|
| 61 |
}
|
| 62 |
```
|
| 63 |
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework.
|
| 67 |
-
Here is some example code for calling this intrinsic from Mellea:
|
| 68 |
-
```
|
| 69 |
-
from mellea.backends.huggingface import LocalHFBackend
|
| 70 |
-
from mellea.stdlib.base import ChatContext, Document
|
| 71 |
-
from mellea.stdlib.chat import Message
|
| 72 |
-
from mellea.stdlib.intrinsics import rag
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
backend = LocalHFBackend(model_id="ibm-granite/granite-3.3-2b-instruct")
|
| 76 |
-
context = ChatContext()
|
| 77 |
-
question = "Who is the CEO of Microsoft?"
|
| 78 |
-
document = Document(
|
| 79 |
-
# Document text does not say who is the CEO.
|
| 80 |
-
"Microsoft Corporation is an American multinational corporation and technology "
|
| 81 |
-
"conglomerate headquartered in Redmond, Washington.[2] Founded in 1975, the "
|
| 82 |
-
"company became influential in the rise of personal computers through software "
|
| 83 |
-
"like Windows, and the company has since expanded to Internet services, cloud "
|
| 84 |
-
"computing, video gaming and other fields. Microsoft is the largest software "
|
| 85 |
-
"maker, one of the most valuable public U.S. companies,[a] and one of the most "
|
| 86 |
-
"valuable brands globally."
|
| 87 |
-
)
|
| 88 |
-
|
| 89 |
-
result = rag.check_context_relevance(question, document, context, backend)
|
| 90 |
-
print(f"Result of context relevance check: {result}")
|
| 91 |
-
```
|
| 92 |
|
| 93 |
|
| 94 |
## Training Details
|
|
@@ -106,7 +80,7 @@ The LoRA adapter was fine-tuned using PEFT under the following regime: rank = 32
|
|
| 106 |
|
| 107 |
Our model was evaluated against several existing benchmarks to measure its ability to determine whether document is relevant or not to the last user question in a multi-turn conversation.
|
| 108 |
|
| 109 |
-
### Forming evaluation datasets for context
|
| 110 |
|
| 111 |
For each of the datasets, we created pairs of document and question pairs and generated a label (based on the available labels in the dataset) as to whether the document is relevant or useful to answering the question.
|
| 112 |
|
|
@@ -117,62 +91,41 @@ Given the dataset groundtruth and the LoRA outputs, we then calculated the preci
|
|
| 117 |
|
| 118 |
### Evaluation Datasets
|
| 119 |
|
| 120 |
-
We evaluated on the following
|
| 121 |
|
| 122 |
-
-
|
| 123 |
-
|
|
|
|
|
|
|
|
|
|
| 124 |
|
| 125 |
-
- SAGE RAG benchmarks used in evaluating Granite models: [FinanceBench](https://arxiv.org/abs/2311.11944), [BioASQ](https://huggingface.co/datasets/enelpol/rag-mini-bioasq), [Open Australian Legal Corpus](https://huggingface.co/datasets/isaacus/open-australian-legal-qa).
|
| 126 |
-
- Negative labels created by random pairings.
|
| 127 |
|
| 128 |
-
|
| 129 |
-
- Pre-existing negative labels which were obtained via prompting a separate model over the original datasets.
|
| 130 |
-
|
| 131 |
-
- Question Answering datasets, that [implicit-fact retrieval](https://arxiv.org/abs/2409.14924): [CLAPNQ](https://arxiv.org/abs/2404.02103) and [Drop](https://aclanthology.org/N19-1246/).
|
| 132 |
-
- Pre-existing negative labels.
|
| 133 |
|
| 134 |
-
|
| 135 |
-
|
| 136 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 137 |
|
| 138 |
-
|
|
|
|
| 139 |
|
| 140 |
-
We evaluated the LoRA adapter against `meta-llama/Llama-3.3-70B-Instruct` and the base Granite model, `ibm-granite/granite-3.3-8b-instruct`. For Llama, we prompted the model using the prompt [proven to perform best for context relevancy](https://arxiv.org/abs/2502.13908).
|
| 141 |
-
|
| 142 |
-
#### Precision for relevant labels
|
| 143 |
-
|
| 144 |
-
| Benchmark Type | Dataset | meta-llama/Llama-3.3-70B-Instruct | ibm-granite/granite-guardian-3.1-5b | ibm-granite/granite-3.3-8b-instruct | Three-label LoRA | Delta from base |
|
| 145 |
-
|-----------------------------|-----------------------------------------|-----------------------------------|-------------------------------------|-------------------------------------|----------------------------------------------------------|-----------------|
|
| 146 |
-
| MTRAG | MTRAG (RAG) | 0.93 | 0.90 | 0.93 | 0.96 | 0.034 |
|
| 147 |
-
| MTRAG | MTRAG Clean (RAG) (71 fixed) | 0.95 | 0.92 | 0.94 | 0.97 | 0.03 |
|
| 148 |
-
| MTRAG | MTRAG (Reference + RAG) | 0.87 | 0.83 | 0.86 | 0.92 | 0.06 |
|
| 149 |
-
| MTRAG | MTRAG (Reference) | 1 | 1 | 1 | 1 | 0 |
|
| 150 |
-
| QA, Implicit Fact Retrieval | PrimeQA/clapnq | 0.97 | 0.31 | 0.69 | 0.99 | 0.30 |
|
| 151 |
-
| QA, Implicit Fact Retrieval | ucinlp/drop | 0.86 | 0.35 | 0.73 | 0.94 | 0.21 |
|
| 152 |
-
| Sage Benchmarks + BeIR | PatronusAI/financebench | 0.92 | 0.44 | 0.66 | 0.99 | 0.33 |
|
| 153 |
-
| Sage Benchmarks + BeIR | enelpol/rag-mini-bioasq | 0.98 | 0.83 | 0.94 | 0.99 | 0.05 |
|
| 154 |
-
| Sage Benchmarks + BeIR | isaacus/open-australian-legal-qa | 0.93 | 0.08 | 0.60 | 1 | 0.40 |
|
| 155 |
-
| Granite Guardian | HotpotQA + SquadV2 | 0.54 | 0.13 | 0.54 | 0.65 | 0.11 |
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
#### Recall for irrelevant labels:
|
| 159 |
-
|
| 160 |
-
| Benchmark Type | Dataset | meta-llama/Llama-3.3-70B-Instruct | ibm-granite/granite-guardian-3.1-5b | ibm-granite/granite-3.3-8b-instruct | Three-label LoRA | Delta from base |
|
| 161 |
-
|-----------------------------|-----------------------------------------|-----------------------------------|-------------------------------------|-------------------------------------|------------------|-----------------|
|
| 162 |
-
| MTRAG | MTRAG (RAG) | 0.41 | 0.21 | 0.29 | 0.73 | 0.44 |
|
| 163 |
-
| MTRAG | MTRAG Clean (RAG) (71 fixed) | 0.44 | 0.21 | 0.33 | 0.74 | 0.41 |
|
| 164 |
-
| MTRAG | MTRAG (Reference + RAG) | 0.39 | 0.21 | 0.29 | 0.73 | 0.44 |
|
| 165 |
-
| MTRAG | MTRAG (Reference) | 0 | 0 | 0 | 0 | 0 |
|
| 166 |
-
| QA, Implicit Fact Retrieval | PrimeQA/clapnq | 0.97 | 0 | 0.65 | 0.99 | 0.34 |
|
| 167 |
-
| QA, Implicit Fact Retrieval | ucinlp/drop | 0.84 | 0.02 | 0.64 | 0.94 | 0.30 |
|
| 168 |
-
| Sage Benchmarks + BeIR | PatronusAI/financebench | 0.89 | 0 | 0.34 | 0.99 | 0.65 |
|
| 169 |
-
| Sage Benchmarks + BeIR | enelpol/rag-mini-bioasq | 0.87 | 0 | 0.53 | 0.99 | 0.46 |
|
| 170 |
-
| Sage Benchmarks + BeIR | isaacus/open-australian-legal-qa | 0.92 | 0 | 0.36 | 1.0 | 0.64 |
|
| 171 |
-
| Granite Guardian | HotpotQA + SquadV2 | 0.16 | 0.08 | 0.15 | 0.35 | 0.20 |
|
| 172 |
|
| 173 |
## Contact
|
| 174 |
|
| 175 |
-
[
|
| 176 |
|
| 177 |
### Framework versions
|
| 178 |
|
|
|
|
| 11 |
|
| 12 |
## Model Summary
|
| 13 |
|
| 14 |
+
This is a RAG-specific family of intrinsics fine-tuned for the context relevance task:
|
| 15 |
|
| 16 |
Given (1) a document and (2) a multi-turn conversation between a user and an AI assistant, identify whether the document is relevant (including partially relevant) and useful to answering the last user question.
|
| 17 |
While this adapter is general purpose, it is especially effective in RAG settings right after the retrieval model's step where the adapter can be used to identify documents or passages that may mislead or harm the downstream generator model's response generation.
|
| 18 |
|
| 19 |
- **Developer:** IBM Research
|
| 20 |
+
- **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granitelib-rag-r1.0/tree/main/context_relevance/granite-4.0-micro)
|
| 21 |
and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
|
| 22 |
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
|
| 23 |
|
| 24 |
## Intended use
|
| 25 |
This is a family of intrinsincs that enables classification of relevant/irrelevant documents for the final user query in a multi-turn conversation. The model has been trained to determine whether a document is useful and relevant for answering the last user question in a multi-turn conversation.
|
| 26 |
|
| 27 |
+
The classification output from the context relevance model can be used in several downstream applications, including but not limited to:
|
| 28 |
|
| 29 |
- Filter out irrelevant and misleading documents before sending them to a generator model in a RAG setting. By removing irrelevant documents upfront, the downsteam generator model reduces the likelihood of outputing misleading responses. Moreover, removing irrelevant documents or passages can help reduce eat up downstream context-length limits , thereby improving inference time.
|
| 30 |
|
|
|
|
| 35 |
1. A conversation formatted using the chat template
|
| 36 |
2. The final user query extracted from the conversation
|
| 37 |
3. A document to evaluate for relevance
|
| 38 |
+
4. A special context relevance invocation prompt
|
| 39 |
|
| 40 |
The model uses a specific format with separate roles for each component:
|
| 41 |
- Conversation: Applied via `tokenizer.apply_chat_template()`
|
|
|
|
| 61 |
}
|
| 62 |
```
|
| 63 |
|
| 64 |
+
**Usage Example**:
|
| 65 |
+
The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework, following the examples [here](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
|
| 67 |
|
| 68 |
## Training Details
|
|
|
|
| 80 |
|
| 81 |
Our model was evaluated against several existing benchmarks to measure its ability to determine whether document is relevant or not to the last user question in a multi-turn conversation.
|
| 82 |
|
| 83 |
+
### Forming evaluation datasets for context relevance
|
| 84 |
|
| 85 |
For each of the datasets, we created pairs of document and question pairs and generated a label (based on the available labels in the dataset) as to whether the document is relevant or useful to answering the question.
|
| 86 |
|
|
|
|
| 91 |
|
| 92 |
### Evaluation Datasets
|
| 93 |
|
| 94 |
+
We evaluated on the following five datasets:
|
| 95 |
|
| 96 |
+
- [CLAPNQ](https://arxiv.org/abs/2404.02103)
|
| 97 |
+
- [RAG MINI BIOASQ](https://huggingface.co/datasets/enelpol/rag-mini-bioasq)
|
| 98 |
+
- [Open Australian Legal QA](https://huggingface.co/datasets/isaacus/open-australian-legal-qa)
|
| 99 |
+
- [FinanceBench](https://arxiv.org/abs/2311.11944)
|
| 100 |
+
- [UCINLP DROP](https://aclanthology.org/N19-1246/)
|
| 101 |
|
|
|
|
|
|
|
| 102 |
|
| 103 |
+
### Evaluation Results
|
|
|
|
|
|
|
|
|
|
|
|
|
| 104 |
|
| 105 |
+
We evaluated the LoRA adapters against various frontier models on the five evaluation datasets. Here, we create strong prompt instruction (three labels) for frontier models, but no instruction for LoRA adapters.
|
| 106 |
+
The results show average F1 scores for relevant and irrelevant classifications, along with overall accuracy.
|
| 107 |
|
| 108 |
+
| Model | Avg. F1 (Relevant) | Avg. F1 (Irrelevant) | Accuracy |
|
| 109 |
+
|-------|-------------------|---------------------|----------|
|
| 110 |
+
| **Frontier Models (Out of Box)** | | | |
|
| 111 |
+
| Granite 3.3-8b Instruct | 94.9% | 92.3% | 94.7% |
|
| 112 |
+
| Granite 4 micro | 93.8% | 88.6% | 93.1% |
|
| 113 |
+
| GPT-OSS-20b | 90.1% | 85.7% | 90.0% |
|
| 114 |
+
| GPT-OSS-120b | 93.7% | 91.4% | 93.6% |
|
| 115 |
+
| GPT 4o mini | 93.7% | 90.8% | 93.6% |
|
| 116 |
+
| GPT 4o | 95.6% | 93.9% | 95.7% |
|
| 117 |
+
| GPT 5 mini | 95.8% | 93.9% | 95.6% |
|
| 118 |
+
| **Trained LoRAs** | | | |
|
| 119 |
+
| Granite 4 micro-context-relevance-LoRA | 93.2% | 87.1% | 92.1% |
|
| 120 |
+
| Granite 4 micro-context-relevance-aLoRA | 87.0% | 72.8% | 83.4% |
|
| 121 |
|
| 122 |
+
**Takeaways:**
|
| 123 |
+
- LoRA without instruction is comparable to frontier models with strong instruction
|
| 124 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 125 |
|
| 126 |
## Contact
|
| 127 |
|
| 128 |
+
[Lihong He](mailto:lihong.he@ibm.com)
|
| 129 |
|
| 130 |
### Framework versions
|
| 131 |
|