Instructions to use ibm-granite/granitelib-rag-gpt-oss-r1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Granite Library
How to use ibm-granite/granitelib-rag-gpt-oss-r1.0 with Granite Library:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
|
Download answerability/README.md from ibm-granite/granitelib-rag-gpt-oss-r1.0: direct link, hf CLI and curl.
- Browser
- Download file 11.8 kB
-
https://huggingface.co/ibm-granite/granitelib-rag-gpt-oss-r1.0/resolve/323ed628c3feba319c4a58219402155a3c4e4cd3/answerability/README.md
- Command line
-
hf download hf://ibm-granite/granitelib-rag-gpt-oss-r1.0@323ed628c3feba319c4a58219402155a3c4e4cd3/answerability/README.md
-
curl -L -o README.md https://huggingface.co/ibm-granite/granitelib-rag-gpt-oss-r1.0/resolve/323ed628c3feba319c4a58219402155a3c4e4cd3/answerability/README.md
11.8 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| library_name: peft | |
| library_name: transformers | |
| # Intrinsics for Answerability Classification | |
| ## Model Summary | |
| This is a RAG-specific family of intrinsics fine-tuned for binary answerability | |
| classification task. The model takes as input a multi-turn conversation and a | |
| set of documents, and classifies whether the user's final query is answerable or | |
| unanswerable based on the available information in the documents. | |
| We provide two intrinsics implemented as LoRA adapters (LoRA/aLoRA) trained over | |
| Granite-3.3-2b-instruct, Granite-3.3-8b-instruct, and GPT-OSS 20b. | |
| - **Developer:** IBM Research | |
| - **Model type:** LoRA and aLoRA adapter for | |
| [ibm-granite/granite-3.3-2b-instruct](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct), | |
| [ibm-granite/granite-3.3-8b-instruct](https://huggingface.co/ibm-granite/granite-3.3-8b-instruct), | |
| and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b) | |
| - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) | |
| ## Intended use | |
| This is a family of intrinsincs that enables answerability classification for | |
| the final user query in a multi-turn conversation, with respect to a set of | |
| provided documents. The model is trained to determine whether the last user | |
| query is answerable or unanswerable, based solely on the information present in | |
| the documents. This makes it suitable for applications involving RAG and | |
| document-grounded chatbots, where knowing whether sufficient information exists | |
| to answer a query is crucial. The classification output from the answerability | |
| model can be used in several downstream applications, including but not limited | |
| to: | |
| - Filter out unanswerable questions before sending them to generation in RAG | |
| setting. By classifying a query as unanswerable upfront, the system can prevent | |
| hallucinated or misleading responses. | |
| - Re-query the retriever to get more | |
| relevant documents. If a query is initially deemed unanswerable, the retriever | |
| can be re-invoked with alternate formulations to fetch more relevant documents. | |
| **Model input**: The input to the answerability intrinsic is an | |
| OpenAI-compatible chat completion request, containing a list of conversation | |
| turns that can alternate between the `user` and `assistant` role and ending with | |
| a `user` turn, as well as list of documents. | |
| **Model output**: The output of the answerability intrinsic is the result of the | |
| original chat completion request formatted as a JSON object containing the | |
| answerability likelihood score. | |
| Please see the code snippets in the Quickstart Example section below for | |
| examples that illustrate the intrinsic's input/output. | |
| ## Quickstart Example | |
| The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework. | |
| Here is some example code for calling this intrinsic from Mellea: | |
| ``` | |
| from mellea.backends.huggingface import LocalHFBackend | |
| from mellea.stdlib.base import ChatContext, Document | |
| from mellea.stdlib.chat import Message | |
| from mellea.stdlib.intrinsics import rag | |
| backend = LocalHFBackend(model_id="ibm-granite/granite-3.3-2b-instruct") | |
| context = ChatContext().add(Message("assistant", "Hello there, how can I help you?")) | |
| next_user_turn = "What is the square root of 4?" | |
| documents_answerable = [Document("The square root of 4 is 2.")] | |
| documents_unanswerable = [Document("The square root of 8 is not 2.")] | |
| result = rag.check_answerability(next_user_turn, documents_answerable, context, backend) | |
| print(f"Result of answerability check when answer is in documents: {result}") | |
| result = rag.check_answerability( | |
| next_user_turn, documents_unanswerable, context, backend | |
| ) | |
| print(f"Result of answerability check when answer is not in documents: {result}") | |
| ``` | |
| ## Training Details | |
| ### Training Data | |
| The training data uses the publicly available Government corpus from | |
| [MT-RAG](https://arxiv.org/pdf/2501.03468) as the source of documents. Based on | |
| this corpus, we constructed a dataset consisting of a mix of human-created and | |
| synthetically generated multi-turn conversations. It includes two types of | |
| examples: (1) Answerable queries, where the final user question can be answered | |
| based on the provided documents. These examples teach the adapter to recognize | |
| when sufficient information is present to support an answer. (2) Unanswerable | |
| queries, where the documents lack the necessary information to answer the final | |
| user query. We used Mixtral as an automatic judge to validate the answerability | |
| labels and filter out noisy samples. | |
| #### Training Hyperparameters | |
| The LoRA adapter was fine-tuned using PEFT under the following regime: rank = | |
| 32, learning rate = 5e-6, number of epochs = 25, with early stopping based on | |
| validation set, and 90/10 split between training and validation. | |
| ## Evaluation | |
| ### Answerability Classification | |
| We evaluated the model on binary answerability classification using MT-RAG | |
| Benchmark. In this setting, the model is given the full multi-turn conversation | |
| history along with the supporting documents. This benchmark evaluates the | |
| model's ability to assess answerability when the final user query can also | |
| depend on prior turns for context. The following table presents results | |
| comparing baselines and frontier models with task-specific answerability | |
| intrinsics on the answerability classification task on MT-RAG data. The LoRAs | |
| consistently outperform frontier models, converging near \~90% accuracy | |
| regardless of base model size. Even small models like Granite 3.3-2B, once | |
| fine-tuned, match or surpass much larger models, including GPT-4o. The | |
| difference between LoRA and aLoRA is minimal, indicating both are effective | |
| fine-tuning strategies. | |
| | | Models | Unanswerable F1 | Answerable F1 | Classification Accuracy | Weighted F1 | | |
| |:--------------------------------------------:|:----------------------------------------------:|:--------------------------:|:---------------------------:|:-------------------------------------:|:-------------------------:| | |
| | Baselines | BigBird (pre-trained embeddings) w/ MLP | 73.4 | 65.2 | 69.8 | 69.6 | | |
| | | llama2-7b as classifier (Full SFT) | 88.2 | 85.9 | 87.1 | 87.1 | | |
| | Frontier Models out-of-the-box | Granite 3.3-2b-instruct | 48.7 | 70.4 | 62.4 | 58.7 | | |
| | | Granite 3.3-8b-instruct | 62.8 | 65.2 | 64.5 | 63.9 | | |
| | | GPT-OSS-20b | 77.3 | 58.3 | 70.7 | 68.5 | | |
| | | GPT-OSS-120b | 70.2 | 68.9 | 69.8 | 69.6 | | |
| | | GPT4o-mini | 82.7 | 78.1 | 80.8 | 80.6 | | |
| | | GPT4o | 85.7 | 77.5 | 82.5 | 81.9 | | |
| | Trained LoRAs/aLoRAs | Granite 3.3-2b LoRA | 91.2 | 89.6 | 90.4 | 90.5 | | |
| | | Granite 3.3-8b LoRA | 91.1 | 90.3 | 90.6 | 90.7 | | |
| | | GPT-OSS-20b LoRA | 91.6 | 89.8 | 90.8 | 90.8 | | |
| | | Granite 3.3-2b aLoRA | 89.8 | 88.6 | 89.1 | 89.2 | | |
| | | Granite 3.3-8b aLoRA | 90.1 | 89.6 | 89.5 | 89.9 | | |
| | | GPT-OSS-20b aLoRA | 90.4 | 88.6 | 89.6 | 89.6 | | |
| ### Comparing the Answerability Intrinsics vs. Vanilla Granite Models for Answer Quality | |
| We compare the performance of Granite 3.3-2b, Granite 3.3-8b Instruct | |
| vs. answerability intrinsics implemented as LoRA adapters on a subset of MT-RAG | |
| Benchmark. In this setup, each query is paired with only 5 retrieved passages as | |
| context. | |
| - Answerability Classification Performance: The answerability intrinsics | |
| outperform the vanilla model in overall F1 on both answerables and | |
| unanswerables. The answerability intrinsics achieves higher recall on | |
| unanswerable queries, making it better at identifying questions that should | |
| not be answered. However, this comes at the cost of lower recall on answerable | |
| queries. | |
| - Joint Answerability-Faithfulness Score computed as: \> = 1 (if model | |
| prediction = IDK/unanswerable ∩ ground truth = unanswerable) | |
| > = RAGAS Faithfulness (if model prediction = non-IDK/answerable ∩ ground | |
| > truth = answerable) | |
| > = 0 (otherwise) | |
| This score rewards the model for correctly abstaining on unanswerable queries | |
| (full credit) and for providing faithful answers on answerable queries | |
| (partial credit based on RAGAS Faithfulness). No credit is given for incorrect | |
| or unfaithful predictions. | |
| The answerability intrinsics for granite-2b and granite-8b achieves 8% and 13% | |
| lifts on this metric respectively. This rewards the model for correctly | |
| abstaining on unanswerable queries and for being faithful when it chooses to | |
| answer. | |
| | | F1 Score Unanswerable | F1 Score Answerable | Recall Unanswerable | Recall Answerable | Joint Answerability- Faithfulness Score | | |
| |:-----------------------:|:---------------------:|:-------------------:|:-------------------:|:-----------------:|:---------------------------------------:| | |
| | Granite 3.3-2b Instruct | 13 | 77 | 7 | 99 | 48 | | |
| | Granite 3.3-2b LoRA | 48 | 78 | 37 | 89 | 56 | | |
| | Granite 3.3-8b Instruct | 17 | 77 | 10 | 99 | 49 | | |
| | Granite 3.3-8b LoRA | 65 | 81 | 60 | 86 | 62 | | |
| ## Model Card Authors | |
| [Vraj Shah](mailto:vraj@ibm.com) | |
| ### Framework versions | |
| - PEFT 0.14.0 | |