frreiss's picture
Initial import
323ed62
|
Raw History Blame
11.8 kB
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: peft
library_name: transformers
---
# Intrinsics for Answerability Classification
## Model Summary
This is a RAG-specific family of intrinsics fine-tuned for binary answerability
classification task. The model takes as input a multi-turn conversation and a
set of documents, and classifies whether the user's final query is answerable or
unanswerable based on the available information in the documents.
We provide two intrinsics implemented as LoRA adapters (LoRA/aLoRA) trained over
Granite-3.3-2b-instruct, Granite-3.3-8b-instruct, and GPT-OSS 20b.
- **Developer:** IBM Research
- **Model type:** LoRA and aLoRA adapter for
[ibm-granite/granite-3.3-2b-instruct](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct),
[ibm-granite/granite-3.3-8b-instruct](https://huggingface.co/ibm-granite/granite-3.3-8b-instruct),
and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
## Intended use
This is a family of intrinsincs that enables answerability classification for
the final user query in a multi-turn conversation, with respect to a set of
provided documents. The model is trained to determine whether the last user
query is answerable or unanswerable, based solely on the information present in
the documents. This makes it suitable for applications involving RAG and
document-grounded chatbots, where knowing whether sufficient information exists
to answer a query is crucial. The classification output from the answerability
model can be used in several downstream applications, including but not limited
to:
- Filter out unanswerable questions before sending them to generation in RAG
setting. By classifying a query as unanswerable upfront, the system can prevent
hallucinated or misleading responses.
- Re-query the retriever to get more
relevant documents. If a query is initially deemed unanswerable, the retriever
can be re-invoked with alternate formulations to fetch more relevant documents.
**Model input**: The input to the answerability intrinsic is an
OpenAI-compatible chat completion request, containing a list of conversation
turns that can alternate between the `user` and `assistant` role and ending with
a `user` turn, as well as list of documents.
**Model output**: The output of the answerability intrinsic is the result of the
original chat completion request formatted as a JSON object containing the
answerability likelihood score.
Please see the code snippets in the Quickstart Example section below for
examples that illustrate the intrinsic's input/output.
## Quickstart Example
The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework.
Here is some example code for calling this intrinsic from Mellea:
```
from mellea.backends.huggingface import LocalHFBackend
from mellea.stdlib.base import ChatContext, Document
from mellea.stdlib.chat import Message
from mellea.stdlib.intrinsics import rag
backend = LocalHFBackend(model_id="ibm-granite/granite-3.3-2b-instruct")
context = ChatContext().add(Message("assistant", "Hello there, how can I help you?"))
next_user_turn = "What is the square root of 4?"
documents_answerable = [Document("The square root of 4 is 2.")]
documents_unanswerable = [Document("The square root of 8 is not 2.")]
result = rag.check_answerability(next_user_turn, documents_answerable, context, backend)
print(f"Result of answerability check when answer is in documents: {result}")
result = rag.check_answerability(
next_user_turn, documents_unanswerable, context, backend
)
print(f"Result of answerability check when answer is not in documents: {result}")
```
## Training Details
### Training Data
The training data uses the publicly available Government corpus from
[MT-RAG](https://arxiv.org/pdf/2501.03468) as the source of documents. Based on
this corpus, we constructed a dataset consisting of a mix of human-created and
synthetically generated multi-turn conversations. It includes two types of
examples: (1) Answerable queries, where the final user question can be answered
based on the provided documents. These examples teach the adapter to recognize
when sufficient information is present to support an answer. (2) Unanswerable
queries, where the documents lack the necessary information to answer the final
user query. We used Mixtral as an automatic judge to validate the answerability
labels and filter out noisy samples.
#### Training Hyperparameters
The LoRA adapter was fine-tuned using PEFT under the following regime: rank =
32, learning rate = 5e-6, number of epochs = 25, with early stopping based on
validation set, and 90/10 split between training and validation.
## Evaluation
### Answerability Classification
We evaluated the model on binary answerability classification using MT-RAG
Benchmark. In this setting, the model is given the full multi-turn conversation
history along with the supporting documents. This benchmark evaluates the
model's ability to assess answerability when the final user query can also
depend on prior turns for context. The following table presents results
comparing baselines and frontier models with task-specific answerability
intrinsics on the answerability classification task on MT-RAG data. The LoRAs
consistently outperform frontier models, converging near \~90% accuracy
regardless of base model size. Even small models like Granite 3.3-2B, once
fine-tuned, match or surpass much larger models, including GPT-4o. The
difference between LoRA and aLoRA is minimal, indicating both are effective
fine-tuning strategies.
| | Models | Unanswerable F1 | Answerable F1 | Classification Accuracy | Weighted F1 |
|:--------------------------------------------:|:----------------------------------------------:|:--------------------------:|:---------------------------:|:-------------------------------------:|:-------------------------:|
| Baselines | BigBird (pre-trained embeddings) w/ MLP | 73.4 | 65.2 | 69.8 | 69.6 |
| | llama2-7b as classifier (Full SFT) | 88.2 | 85.9 | 87.1 | 87.1 |
| Frontier Models out-of-the-box | Granite 3.3-2b-instruct | 48.7 | 70.4 | 62.4 | 58.7 |
| | Granite 3.3-8b-instruct | 62.8 | 65.2 | 64.5 | 63.9 |
| | GPT-OSS-20b | 77.3 | 58.3 | 70.7 | 68.5 |
| | GPT-OSS-120b | 70.2 | 68.9 | 69.8 | 69.6 |
| | GPT4o-mini | 82.7 | 78.1 | 80.8 | 80.6 |
| | GPT4o | 85.7 | 77.5 | 82.5 | 81.9 |
| Trained LoRAs/aLoRAs | Granite 3.3-2b LoRA | 91.2 | 89.6 | 90.4 | 90.5 |
| | Granite 3.3-8b LoRA | 91.1 | 90.3 | 90.6 | 90.7 |
| | GPT-OSS-20b LoRA | 91.6 | 89.8 | 90.8 | 90.8 |
| | Granite 3.3-2b aLoRA | 89.8 | 88.6 | 89.1 | 89.2 |
| | Granite 3.3-8b aLoRA | 90.1 | 89.6 | 89.5 | 89.9 |
| | GPT-OSS-20b aLoRA | 90.4 | 88.6 | 89.6 | 89.6 |
### Comparing the Answerability Intrinsics vs. Vanilla Granite Models for Answer Quality
We compare the performance of Granite 3.3-2b, Granite 3.3-8b Instruct
vs. answerability intrinsics implemented as LoRA adapters on a subset of MT-RAG
Benchmark. In this setup, each query is paired with only 5 retrieved passages as
context.
- Answerability Classification Performance: The answerability intrinsics
outperform the vanilla model in overall F1 on both answerables and
unanswerables. The answerability intrinsics achieves higher recall on
unanswerable queries, making it better at identifying questions that should
not be answered. However, this comes at the cost of lower recall on answerable
queries.
- Joint Answerability-Faithfulness Score computed as: \> = 1 (if model
prediction = IDK/unanswerable ∩ ground truth = unanswerable)
> = RAGAS Faithfulness (if model prediction = non-IDK/answerable ∩ ground
> truth = answerable)
> = 0 (otherwise)
This score rewards the model for correctly abstaining on unanswerable queries
(full credit) and for providing faithful answers on answerable queries
(partial credit based on RAGAS Faithfulness). No credit is given for incorrect
or unfaithful predictions.
The answerability intrinsics for granite-2b and granite-8b achieves 8% and 13%
lifts on this metric respectively. This rewards the model for correctly
abstaining on unanswerable queries and for being faithful when it chooses to
answer.
| | F1 Score Unanswerable | F1 Score Answerable | Recall Unanswerable | Recall Answerable | Joint Answerability- Faithfulness Score |
|:-----------------------:|:---------------------:|:-------------------:|:-------------------:|:-----------------:|:---------------------------------------:|
| Granite 3.3-2b Instruct | 13 | 77 | 7 | 99 | 48 |
| Granite 3.3-2b LoRA | 48 | 78 | 37 | 89 | 56 |
| Granite 3.3-8b Instruct | 17 | 77 | 10 | 99 | 49 |
| Granite 3.3-8b LoRA | 65 | 81 | 60 | 86 | 62 |
## Model Card Authors
[Vraj Shah](mailto:vraj@ibm.com)
### Framework versions
- PEFT 0.14.0