Granite Library
Safetensors
GGUF
English

Update answerability model card with granite 4 micro lora/alora eval

#12
Files changed (1) hide show
  1. answerability/answerability.md +143 -0
answerability/answerability.md ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ pipeline_tag: text-generation
6
+ library_name: peft
7
+ library_name: transformers
8
+ ---
9
+
10
+ # Intrinsics for Answerability Classification
11
+
12
+ ## Model Summary
13
+ This is a RAG-specific family of intrinsics fine-tuned for binary answerability
14
+ classification task. The model takes as input a multi-turn conversation and a
15
+ set of documents, and classifies whether the user's final query is answerable or
16
+ unanswerable based on the available information in the documents.
17
+
18
+ We provide answerability intrinsics implemented as LoRA adapters trained over
19
+ Granite-4.0-micro and GPT-OSS 20b.
20
+
21
+ - **Developer:** IBM Research
22
+ - **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granite-4.0-micro) and [openai/gpt-oss-20b](https://huggingface.co/openai/gpt-oss-20b)
23
+ - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
24
+
25
+ ## Intended use
26
+ This is a family of intrinsincs that enables answerability classification for
27
+ the final user query in a multi-turn conversation, with respect to a set of
28
+ provided documents. The model is trained to determine whether the last user
29
+ query is answerable or unanswerable, based solely on the information present in
30
+ the documents. This makes it suitable for applications involving RAG and
31
+ document-grounded chatbots, where knowing whether sufficient information exists
32
+ to answer a query is crucial. The classification output from the answerability
33
+ model can be used in several downstream applications, including but not limited
34
+ to:
35
+ - Filter out unanswerable questions before sending them to generation in RAG
36
+ setting. By classifying a query as unanswerable upfront, the system can prevent
37
+ hallucinated or misleading responses.
38
+ - Re-query the retriever to get more
39
+ relevant documents. If a query is initially deemed unanswerable, the retriever
40
+ can be re-invoked with alternate formulations to fetch more relevant documents.
41
+
42
+ **Intrinsic input**: The input to the answerability intrinsic is an
43
+ OpenAI-compatible chat completion request, containing a list of conversation
44
+ turns that can alternate between the `user` and `assistant` role and ending with
45
+ a `user` turn, as well as list of documents.
46
+
47
+ **Intrinsic output**: The output of the answerability intrinsic is the result of the
48
+ original chat completion request formatted as a JSON object as follows:
49
+ ```json
50
+ {
51
+ "answerability_likelihood": <float>
52
+ }
53
+ ```
54
+
55
+ ### Example
56
+
57
+ **Input conversation:**
58
+
59
+ | Role | Message |
60
+ |------|---------|
61
+ | assistant | Hello there, how can I help you? |
62
+ | user | What is the square root of 4? |
63
+
64
+ **Input documents (answerable case):**
65
+ - Document 1: "The square root of 4 is 2."
66
+
67
+ **Output (answerable):**
68
+ ```json
69
+ {
70
+ "answerability_likelihood": 0.9999646429576308
71
+ }
72
+ ```
73
+
74
+ **Input documents (unanswerable case):**
75
+ - Document 1: "The square root of 8 is not 2."
76
+
77
+ **Output (unanswerable):**
78
+ ```json
79
+ {
80
+ "answerability_likelihood": 0.0001234567890123
81
+ }
82
+ ```
83
+
84
+ ## Usage Examples
85
+
86
+ The recommended way to call this intrinsic is through the [Mellea](https://mellea.ai) framework. For detailed examples on how to use this and other intrinsics, please refer to the [Mellea intrinsics examples](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
87
+
88
+
89
+ ## Training Details
90
+
91
+ ### Training Data
92
+
93
+ The training data uses the publicly available Government corpus from
94
+ [MT-RAG](https://arxiv.org/pdf/2501.03468) as the source of documents. Based on
95
+ this corpus, we constructed a dataset consisting of a mix of human-created and
96
+ synthetically generated multi-turn conversations. It includes two types of
97
+ examples: (1) Answerable queries, where the final user question can be answered
98
+ based on the provided documents. These examples teach the adapter to recognize
99
+ when sufficient information is present to support an answer. (2) Unanswerable
100
+ queries, where the documents lack the necessary information to answer the final
101
+ user query. We used Mixtral as an automatic judge to validate the answerability
102
+ labels and filter out noisy samples.
103
+
104
+ #### Training Hyperparameters
105
+
106
+ The LoRA adapter was fine-tuned using PEFT under the following regime: rank =
107
+ 32, learning rate = 5e-6, number of epochs = 25, with early stopping based on
108
+ validation set, and 90/10 split between training and validation.
109
+
110
+ ## Evaluation: Answerability Classification
111
+
112
+ We evaluated the model on binary answerability classification using MT-RAG
113
+ Benchmark. In this setting, the model is given the full multi-turn conversation
114
+ history along with the supporting documents. This benchmark evaluates the
115
+ model's ability to assess answerability when the final user query can also
116
+ depend on prior turns for context. The following table presents results
117
+ comparing baselines and frontier models with task-specific answerability
118
+ intrinsics on the answerability classification task on MT-RAG data. The LoRAs
119
+ consistently outperform frontier models, converging near \~90% accuracy
120
+ regardless of base model size. Even small models like Granite 4.0-micro, once
121
+ fine-tuned, match or surpass much larger models, including GPT-4o.
122
+
123
+ | | Models | Unanswerable F1 | Answerable F1 | Classification Accuracy | Weighted F1 |
124
+ |:--------------------------------------------:|:----------------------------------------------:|:--------------------------:|:---------------------------:|:-------------------------------------:|:-------------------------:|
125
+ | Baselines | BigBird (pre-trained embeddings) w/ MLP | 73.4 | 65.2 | 69.8 | 69.6 |
126
+ | | llama2-7b as classifier (Full SFT) | 88.2 | 85.9 | 87.1 | 87.1 |
127
+ | Frontier Models out-of-the-box | GPT-OSS-20b | 77.3 | 58.3 | 70.7 | 68.5 |
128
+ | | GPT-OSS-120b | 70.2 | 68.9 | 69.8 | 69.6 |
129
+ | | GPT4o-mini | 82.7 | 78.1 | 80.8 | 80.6 |
130
+ | | GPT4o | 85.7 | 77.5 | 82.5 | 81.9 |
131
+ | Trained LoRAs/aLoRAs | Granite 4.0-micro LoRA | 90.9 | 90.0 | 90.4 | 90.5 |
132
+ | | GPT-OSS-20b LoRA | 91.6 | 89.8 | 90.8 | 90.8 |
133
+ | | Granite 4.0-micro aLoRA | 90.0 | 89.4 | 89.6 | 89.7 |
134
+ | | GPT-OSS-20b aLoRA | 90.4 | 88.6 | 89.6 | 89.6 |
135
+
136
+
137
+ ## Model Card Authors
138
+
139
+ [Vraj Shah](mailto:vraj@ibm.com)
140
+
141
+ ### Framework versions
142
+
143
+ - PEFT 0.14.0