Granite Library
Safetensors
GGUF
English
lucianpopa commited on
Commit
6d1bf67
·
verified ·
1 Parent(s): c021f94

Update query_clarification/README.md

Browse files
Files changed (1) hide show
  1. query_clarification/README.md +25 -24
query_clarification/README.md CHANGED
@@ -11,7 +11,7 @@ library_name: transformers
11
 
12
  ## Model Summary
13
 
14
- **Query Clarification** is an adpater designed for conversational use cases where user queries may be ill-formed, unclear, or have multiple valid interpretations based on the underlying system or content. Use cases include RAG systems, routing, and tool invocation scenarios. We provide experimental results showing that the adapter achieves significantly higher classification accuracy than prompted out-of-the-box models, including frontier models such as gpt-4o. The adapter released here works with the IBM granite-4.0-micro model and is specifically fine-tuned for the following task
15
 
16
  Given a multi-turn conversation between a user and an AI assistant (and optionally relevant content such as
17
  RAG documents), detect whether the last user query is underspecified (no clear interpretation or multiple
@@ -20,7 +20,7 @@ library_name: transformers
20
  - **Developer:** IBM Research
21
  - **HF Collection:** [Granite Libraries](https://huggingface.co/collections/ibm-granite/granite-libraries)
22
  - **GitHub Repository:** https://github.com/ibm-granite
23
- - **Release Date:** March 18th, 2026
24
  - **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granite-4.0-micro)
25
  - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
26
 
@@ -68,7 +68,7 @@ In this example, the adapter detects that the user's question is underspecified
68
 
69
  ## Usage Examples
70
 
71
- The recommended way to call this adapter is through the [Mellea](https://mellea.ai) framework. For detailed examples on how to use this and other adapter, please refer to the [Mellea examples](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
72
 
73
 
74
  ## Training Details
@@ -81,6 +81,7 @@ The training dataset consists of:
81
  - Human-generated underspecified questions created by annotators over content from Gov and ClapNQ domains
82
  - Multi-turn conversations with underspecified final turns
83
  - Synthetically generated clarification requests using target patterns (hedging with answers, hedging list, open domain)
 
84
 
85
  The training dataset is proprietary and was obtained in combination with a third-party company who contracted the human annotators.
86
 
@@ -105,9 +106,9 @@ We evaluate the adapter on two dimensions: (1) classification accuracy in detect
105
  ### Test Set
106
 
107
  The evaluation uses a test set containing:
108
- - 80 positive examples: conversations where the last user question is underspecified
109
- - 150 random negatives: conversations where the last user question is clear
110
- - 50 hard negatives: conversations intended to be underspecified but adjudicated as clear by humans
111
 
112
  ### Classification Results
113
 
@@ -115,39 +116,39 @@ The evaluation uses a test set containing:
115
 
116
  **Trained LoRA:**
117
 
118
- | Model | Underspecified (80) | CLEAR - random (150) | CLEAR - hard (50) | Overall Accuracy |
119
  |-------|---------------------|----------------------|-------------------|------------------|
120
- | Granite-4.0-micro LoRA | 76.25% | 99.33% | 80% | **89.3%** |
121
 
122
  **OOB models (prompted, zero-shot):**
123
 
124
- | Model | Underspecified (80) | CLEAR - random (150) | CLEAR - hard (50) | Overall Accuracy |
125
  |-------|---------------------|----------------------|-------------------|------------------|
126
- | gpt-oss-20b | 85% | 78% | 62% | 77.14% |
127
- | GPT 4o | 60% | 85.3% | 48% | 71.42% |
128
- | Granite-4.0-micro | 97.5% | 5.3% | 0% | 30.71% |
129
 
130
- The LoRA achieves higher overall accuracy due to its ability to correctly identify clear queries, while prompted models tend to over-clarify.
131
 
132
  ### Clarification Quality Results
133
 
134
- For queries identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The numbers in parentheses indicate how many of the 80 underspecified records each model correctly identified and generated clarifications for—quality metrics are measured only on those data points.
135
 
136
  **Trained LoRA:**
137
 
138
  | Model | Precision | Recall | F1 |
139
  |-------|-----------|--------|-----|
140
- | Granite-4.0-micro LoRA (61/80) | 90.4% | 80.1% | 84.9% |
141
 
142
  **OOB models (prompted, zero-shot):**
143
 
144
  | Model | Precision | Recall | F1 |
145
  |-------|-----------|--------|-----|
146
- | GPT 4o (48/80) | 91.9% | 83.4% | 87.4% |
147
- | gpt-oss-20b (68/80) | 86.9% | 66.9% | 75.6% |
148
- | Granite-4.0-micro (78/80) | 74.9% | 58.0% | 65.4% |
149
 
150
- The LoRA achieves clarification quality close to GPT-4o while maintaining much better classification accuracy and evaluating on more data points.
151
 
152
 
153
  ### Adapter Details
@@ -161,16 +162,16 @@ The LoRA achieves clarification quality close to GPT-4o while maintaining much b
161
  | **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
162
 
163
  **Infrastructure:**
164
- We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 8 H100 GPUs.
165
 
166
  **Ethical Considerations & Limitations:**
167
  The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
168
 
169
  ## Resources
170
 
171
- - ⭐️ Learn about the latest updates with Granite: https://www.ibm.com/granite
172
- - 📄 Get started with tutorials, best practices, and prompt engineering advice: https://www.ibm.com/granite/docs/
173
- - 💡 Learn about the latest Granite learning resources: https://github.com/ibm-granite/granite-guardian/tree/main/cookbooks
174
 
175
  ## Contact
176
- [Lucian Popa](mailto:lpopa@us.ibm.com)
 
11
 
12
  ## Model Summary
13
 
14
+ **Query Clarification** is an adapter designed for conversational use cases where user queries may be ill-formed, unclear, or have multiple valid interpretations based on the underlying system or content. Use cases include RAG systems, routing, and tool invocation scenarios. We provide experimental results showing that the adapter achieves significantly higher classification accuracy than prompted out-of-the-box models, including frontier models such as gpt-4o. The adapter released here works with the IBM granite-4.0-micro model and is specifically fine-tuned for the following task:
15
 
16
  Given a multi-turn conversation between a user and an AI assistant (and optionally relevant content such as
17
  RAG documents), detect whether the last user query is underspecified (no clear interpretation or multiple
 
20
  - **Developer:** IBM Research
21
  - **HF Collection:** [Granite Libraries](https://huggingface.co/collections/ibm-granite/granite-libraries)
22
  - **GitHub Repository:** https://github.com/ibm-granite
23
+ - **Release Date:** April 2026
24
  - **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granite-4.0-micro)
25
  - **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
26
 
 
68
 
69
  ## Usage Examples
70
 
71
+ The recommended way to call this adapter is through the [Mellea](https://mellea.ai) framework. For detailed examples on how to use this and other adapters, please refer to the [Mellea examples](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
72
 
73
 
74
  ## Training Details
 
81
  - Human-generated underspecified questions created by annotators over content from Gov and ClapNQ domains
82
  - Multi-turn conversations with underspecified final turns
83
  - Synthetically generated clarification requests using target patterns (hedging with answers, hedging list, open domain)
84
+ - Conversation-only examples (without document passages) targeting pre-retrieval scenarios
85
 
86
  The training dataset is proprietary and was obtained in combination with a third-party company who contracted the human annotators.
87
 
 
106
  ### Test Set
107
 
108
  The evaluation uses a test set containing:
109
+ - 104 positive examples: conversations where the last user question is underspecified (generated by a combination of humans and AI model -- mistralai/Mistral-Large-3-675B-Instruct-2512)
110
+ - 200 random negatives: randomly sampled, pre-existing conversations where the last user question is clear
111
+ - 79 hard negatives: conversations where the last user question was intended to be underspecified (generated by an AI model -- mistralai/Mistral-Large-3-675B-Instruct-2512) but adjudicated as clear by humans
112
 
113
  ### Classification Results
114
 
 
116
 
117
  **Trained LoRA:**
118
 
119
+ | Model | Underspecified (104) | CLEAR - random (200) | CLEAR - hard (79) | Overall Accuracy |
120
  |-------|---------------------|----------------------|-------------------|------------------|
121
+ | Granite-4.0-micro LoRA | **84.6%** | **98.5%** | **87.3%** | **92.4%** |
122
 
123
  **OOB models (prompted, zero-shot):**
124
 
125
+ | Model | Underspecified (104) | CLEAR - random (200) | CLEAR - hard (79) | Overall Accuracy |
126
  |-------|---------------------|----------------------|-------------------|------------------|
127
+ | gpt-4o | 48.1% | 90.5% | 92.4% | 79.4% |
128
+ | gpt-oss-20b | 90.4% | 72.5% | 58.2% | 74.4% |
129
+ | Granite-4.0-micro | 25.0% | 85.0% | 63.3% | 64.2% |
130
 
131
+ The LoRA achieves the highest overall accuracy by balancing high recall on underspecified queries with strong specificity on clear queries. Prompted models tend to err in one direction: gpt-oss-20b over-clarifies (high underspecified recall but low CLEAR accuracy), while gpt-4o and base Granite under-clarify (low underspecified recall).
132
 
133
  ### Clarification Quality Results
134
 
135
+ For queries correctly identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The judge compares the entities and options surfaced in the generated clarification against those in the gold-standard reference: precision measures what fraction of the generated options are correct, and recall measures what fraction of the gold options are covered. Additionally, the numbers in parentheses (near the model name) indicate how many of the 104 underspecified records each model correctly identified and generated clarifications for -- the quality metrics are measured only on those data points.
136
 
137
  **Trained LoRA:**
138
 
139
  | Model | Precision | Recall | F1 |
140
  |-------|-----------|--------|-----|
141
+ | Granite-4.0-micro LoRA (88/104) | **91.7%** | **87.5%** | **89.5%** |
142
 
143
  **OOB models (prompted, zero-shot):**
144
 
145
  | Model | Precision | Recall | F1 |
146
  |-------|-----------|--------|-----|
147
+ | gpt-oss-20b (93/104) | 93.3% | 70.7% | 80.5% |
148
+ | gpt-4o (50/104) | 81.3% | 69.8% | 75.1% |
149
+ | Granite-4.0-micro (26/104) | 75.0% | 45.7% | 56.8% |
150
 
151
+ The LoRA achieves the highest F1 (89.5%) across the largest number of correctly classified records (88/104), demonstrating both strong classification and high generation quality.
152
 
153
 
154
  ### Adapter Details
 
162
  | **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
163
 
164
  **Infrastructure:**
165
+ We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 1 H100 GPU.
166
 
167
  **Ethical Considerations & Limitations:**
168
  The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
169
 
170
  ## Resources
171
 
172
+ - Learn about the latest updates with Granite: https://www.ibm.com/granite
173
+ - Get started with tutorials, best practices, and prompt engineering advice: https://www.ibm.com/granite/docs/
174
+ - Learn about the latest Granite learning resources: https://github.com/ibm-granite/granite-guardian/tree/main/cookbooks
175
 
176
  ## Contact
177
+ [Lucian Popa](mailto:lpopa@us.ibm.com)