Update query_clarification/README.md
Browse files- query_clarification/README.md +25 -24
query_clarification/README.md
CHANGED
|
@@ -11,7 +11,7 @@ library_name: transformers
|
|
| 11 |
|
| 12 |
## Model Summary
|
| 13 |
|
| 14 |
-
**Query Clarification** is an
|
| 15 |
|
| 16 |
Given a multi-turn conversation between a user and an AI assistant (and optionally relevant content such as
|
| 17 |
RAG documents), detect whether the last user query is underspecified (no clear interpretation or multiple
|
|
@@ -20,7 +20,7 @@ library_name: transformers
|
|
| 20 |
- **Developer:** IBM Research
|
| 21 |
- **HF Collection:** [Granite Libraries](https://huggingface.co/collections/ibm-granite/granite-libraries)
|
| 22 |
- **GitHub Repository:** https://github.com/ibm-granite
|
| 23 |
-
- **Release Date:**
|
| 24 |
- **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granite-4.0-micro)
|
| 25 |
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
|
| 26 |
|
|
@@ -68,7 +68,7 @@ In this example, the adapter detects that the user's question is underspecified
|
|
| 68 |
|
| 69 |
## Usage Examples
|
| 70 |
|
| 71 |
-
The recommended way to call this adapter is through the [Mellea](https://mellea.ai) framework. For detailed examples on how to use this and other
|
| 72 |
|
| 73 |
|
| 74 |
## Training Details
|
|
@@ -81,6 +81,7 @@ The training dataset consists of:
|
|
| 81 |
- Human-generated underspecified questions created by annotators over content from Gov and ClapNQ domains
|
| 82 |
- Multi-turn conversations with underspecified final turns
|
| 83 |
- Synthetically generated clarification requests using target patterns (hedging with answers, hedging list, open domain)
|
|
|
|
| 84 |
|
| 85 |
The training dataset is proprietary and was obtained in combination with a third-party company who contracted the human annotators.
|
| 86 |
|
|
@@ -105,9 +106,9 @@ We evaluate the adapter on two dimensions: (1) classification accuracy in detect
|
|
| 105 |
### Test Set
|
| 106 |
|
| 107 |
The evaluation uses a test set containing:
|
| 108 |
-
-
|
| 109 |
-
-
|
| 110 |
-
-
|
| 111 |
|
| 112 |
### Classification Results
|
| 113 |
|
|
@@ -115,39 +116,39 @@ The evaluation uses a test set containing:
|
|
| 115 |
|
| 116 |
**Trained LoRA:**
|
| 117 |
|
| 118 |
-
| Model | Underspecified (
|
| 119 |
|-------|---------------------|----------------------|-------------------|------------------|
|
| 120 |
-
| Granite-4.0-micro LoRA |
|
| 121 |
|
| 122 |
**OOB models (prompted, zero-shot):**
|
| 123 |
|
| 124 |
-
| Model | Underspecified (
|
| 125 |
|-------|---------------------|----------------------|-------------------|------------------|
|
| 126 |
-
| gpt-
|
| 127 |
-
|
|
| 128 |
-
| Granite-4.0-micro |
|
| 129 |
|
| 130 |
-
The LoRA achieves
|
| 131 |
|
| 132 |
### Clarification Quality Results
|
| 133 |
|
| 134 |
-
For queries identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The numbers in parentheses indicate how many of the
|
| 135 |
|
| 136 |
**Trained LoRA:**
|
| 137 |
|
| 138 |
| Model | Precision | Recall | F1 |
|
| 139 |
|-------|-----------|--------|-----|
|
| 140 |
-
| Granite-4.0-micro LoRA (
|
| 141 |
|
| 142 |
**OOB models (prompted, zero-shot):**
|
| 143 |
|
| 144 |
| Model | Precision | Recall | F1 |
|
| 145 |
|-------|-----------|--------|-----|
|
| 146 |
-
|
|
| 147 |
-
| gpt-
|
| 148 |
-
| Granite-4.0-micro (
|
| 149 |
|
| 150 |
-
The LoRA achieves
|
| 151 |
|
| 152 |
|
| 153 |
### Adapter Details
|
|
@@ -161,16 +162,16 @@ The LoRA achieves clarification quality close to GPT-4o while maintaining much b
|
|
| 161 |
| **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
|
| 162 |
|
| 163 |
**Infrastructure:**
|
| 164 |
-
We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using
|
| 165 |
|
| 166 |
**Ethical Considerations & Limitations:**
|
| 167 |
The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
|
| 168 |
|
| 169 |
## Resources
|
| 170 |
|
| 171 |
-
-
|
| 172 |
-
-
|
| 173 |
-
-
|
| 174 |
|
| 175 |
## Contact
|
| 176 |
-
[Lucian Popa](mailto:lpopa@us.ibm.com)
|
|
|
|
| 11 |
|
| 12 |
## Model Summary
|
| 13 |
|
| 14 |
+
**Query Clarification** is an adapter designed for conversational use cases where user queries may be ill-formed, unclear, or have multiple valid interpretations based on the underlying system or content. Use cases include RAG systems, routing, and tool invocation scenarios. We provide experimental results showing that the adapter achieves significantly higher classification accuracy than prompted out-of-the-box models, including frontier models such as gpt-4o. The adapter released here works with the IBM granite-4.0-micro model and is specifically fine-tuned for the following task:
|
| 15 |
|
| 16 |
Given a multi-turn conversation between a user and an AI assistant (and optionally relevant content such as
|
| 17 |
RAG documents), detect whether the last user query is underspecified (no clear interpretation or multiple
|
|
|
|
| 20 |
- **Developer:** IBM Research
|
| 21 |
- **HF Collection:** [Granite Libraries](https://huggingface.co/collections/ibm-granite/granite-libraries)
|
| 22 |
- **GitHub Repository:** https://github.com/ibm-granite
|
| 23 |
+
- **Release Date:** April 2026
|
| 24 |
- **Model type:** LoRA adapter for [ibm-granite/granite-4.0-micro](https://huggingface.co/ibm-granite/granite-4.0-micro)
|
| 25 |
- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
|
| 26 |
|
|
|
|
| 68 |
|
| 69 |
## Usage Examples
|
| 70 |
|
| 71 |
+
The recommended way to call this adapter is through the [Mellea](https://mellea.ai) framework. For detailed examples on how to use this and other adapters, please refer to the [Mellea examples](https://github.com/generative-computing/mellea/tree/main/docs/examples/intrinsics).
|
| 72 |
|
| 73 |
|
| 74 |
## Training Details
|
|
|
|
| 81 |
- Human-generated underspecified questions created by annotators over content from Gov and ClapNQ domains
|
| 82 |
- Multi-turn conversations with underspecified final turns
|
| 83 |
- Synthetically generated clarification requests using target patterns (hedging with answers, hedging list, open domain)
|
| 84 |
+
- Conversation-only examples (without document passages) targeting pre-retrieval scenarios
|
| 85 |
|
| 86 |
The training dataset is proprietary and was obtained in combination with a third-party company who contracted the human annotators.
|
| 87 |
|
|
|
|
| 106 |
### Test Set
|
| 107 |
|
| 108 |
The evaluation uses a test set containing:
|
| 109 |
+
- 104 positive examples: conversations where the last user question is underspecified (generated by a combination of humans and AI model -- mistralai/Mistral-Large-3-675B-Instruct-2512)
|
| 110 |
+
- 200 random negatives: randomly sampled, pre-existing conversations where the last user question is clear
|
| 111 |
+
- 79 hard negatives: conversations where the last user question was intended to be underspecified (generated by an AI model -- mistralai/Mistral-Large-3-675B-Instruct-2512) but adjudicated as clear by humans
|
| 112 |
|
| 113 |
### Classification Results
|
| 114 |
|
|
|
|
| 116 |
|
| 117 |
**Trained LoRA:**
|
| 118 |
|
| 119 |
+
| Model | Underspecified (104) | CLEAR - random (200) | CLEAR - hard (79) | Overall Accuracy |
|
| 120 |
|-------|---------------------|----------------------|-------------------|------------------|
|
| 121 |
+
| Granite-4.0-micro LoRA | **84.6%** | **98.5%** | **87.3%** | **92.4%** |
|
| 122 |
|
| 123 |
**OOB models (prompted, zero-shot):**
|
| 124 |
|
| 125 |
+
| Model | Underspecified (104) | CLEAR - random (200) | CLEAR - hard (79) | Overall Accuracy |
|
| 126 |
|-------|---------------------|----------------------|-------------------|------------------|
|
| 127 |
+
| gpt-4o | 48.1% | 90.5% | 92.4% | 79.4% |
|
| 128 |
+
| gpt-oss-20b | 90.4% | 72.5% | 58.2% | 74.4% |
|
| 129 |
+
| Granite-4.0-micro | 25.0% | 85.0% | 63.3% | 64.2% |
|
| 130 |
|
| 131 |
+
The LoRA achieves the highest overall accuracy by balancing high recall on underspecified queries with strong specificity on clear queries. Prompted models tend to err in one direction: gpt-oss-20b over-clarifies (high underspecified recall but low CLEAR accuracy), while gpt-4o and base Granite under-clarify (low underspecified recall).
|
| 132 |
|
| 133 |
### Clarification Quality Results
|
| 134 |
|
| 135 |
+
For queries correctly identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The judge compares the entities and options surfaced in the generated clarification against those in the gold-standard reference: precision measures what fraction of the generated options are correct, and recall measures what fraction of the gold options are covered. Additionally, the numbers in parentheses (near the model name) indicate how many of the 104 underspecified records each model correctly identified and generated clarifications for -- the quality metrics are measured only on those data points.
|
| 136 |
|
| 137 |
**Trained LoRA:**
|
| 138 |
|
| 139 |
| Model | Precision | Recall | F1 |
|
| 140 |
|-------|-----------|--------|-----|
|
| 141 |
+
| Granite-4.0-micro LoRA (88/104) | **91.7%** | **87.5%** | **89.5%** |
|
| 142 |
|
| 143 |
**OOB models (prompted, zero-shot):**
|
| 144 |
|
| 145 |
| Model | Precision | Recall | F1 |
|
| 146 |
|-------|-----------|--------|-----|
|
| 147 |
+
| gpt-oss-20b (93/104) | 93.3% | 70.7% | 80.5% |
|
| 148 |
+
| gpt-4o (50/104) | 81.3% | 69.8% | 75.1% |
|
| 149 |
+
| Granite-4.0-micro (26/104) | 75.0% | 45.7% | 56.8% |
|
| 150 |
|
| 151 |
+
The LoRA achieves the highest F1 (89.5%) across the largest number of correctly classified records (88/104), demonstrating both strong classification and high generation quality.
|
| 152 |
|
| 153 |
|
| 154 |
### Adapter Details
|
|
|
|
| 162 |
| **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
|
| 163 |
|
| 164 |
**Infrastructure:**
|
| 165 |
+
We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 1 H100 GPU.
|
| 166 |
|
| 167 |
**Ethical Considerations & Limitations:**
|
| 168 |
The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
|
| 169 |
|
| 170 |
## Resources
|
| 171 |
|
| 172 |
+
- Learn about the latest updates with Granite: https://www.ibm.com/granite
|
| 173 |
+
- Get started with tutorials, best practices, and prompt engineering advice: https://www.ibm.com/granite/docs/
|
| 174 |
+
- Learn about the latest Granite learning resources: https://github.com/ibm-granite/granite-guardian/tree/main/cookbooks
|
| 175 |
|
| 176 |
## Contact
|
| 177 |
+
[Lucian Popa](mailto:lpopa@us.ibm.com)
|