Update the numbers to reflect the LoRA with json object as output format (as opposed to plain text, which was for a previous version).
#33
by lucianpopa - opened
query_clarification/README.md
CHANGED
|
@@ -117,7 +117,7 @@ The evaluation uses a test set containing:
|
|
| 117 |
|
| 118 |
| Model | Underspecified (80) | CLEAR - random (150) | CLEAR - hard (50) | Overall Accuracy |
|
| 119 |
|-------|---------------------|----------------------|-------------------|------------------|
|
| 120 |
-
| Granite-4.0-micro LoRA |
|
| 121 |
|
| 122 |
**OOB models (prompted, zero-shot):**
|
| 123 |
|
|
@@ -127,17 +127,17 @@ The evaluation uses a test set containing:
|
|
| 127 |
| GPT 4o | 60% | 85.3% | 48% | 71.42% |
|
| 128 |
| Granite-4.0-micro | 97.5% | 5.3% | 0% | 30.71% |
|
| 129 |
|
| 130 |
-
The LoRA achieves
|
| 131 |
|
| 132 |
### Clarification Quality Results
|
| 133 |
|
| 134 |
-
For queries
|
| 135 |
|
| 136 |
**Trained LoRA:**
|
| 137 |
|
| 138 |
| Model | Precision | Recall | F1 |
|
| 139 |
|-------|-----------|--------|-----|
|
| 140 |
-
| Granite-4.0-micro LoRA (
|
| 141 |
|
| 142 |
**OOB models (prompted, zero-shot):**
|
| 143 |
|
|
@@ -161,7 +161,7 @@ The LoRA achieves clarification quality close to GPT-4o while maintaining much b
|
|
| 161 |
| **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
|
| 162 |
|
| 163 |
**Infrastructure:**
|
| 164 |
-
We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 8
|
| 165 |
|
| 166 |
**Ethical Considerations & Limitations:**
|
| 167 |
The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
|
|
|
|
| 117 |
|
| 118 |
| Model | Underspecified (80) | CLEAR - random (150) | CLEAR - hard (50) | Overall Accuracy |
|
| 119 |
|-------|---------------------|----------------------|-------------------|------------------|
|
| 120 |
+
| Granite-4.0-micro LoRA | 76.25% | 99.33% | 80% | **89.3%** |
|
| 121 |
|
| 122 |
**OOB models (prompted, zero-shot):**
|
| 123 |
|
|
|
|
| 127 |
| GPT 4o | 60% | 85.3% | 48% | 71.42% |
|
| 128 |
| Granite-4.0-micro | 97.5% | 5.3% | 0% | 30.71% |
|
| 129 |
|
| 130 |
+
The LoRA achieves higher overall accuracy due to its ability to correctly identify clear queries, while prompted models tend to over-clarify.
|
| 131 |
|
| 132 |
### Clarification Quality Results
|
| 133 |
|
| 134 |
+
For queries identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The numbers in parentheses indicate how many of the 80 underspecified records each model correctly identified and generated clarifications for—quality metrics are measured only on those data points.
|
| 135 |
|
| 136 |
**Trained LoRA:**
|
| 137 |
|
| 138 |
| Model | Precision | Recall | F1 |
|
| 139 |
|-------|-----------|--------|-----|
|
| 140 |
+
| Granite-4.0-micro LoRA (61/80) | 90.4% | 80.1% | 84.9% |
|
| 141 |
|
| 142 |
**OOB models (prompted, zero-shot):**
|
| 143 |
|
|
|
|
| 161 |
| **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
|
| 162 |
|
| 163 |
**Infrastructure:**
|
| 164 |
+
We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 8 H100 GPUs.
|
| 165 |
|
| 166 |
**Ethical Considerations & Limitations:**
|
| 167 |
The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
|