Granite Library
Safetensors
GGUF
English

Update the numbers to reflect the LoRA with json object as output format (as opposed to plain text, which was for a previous version).

#33
Files changed (1) hide show
  1. query_clarification/README.md +5 -5
query_clarification/README.md CHANGED
@@ -117,7 +117,7 @@ The evaluation uses a test set containing:
117
 
118
  | Model | Underspecified (80) | CLEAR - random (150) | CLEAR - hard (50) | Overall Accuracy |
119
  |-------|---------------------|----------------------|-------------------|------------------|
120
- | Granite-4.0-micro LoRA | 81.25% | 100% | 86% | **92.14%** |
121
 
122
  **OOB models (prompted, zero-shot):**
123
 
@@ -127,17 +127,17 @@ The evaluation uses a test set containing:
127
  | GPT 4o | 60% | 85.3% | 48% | 71.42% |
128
  | Granite-4.0-micro | 97.5% | 5.3% | 0% | 30.71% |
129
 
130
- The LoRA achieves significantly higher overall accuracy due to its ability to correctly identify clear queries, while prompted models tend to over-clarify.
131
 
132
  ### Clarification Quality Results
133
 
134
- For queries correctly identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The numbers in parentheses indicate how many of the 80 underspecified records each model correctly identified and generated clarifications for—quality metrics are measured only on those data points.
135
 
136
  **Trained LoRA:**
137
 
138
  | Model | Precision | Recall | F1 |
139
  |-------|-----------|--------|-----|
140
- | Granite-4.0-micro LoRA (65/80) | 88.4% | 78.3% | 83.0% |
141
 
142
  **OOB models (prompted, zero-shot):**
143
 
@@ -161,7 +161,7 @@ The LoRA achieves clarification quality close to GPT-4o while maintaining much b
161
  | **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
162
 
163
  **Infrastructure:**
164
- We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 8 A100 GPUs.
165
 
166
  **Ethical Considerations & Limitations:**
167
  The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.
 
117
 
118
  | Model | Underspecified (80) | CLEAR - random (150) | CLEAR - hard (50) | Overall Accuracy |
119
  |-------|---------------------|----------------------|-------------------|------------------|
120
+ | Granite-4.0-micro LoRA | 76.25% | 99.33% | 80% | **89.3%** |
121
 
122
  **OOB models (prompted, zero-shot):**
123
 
 
127
  | GPT 4o | 60% | 85.3% | 48% | 71.42% |
128
  | Granite-4.0-micro | 97.5% | 5.3% | 0% | 30.71% |
129
 
130
+ The LoRA achieves higher overall accuracy due to its ability to correctly identify clear queries, while prompted models tend to over-clarify.
131
 
132
  ### Clarification Quality Results
133
 
134
+ For queries identified as underspecified, we measure the quality of generated clarification requests using an LLM judge (Llama-3.3-70b). The numbers in parentheses indicate how many of the 80 underspecified records each model correctly identified and generated clarifications for—quality metrics are measured only on those data points.
135
 
136
  **Trained LoRA:**
137
 
138
  | Model | Precision | Recall | F1 |
139
  |-------|-----------|--------|-----|
140
+ | Granite-4.0-micro LoRA (61/80) | 90.4% | 80.1% | 84.9% |
141
 
142
  **OOB models (prompted, zero-shot):**
143
 
 
161
  | **Target Modules** | q_proj, k_proj, v_proj, o_proj, input_linear, output_linear |
162
 
163
  **Infrastructure:**
164
+ We trained the query clarification granite-4.0-micro LoRA adapter on IBM's Vela cluster using 8 H100 GPUs.
165
 
166
  **Ethical Considerations & Limitations:**
167
  The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.