Granite Library
Safetensors
GGUF
English
frreiss's picture
Update query_rewrite/README.md (#32)
76625c1
|
Raw History Blame
9.29 kB

Query Rewrite

Model Summary

Query re-write is a general-purpose adapter designed for conversational use cases that require rewriting a user query, for example, before accessing a database, or before routing to other APIs or tools. The adapter uses only the dialog turns with what is being said between the user and assistant. We are providing experimental results on a specialized enterprise setup, showing that the adapter performance is significantly higher than when prompting out-of-the-box models, including open-source models such as gpt-oss as well as frontier models such as gpt-4o. Query rewrite is implemented as part of a family of adapters that are fine-tuned specifically for the following task:

Given a multi-turn conversation between a user and an AI assistant, rewrite the last
user utterance (query) by transforming it (only if necessary) into an equivalent version that
is standalone and can be understood by itself (without the conversation).

We have created two implementations of LoRA adapters trained over granite-4.0-micro and gpt-oss-20b, respectively. This is the model card for the LoRA adapter trained over granite-4.0-micro. The LoRA adapter trained over gpt-oss-20b can be found here.

Intended use

The adapter gives the ability to rewrite the last user query in a multi-turn conversation. Typically, the rewrite is a form of expansion that inlines into the query any implicit references that are made to entities, concepts, or even parts of the conversation that occur in the previous turns (either by the user or the AI assistant). Such expansion can include coreference resolution (i.e., replacement of pronouns with the actual entities), handling of ellipsis, which is the common linguistic phenomenon where parts of a sentence or phrase are omitted by the user, but can be understood from the context (i.e., for whom, of what, with respect to something discussed above, etc.).

As a result of the expansion, the query becomes a standalone query, still equivalent in meaning to what the user asked in the last turn. The rewritten query can be sent to downstream tasks as a better replacement for the original user query, and without the need for (a potentially very long) context.

Adapter input: The input to the query rewrite adapter is an OpenAI-compatible chat completion request, containing a list of conversation turns that can alternate between the user and assistant role, and ending with a user turn. The last user turn in the list is assumed to be the query that needs to be rewritten (if not already standalone).

Adapter output: The output of the query rewrite adapter is the result of the original chat completion request, formatted as a JSON object as follows:

{
  "rewritten_question": <Rewritten last user question>
}

Example

Input conversation:

Role Message
assistant Welcome to pet questions!
user I have two pets, a dog named Rex and a cat named Lucy. Rex spends a lot of time in the backyard and outdoors, and Lucy is always inside.
assistant Sounds good! Rex must love exploring outside, while Lucy probably enjoys her cozy indoor life.
user But is he more likely to get fleas because of that?

Output:

{
  "rewritten_question": "Is Rex more likely to get fleas because he spends a lot of time in the backyard and outdoors?"
}

In this example, the adapter resolves the pronoun "he" to "Rex" and expands "because of that" to explicitly reference spending time outdoors, making the query standalone and understandable without the conversation history.

Usage Examples

The recommended way to call this adapter is through the Mellea framework. For detailed examples on how to use this and other adapter, please refer to the Mellea examples.

Training Details

The training data contains both: 1) standalone examples, which teach the adapter to refrain from rewriting user questions that are already standalone, and 2) non-standalone examples containing a diversity of patterns that are used to teach the adapter to expand the user turn so that it becomes standalone.

Training Data

The training data uses the publicly available Cloud corpus of technical documentation pages from MT-RAG. Based on this corpus of documents, we constructed a dataset consisting of high-quality, human-created conversations, where the last turn of the conversation comes into versions: non-standalone version, and corresponding standalone version. In addition, we have also used a synthetically generated set of training examples, to maximize the diversity across a variety of patterns.

The training dataset is proprietary and was obtained in combination with a third-party company that contracted the human annotators.

Robustness to System Prompts

Different researchers or practitioners may use various system prompts tailored to their specific use cases. For the IBM granite versions of the adapter, to enhance robustness against these variations, we generated three distinct versions of each training sample, each paired with a different system prompt. This expanded and diversified training dataset is then used to train the LoRA adapters, improving their ability to handle diverse prompt styles effectively.

System Prompts Used (specifically for IBM granite base models):

Version 1: <|start_of_role|>system<|end_of_role|> Knowledge Cutoff Date: April 2024. Today's Date: May 20, 2025. You are Granite, developed by IBM. You are a helpful AI assistant. <|end_of_text|>

Version 2: <|start_of_role|>system<|end_of_role|> Knowledge Cutoff Date: April 2024. Today's Date: May 20, 2025. You are Granite, developed by IBM. Write the response to the user's input by strictly aligning with the facts in the provided documents. If the information needed to answer the question is not available in the documents, inform the user that the question cannot be answered based on the available data.<|end_of_text|>

Version 3: An empty system prompt (no instructions provided).

This approach ensures that the adapters remain effective and reliable across varying system prompt formats commonly encountered in real-world applications.

Training Hyperparameters

The adapters, both for granite and gpt-oss, were fine-tuned using PEFT under the following regime: rank = 32, learning rate = 1e-5, number of epochs = 25 with early stopping.

Evaluation

Here we evaluate the quality of the rewritten queries themselves on an enterprise internal dataset. This is a dataset with two turn conversations, where the last user turn may or may not be standalone. We do have the gold rewritten queries, and also make use of a LLM judge (with Llama-3.3-70b as the model) with a specific prompt to check whether the model-generate rewriting is equivalent to the gold rewriting. This is a challenging benchmark, with the specific requirement that the models minimally change the query with only the additions needed to make the query standalone (if not already standalone). If any changes are not minimal, the judge will penalize the rewritten query.

Results

Trained LoRAs:

Model Accuracy
Granite-4.0-micro-query-rewrite-LoRA 87.1%
GPT-OSS-20b-query-rewrite-LoRA 87.7%

OOB models (prompted, zero-shot):

Model Accuracy
GPT 4o 77.8%
GPT 4o mini 73.4%
Granite-4.0-micro 70.6%
GPT-OSS-120b 67.9%
GPT-OSS-20b 52.9%

Adapter Details

Property LoRA
Base Model ibm-granite/granite-4.0-micro
PEFT Type LORA
Rank (r) 32
Alpha 32
Target Modules q_proj, k_proj, v_proj

Infrastructure: We trained the query rewrite granite-4.0-micro LoRA adapter on IBM's Vela cluster using 8 A100 GPUs.

Ethical Considerations & Limitations: The model's outputs are not guaranteed to be factually accurate or complete. All outputs should be independently validated before use in decision-making or downstream applications. The model has been trained and evaluated on English data only.

Resources

Contact

Lucian Popa