Text Generation
Transformers
Safetensors
English
qwen3
proactive-agents
information-seeking
question-asking
agents
multi-hop-qa
tau2-bench
dpo
conversational
text-generation-inference
Instructions to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dolev31/ProactiveInquirer-Qwen3-8B-Merged") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dolev31/ProactiveInquirer-Qwen3-8B-Merged") model = AutoModelForCausalLM.from_pretrained("dolev31/ProactiveInquirer-Qwen3-8B-Merged", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dolev31/ProactiveInquirer-Qwen3-8B-Merged" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-Merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-Merged
- SGLang
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dolev31/ProactiveInquirer-Qwen3-8B-Merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-Merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dolev31/ProactiveInquirer-Qwen3-8B-Merged" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-Merged", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use dolev31/ProactiveInquirer-Qwen3-8B-Merged with Docker Model Runner:
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-Merged
File size: 6,548 Bytes
b6e8077 ecea20b b6e8077 217f99d b6e8077 ecea20b b6e8077 45b629e b6e8077 45b629e b6e8077 217f99d b6e8077 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | ---
license: apache-2.0
base_model: Qwen/Qwen3-8B
library_name: transformers
pipeline_tag: text-generation
language:
- en
tags:
- proactive-agents
- information-seeking
- question-asking
- agents
- multi-hop-qa
- tau2-bench
- dpo
- qwen3
datasets:
- dgslibisey/MuSiQue
- ChilleD/StrategyQA
- xanhho/2WikiMultihopQA
---
<div align="center">
# ProactiveInquirer-Qwen3-8B-Merged
<img src="assets/title-card.png" width="100%" alt="Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents">
[Ido Levy](https://scholar.google.com/citations?user=Ok_7M80AAAAJ)<sup>1,2</sup> · [Asaf Yehudai](https://scholar.google.com/citations?user=FprEf4oAAAAJ)<sup>1</sup> · [Segev Shlomov](https://scholar.google.com/citations?user=hhtOihkAAAAJ)<sup>1</sup> · [Asaf Adi](https://www.linkedin.com/in/asaf-adi/)<sup>1</sup> · [Leshem Choshen](https://borgr.github.io/)<sup>1,2</sup><br>
<sup>1</sup>IBM <sup>2</sup>Weizmann Institute of Science
[](https://dolev31.github.io/ProactiveInquirer/)
[](https://arxiv.org/abs/2609.37236)
[](https://github.com/dolev31/ProactiveInquirer)
[](https://www.apache.org/licenses/LICENSE-2.0)
<video controls autoplay muted loop playsinline poster="https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B/resolve/main/assets/animation-poster.webp" src="https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B/resolve/main/assets/animation.mp4" width="100%"></video>
<sub>▶ The paper's example, step by step (22 seconds): the questioner trained with Q&D finds the account, the order with the boots and the size-8 boots, and the task is completed.</sub>
</div>
The trained **questioner** from *Asking for What Was Never Requested: Horizontal and Vertical
Proactivity in Agents*, with its LoRA adapter merged into [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
It is a standard full-weight model: it loads without PEFT and serves with vLLM, SGLang or TGI like any
Qwen3-8B.
- **The adapter**, with the results, the training details, the limitations and a complete two-turn
example: [dolev31/ProactiveInquirer-Qwen3-8B](https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B).
- **Quantized** for llama.cpp, Ollama and LM Studio:
[dolev31/ProactiveInquirer-Qwen3-8B-GGUF](https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF).
This is training seed 1, the adapter at the root of the adapter repository. The merge ran in float32
and the weights are stored in bfloat16. On the adapter card's two-turn example, greedy decoding with
this model returns the adapter's output character for character.
## Results
The results are the trained questioner's, as the paper reports them: see the
[adapter card's Results](https://huggingface.co/dolev31/ProactiveInquirer-Qwen3-8B#results). This merged model reproduces the adapter's output on that card's example, as the paragraph
above says.
## How to use it
The questioner reads the prompt template it was trained on, in [`prompts/`](prompts/), and replies
with one JSON action per step: `{"action": "ask", "question": ...}` or `{"action": "stop", ...}`.
Keep Qwen3's thinking off, as in training.
```python
import re
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "dolev31/ProactiveInquirer-Qwen3-8B-Merged"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(REPO, dtype=torch.bfloat16, device_map="auto")
template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()
def next_action(**state):
fields = dict(state, user_channel=placebo.strip())
prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
ids = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
).to(model.device)
out = model.generate(**ids, max_new_tokens=200, do_sample=False)
return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)
print(next_action(
question="Who was the spouse of the director of the film The Great Flamarion?",
instructions="Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
"queries against that pool before answering; several paragraphs are distractors, and the answer "
"usually requires composing facts from more than one of them.",
evidence="(nothing retrieved yet)", draft="(no draft yet)", history="(nothing asked yet)",
))
# {"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}
```
With vLLM, serve it and send the filled template as the user message, with thinking off:
```bash
vllm serve dolev31/ProactiveInquirer-Qwen3-8B-Merged
# request body: {"messages": [{"role": "user", "content": "<the filled template>"}],
# "chat_template_kwargs": {"enable_thinking": false}, "temperature": 0}
```
## Limitations
- The questioner's own limitations, from the paper: it has learned what to ask more readily than when to
stop, the extra evidence it finds does not yet translate into better final answers, and its user-facing
results come from a simulated customer, not from real people.
- It is a component inside an agent, meant to be called with its prompt template. It is not a chat
assistant, and it was trained and evaluated in English.
- This is one training seed (seed 1), merged in float32 and stored in bfloat16. Its equality with the
adapter was checked on the card's example, not on a benchmark.
## Citation
```bibtex
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint arXiv:2609.37236},
url = {https://arxiv.org/abs/2609.37236},
year = {2026}
}
```
## License
Apache-2.0, like the base model [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
|