FastFlowLM commited on
Commit
3bd01ec
·
verified ·
1 Parent(s): 070b3b9

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +353 -49
README.md CHANGED
@@ -1,69 +1,373 @@
1
  ---
2
- license: mit
3
  language:
4
  - en
5
- library_name: transformers
 
6
  tags:
7
- - qwen
8
  - qwen3
9
- - qwen3-8b
10
- - text-generation
11
- - AMD
12
- - Ryzen
13
- - NPU
14
- pipeline_tag: text-generation
15
- base_model:
16
- - Qwen/Qwen3-7B-Instruct
17
  ---
18
 
19
- # 🐉 Qwen3 8B – Optimized for FastFlowLM on AMD Ryzen™ AI NPU (XDNA2 Only)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
 
21
- ## Model Summary
22
- This model is based on **Qwen3 8B Instruct** from Alibaba Cloud. It uses the original Qwen3-7B architecture and weights, potentially with enhancements such as quantization or kernel-level acceleration for NPU efficiency via the FastFlowLM runtime.
 
 
23
 
24
- > ✅ **Licensed under the permissive MIT License.**
 
25
 
26
- ## 📝 License & Usage Terms
 
27
 
28
- ### Base Model License
29
- - Released under the MIT License by Alibaba Cloud:
30
- 👉 https://huggingface.co/Qwen/Qwen3-7B-Instruct
31
 
32
- - Permissions include:
33
- - Free commercial and non-commercial use
34
- - Permission to modify and redistribute
35
- - Attribution not required, but appreciated
36
 
37
- ### Redistribution Notice
38
- - This repository **does not** contain original base weights.
39
- - You must acquire the official weights from Qwen’s Hugging Face page:
40
- 👉 https://huggingface.co/Qwen/Qwen3-7B-Instruct
41
 
42
- ### If Fine-tuned
43
- If this model has been modified (e.g., quantized, fine-tuned):
44
 
45
- - **Base Model License**: MIT
46
- - **Derivative Weights License**: [e.g., MIT, CC-BY-NC-4.0, custom]
47
- - **Training Dataset License(s)**:
48
- - [Dataset A] – [license]
49
- - [Dataset B] – [license]
50
 
51
- Ensure all datasets used are appropriately licensed for redistribution and use.
52
 
53
- ## Intended Use
54
- - **Recommended For**: High-performance on-device inference, private LLM workloads, local chat assistants, research
55
- - **Not Recommended For**: Critical systems or regulated applications without thorough evaluation
56
 
57
- ## Limitations & Risks
58
- - Large model may require tuning for low-latency use
59
- - Possibility of hallucination, toxicity, or factual errors
60
- - May encode pretraining biases
61
 
62
- ## Citation
63
- ```bibtex
64
- @misc{qwen32024,
65
- title={Qwen3: Smaller, Smarter, and More Open},
66
- author={Alibaba Cloud},
67
- year={2024},
68
- url={https://huggingface.co/Qwen}
69
  ```
 
 
 
 
 
 
 
 
 
1
  ---
2
+ base_model: Qwen/Qwen3-8B
3
  language:
4
  - en
5
+ license_link: https://huggingface.co/Qwen/Qwen3-8B/blob/main/LICENSE
6
+ license: apache-2.0
7
  tags:
 
8
  - qwen3
9
+ - qwen
10
+ - unsloth
11
+ - transformers
 
 
 
 
 
12
  ---
13
 
14
+ # To Switch Between Thinking and Non-Thinking
15
+ If you are using llama.cpp, Ollama, Open WebUI etc., you can add `/think` and `/no_think` to user prompts or system messages to switch the model's thinking mode from turn to turn. The model will follow the most recent instruction in multi-turn conversations.
16
+
17
+ Here is an example of multi-turn conversation:
18
+
19
+ ```
20
+ > Who are you /no_think
21
+
22
+ <think>
23
+
24
+ </think>
25
+
26
+ I am Qwen, a large-scale language model developed by Alibaba Cloud. [...]
27
+
28
+ > How many 'r's are in 'strawberries'? /think
29
+
30
+ <think>
31
+ Okay, let's see. The user is asking how many times the letter 'r' appears in the word "strawberries". [...]
32
+ </think>
33
+
34
+ The word strawberries contains 3 instances of the letter r. [...]
35
+ ```
36
+
37
+ # Qwen3-8B
38
+
39
+ ## Qwen3 Highlights
40
+
41
+ Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
42
+
43
+ - **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios.
44
+ - **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning.
45
+ - **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience.
46
+ - **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks.
47
+ - **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**.
48
+
49
+ ## Model Overview
50
+
51
+ **Qwen3-8B** has the following features:
52
+ - Type: Causal Language Models
53
+ - Training Stage: Pretraining & Post-training
54
+ - Number of Parameters: 8.2B
55
+ - Number of Paramaters (Non-Embedding): 6.95B
56
+ - Number of Layers: 36
57
+ - Number of Attention Heads (GQA): 32 for Q and 8 for KV
58
+ - Context Length: 32,768 natively and [131,072 tokens with YaRN](#processing-long-texts).
59
+
60
+ For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our [blog](https://qwenlm.github.io/blog/qwen3/), [GitHub](https://github.com/QwenLM/Qwen3), and [Documentation](https://qwen.readthedocs.io/en/latest/).
61
+
62
+ ## Quickstart
63
+
64
+ The code of Qwen3 has been in the latest Hugging Face `transformers` and we advise you to use the latest version of `transformers`.
65
+
66
+ With `transformers<4.51.0`, you will encounter the following error:
67
+ ```
68
+ KeyError: 'qwen3'
69
+ ```
70
+
71
+ The following contains a code snippet illustrating how to use the model generate content based on given inputs.
72
+ ```python
73
+ from transformers import AutoModelForCausalLM, AutoTokenizer
74
+
75
+ model_name = "Qwen/Qwen3-8B"
76
+
77
+ # load the tokenizer and the model
78
+ tokenizer = AutoTokenizer.from_pretrained(model_name)
79
+ model = AutoModelForCausalLM.from_pretrained(
80
+ model_name,
81
+ torch_dtype="auto",
82
+ device_map="auto"
83
+ )
84
+
85
+ # prepare the model input
86
+ prompt = "Give me a short introduction to large language model."
87
+ messages = [
88
+ {"role": "user", "content": prompt}
89
+ ]
90
+ text = tokenizer.apply_chat_template(
91
+ messages,
92
+ tokenize=False,
93
+ add_generation_prompt=True,
94
+ enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
95
+ )
96
+ model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
97
+
98
+ # conduct text completion
99
+ generated_ids = model.generate(
100
+ **model_inputs,
101
+ max_new_tokens=32768
102
+ )
103
+ output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
104
+
105
+ # parsing thinking content
106
+ try:
107
+ # rindex finding 151668 (</think>)
108
+ index = len(output_ids) - output_ids[::-1].index(151668)
109
+ except ValueError:
110
+ index = 0
111
+
112
+ thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
113
+ content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")
114
+
115
+ print("thinking content:", thinking_content)
116
+ print("content:", content)
117
+ ```
118
+
119
+ For deployment, you can use `vllm>=0.8.5` or `sglang>=0.4.5.post2` to create an OpenAI-compatible API endpoint:
120
+ - vLLM:
121
+ ```shell
122
+ vllm serve Qwen/Qwen3-8B --enable-reasoning --reasoning-parser deepseek_r1
123
+ ```
124
+ - SGLang:
125
+ ```shell
126
+ python -m sglang.launch_server --model-path Qwen/Qwen3-8B --reasoning-parser deepseek-r1
127
+ ```
128
+
129
+ ## Switching Between Thinking and Non-Thinking Mode
130
+
131
+ > [!TIP]
132
+ > The `enable_thinking` switch is also available in APIs created by vLLM and SGLang.
133
+ > Please refer to our documentation for [vLLM](https://qwen.readthedocs.io/en/latest/deployment/vllm.html#thinking-non-thinking-modes) and [SGLang](https://qwen.readthedocs.io/en/latest/deployment/sglang.html#thinking-non-thinking-modes) users.
134
+
135
+ ### `enable_thinking=True`
136
+
137
+ By default, Qwen3 has thinking capabilities enabled, similar to QwQ-32B. This means the model will use its reasoning abilities to enhance the quality of generated responses. For example, when explicitly setting `enable_thinking=True` or leaving it as the default value in `tokenizer.apply_chat_template`, the model will engage its thinking mode.
138
+
139
+ ```python
140
+ text = tokenizer.apply_chat_template(
141
+ messages,
142
+ tokenize=False,
143
+ add_generation_prompt=True,
144
+ enable_thinking=True # True is the default value for enable_thinking
145
+ )
146
+ ```
147
+
148
+ In this mode, the model will generate think content wrapped in a `<think>...</think>` block, followed by the final response.
149
+
150
+ > [!NOTE]
151
+ > For thinking mode, use `Temperature=0.6`, `TopP=0.95`, `TopK=20`, and `MinP=0` (the default setting in `generation_config.json`). **DO NOT use greedy decoding**, as it can lead to performance degradation and endless repetitions. For more detailed guidance, please refer to the [Best Practices](#best-practices) section.
152
+
153
+
154
+ ### `enable_thinking=False`
155
+
156
+ We provide a hard switch to strictly disable the model's thinking behavior, aligning its functionality with the previous Qwen2.5-Instruct models. This mode is particularly useful in scenarios where disabling thinking is essential for enhancing efficiency.
157
+
158
+ ```python
159
+ text = tokenizer.apply_chat_template(
160
+ messages,
161
+ tokenize=False,
162
+ add_generation_prompt=True,
163
+ enable_thinking=False # Setting enable_thinking=False disables thinking mode
164
+ )
165
+ ```
166
+
167
+ In this mode, the model will not generate any think content and will not include a `<think>...</think>` block.
168
+
169
+ > [!NOTE]
170
+ > For non-thinking mode, we suggest using `Temperature=0.7`, `TopP=0.8`, `TopK=20`, and `MinP=0`. For more detailed guidance, please refer to the [Best Practices](#best-practices) section.
171
+
172
+ ### Advanced Usage: Switching Between Thinking and Non-Thinking Modes via User Input
173
+
174
+ We provide a soft switch mechanism that allows users to dynamically control the model's behavior when `enable_thinking=True`. Specifically, you can add `/think` and `/no_think` to user prompts or system messages to switch the model's thinking mode from turn to turn. The model will follow the most recent instruction in multi-turn conversations.
175
+
176
+ Here is an example of a multi-turn conversation:
177
+
178
+ ```python
179
+ from transformers import AutoModelForCausalLM, AutoTokenizer
180
+
181
+ class QwenChatbot:
182
+ def __init__(self, model_name="Qwen/Qwen3-8B"):
183
+ self.tokenizer = AutoTokenizer.from_pretrained(model_name)
184
+ self.model = AutoModelForCausalLM.from_pretrained(model_name)
185
+ self.history = []
186
+
187
+ def generate_response(self, user_input):
188
+ messages = self.history + [{"role": "user", "content": user_input}]
189
+
190
+ text = self.tokenizer.apply_chat_template(
191
+ messages,
192
+ tokenize=False,
193
+ add_generation_prompt=True
194
+ )
195
+
196
+ inputs = self.tokenizer(text, return_tensors="pt")
197
+ response_ids = self.model.generate(**inputs, max_new_tokens=32768)[0][len(inputs.input_ids[0]):].tolist()
198
+ response = self.tokenizer.decode(response_ids, skip_special_tokens=True)
199
+
200
+ # Update history
201
+ self.history.append({"role": "user", "content": user_input})
202
+ self.history.append({"role": "assistant", "content": response})
203
+
204
+ return response
205
+
206
+ # Example Usage
207
+ if __name__ == "__main__":
208
+ chatbot = QwenChatbot()
209
+
210
+ # First input (without /think or /no_think tags, thinking mode is enabled by default)
211
+ user_input_1 = "How many r's in strawberries?"
212
+ print(f"User: {user_input_1}")
213
+ response_1 = chatbot.generate_response(user_input_1)
214
+ print(f"Bot: {response_1}")
215
+ print("----------------------")
216
+
217
+ # Second input with /no_think
218
+ user_input_2 = "Then, how many r's in blueberries? /no_think"
219
+ print(f"User: {user_input_2}")
220
+ response_2 = chatbot.generate_response(user_input_2)
221
+ print(f"Bot: {response_2}")
222
+ print("----------------------")
223
+
224
+ # Third input with /think
225
+ user_input_3 = "Really? /think"
226
+ print(f"User: {user_input_3}")
227
+ response_3 = chatbot.generate_response(user_input_3)
228
+ print(f"Bot: {response_3}")
229
+ ```
230
+
231
+ > **Note**
232
+ > For API compatibility, when `enable_thinking=True`, regardless of whether the user uses `/think` or `/no_think`, the model will always output a block wrapped in `<think>...</think>`. However, the content inside this block may be empty if thinking is disabled.
233
+ > When `enable_thinking=False`, the soft switches are not valid. Regardless of any `/think` or `/no_think` tags input by the user, the model will not generate think content and will not include a `<think>...</think>` block.
234
+
235
+ ## Agentic Use
236
+
237
+ Qwen3 excels in tool calling capabilities. We recommend using [Qwen-Agent](https://github.com/QwenLM/Qwen-Agent) to make the best use of agentic ability of Qwen3. Qwen-Agent encapsulates tool-calling templates and tool-calling parsers internally, greatly reducing coding complexity.
238
+
239
+ To define the available tools, you can use the MCP configuration file, use the integrated tool of Qwen-Agent, or integrate other tools by yourself.
240
+ ```python
241
+ from qwen_agent.agents import Assistant
242
+
243
+ # Define LLM
244
+ llm_cfg = {
245
+ 'model': 'Qwen3-8B',
246
+
247
+ # Use the endpoint provided by Alibaba Model Studio:
248
+ # 'model_type': 'qwen_dashscope',
249
+ # 'api_key': os.getenv('DASHSCOPE_API_KEY'),
250
+
251
+ # Use a custom endpoint compatible with OpenAI API:
252
+ 'model_server': 'http://localhost:8000/v1', # api_base
253
+ 'api_key': 'EMPTY',
254
+
255
+ # Other parameters:
256
+ # 'generate_cfg': {
257
+ # # Add: When the response content is `<think>this is the thought</think>this is the answer;
258
+ # # Do not add: When the response has been separated by reasoning_content and content.
259
+ # 'thought_in_content': True,
260
+ # },
261
+ }
262
+
263
+ # Define Tools
264
+ tools = [
265
+ {'mcpServers': { # You can specify the MCP configuration file
266
+ 'time': {
267
+ 'command': 'uvx',
268
+ 'args': ['mcp-server-time', '--local-timezone=Asia/Shanghai']
269
+ },
270
+ "fetch": {
271
+ "command": "uvx",
272
+ "args": ["mcp-server-fetch"]
273
+ }
274
+ }
275
+ },
276
+ 'code_interpreter', # Built-in tools
277
+ ]
278
+
279
+ # Define Agent
280
+ bot = Assistant(llm=llm_cfg, function_list=tools)
281
+
282
+ # Streaming generation
283
+ messages = [{'role': 'user', 'content': 'https://qwenlm.github.io/blog/ Introduce the latest developments of Qwen'}]
284
+ for responses in bot.run(messages=messages):
285
+ pass
286
+ print(responses)
287
+ ```
288
+
289
+ ## Processing Long Texts
290
+
291
+ Qwen3 natively supports context lengths of up to 32,768 tokens. For conversations where the total length (including both input and output) significantly exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively. We have validated the model's performance on context lengths of up to 131,072 tokens using the [YaRN](https://arxiv.org/abs/2309.00071) method.
292
+
293
+ YaRN is currently supported by several inference frameworks, e.g., `transformers` and `llama.cpp` for local use, `vllm` and `sglang` for deployment. In general, there are two approaches to enabling YaRN for supported frameworks:
294
+
295
+ - Modifying the model files:
296
+ In the `config.json` file, add the `rope_scaling` fields:
297
+ ```json
298
+ {
299
+ ...,
300
+ "rope_scaling": {
301
+ "type": "yarn",
302
+ "factor": 4.0,
303
+ "original_max_position_embeddings": 32768
304
+ }
305
+ }
306
+ ```
307
+ For `llama.cpp`, you need to regenerate the GGUF file after the modification.
308
+
309
+ - Passing command line arguments:
310
+
311
+ For `vllm`, you can use
312
+ ```shell
313
+ vllm serve ... --rope-scaling '{"type":"yarn","factor":4.0,"original_max_position_embeddings":32768}' --max-model-len 131072
314
+ ```
315
+
316
+ For `sglang`, you can use
317
+ ```shell
318
+ python -m sglang.launch_server ... --json-model-override-args '{"rope_scaling":{"type":"yarn","factor":4.0,"original_max_position_embeddings":32768}}'
319
+ ```
320
+
321
+ For `llama-server` from `llama.cpp`, you can use
322
+ ```shell
323
+ llama-server ... --rope-scaling yarn --rope-scale 4 --yarn-orig-ctx 32768
324
+ ```
325
+
326
+ > [!IMPORTANT]
327
+ > If you encounter the following warning
328
+ > ```
329
+ > Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'original_max_position_embeddings'}
330
+ > ```
331
+ > please upgrade `transformers>=4.51.0`.
332
 
333
+ > [!NOTE]
334
+ > All the notable open-source frameworks implement static YaRN, which means the scaling factor remains constant regardless of input length, **potentially impacting performance on shorter texts.**
335
+ > We advise adding the `rope_scaling` configuration only when processing long contexts is required.
336
+ > It is also recommended to modify the `factor` as needed. For example, if the typical context length for your application is 65,536 tokens, it would be better to set `factor` as 2.0.
337
 
338
+ > [!NOTE]
339
+ > The default `max_position_embeddings` in `config.json` is set to 40,960. This allocation includes reserving 32,768 tokens for outputs and 8,192 tokens for typical prompts, which is sufficient for most scenarios involving short text processing. If the average context length does not exceed 32,768 tokens, we do not recommend enabling YaRN in this scenario, as it may potentially degrade model performance.
340
 
341
+ > [!TIP]
342
+ > The endpoint provided by Alibaba Model Studio supports dynamic YaRN by default and no extra configuration is needed.
343
 
344
+ ## Best Practices
 
 
345
 
346
+ To achieve optimal performance, we recommend the following settings:
 
 
 
347
 
348
+ 1. **Sampling Parameters**:
349
+ - For thinking mode (`enable_thinking=True`), use `Temperature=0.6`, `TopP=0.95`, `TopK=20`, and `MinP=0`. **DO NOT use greedy decoding**, as it can lead to performance degradation and endless repetitions.
350
+ - For non-thinking mode (`enable_thinking=False`), we suggest using `Temperature=0.7`, `TopP=0.8`, `TopK=20`, and `MinP=0`.
351
+ - For supported frameworks, you can adjust the `presence_penalty` parameter between 0 and 2 to reduce endless repetitions. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.
352
 
353
+ 2. **Adequate Output Length**: We recommend using an output length of 32,768 tokens for most queries. For benchmarking on highly complex problems, such as those found in math and programming competitions, we suggest setting the max output length to 38,912 tokens. This provides the model with sufficient space to generate detailed and comprehensive responses, thereby enhancing its overall performance.
 
354
 
355
+ 3. **Standardize Output Format**: We recommend using prompts to standardize model outputs when benchmarking.
356
+ - **Math Problems**: Include "Please reason step by step, and put your final answer within \boxed{}." in the prompt.
357
+ - **Multiple-Choice Questions**: Add the following JSON structure to the prompt to standardize responses: "Please show your choice in the `answer` field with only the choice letter, e.g., `"answer": "C"`."
 
 
358
 
359
+ 4. **No Thinking Content in History**: In multi-turn conversations, the historical model output should only include the final output part and does not need to include the thinking content. It is implemented in the provided chat template in Jinja2. However, for frameworks that do not directly use the Jinja2 chat template, it is up to the developers to ensure that the best practice is followed.
360
 
361
+ ### Citation
 
 
362
 
363
+ If you find our work helpful, feel free to give us a cite.
 
 
 
364
 
 
 
 
 
 
 
 
365
  ```
366
+ @misc{qwen3,
367
+ title = {Qwen3},
368
+ url = {https://qwenlm.github.io/blog/qwen3/},
369
+ author = {Qwen Team},
370
+ month = {April},
371
+ year = {2025}
372
+ }
373
+ ```