Text Generation
MLX
Safetensors
Russian
English
qwen3
conversational
8-bit precision
Bogdan commited on
Commit
ffbebff
·
verified ·
1 Parent(s): c568279

Upload folder using huggingface_hub

Browse files
README.md CHANGED
@@ -3,86 +3,41 @@ license: apache-2.0
3
  datasets:
4
  - dichspace/darulm
5
  - HuggingFaceFW/fineweb-2
6
- - RefalMachine/ruadapt_instruct_2507
7
  language:
8
  - ru
9
  - en
10
- base_model: RefalMachine/RuadaptQwen3-4B-Instruct
11
- pipeline_tag: text-generation
12
  tags:
13
  - mlx
14
- library_name: mlx
15
  ---
16
 
17
- # RuadaptQwen3-4B-Instruct-4bit MLX
18
-
19
- This is a 4-bit quantized MLX version of [RefalMachine/RuadaptQwen3-4B-Instruct](https://huggingface.co/RefalMachine/RuadaptQwen3-4B-Instruct), optimized for Apple Silicon devices using the MLX framework.
20
-
21
- ## Performance Metrics
22
-
23
- Model was tested using LM Studio as inference provider on a binned M4 Max MacBook Pro 14" with the following specifications:
24
- - 32-core GPU
25
- - 36 GB unified memory
26
- - Full GPU offload enabled
27
-
28
- ### Test prompt:
29
- ```
30
- напиши стих в стиле Евгения Онегина о бесконечно вечном
31
- ```
32
-
33
- ### Results:
34
 
35
- **4-bit MLX (this model):**
36
- - Tokens generated: 287
37
- - Time to First Token (TTFT): 0.14s
38
- - Throughput: up to 118.93 tok/s
39
 
40
- **Q4_0 GGUF** ([RefalMachine/RuadaptQwen3-4B-Instruct-GGUF](https://huggingface.co/RefalMachine/RuadaptQwen3-4B-Instruct-GGUF)):
41
- - Tokens generated: 233
42
- - TTFT: 0.16s
43
- - Throughput: up to 97.52 tok/s
44
-
45
- **Note:** The MLX 4-bit quantization demonstrates up to ~22% higher throughput compared to Q4_0 GGUF on Apple Silicon in this simple test.
46
-
47
- ## Usage
48
-
49
- Install MLX and mlx-lm:
50
 
51
  ```bash
52
- pip install mlx mlx-lm
53
  ```
54
 
55
- Use the model:
56
-
57
  ```python
58
  from mlx_lm import load, generate
59
 
60
- model, tokenizer = load("Bogdan01m/RuadaptQwen3-4B-Instruct-4bit")
61
- response = generate(model, tokenizer, prompt="Привет! Как дела?", verbose=True)
62
- ```
63
-
64
- ## Recommended Generation Parameters
65
-
66
- For more stable results, it is recommended to use low temperatures 0.0-0.3, top_p in the range from 0.85 to 0.95 and repetition_penalty 1.05.
67
-
68
- ## Original Model
69
 
70
- This model is a quantized version of [RefalMachine/RuadaptQwen3-4B-Instruct](https://huggingface.co/RefalMachine/RuadaptQwen3-4B-Instruct).
71
 
72
- ## Citation
 
 
 
 
73
 
74
- ```bibtex
75
- @article{tikhomirov2024facilitating,
76
- title={Facilitating Large Language Model Russian Adaptation with Learned Embedding Propagation},
77
- author={Tikhomirov, Mikhail and Chernyshov, Daniil},
78
- journal={Journal of Language and Education},
79
- volume={10},
80
- number={4},
81
- pages={130--145},
82
- year={2024}
83
- }
84
  ```
85
-
86
- ## Important
87
-
88
- The model's answers do not reflect the authors' opinions; they merely reproduce the knowledge obtained from data at all training stages. The model is based on a third-party pretrained model. Use with caution.
 
3
  datasets:
4
  - dichspace/darulm
5
  - HuggingFaceFW/fineweb-2
6
+ - RefalMachine/hybrid_reasoning_dataset_ru
7
  language:
8
  - ru
9
  - en
10
+ base_model: RefalMachine/RuadaptQwen3-32B-Instruct
11
+ library_name: mlx
12
  tags:
13
  - mlx
14
+ pipeline_tag: text-generation
15
  ---
16
 
17
+ # Bogdan01m/RuadaptQwen3-32B-Instruct-MLX-8bit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
+ This model [Bogdan01m/RuadaptQwen3-32B-Instruct-MLX-8bit](https://huggingface.co/Bogdan01m/RuadaptQwen3-32B-Instruct-MLX-8bit) was
20
+ converted to MLX format from [RefalMachine/RuadaptQwen3-32B-Instruct](https://huggingface.co/RefalMachine/RuadaptQwen3-32B-Instruct)
21
+ using mlx-lm version **0.28.3**.
 
22
 
23
+ ## Use with mlx
 
 
 
 
 
 
 
 
 
24
 
25
  ```bash
26
+ pip install mlx-lm
27
  ```
28
 
 
 
29
  ```python
30
  from mlx_lm import load, generate
31
 
32
+ model, tokenizer = load("Bogdan01m/RuadaptQwen3-32B-Instruct-MLX-8bit")
 
 
 
 
 
 
 
 
33
 
34
+ prompt = "hello"
35
 
36
+ if tokenizer.chat_template is not None:
37
+ messages = [{"role": "user", "content": prompt}]
38
+ prompt = tokenizer.apply_chat_template(
39
+ messages, add_generation_prompt=True
40
+ )
41
 
42
+ response = generate(model, tokenizer, prompt=prompt, verbose=True)
 
 
 
 
 
 
 
 
 
43
  ```
 
 
 
 
chat_template.jinja CHANGED
@@ -1,42 +1,72 @@
1
  {%- if tools %}
2
  {{- '<|im_start|>system\n' }}
3
- {%- if messages[0]['role'] == 'system' %}
4
- {{- messages[0]['content'] }}
5
- {%- else %}
6
- {{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
7
  {%- endif %}
8
- {{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
9
  {%- for tool in tools %}
10
  {{- "\n" }}
11
  {{- tool | tojson }}
12
  {%- endfor %}
13
  {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
14
  {%- else %}
15
- {%- if messages[0]['role'] == 'system' %}
16
- {{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
17
  {%- endif %}
18
  {%- endif %}
 
 
 
 
 
 
 
 
19
  {%- for message in messages %}
20
- {%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
21
  {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
22
  {%- elif message.role == "assistant" %}
23
- {{- '<|im_start|>' + message.role }}
24
- {%- if message.content %}
25
- {{- '\n' + message.content }}
 
 
 
 
 
 
26
  {%- endif %}
27
- {%- for tool_call in message.tool_calls %}
28
- {%- if tool_call.function is defined %}
29
- {%- set tool_call = tool_call.function %}
 
 
30
  {%- endif %}
31
- {{- '\n<tool_call>\n{"name": "' }}
32
- {{- tool_call.name }}
33
- {{- '", "arguments": ' }}
34
- {{- tool_call.arguments | tojson }}
35
- {{- '}\n</tool_call>' }}
36
- {%- endfor %}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
  {{- '<|im_end|>\n' }}
38
  {%- elif message.role == "tool" %}
39
- {%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
40
  {{- '<|im_start|>user' }}
41
  {%- endif %}
42
  {{- '\n<tool_response>\n' }}
@@ -49,4 +79,7 @@
49
  {%- endfor %}
50
  {%- if add_generation_prompt %}
51
  {{- '<|im_start|>assistant\n' }}
52
- {%- endif %}
 
 
 
 
1
  {%- if tools %}
2
  {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {{- messages[0].content + '\n\n' }}
 
 
5
  {%- endif %}
6
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
7
  {%- for tool in tools %}
8
  {{- "\n" }}
9
  {{- tool | tojson }}
10
  {%- endfor %}
11
  {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
12
  {%- else %}
13
+ {%- if messages[0].role == 'system' %}
14
+ {{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
15
  {%- endif %}
16
  {%- endif %}
17
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
18
+ {%- for message in messages[::-1] %}
19
+ {%- set index = (messages|length - 1) - loop.index0 %}
20
+ {%- if ns.multi_step_tool and message.role == "user" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
21
+ {%- set ns.multi_step_tool = false %}
22
+ {%- set ns.last_query_index = index %}
23
+ {%- endif %}
24
+ {%- endfor %}
25
  {%- for message in messages %}
26
+ {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
27
  {{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
28
  {%- elif message.role == "assistant" %}
29
+ {%- set content = message.content %}
30
+ {%- set reasoning_content = '' %}
31
+ {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
32
+ {%- set reasoning_content = message.reasoning_content %}
33
+ {%- else %}
34
+ {%- if '</think>' in message.content %}
35
+ {%- set content = message.content.split('</think>')[-1].lstrip('\n') %}
36
+ {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
37
+ {%- endif %}
38
  {%- endif %}
39
+ {%- if loop.index0 > ns.last_query_index %}
40
+ {%- if loop.last or (not loop.last and reasoning_content) %}
41
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
42
+ {%- else %}
43
+ {{- '<|im_start|>' + message.role + '\n' + content }}
44
  {%- endif %}
45
+ {%- else %}
46
+ {{- '<|im_start|>' + message.role + '\n' + content }}
47
+ {%- endif %}
48
+ {%- if message.tool_calls %}
49
+ {%- for tool_call in message.tool_calls %}
50
+ {%- if (loop.first and content) or (not loop.first) %}
51
+ {{- '\n' }}
52
+ {%- endif %}
53
+ {%- if tool_call.function %}
54
+ {%- set tool_call = tool_call.function %}
55
+ {%- endif %}
56
+ {{- '<tool_call>\n{"name": "' }}
57
+ {{- tool_call.name }}
58
+ {{- '", "arguments": ' }}
59
+ {%- if tool_call.arguments is string %}
60
+ {{- tool_call.arguments }}
61
+ {%- else %}
62
+ {{- tool_call.arguments | tojson }}
63
+ {%- endif %}
64
+ {{- '}\n</tool_call>' }}
65
+ {%- endfor %}
66
+ {%- endif %}
67
  {{- '<|im_end|>\n' }}
68
  {%- elif message.role == "tool" %}
69
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
70
  {{- '<|im_start|>user' }}
71
  {%- endif %}
72
  {{- '\n<tool_response>\n' }}
 
79
  {%- endfor %}
80
  {%- if add_generation_prompt %}
81
  {{- '<|im_start|>assistant\n' }}
82
+ {%- if enable_thinking is defined and enable_thinking is false %}
83
+ {{- '<think>\n\n</think>\n\n' }}
84
+ {%- endif %}
85
+ {%- endif %}
config.json CHANGED
@@ -7,31 +7,31 @@
7
  "eos_token_id": 146215,
8
  "head_dim": 128,
9
  "hidden_act": "silu",
10
- "hidden_size": 2560,
11
  "initializer_range": 0.02,
12
- "intermediate_size": 9728,
13
- "max_position_embeddings": 262144,
14
- "max_window_layers": 36,
15
  "model_type": "qwen3",
16
- "num_attention_heads": 32,
17
- "num_hidden_layers": 36,
18
  "num_key_value_heads": 8,
19
  "pad_token_id": 146213,
20
  "quantization": {
21
  "group_size": 64,
22
- "bits": 4,
23
  "mode": "affine"
24
  },
25
  "quantization_config": {
26
  "group_size": 64,
27
- "bits": 4,
28
  "mode": "affine"
29
  },
30
  "rms_norm_eps": 1e-06,
31
  "rope_scaling": null,
32
- "rope_theta": 5000000,
33
  "sliding_window": null,
34
- "tie_word_embeddings": true,
35
  "torch_dtype": "bfloat16",
36
  "transformers_version": "4.51.3",
37
  "use_cache": true,
 
7
  "eos_token_id": 146215,
8
  "head_dim": 128,
9
  "hidden_act": "silu",
10
+ "hidden_size": 5120,
11
  "initializer_range": 0.02,
12
+ "intermediate_size": 25600,
13
+ "max_position_embeddings": 40960,
14
+ "max_window_layers": 64,
15
  "model_type": "qwen3",
16
+ "num_attention_heads": 64,
17
+ "num_hidden_layers": 64,
18
  "num_key_value_heads": 8,
19
  "pad_token_id": 146213,
20
  "quantization": {
21
  "group_size": 64,
22
+ "bits": 8,
23
  "mode": "affine"
24
  },
25
  "quantization_config": {
26
  "group_size": 64,
27
+ "bits": 8,
28
  "mode": "affine"
29
  },
30
  "rms_norm_eps": 1e-06,
31
  "rope_scaling": null,
32
+ "rope_theta": 1000000,
33
  "sliding_window": null,
34
+ "tie_word_embeddings": false,
35
  "torch_dtype": "bfloat16",
36
  "transformers_version": "4.51.3",
37
  "use_cache": true,
generation_config.json CHANGED
@@ -2,8 +2,8 @@
2
  "do_sample": true,
3
  "eos_token_id": 146215,
4
  "pad_token_id": 146213,
5
- "temperature": 0.7,
6
  "top_k": 20,
7
- "top_p": 0.8,
8
  "transformers_version": "4.51.3"
9
  }
 
2
  "do_sample": true,
3
  "eos_token_id": 146215,
4
  "pad_token_id": 146213,
5
+ "temperature": 0.6,
6
  "top_k": 20,
7
+ "top_p": 0.95,
8
  "transformers_version": "4.51.3"
9
  }
model-00001-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1f517f2a218f6f6a616eb3a5df110283f5924019b95af2b552f61146520d5615
3
+ size 5319142940
model-00002-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:956770fc5f486bb7ffc990b78ec13e384d592ee6522a70e8b494a0f282ce2105
3
+ size 5364709259
model-00003-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:47d6511aaa939cd61e8cb8393d3a863795b20be012b44593823045fae2d7978e
3
+ size 5367638894
model-00004-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:acf4ff28cc7e11d233f74c73e21dfa62160b2755e98af85b864142c09b675d96
3
+ size 5328316006
model-00005-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cdea6a2216e1d7452cb9f152e16851722ba7fbcaee366d235e980f1d54d7136f
3
+ size 5364709271
model-00006-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:69389761d8fb8d889af2d6e1b3bb7b381a7e3c59ba2669dfe7619640ddf24373
3
+ size 5367638918
model-00007-of-00007.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e2e574194ef39c869e912272096d3c3d0e5f656af51dd1dd10492c620ebcd46e
3
+ size 2636664420
model.safetensors.index.json CHANGED
The diff for this file is too large to render. See raw diff
 
special_tokens_map.json CHANGED
@@ -14,13 +14,6 @@
14
  "<|image_pad|>",
15
  "<|video_pad|>"
16
  ],
17
- "bos_token": {
18
- "content": "<|endoftext|>",
19
- "lstrip": false,
20
- "normalized": false,
21
- "rstrip": false,
22
- "single_word": false
23
- },
24
  "eos_token": {
25
  "content": "<|im_end|>",
26
  "lstrip": false,
 
14
  "<|image_pad|>",
15
  "<|video_pad|>"
16
  ],
 
 
 
 
 
 
 
17
  "eos_token": {
18
  "content": "<|im_end|>",
19
  "lstrip": false,
tokenizer.json CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5408c0e76d30d5f9348593818dcaafdd45dbe64ba19fcd5fb92c2677c15f4cf1
3
- size 12380040
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e5259c1df4795303715f22c2d9d9ca43fe8e25a4909342df3b90b56bc19a2f42
3
+ size 12380042
tokenizer_config.json CHANGED
@@ -200,7 +200,7 @@
200
  "normalized": false,
201
  "rstrip": false,
202
  "single_word": false,
203
- "special": true
204
  },
205
  "146238": {
206
  "content": "</think>",
@@ -208,7 +208,7 @@
208
  "normalized": false,
209
  "rstrip": false,
210
  "single_word": false,
211
- "special": true
212
  },
213
  "146239": {
214
  "content": "<|free_token1|>",
@@ -394,7 +394,7 @@
394
  "<|image_pad|>",
395
  "<|video_pad|>"
396
  ],
397
- "bos_token": "<|endoftext|>",
398
  "clean_up_tokenization_spaces": false,
399
  "eos_token": "<|im_end|>",
400
  "errors": "replace",
 
200
  "normalized": false,
201
  "rstrip": false,
202
  "single_word": false,
203
+ "special": false
204
  },
205
  "146238": {
206
  "content": "</think>",
 
208
  "normalized": false,
209
  "rstrip": false,
210
  "single_word": false,
211
+ "special": false
212
  },
213
  "146239": {
214
  "content": "<|free_token1|>",
 
394
  "<|image_pad|>",
395
  "<|video_pad|>"
396
  ],
397
+ "bos_token": null,
398
  "clean_up_tokenization_spaces": false,
399
  "eos_token": "<|im_end|>",
400
  "errors": "replace",