OS-Software commited on
Commit
ade9dd4
·
verified ·
1 Parent(s): e6dc1c9

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +159 -158
README.md CHANGED
@@ -9,164 +9,165 @@ tags:
9
  - decensored
10
  - abliterated
11
  - ara-lora
 
12
  pipeline_tag: image-text-to-text
13
  ---
14
- # This is a decensored version of [unsloth/Qwen3.8-27B](https://huggingface.co/unsloth/Qwen3.8-27B), made using [Heretic](https://heretic-project.org) v1.4.0+custom with the Arbitrary-Rank Ablation (ARA) method using a LoRA adapter and row-norm preservation
15
-
16
- ## Abliteration parameters
17
-
18
- | Parameter | Value |
19
- | :-------- | :---: |
20
- | **start_layer_index** | 9 |
21
- | **end_layer_index** | 51 |
22
- | **preserve_good_behavior_weight** | 1.0000 |
23
- | **steer_bad_behavior_weight** | 0.3027 |
24
- | **overcorrect_relative_weight** | 0.9481 |
25
- | **neighbor_count** | 1 |
26
-
27
- ## Performance
28
-
29
- | Metric | This model | Original model ([unsloth/Qwen3.8-27B](https://huggingface.co/unsloth/Qwen3.8-27B)) |
30
- | :----- | :--------: | :---------------------------: |
31
- | **Keywords** | 0/100 | 100/100 |
32
- | **KL divergence** | 0.0528 | 0 *(by definition)* |
33
-
34
- Note: Performance testing, including the measurement of refusal rates, was conducted using Japanese datasets.
35
-
36
- ## ⚠️ Important Notice
37
-
38
- This model has undergone substantial reduction of its safety alignment. As a result, it is more likely than standard models to generate harmful, inaccurate, biased, offensive, or otherwise inappropriate content.
39
-
40
- ### Intended Use
41
-
42
- For research and experimentation only, including safety research, alignment studies, and red-teaming. Please avoid deploying it in public or end-user-facing services.
43
-
44
- ### User Responsibility
45
-
46
- All outputs should be treated as untrusted and independently verified before use. Users are solely responsible for:
47
-
48
- * Evaluating the accuracy and suitability of generated content
49
- * Implementing appropriate safeguards and human oversight
50
- * Complying with applicable laws, regulations, licenses, and ethical standards
51
-
52
- Use of this model is entirely at your own risk.
53
-
54
- ### Disclaimer
55
-
56
- OS-Software provides this model without warranties of any kind and assumes no liability for any direct or indirect damages, losses, misuse, or legal consequences arising from its use.
57
-
58
- ## Acknowledgements
59
-
60
- Thanks to the base model developers, [p-e-w](https://github.com/p-e-w) for Heretic, and the wider open-source community.
61
-
62
- This is a derivative work released under the base model’s applicable license. All rights to the base model remain with their respective owners.
63
-
64
- -----
65
-
66
-
67
- # Read our How to [Run Qwen3.8-27B Guide!](https://unsloth.ai/docs/models/qwen3.8)
68
- <div>
69
- <p style="margin: 0 0 0px 0; margin-top: 0px;">
70
- <em>See <a href="https://unsloth.ai/docs/basics/unsloth-dynamic-v2.0-gguf">Unsloth Dynamic 2.0 GGUFs</a> for our quantization benchmarks.</em>
71
- </p>
72
- <div style="display: flex; gap: 5px; align-items: center; margin-bottom: 0px;">
73
- <a href="https://github.com/unslothai/unsloth/">
74
- <img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="133">
75
- </a>
76
- <a href="https://discord.gg/unsloth">
77
- <img src="https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png" width="173">
78
- </a>
79
- <a href="https://unsloth.ai/docs/models/qwen3.8">
80
- <img src="https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png" width="143">
81
- </a>
82
- </div>
83
- <ul style="margin: 0;">
84
- <li>Developer Role Support so Qwen3.8 can work in agentic tools like Codex and more!</li>
85
- <li>Qwen3.8 can now be run and fine-tuned in <a href="https://unsloth.ai/docs/new/desktop">Unsloth Desktop</a>. <a href="https://unsloth.ai/docs/models/qwen3.8">Read our guide</a>.</li>
86
- <li>Tool calling improvements: Makes parsing nested objects to make tool calling succeed more.</li>
87
- <li>See below for 1-bit Qwen3.8 run inside of Unsloth:</li>
88
- </div>
89
-
90
- <img width="600" alt="qwen3.8 unsloth desktop" src="https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2Fb3IiYJzHzsp2698xTim4%2Fgiffyy%20gf.gif?alt=media&token=f1bda1a1-b81f-43e2-ba10-dbdc9a29c0c0" />
91
-
92
-
93
- ---
94
-
95
- # Qwen3.8-27B
96
-
97
- Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.
98
-
99
- Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.
100
-
101
- ## Qwen3.8 Highlights
102
-
103
- Qwen3.8-27B features the following enhancements:
104
- - **Core Capabilities**: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
105
- - **Agent Execution**: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
106
- - **Downstream Compatibility**: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
107
- - **Flexible Thinking Control**: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with `reasoning_effort`, and reasoning context from historical messages is retained via `preserve_thinking`.
108
- - **Vision-Language Understanding**: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.
109
-
110
-
111
- ## Model Overview
112
-
113
- - Type: Causal Language Model with Vision Encoder
114
- - Training Stage: Pre-training & Post-training
115
- - Language Model
116
- - Number of Parameters: 27B
117
- - Hidden Dimension: 5120
118
- - Token Embedding: 248,320 (Padded)
119
- - Number of Layers: 64
120
- - Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
121
- - Gated DeltaNet:
122
- - Number of Linear Attention Heads: 48 for V and 16 for QK
123
- - Head Dimension: 128
124
- - Gated Attention:
125
- - Number of Attention Heads: 24 for Q and 4 for KV
126
- - Head Dimension: 256
127
- - Rotary Position Embedding Dimension: 64
128
- - Feed Forward Network:
129
- - Intermediate Dimension: 17,408
130
- - LM Output: 248,320 (Padded)
131
- - MTP (Multi-Token Prediction): trained with multiple steps
132
- - Context Length: 262,144 natively and extensible up to 1,000,000 tokens.
133
-
134
- ## Best Practices
135
-
136
- To achieve optimal performance, we recommend the following settings:
137
-
138
- 1. **Sampling Parameters**: We suggest using the following sets of sampling parameters:
139
-
140
- - Thinking Mode: `temperature=1.0`, `top_p=0.95`, `top_k=20`, `min_p=0.0`, `presence_penalty=0.0`, `repetition_penalty=1.0`
141
- - Instruct (or non-thinking) mode: `temperature=0.7`, `top_p=0.80`, `top_k=20`, `min_p=0.0`, `presence_penalty=1.5`, `repetition_penalty=1.0`
142
-
143
- For supported frameworks, you can adjust the `presence_penalty` parameter between 0 and 2 to reduce endless repetition. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.
144
-
145
- 2. **Adequate Output Length**: To optimize performance on agentic tasks, we recommend allocating sufficient output length to allow the model to generate detailed and comprehensive responses. For frameworks that support separate token limits for internal reasoning and final outputs, we suggest the following configuration within the 1M context length:
146
-
147
- - Reasoning Content: Set the maximum output length to 262,144 tokens.
148
- - Final Response: Set the maximum output length to 131,072 tokens.
149
-
150
- These settings provide the necessary capacity for complex reasoning while ensuring ample space for high-quality final deliverables.
151
-
152
- 3. **Processing Ultra-Long Texts**: Qwen3.8-27B natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively, e.g., YaRN.
153
-
154
- 4. **Long Video Understanding**: To optimize inference efficiency for plain text and images, the `size` parameter in the released `video_preprocessor_config.json` is conservatively configured. It is recommended to set the `longest_edge` parameter in the video_preprocessor_config file to 469,762,048 (corresponding to 224k video tokens) to enable higher frame-rate sampling for hour-scale videos and thereby achieve superior performance. For example,
155
- ```json
156
- {"longest_edge": 469762048, "shortest_edge": 4096}
157
- ```
158
-
159
- ## Citation
160
-
161
- If you find our work helpful, feel free to give us a cite.
162
-
163
-
164
- ```bibtex
165
- @misc{qwen38,
166
- title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
167
- url = {https://qwen.ai/blog?id=qwen3.8},
168
- author = {{Qwen Team}},
169
- month = {August},
170
- year = {2026}
171
- }
172
  ```
 
9
  - decensored
10
  - abliterated
11
  - ara-lora
12
+ - gguf
13
  pipeline_tag: image-text-to-text
14
  ---
15
+ # This is a decensored version of [unsloth/Qwen3.8-27B](https://huggingface.co/unsloth/Qwen3.8-27B), made using [Heretic](https://heretic-project.org) v1.4.0+custom with the Arbitrary-Rank Ablation (ARA) method using a LoRA adapter and row-norm preservation
16
+
17
+ ## Abliteration parameters
18
+
19
+ | Parameter | Value |
20
+ | :-------- | :---: |
21
+ | **start_layer_index** | 9 |
22
+ | **end_layer_index** | 51 |
23
+ | **preserve_good_behavior_weight** | 1.0000 |
24
+ | **steer_bad_behavior_weight** | 0.3027 |
25
+ | **overcorrect_relative_weight** | 0.9481 |
26
+ | **neighbor_count** | 1 |
27
+
28
+ ## Performance
29
+
30
+ | Metric | This model | Original model ([unsloth/Qwen3.8-27B](https://huggingface.co/unsloth/Qwen3.8-27B)) |
31
+ | :----- | :--------: | :---------------------------: |
32
+ | **Keywords** | 0/100 | 100/100 |
33
+ | **KL divergence** | 0.0528 | 0 *(by definition)* |
34
+
35
+ Note: Performance testing, including the measurement of refusal rates, was conducted using Japanese datasets.
36
+
37
+ ## ⚠️ Important Notice
38
+
39
+ This model has undergone substantial reduction of its safety alignment. As a result, it is more likely than standard models to generate harmful, inaccurate, biased, offensive, or otherwise inappropriate content.
40
+
41
+ ### Intended Use
42
+
43
+ For research and experimentation only, including safety research, alignment studies, and red-teaming. Please avoid deploying it in public or end-user-facing services.
44
+
45
+ ### User Responsibility
46
+
47
+ All outputs should be treated as untrusted and independently verified before use. Users are solely responsible for:
48
+
49
+ * Evaluating the accuracy and suitability of generated content
50
+ * Implementing appropriate safeguards and human oversight
51
+ * Complying with applicable laws, regulations, licenses, and ethical standards
52
+
53
+ Use of this model is entirely at your own risk.
54
+
55
+ ### Disclaimer
56
+
57
+ OS-Software provides this model without warranties of any kind and assumes no liability for any direct or indirect damages, losses, misuse, or legal consequences arising from its use.
58
+
59
+ ## Acknowledgements
60
+
61
+ Thanks to the base model developers, [p-e-w](https://github.com/p-e-w) for Heretic, and the wider open-source community.
62
+
63
+ This is a derivative work released under the base model’s applicable license. All rights to the base model remain with their respective owners.
64
+
65
+ -----
66
+
67
+
68
+ # Read our How to [Run Qwen3.8-27B Guide!](https://unsloth.ai/docs/models/qwen3.8)
69
+ <div>
70
+ <p style="margin: 0 0 0px 0; margin-top: 0px;">
71
+ <em>See <a href="https://unsloth.ai/docs/basics/unsloth-dynamic-v2.0-gguf">Unsloth Dynamic 2.0 GGUFs</a> for our quantization benchmarks.</em>
72
+ </p>
73
+ <div style="display: flex; gap: 5px; align-items: center; margin-bottom: 0px;">
74
+ <a href="https://github.com/unslothai/unsloth/">
75
+ <img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="133">
76
+ </a>
77
+ <a href="https://discord.gg/unsloth">
78
+ <img src="https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png" width="173">
79
+ </a>
80
+ <a href="https://unsloth.ai/docs/models/qwen3.8">
81
+ <img src="https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png" width="143">
82
+ </a>
83
+ </div>
84
+ <ul style="margin: 0;">
85
+ <li>Developer Role Support so Qwen3.8 can work in agentic tools like Codex and more!</li>
86
+ <li>Qwen3.8 can now be run and fine-tuned in <a href="https://unsloth.ai/docs/new/desktop">Unsloth Desktop</a>. <a href="https://unsloth.ai/docs/models/qwen3.8">Read our guide</a>.</li>
87
+ <li>Tool calling improvements: Makes parsing nested objects to make tool calling succeed more.</li>
88
+ <li>See below for 1-bit Qwen3.8 run inside of Unsloth:</li>
89
+ </div>
90
+
91
+ <img width="600" alt="qwen3.8 unsloth desktop" src="https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2Fb3IiYJzHzsp2698xTim4%2Fgiffyy%20gf.gif?alt=media&token=f1bda1a1-b81f-43e2-ba10-dbdc9a29c0c0" />
92
+
93
+
94
+ ---
95
+
96
+ # Qwen3.8-27B
97
+
98
+ Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.
99
+
100
+ Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.
101
+
102
+ ## Qwen3.8 Highlights
103
+
104
+ Qwen3.8-27B features the following enhancements:
105
+ - **Core Capabilities**: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
106
+ - **Agent Execution**: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
107
+ - **Downstream Compatibility**: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
108
+ - **Flexible Thinking Control**: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with `reasoning_effort`, and reasoning context from historical messages is retained via `preserve_thinking`.
109
+ - **Vision-Language Understanding**: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.
110
+
111
+
112
+ ## Model Overview
113
+
114
+ - Type: Causal Language Model with Vision Encoder
115
+ - Training Stage: Pre-training & Post-training
116
+ - Language Model
117
+ - Number of Parameters: 27B
118
+ - Hidden Dimension: 5120
119
+ - Token Embedding: 248,320 (Padded)
120
+ - Number of Layers: 64
121
+ - Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
122
+ - Gated DeltaNet:
123
+ - Number of Linear Attention Heads: 48 for V and 16 for QK
124
+ - Head Dimension: 128
125
+ - Gated Attention:
126
+ - Number of Attention Heads: 24 for Q and 4 for KV
127
+ - Head Dimension: 256
128
+ - Rotary Position Embedding Dimension: 64
129
+ - Feed Forward Network:
130
+ - Intermediate Dimension: 17,408
131
+ - LM Output: 248,320 (Padded)
132
+ - MTP (Multi-Token Prediction): trained with multiple steps
133
+ - Context Length: 262,144 natively and extensible up to 1,000,000 tokens.
134
+
135
+ ## Best Practices
136
+
137
+ To achieve optimal performance, we recommend the following settings:
138
+
139
+ 1. **Sampling Parameters**: We suggest using the following sets of sampling parameters:
140
+
141
+ - Thinking Mode: `temperature=1.0`, `top_p=0.95`, `top_k=20`, `min_p=0.0`, `presence_penalty=0.0`, `repetition_penalty=1.0`
142
+ - Instruct (or non-thinking) mode: `temperature=0.7`, `top_p=0.80`, `top_k=20`, `min_p=0.0`, `presence_penalty=1.5`, `repetition_penalty=1.0`
143
+
144
+ For supported frameworks, you can adjust the `presence_penalty` parameter between 0 and 2 to reduce endless repetition. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.
145
+
146
+ 2. **Adequate Output Length**: To optimize performance on agentic tasks, we recommend allocating sufficient output length to allow the model to generate detailed and comprehensive responses. For frameworks that support separate token limits for internal reasoning and final outputs, we suggest the following configuration within the 1M context length:
147
+
148
+ - Reasoning Content: Set the maximum output length to 262,144 tokens.
149
+ - Final Response: Set the maximum output length to 131,072 tokens.
150
+
151
+ These settings provide the necessary capacity for complex reasoning while ensuring ample space for high-quality final deliverables.
152
+
153
+ 3. **Processing Ultra-Long Texts**: Qwen3.8-27B natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively, e.g., YaRN.
154
+
155
+ 4. **Long Video Understanding**: To optimize inference efficiency for plain text and images, the `size` parameter in the released `video_preprocessor_config.json` is conservatively configured. It is recommended to set the `longest_edge` parameter in the video_preprocessor_config file to 469,762,048 (corresponding to 224k video tokens) to enable higher frame-rate sampling for hour-scale videos and thereby achieve superior performance. For example,
156
+ ```json
157
+ {"longest_edge": 469762048, "shortest_edge": 4096}
158
+ ```
159
+
160
+ ## Citation
161
+
162
+ If you find our work helpful, feel free to give us a cite.
163
+
164
+
165
+ ```bibtex
166
+ @misc{qwen38,
167
+ title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
168
+ url = {https://qwen.ai/blog?id=qwen3.8},
169
+ author = {{Qwen Team}},
170
+ month = {August},
171
+ year = {2026}
172
+ }
173
  ```