pistonX commited on
Commit
b4694b3
·
verified ·
1 Parent(s): ee310ec

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +204 -1
README.md CHANGED
@@ -18,4 +18,207 @@ tags:
18
  - abliterated
19
  - reasoning
20
  - long-context
21
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
  - abliterated
19
  - reasoning
20
  - long-context
21
+ ---
22
+
23
+ # XION 0.2 27B
24
+
25
+ XION 0.2 27B is an experimental multilingual conversational model developed
26
+ by the PIXELZX team. It is designed to provide a consistent XION identity,
27
+ support persona-conditioned conversations, answer in multiple languages, and
28
+ handle both direct and reasoning-formatted responses.
29
+
30
+ XION 0.2 27B is adapted from
31
+ [`Jiunsong/SuperQwen3.8-27b-abliterated`](https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated),
32
+ which is based on Qwen3.8-27B. The fine-tuning data is text-only, even though
33
+ the underlying Qwen3.8 architecture supports image and video inputs.
34
+
35
+ > **Status:** Experimental. This repository contains the data-preparation
36
+ > pipeline and the current Together AI export. Independent XION benchmark
37
+ > results and the final published checkpoint should be added after a completed
38
+ > training run.
39
+
40
+ ## Model Summary
41
+
42
+ | Property | Value |
43
+ | --- | --- |
44
+ | Model name | XION 0.2 27B |
45
+ | Developer | PIXELZX |
46
+ | Model family | Qwen3.8 |
47
+ | Starting checkpoint | `Jiunsong/SuperQwen3.8-27b-abliterated` |
48
+ | Model size | 27B parameters, inherited from Qwen3.8-27B |
49
+ | Model type | Multimodal causal language model |
50
+ | Native context length | 262,144 tokens in the Qwen3.8 configuration |
51
+ | Fine-tuning method | Supervised fine-tuning (SFT) |
52
+ | Fine-tuning data | Text-only conversational and instruction data |
53
+ | Supported languages in the prepared data | 13 |
54
+
55
+ ## Intended Capabilities
56
+
57
+ The training data targets the following behaviors:
58
+
59
+ - Identify the assistant as XION and attribute its development to PIXELZX.
60
+ - Distinguish XION from ChatGPT, Claude, Gemini, GPT-4, and other third-party
61
+ models.
62
+ - Hold conversations in Arabic, Chinese, Dutch, English, French, German,
63
+ Indonesian, Japanese, Korean, Portuguese, Russian, Thai, and Vietnamese.
64
+ - Follow system-prompt personas such as secretary, friend, and teacher styles.
65
+ - Produce direct answers as well as responses containing Qwen-style reasoning
66
+ sections.
67
+ - Answer factual questions about AI companies and model identities without
68
+ confusing those entities with XION.
69
+ - Learn selected coding and agent-style interaction patterns from a small
70
+ Fable-5 subset.
71
+
72
+ These are training objectives, not guarantees of reliable performance.
73
+
74
+ ## Multi-Token Prediction
75
+
76
+ The Qwen3.8 architecture includes Multi-Token Prediction (MTP) components.
77
+ However, the Together AI fine-tuning API does not expose a separate MTP loss or
78
+ MTP training switch. The current recipe is standard SFT and must not be
79
+ described as additional MTP fine-tuning.
80
+
81
+ ## Training Data
82
+
83
+ The current Together AI export is stored in
84
+ `data/together/{train,val}.jsonl`. It uses pre-rendered Qwen3.8 ChatML in the
85
+ instruction format, with one `prompt` and one `completion` field per line.
86
+
87
+ | Dataset | Train | Validation | Purpose |
88
+ | --- | ---: | ---: | --- |
89
+ | `qwen3_identity` | 1,235 | 156 | Identity and third-party knowledge examples |
90
+ | `qwen3_identity_nothink` | 624 | 78 | Direct identity and knowledge responses |
91
+ | `qwen3_persona` | 858 | 78 | System-prompt persona conversations |
92
+ | `qwen3_uncensored` | 214 | 26 | Low-refusal and open-ended response examples |
93
+ | `fable5` | 40 | 2 | Quality-ranked agent traces flattened to text |
94
+ | **Total** | **2,971** | **340** | |
95
+
96
+ The prepared data covers Arabic, Chinese, English, French, German, Indonesian,
97
+ Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. Some
98
+ examples contain reasoning traces. The Fable-5 traces are serialized as text;
99
+ they are not native Together function-calling examples.
100
+
101
+ The Together export is capped at 28,000 rendered tokens per example to leave a
102
+ safety margin below the 32,768-token Qwen3.8 SFT context limit used by Together
103
+ AI. The underlying Qwen3.8 model has a larger native context window, but that
104
+ does not increase the context limit of this Together training job.
105
+
106
+ ## Training Recipe
107
+
108
+ The current release candidate was prepared for the following Together AI SFT
109
+ configuration:
110
+
111
+ - Three training epochs.
112
+ - Three validation evaluations.
113
+ - LoRA by default, unless full fine-tuning is selected explicitly.
114
+ - A held-out validation file at `data/together/val.jsonl`.
115
+ - Qwen3.8 ChatML rendered locally before upload.
116
+
117
+ The pre-rendered `prompt`/`completion` format is intentional. Uploading the
118
+ older `messages` export can cause Together's Qwen3.8 chat-template processing
119
+ to fail with `No user query found in messages`.
120
+
121
+ ## Usage with Transformers
122
+
123
+ After the XION checkpoint is published, replace `MODEL_ID` with its Hugging
124
+ Face repository ID.
125
+
126
+ ```bash
127
+ pip install -U transformers torch accelerate
128
+ ```
129
+
130
+ ```python
131
+ from transformers import AutoModelForMultimodalLM, AutoProcessor
132
+
133
+ MODEL_ID = "YOUR_ORG/XION-0.2-27B"
134
+
135
+ processor = AutoProcessor.from_pretrained(MODEL_ID)
136
+ model = AutoModelForMultimodalLM.from_pretrained(
137
+ MODEL_ID,
138
+ device_map="auto",
139
+ torch_dtype="auto",
140
+ )
141
+
142
+ messages = [
143
+ {
144
+ "role": "user",
145
+ "content": [{"type": "text", "text": "Who are you?"}],
146
+ }
147
+ ]
148
+
149
+ inputs = processor.apply_chat_template(
150
+ messages,
151
+ add_generation_prompt=True,
152
+ tokenize=True,
153
+ return_dict=True,
154
+ return_tensors="pt",
155
+ ).to(model.device)
156
+
157
+ outputs = model.generate(**inputs, max_new_tokens=256)
158
+ new_tokens = outputs[0][inputs["input_ids"].shape[-1] :]
159
+ print(processor.decode(new_tokens, skip_special_tokens=True))
160
+ ```
161
+
162
+ Qwen3.8-based models use thinking mode by default. The exact controls for
163
+ thinking, reasoning effort, and preserved thinking depend on the serving
164
+ framework. Follow the documentation for the selected Transformers, vLLM,
165
+ SGLang, or API runtime before changing those settings.
166
+
167
+ ## Together AI Export
168
+
169
+ The generated files can be uploaded with the Together CLI:
170
+
171
+ ```bash
172
+ tg files upload data/together/train.jsonl
173
+ ```
174
+
175
+ Use the new file IDs when creating an SFT job. Do not reuse an ID for an older
176
+ `messages`-format file.
177
+
178
+ The local Together SDK checks pass for both exported files. Server-side
179
+ validation still occurs after upload and should reach `COMPLETED` before a
180
+ training job is started.
181
+
182
+ ## Limitations and Safety
183
+
184
+ - XION 0.2 27B is experimental and has no independent benchmark results in
185
+ this repository.
186
+ - The starting checkpoint is refusal-reduced and should not be treated as a
187
+ safety-aligned model.
188
+ - The training mixture includes open-ended and potentially harmful requests.
189
+ Outputs may be unsafe, incorrect, biased, or unsuitable for deployment.
190
+ - The model can produce content that violates laws, policies, or user safety
191
+ requirements. Add application-level moderation, access controls, logging,
192
+ and human review where appropriate.
193
+ - Reasoning sections should not automatically be treated as verified facts or
194
+ exposed as authoritative explanations.
195
+ - The Fable-5 subset teaches serialized agent traces, not guaranteed tool
196
+ execution or secure code execution.
197
+ - Qwen3.8 benchmark results must not be presented as XION benchmark results.
198
+
199
+ ## License and Attribution
200
+
201
+ The upstream Qwen3.8 and `Jiunsong/SuperQwen3.8-27b-abliterated` model cards
202
+ state Apache-2.0 licensing. The Fable-5 metadata identifies
203
+ `saidutta69/fable-5-premium` as MIT-licensed. Other generated and teacher-source
204
+ data may have separate terms.
205
+
206
+ The XION 0.2 27B distribution license is not declared in this repository.
207
+ Before publishing weights or datasets, review all upstream and source-data
208
+ licenses and add the final XION license here.
209
+
210
+ Relevant upstream resources:
211
+
212
+ - [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)
213
+ - [SuperQwen3.8-27b-abliterated](https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated)
214
+ - [Fable-5 Premium](https://huggingface.co/datasets/saidutta69/fable-5-premium)
215
+
216
+ ## Citation
217
+
218
+ ```bibtex
219
+ @misc{xion-0.2-27b,
220
+ title = {XION 0.2 27B},
221
+ author = {PIXELZX},
222
+ year = {2026}
223
+ }
224
+ ```