pistonX commited on
Commit
138b9ec
·
verified ·
1 Parent(s): 927a204

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +0 -127
README.md CHANGED
@@ -50,133 +50,6 @@ the underlying Qwen3.8 architecture supports image and video inputs.
50
  | Fine-tuning data | Text-only conversational and instruction data |
51
  | Supported languages in the prepared data | 13 |
52
 
53
- ## Intended Capabilities
54
-
55
- The training data targets the following behaviors:
56
-
57
- - Identify the assistant as XION and attribute its development to PIXELZX.
58
- - Distinguish XION from ChatGPT, Claude, Gemini, GPT-4, and other third-party
59
- models.
60
- - Hold conversations in Arabic, Chinese, Dutch, English, French, German,
61
- Indonesian, Japanese, Korean, Portuguese, Russian, Thai, and Vietnamese.
62
- - Follow system-prompt personas such as secretary, friend, and teacher styles.
63
- - Produce direct answers as well as responses containing Qwen-style reasoning
64
- sections.
65
- - Answer factual questions about AI companies and model identities without
66
- confusing those entities with XION.
67
- - Learn selected coding and agent-style interaction patterns from a small
68
- Fable-5 subset.
69
-
70
- These are training objectives, not guarantees of reliable performance.
71
-
72
- ## Multi-Token Prediction
73
-
74
- The Qwen3.8 architecture includes Multi-Token Prediction (MTP) components.
75
- However, the Together AI fine-tuning API does not expose a separate MTP loss or
76
- MTP training switch. The current recipe is standard SFT and must not be
77
- described as additional MTP fine-tuning.
78
-
79
- ## Training Data
80
-
81
- The current Together AI export is stored in
82
- `data/together/{train,val}.jsonl`. It uses pre-rendered Qwen3.8 ChatML in the
83
- instruction format, with one `prompt` and one `completion` field per line.
84
-
85
- | Dataset | Train | Validation | Purpose |
86
- | --- | ---: | ---: | --- |
87
- | `qwen3_identity` | 1,235 | 156 | Identity and third-party knowledge examples |
88
- | `qwen3_identity_nothink` | 624 | 78 | Direct identity and knowledge responses |
89
- | `qwen3_persona` | 858 | 78 | System-prompt persona conversations |
90
- | `qwen3_uncensored` | 214 | 26 | Low-refusal and open-ended response examples |
91
- | `fable5` | 40 | 2 | Quality-ranked agent traces flattened to text |
92
- | **Total** | **2,971** | **340** | |
93
-
94
- The prepared data covers Arabic, Chinese, English, French, German, Indonesian,
95
- Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. Some
96
- examples contain reasoning traces. The Fable-5 traces are serialized as text;
97
- they are not native Together function-calling examples.
98
-
99
- The Together export is capped at 28,000 rendered tokens per example to leave a
100
- safety margin below the 32,768-token Qwen3.8 SFT context limit used by Together
101
- AI. The underlying Qwen3.8 model has a larger native context window, but that
102
- does not increase the context limit of this Together training job.
103
-
104
- ## Training Recipe
105
-
106
- The current release candidate was prepared for the following Together AI SFT
107
- configuration:
108
-
109
- - Three training epochs.
110
- - Three validation evaluations.
111
- - LoRA by default, unless full fine-tuning is selected explicitly.
112
- - A held-out validation file at `data/together/val.jsonl`.
113
- - Qwen3.8 ChatML rendered locally before upload.
114
-
115
- The pre-rendered `prompt`/`completion` format is intentional. Uploading the
116
- older `messages` export can cause Together's Qwen3.8 chat-template processing
117
- to fail with `No user query found in messages`.
118
-
119
- ## Usage with Transformers
120
-
121
- After the XION checkpoint is published, replace `MODEL_ID` with its Hugging
122
- Face repository ID.
123
-
124
- ```bash
125
- pip install -U transformers torch accelerate
126
- ```
127
-
128
- ```python
129
- from transformers import AutoModelForMultimodalLM, AutoProcessor
130
-
131
- MODEL_ID = "YOUR_ORG/XION-0.2-27B"
132
-
133
- processor = AutoProcessor.from_pretrained(MODEL_ID)
134
- model = AutoModelForMultimodalLM.from_pretrained(
135
- MODEL_ID,
136
- device_map="auto",
137
- torch_dtype="auto",
138
- )
139
-
140
- messages = [
141
- {
142
- "role": "user",
143
- "content": [{"type": "text", "text": "Who are you?"}],
144
- }
145
- ]
146
-
147
- inputs = processor.apply_chat_template(
148
- messages,
149
- add_generation_prompt=True,
150
- tokenize=True,
151
- return_dict=True,
152
- return_tensors="pt",
153
- ).to(model.device)
154
-
155
- outputs = model.generate(**inputs, max_new_tokens=256)
156
- new_tokens = outputs[0][inputs["input_ids"].shape[-1] :]
157
- print(processor.decode(new_tokens, skip_special_tokens=True))
158
- ```
159
-
160
- Qwen3.8-based models use thinking mode by default. The exact controls for
161
- thinking, reasoning effort, and preserved thinking depend on the serving
162
- framework. Follow the documentation for the selected Transformers, vLLM,
163
- SGLang, or API runtime before changing those settings.
164
-
165
- ## Together AI Export
166
-
167
- The generated files can be uploaded with the Together CLI:
168
-
169
- ```bash
170
- tg files upload data/together/train.jsonl
171
- ```
172
-
173
- Use the new file IDs when creating an SFT job. Do not reuse an ID for an older
174
- `messages`-format file.
175
-
176
- The local Together SDK checks pass for both exported files. Server-side
177
- validation still occurs after upload and should reach `COMPLETED` before a
178
- training job is started.
179
-
180
  ## Limitations and Safety
181
 
182
  - XION 0.2 27B is experimental and has no independent benchmark results in
 
50
  | Fine-tuning data | Text-only conversational and instruction data |
51
  | Supported languages in the prepared data | 13 |
52
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
53
  ## Limitations and Safety
54
 
55
  - XION 0.2 27B is experimental and has no independent benchmark results in