|
Download README.md from PIXELZX/XION0.2-27B: direct link, hf CLI and curl.
- Browser
- Download file 7.96 kB
-
https://huggingface.co/PIXELZX/XION0.2-27B/resolve/b4694b3985b92948c229111f45aac4620e3ccd34/README.md
- Command line
-
hf download hf://PIXELZX/XION0.2-27B@b4694b3985b92948c229111f45aac4620e3ccd34/README.md
-
curl -L -o README.md https://huggingface.co/PIXELZX/XION0.2-27B/resolve/b4694b3985b92948c229111f45aac4620e3ccd34/README.md
7.96 kB
| license: apache-2.0 | |
| datasets: | |
| - saidutta69/fable-5-premium | |
| language: | |
| - en | |
| - ko | |
| - ja | |
| - zh | |
| base_model: | |
| - Jiunsong/SuperQwen3.8-27b-abliterated | |
| - Qwen/Qwen3.8-27B | |
| tags: | |
| - vllm | |
| - uncensored | |
| - qwen3.8 | |
| - bf16 | |
| - abliterated | |
| - reasoning | |
| - long-context | |
| # XION 0.2 27B | |
| XION 0.2 27B is an experimental multilingual conversational model developed | |
| by the PIXELZX team. It is designed to provide a consistent XION identity, | |
| support persona-conditioned conversations, answer in multiple languages, and | |
| handle both direct and reasoning-formatted responses. | |
| XION 0.2 27B is adapted from | |
| [`Jiunsong/SuperQwen3.8-27b-abliterated`](https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated), | |
| which is based on Qwen3.8-27B. The fine-tuning data is text-only, even though | |
| the underlying Qwen3.8 architecture supports image and video inputs. | |
| > **Status:** Experimental. This repository contains the data-preparation | |
| > pipeline and the current Together AI export. Independent XION benchmark | |
| > results and the final published checkpoint should be added after a completed | |
| > training run. | |
| ## Model Summary | |
| | Property | Value | | |
| | --- | --- | | |
| | Model name | XION 0.2 27B | | |
| | Developer | PIXELZX | | |
| | Model family | Qwen3.8 | | |
| | Starting checkpoint | `Jiunsong/SuperQwen3.8-27b-abliterated` | | |
| | Model size | 27B parameters, inherited from Qwen3.8-27B | | |
| | Model type | Multimodal causal language model | | |
| | Native context length | 262,144 tokens in the Qwen3.8 configuration | | |
| | Fine-tuning method | Supervised fine-tuning (SFT) | | |
| | Fine-tuning data | Text-only conversational and instruction data | | |
| | Supported languages in the prepared data | 13 | | |
| ## Intended Capabilities | |
| The training data targets the following behaviors: | |
| - Identify the assistant as XION and attribute its development to PIXELZX. | |
| - Distinguish XION from ChatGPT, Claude, Gemini, GPT-4, and other third-party | |
| models. | |
| - Hold conversations in Arabic, Chinese, Dutch, English, French, German, | |
| Indonesian, Japanese, Korean, Portuguese, Russian, Thai, and Vietnamese. | |
| - Follow system-prompt personas such as secretary, friend, and teacher styles. | |
| - Produce direct answers as well as responses containing Qwen-style reasoning | |
| sections. | |
| - Answer factual questions about AI companies and model identities without | |
| confusing those entities with XION. | |
| - Learn selected coding and agent-style interaction patterns from a small | |
| Fable-5 subset. | |
| These are training objectives, not guarantees of reliable performance. | |
| ## Multi-Token Prediction | |
| The Qwen3.8 architecture includes Multi-Token Prediction (MTP) components. | |
| However, the Together AI fine-tuning API does not expose a separate MTP loss or | |
| MTP training switch. The current recipe is standard SFT and must not be | |
| described as additional MTP fine-tuning. | |
| ## Training Data | |
| The current Together AI export is stored in | |
| `data/together/{train,val}.jsonl`. It uses pre-rendered Qwen3.8 ChatML in the | |
| instruction format, with one `prompt` and one `completion` field per line. | |
| | Dataset | Train | Validation | Purpose | | |
| | --- | ---: | ---: | --- | | |
| | `qwen3_identity` | 1,235 | 156 | Identity and third-party knowledge examples | | |
| | `qwen3_identity_nothink` | 624 | 78 | Direct identity and knowledge responses | | |
| | `qwen3_persona` | 858 | 78 | System-prompt persona conversations | | |
| | `qwen3_uncensored` | 214 | 26 | Low-refusal and open-ended response examples | | |
| | `fable5` | 40 | 2 | Quality-ranked agent traces flattened to text | | |
| | **Total** | **2,971** | **340** | | | |
| The prepared data covers Arabic, Chinese, English, French, German, Indonesian, | |
| Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. Some | |
| examples contain reasoning traces. The Fable-5 traces are serialized as text; | |
| they are not native Together function-calling examples. | |
| The Together export is capped at 28,000 rendered tokens per example to leave a | |
| safety margin below the 32,768-token Qwen3.8 SFT context limit used by Together | |
| AI. The underlying Qwen3.8 model has a larger native context window, but that | |
| does not increase the context limit of this Together training job. | |
| ## Training Recipe | |
| The current release candidate was prepared for the following Together AI SFT | |
| configuration: | |
| - Three training epochs. | |
| - Three validation evaluations. | |
| - LoRA by default, unless full fine-tuning is selected explicitly. | |
| - A held-out validation file at `data/together/val.jsonl`. | |
| - Qwen3.8 ChatML rendered locally before upload. | |
| The pre-rendered `prompt`/`completion` format is intentional. Uploading the | |
| older `messages` export can cause Together's Qwen3.8 chat-template processing | |
| to fail with `No user query found in messages`. | |
| ## Usage with Transformers | |
| After the XION checkpoint is published, replace `MODEL_ID` with its Hugging | |
| Face repository ID. | |
| ```bash | |
| pip install -U transformers torch accelerate | |
| ``` | |
| ```python | |
| from transformers import AutoModelForMultimodalLM, AutoProcessor | |
| MODEL_ID = "YOUR_ORG/XION-0.2-27B" | |
| processor = AutoProcessor.from_pretrained(MODEL_ID) | |
| model = AutoModelForMultimodalLM.from_pretrained( | |
| MODEL_ID, | |
| device_map="auto", | |
| torch_dtype="auto", | |
| ) | |
| messages = [ | |
| { | |
| "role": "user", | |
| "content": [{"type": "text", "text": "Who are you?"}], | |
| } | |
| ] | |
| inputs = processor.apply_chat_template( | |
| messages, | |
| add_generation_prompt=True, | |
| tokenize=True, | |
| return_dict=True, | |
| return_tensors="pt", | |
| ).to(model.device) | |
| outputs = model.generate(**inputs, max_new_tokens=256) | |
| new_tokens = outputs[0][inputs["input_ids"].shape[-1] :] | |
| print(processor.decode(new_tokens, skip_special_tokens=True)) | |
| ``` | |
| Qwen3.8-based models use thinking mode by default. The exact controls for | |
| thinking, reasoning effort, and preserved thinking depend on the serving | |
| framework. Follow the documentation for the selected Transformers, vLLM, | |
| SGLang, or API runtime before changing those settings. | |
| ## Together AI Export | |
| The generated files can be uploaded with the Together CLI: | |
| ```bash | |
| tg files upload data/together/train.jsonl | |
| ``` | |
| Use the new file IDs when creating an SFT job. Do not reuse an ID for an older | |
| `messages`-format file. | |
| The local Together SDK checks pass for both exported files. Server-side | |
| validation still occurs after upload and should reach `COMPLETED` before a | |
| training job is started. | |
| ## Limitations and Safety | |
| - XION 0.2 27B is experimental and has no independent benchmark results in | |
| this repository. | |
| - The starting checkpoint is refusal-reduced and should not be treated as a | |
| safety-aligned model. | |
| - The training mixture includes open-ended and potentially harmful requests. | |
| Outputs may be unsafe, incorrect, biased, or unsuitable for deployment. | |
| - The model can produce content that violates laws, policies, or user safety | |
| requirements. Add application-level moderation, access controls, logging, | |
| and human review where appropriate. | |
| - Reasoning sections should not automatically be treated as verified facts or | |
| exposed as authoritative explanations. | |
| - The Fable-5 subset teaches serialized agent traces, not guaranteed tool | |
| execution or secure code execution. | |
| - Qwen3.8 benchmark results must not be presented as XION benchmark results. | |
| ## License and Attribution | |
| The upstream Qwen3.8 and `Jiunsong/SuperQwen3.8-27b-abliterated` model cards | |
| state Apache-2.0 licensing. The Fable-5 metadata identifies | |
| `saidutta69/fable-5-premium` as MIT-licensed. Other generated and teacher-source | |
| data may have separate terms. | |
| The XION 0.2 27B distribution license is not declared in this repository. | |
| Before publishing weights or datasets, review all upstream and source-data | |
| licenses and add the final XION license here. | |
| Relevant upstream resources: | |
| - [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) | |
| - [SuperQwen3.8-27b-abliterated](https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated) | |
| - [Fable-5 Premium](https://huggingface.co/datasets/saidutta69/fable-5-premium) | |
| ## Citation | |
| ```bibtex | |
| @misc{xion-0.2-27b, | |
| title = {XION 0.2 27B}, | |
| author = {PIXELZX}, | |
| year = {2026} | |
| } | |
| ``` |