--- license: apache-2.0 datasets: - saidutta69/fable-5-premium language: - en - ko - ja - zh base_model: - Jiunsong/SuperQwen3.8-27b-abliterated - Qwen/Qwen3.8-27B tags: - vllm - uncensored - qwen3.8 - bf16 - abliterated - reasoning - long-context --- # XION 0.2 27B XION 0.2 27B is an experimental multilingual conversational model developed by the PIXELZX team. It is designed to provide a consistent XION identity, support persona-conditioned conversations, answer in multiple languages, and handle both direct and reasoning-formatted responses. XION 0.2 27B is adapted from [`Jiunsong/SuperQwen3.8-27b-abliterated`](https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated), which is based on Qwen3.8-27B. The fine-tuning data is text-only, even though the underlying Qwen3.8 architecture supports image and video inputs. > **Status:** Experimental. This repository contains the data-preparation > pipeline and the current Together AI export. Independent XION benchmark > results and the final published checkpoint should be added after a completed > training run. ## Model Summary | Property | Value | | --- | --- | | Model name | XION 0.2 27B | | Developer | PIXELZX | | Model family | Qwen3.8 | | Starting checkpoint | `Jiunsong/SuperQwen3.8-27b-abliterated` | | Model size | 27B parameters, inherited from Qwen3.8-27B | | Model type | Multimodal causal language model | | Native context length | 262,144 tokens in the Qwen3.8 configuration | | Fine-tuning method | Supervised fine-tuning (SFT) | | Fine-tuning data | Text-only conversational and instruction data | | Supported languages in the prepared data | 13 | ## Intended Capabilities The training data targets the following behaviors: - Identify the assistant as XION and attribute its development to PIXELZX. - Distinguish XION from ChatGPT, Claude, Gemini, GPT-4, and other third-party models. - Hold conversations in Arabic, Chinese, Dutch, English, French, German, Indonesian, Japanese, Korean, Portuguese, Russian, Thai, and Vietnamese. - Follow system-prompt personas such as secretary, friend, and teacher styles. - Produce direct answers as well as responses containing Qwen-style reasoning sections. - Answer factual questions about AI companies and model identities without confusing those entities with XION. - Learn selected coding and agent-style interaction patterns from a small Fable-5 subset. These are training objectives, not guarantees of reliable performance. ## Multi-Token Prediction The Qwen3.8 architecture includes Multi-Token Prediction (MTP) components. However, the Together AI fine-tuning API does not expose a separate MTP loss or MTP training switch. The current recipe is standard SFT and must not be described as additional MTP fine-tuning. ## Training Data The current Together AI export is stored in `data/together/{train,val}.jsonl`. It uses pre-rendered Qwen3.8 ChatML in the instruction format, with one `prompt` and one `completion` field per line. | Dataset | Train | Validation | Purpose | | --- | ---: | ---: | --- | | `qwen3_identity` | 1,235 | 156 | Identity and third-party knowledge examples | | `qwen3_identity_nothink` | 624 | 78 | Direct identity and knowledge responses | | `qwen3_persona` | 858 | 78 | System-prompt persona conversations | | `qwen3_uncensored` | 214 | 26 | Low-refusal and open-ended response examples | | `fable5` | 40 | 2 | Quality-ranked agent traces flattened to text | | **Total** | **2,971** | **340** | | The prepared data covers Arabic, Chinese, English, French, German, Indonesian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. Some examples contain reasoning traces. The Fable-5 traces are serialized as text; they are not native Together function-calling examples. The Together export is capped at 28,000 rendered tokens per example to leave a safety margin below the 32,768-token Qwen3.8 SFT context limit used by Together AI. The underlying Qwen3.8 model has a larger native context window, but that does not increase the context limit of this Together training job. ## Training Recipe The current release candidate was prepared for the following Together AI SFT configuration: - Three training epochs. - Three validation evaluations. - LoRA by default, unless full fine-tuning is selected explicitly. - A held-out validation file at `data/together/val.jsonl`. - Qwen3.8 ChatML rendered locally before upload. The pre-rendered `prompt`/`completion` format is intentional. Uploading the older `messages` export can cause Together's Qwen3.8 chat-template processing to fail with `No user query found in messages`. ## Usage with Transformers After the XION checkpoint is published, replace `MODEL_ID` with its Hugging Face repository ID. ```bash pip install -U transformers torch accelerate ``` ```python from transformers import AutoModelForMultimodalLM, AutoProcessor MODEL_ID = "YOUR_ORG/XION-0.2-27B" processor = AutoProcessor.from_pretrained(MODEL_ID) model = AutoModelForMultimodalLM.from_pretrained( MODEL_ID, device_map="auto", torch_dtype="auto", ) messages = [ { "role": "user", "content": [{"type": "text", "text": "Who are you?"}], } ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) new_tokens = outputs[0][inputs["input_ids"].shape[-1] :] print(processor.decode(new_tokens, skip_special_tokens=True)) ``` Qwen3.8-based models use thinking mode by default. The exact controls for thinking, reasoning effort, and preserved thinking depend on the serving framework. Follow the documentation for the selected Transformers, vLLM, SGLang, or API runtime before changing those settings. ## Together AI Export The generated files can be uploaded with the Together CLI: ```bash tg files upload data/together/train.jsonl ``` Use the new file IDs when creating an SFT job. Do not reuse an ID for an older `messages`-format file. The local Together SDK checks pass for both exported files. Server-side validation still occurs after upload and should reach `COMPLETED` before a training job is started. ## Limitations and Safety - XION 0.2 27B is experimental and has no independent benchmark results in this repository. - The starting checkpoint is refusal-reduced and should not be treated as a safety-aligned model. - The training mixture includes open-ended and potentially harmful requests. Outputs may be unsafe, incorrect, biased, or unsuitable for deployment. - The model can produce content that violates laws, policies, or user safety requirements. Add application-level moderation, access controls, logging, and human review where appropriate. - Reasoning sections should not automatically be treated as verified facts or exposed as authoritative explanations. - The Fable-5 subset teaches serialized agent traces, not guaranteed tool execution or secure code execution. - Qwen3.8 benchmark results must not be presented as XION benchmark results. ## License and Attribution The upstream Qwen3.8 and `Jiunsong/SuperQwen3.8-27b-abliterated` model cards state Apache-2.0 licensing. The Fable-5 metadata identifies `saidutta69/fable-5-premium` as MIT-licensed. Other generated and teacher-source data may have separate terms. The XION 0.2 27B distribution license is not declared in this repository. Before publishing weights or datasets, review all upstream and source-data licenses and add the final XION license here. Relevant upstream resources: - [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) - [SuperQwen3.8-27b-abliterated](https://huggingface.co/Jiunsong/SuperQwen3.8-27b-abliterated) - [Fable-5 Premium](https://huggingface.co/datasets/saidutta69/fable-5-premium) ## Citation ```bibtex @misc{xion-0.2-27b, title = {XION 0.2 27B}, author = {PIXELZX}, year = {2026} } ```