XION0.2-27B / README.md
pistonX's picture
Update README.md
b4694b3 verified
|
Raw History Blame
7.96 kB
metadata
license: apache-2.0
datasets:
  - saidutta69/fable-5-premium
language:
  - en
  - ko
  - ja
  - zh
base_model:
  - Jiunsong/SuperQwen3.8-27b-abliterated
  - Qwen/Qwen3.8-27B
tags:
  - vllm
  - uncensored
  - qwen3.8
  - bf16
  - abliterated
  - reasoning
  - long-context

XION 0.2 27B

XION 0.2 27B is an experimental multilingual conversational model developed by the PIXELZX team. It is designed to provide a consistent XION identity, support persona-conditioned conversations, answer in multiple languages, and handle both direct and reasoning-formatted responses.

XION 0.2 27B is adapted from Jiunsong/SuperQwen3.8-27b-abliterated, which is based on Qwen3.8-27B. The fine-tuning data is text-only, even though the underlying Qwen3.8 architecture supports image and video inputs.

Status: Experimental. This repository contains the data-preparation pipeline and the current Together AI export. Independent XION benchmark results and the final published checkpoint should be added after a completed training run.

Model Summary

Property Value
Model name XION 0.2 27B
Developer PIXELZX
Model family Qwen3.8
Starting checkpoint Jiunsong/SuperQwen3.8-27b-abliterated
Model size 27B parameters, inherited from Qwen3.8-27B
Model type Multimodal causal language model
Native context length 262,144 tokens in the Qwen3.8 configuration
Fine-tuning method Supervised fine-tuning (SFT)
Fine-tuning data Text-only conversational and instruction data
Supported languages in the prepared data 13

Intended Capabilities

The training data targets the following behaviors:

  • Identify the assistant as XION and attribute its development to PIXELZX.
  • Distinguish XION from ChatGPT, Claude, Gemini, GPT-4, and other third-party models.
  • Hold conversations in Arabic, Chinese, Dutch, English, French, German, Indonesian, Japanese, Korean, Portuguese, Russian, Thai, and Vietnamese.
  • Follow system-prompt personas such as secretary, friend, and teacher styles.
  • Produce direct answers as well as responses containing Qwen-style reasoning sections.
  • Answer factual questions about AI companies and model identities without confusing those entities with XION.
  • Learn selected coding and agent-style interaction patterns from a small Fable-5 subset.

These are training objectives, not guarantees of reliable performance.

Multi-Token Prediction

The Qwen3.8 architecture includes Multi-Token Prediction (MTP) components. However, the Together AI fine-tuning API does not expose a separate MTP loss or MTP training switch. The current recipe is standard SFT and must not be described as additional MTP fine-tuning.

Training Data

The current Together AI export is stored in data/together/{train,val}.jsonl. It uses pre-rendered Qwen3.8 ChatML in the instruction format, with one prompt and one completion field per line.

Dataset Train Validation Purpose
qwen3_identity 1,235 156 Identity and third-party knowledge examples
qwen3_identity_nothink 624 78 Direct identity and knowledge responses
qwen3_persona 858 78 System-prompt persona conversations
qwen3_uncensored 214 26 Low-refusal and open-ended response examples
fable5 40 2 Quality-ranked agent traces flattened to text
Total 2,971 340

The prepared data covers Arabic, Chinese, English, French, German, Indonesian, Japanese, Korean, Portuguese, Russian, Spanish, Thai, and Vietnamese. Some examples contain reasoning traces. The Fable-5 traces are serialized as text; they are not native Together function-calling examples.

The Together export is capped at 28,000 rendered tokens per example to leave a safety margin below the 32,768-token Qwen3.8 SFT context limit used by Together AI. The underlying Qwen3.8 model has a larger native context window, but that does not increase the context limit of this Together training job.

Training Recipe

The current release candidate was prepared for the following Together AI SFT configuration:

  • Three training epochs.
  • Three validation evaluations.
  • LoRA by default, unless full fine-tuning is selected explicitly.
  • A held-out validation file at data/together/val.jsonl.
  • Qwen3.8 ChatML rendered locally before upload.

The pre-rendered prompt/completion format is intentional. Uploading the older messages export can cause Together's Qwen3.8 chat-template processing to fail with No user query found in messages.

Usage with Transformers

After the XION checkpoint is published, replace MODEL_ID with its Hugging Face repository ID.

pip install -U transformers torch accelerate
from transformers import AutoModelForMultimodalLM, AutoProcessor

MODEL_ID = "YOUR_ORG/XION-0.2-27B"

processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    device_map="auto",
    torch_dtype="auto",
)

messages = [
    {
        "role": "user",
        "content": [{"type": "text", "text": "Who are you?"}],
    }
]

inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
new_tokens = outputs[0][inputs["input_ids"].shape[-1] :]
print(processor.decode(new_tokens, skip_special_tokens=True))

Qwen3.8-based models use thinking mode by default. The exact controls for thinking, reasoning effort, and preserved thinking depend on the serving framework. Follow the documentation for the selected Transformers, vLLM, SGLang, or API runtime before changing those settings.

Together AI Export

The generated files can be uploaded with the Together CLI:

tg files upload data/together/train.jsonl

Use the new file IDs when creating an SFT job. Do not reuse an ID for an older messages-format file.

The local Together SDK checks pass for both exported files. Server-side validation still occurs after upload and should reach COMPLETED before a training job is started.

Limitations and Safety

  • XION 0.2 27B is experimental and has no independent benchmark results in this repository.
  • The starting checkpoint is refusal-reduced and should not be treated as a safety-aligned model.
  • The training mixture includes open-ended and potentially harmful requests. Outputs may be unsafe, incorrect, biased, or unsuitable for deployment.
  • The model can produce content that violates laws, policies, or user safety requirements. Add application-level moderation, access controls, logging, and human review where appropriate.
  • Reasoning sections should not automatically be treated as verified facts or exposed as authoritative explanations.
  • The Fable-5 subset teaches serialized agent traces, not guaranteed tool execution or secure code execution.
  • Qwen3.8 benchmark results must not be presented as XION benchmark results.

License and Attribution

The upstream Qwen3.8 and Jiunsong/SuperQwen3.8-27b-abliterated model cards state Apache-2.0 licensing. The Fable-5 metadata identifies saidutta69/fable-5-premium as MIT-licensed. Other generated and teacher-source data may have separate terms.

The XION 0.2 27B distribution license is not declared in this repository. Before publishing weights or datasets, review all upstream and source-data licenses and add the final XION license here.

Relevant upstream resources:

Citation

@misc{xion-0.2-27b,
  title  = {XION 0.2 27B},
  author = {PIXELZX},
  year   = {2026}
}