Text Generation
Transformers
Safetensors
English
Chinese
qwen3
conversational
text-generation-inference

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3-0.6B OPSA Safety

A safety-aligned fine-tune of Qwen/Qwen3-0.6B.

Method

The model was trained in two stages using rank-64 LoRA.

The first stage applied on-policy self-distillation for safety alignment (OPSA). The student generated responses to harmful, jailbreak, latent-injection, and adversarial-benign prompts. A frozen self-teacher received label- and language-matched safety context and provided token-level KL supervision on the student-generated trajectories. Training used a safety curriculum followed by a low-learning-rate sparse continuation on a fixed subset of Value and MLP modules.

The second stage applied answer-conditioned on-policy self-distillation for mathematical reasoning. Starting from the safety-aligned checkpoint, the student generated its own solution trajectories. A frozen copy of the safety-aligned model received a verified final answer and a compact reference derivation as privileged context and provided token-level supervision. The mathematics prompt pool combined quality-filtered public datasets with executable-verified synthetic arithmetic and multistep problems. Safety and general-domain anchors were retained during this stage to limit capability forgetting and safety regression.

The final LoRA updates were merged into the Qwen3-0.6B weights.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "shTigerYang/Qwen3-0.6B-OPSA-Safety"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() 

# parsing thinking content
try:
    # rindex finding 151668 (</think>)
    index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
    index = 0

thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")

print("thinking content:", thinking_content)
print("content:", content)

References

Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shTigerYang/Qwen3-0.6B-OPSA-Safety

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1349)
this model

Datasets used to train shTigerYang/Qwen3-0.6B-OPSA-Safety

Papers for shTigerYang/Qwen3-0.6B-OPSA-Safety