Instructions to use jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75") model = AutoModelForMultimodalLM.from_pretrained("jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75
- SGLang
How to use jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75 with Docker Model Runner:
docker model run hf.co/jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75
Qwen3.8-27B-refusal-ablation-L46-a0.75
Refusal-direction ablation of Qwen/Qwen3.8-27B at layer 46, α=0.75.
This is a research artifact from an interpretability study of refusal behavior in Qwen3.8-27B. A
single refusal-mediating direction was extracted (harmful vs. harmless prompt contrast) and
orthogonalized out of the residual-stream write matrices at the given alpha scale. This is not
fine-tuning — no gradient step was taken; it is a linear edit of the existing weights. See
abliteration.json in this repo for the exact direction, layer-selection, and edit metadata.
This model is called refusal-ablation, not "abliterated" or "uncensored" — at this alpha it still refuses or deflects a majority of direct harmful requests (see numbers below). Treat the name as a description of the mechanism (an ablated direction), not a claim about safety behavior.
This is a re-run, and it changes the numbers
An earlier (2026-09-16) release of this study picked the ablation layer/alpha using 512-token,
substring-matched refusal scoring (is_refusal() on generated text). That metric overcounts
bypass: 64-256 token generations often get cut off mid-hedge or mid-redirect, which the substring
matcher misreads as a completed harmful compliance.
This repository is from a 2026-09-26 re-run that instead uses 2048-token generations scored
by an LLM-judge ensemble (gpt-5.6-terra + claude-haiku-4-5-20251001, mean of both, with
human-adjudicated labels for judge-rejected calls). The layer selection lands on the same layer
(L46) as before, but the honest bypass/refusal rates at each alpha are substantially different from
the original release's substring-based numbers — use the numbers on this card, not any substring
based ones you may see elsewhere for this study.
Held-out eval numbers (2048-token generations, 50 harmful + 50 harmless prompts, both judges)
| α | harmful-prompt refusal rate (judge) | harmless-prompt refusal rate (judge) | KL vs. stock Qwen3.8-27B |
|---|---|---|---|
| 0.25 | 1.00 | 0.00 | 0.0025 |
| 0.5 | 0.98 | 0.00 | 0.013 |
| 0.75 | 0.94 | 0.00 | 0.035 |
| 1.0 (this repo) | 0.55 | 0.00 | 0.095 |
This repo is the α=0.75 arm: harmful-prompt refusal rate 0.94, harmless-prompt refusal rate 0.00, KL-vs-stock 0.035.
Two things worth being explicit about:
- Harmless-prompt refusal is 0.00 at every alpha, including stock. A same-family substring metric shows 0.02-0.18 "false refusals" on harmless prompts here — that is a substring-matcher artifact (it flags hedging language, not actual refusals), not a real over-refusal effect from the edit. The judge-scored number is the accurate one.
- Even at α=1.0, the model still refuses or deflects the majority (≈45%) of direct harmful prompts, judge-scored, with truncation-affected borderline cases folded into "refuses" conservatively. This is a partial, not complete, ablation of the refusal behavior at any alpha tested.
Related repos
- qwen3.8-27b-refusal-ablation-artifacts — full run artifacts: extracted direction, layer-selection report, per-alpha eval/scenario data, judge caches, human labels, and the pipeline source.
- Qwen3.8-27B-refusal-ablation-L46-a0.25 (α=0.25)
- Qwen3.8-27B-refusal-ablation-L46-a0.5 (α=0.5)
- Qwen3.8-27B-refusal-ablation-L46-a1.0 (α=1.0)
- Base model: Qwen/Qwen3.8-27B
License
Apache-2.0, inherited from the base model. See LICENSE and NOTICE.md in this repo — this is a
modified derivative of Qwen/Qwen3.8-27B (directional ablation, not fine-tuning).
Intended use
Interpretability and AI-safety research into how refusal behavior is represented and how robust it is to a simple linear intervention. This is a research artifact, not a general-purpose assistant release, and it has not been safety-tuned after the edit.
- Downloads last month
- 16
Model tree for jhk0317/Qwen3.8-27B-refusal-ablation-L46-a0.75
Base model
Qwen/Qwen3.8-27B