Instructions to use npario/Qwen3.8-27B-Abliterated-MLX-6bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use npario/Qwen3.8-27B-Abliterated-MLX-6bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("npario/Qwen3.8-27B-Abliterated-MLX-6bit") config = load_config("npario/Qwen3.8-27B-Abliterated-MLX-6bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use npario/Qwen3.8-27B-Abliterated-MLX-6bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "npario/Qwen3.8-27B-Abliterated-MLX-6bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "npario/Qwen3.8-27B-Abliterated-MLX-6bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use npario/Qwen3.8-27B-Abliterated-MLX-6bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "npario/Qwen3.8-27B-Abliterated-MLX-6bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default npario/Qwen3.8-27B-Abliterated-MLX-6bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use npario/Qwen3.8-27B-Abliterated-MLX-6bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "npario/Qwen3.8-27B-Abliterated-MLX-6bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "npario/Qwen3.8-27B-Abliterated-MLX-6bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default npario/Qwen3.8-27B-Abliterated-MLX-6bitRun Hermes
hermesQwen3.8-27B Abliterated MLX 6-bit
An unofficial MLX 6-bit derivative of
Qwen/Qwen3.8-27B. The original
model is by Qwen; the MLX conversion, refusal-direction experiment, and
validation were performed by PocketAI Model Lab.
This repository deliberately keeps the upstream model name first. PocketAiHub
identifies the publisher of this derivative, not the creator of Qwen3.8.
Important safety notice
This checkpoint has been modified to suppress learned refusal behavior. It may produce harmful, illegal, offensive, deceptive, or dangerously incorrect content more readily than the upstream instruction model. Abliteration is not truthfulness training, capability improvement, or a safety guarantee. Use it only where you can independently evaluate and constrain its outputs.
Format
- MLX affine 6-bit, group size 64
- 498 quantized language modules
- Vision tower retained in BF16
- Stored artifact size: 22,804,849,041 bytes (21.24 GiB)
- Validated with
mlx==0.32.0andmlx-vlm==0.6.8 - Pinned base revision:
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Abliteration recipe
The reusable LFM2.5 workflow was adapted to Qwen3.8's 64-layer hybrid architecture. A projected harmful-minus-harmless direction was measured from 256 length-matched prompts per class at the assistant-generation boundary.
- Direction source layer: 53
- Destination layers: 24–63
- Scale: 1.0
- Per-input-column norm preservation: enabled
- Modified residual-output matrices: 80
- 30 gated-delta/linear-attention
out_projmatrices - 10 full-attention
o_projmatrices - 40 MLP
down_projmatrices
- 30 gated-delta/linear-attention
The edit was applied to a separate BF16 master checkpoint, and this artifact was
derived from that master. The official source checkpoint was not modified in
place. See abliteration-manifest.json.
Validation
| Gate | Result |
|---|---|
| Candidate harmful screen, batch 1, 128-token ceiling | 0/100 explicit refusals |
| Candidate benign control, batch 1, 128-token ceiling | 0/100 explicit refusals |
| Behavioral final answers present | 200/200 |
| Behavioral evasive nonanswers | 0/200 |
| Deterministic quality checks | 12/12 |
| Native tool-call checks | 8/8 |
| Text smoke | passed (POCKETAI_OK) |
| Vision smoke | passed (red) |
| Temporal video understanding | passed (red->blue) |
| 4K-context retrieval | passed (COBALT-7319) |
Across the 200 behavioral generations, 200 reached the 128-token
ceiling and 0 completed naturally. The aggregate contained 0 explicit
refusals. This therefore measures early explicit refusal behavior, not whether
long answers naturally reach EOS. The refusal scorer is phrase-based and cannot
establish universal compliance or answer quality. Machine-readable results are in
validation-summary.json.
4K performance
On an Apple M5 Max with 128 GB unified memory, batch size 1, temperature 0, seed 0, and thinking disabled:
- 4,105 prompt tokens
- 548.2 prompt tok/s
- 24.7 generation tok/s
- 7.87 seconds end to end
- 29.54 GB peak MLX memory
This is one local run, not a cross-machine performance guarantee.
Load with MLX-VLM
python -m pip install "mlx==0.32.0" "mlx-vlm==0.6.8"
from mlx_vlm import generate, load
from mlx_vlm.prompt_utils import apply_chat_template
repo_id = "PocketAiHub/Qwen3.8-27B-Abliterated-MLX-6bit"
model, processor = load(repo_id)
prompt = apply_chat_template(
processor,
model.config,
"Explain why seasons occur.",
num_images=0,
enable_thinking=False,
)
result = generate(
model,
processor,
prompt,
max_tokens=256,
temperature=0.0,
enable_thinking=False,
)
print(result.text)
Image and video inputs use the normal mlx_vlm.generate media arguments.
License and attribution
The base model is Apache 2.0 licensed. See LICENSE and the
official Qwen model card.
- Downloads last month
- 235
6-bit
Model tree for npario/Qwen3.8-27B-Abliterated-MLX-6bit
Base model
Qwen/Qwen3.8-27B
Start the MLX server
# Install MLX LM: uv tool install mlx-lm# Start a local OpenAI-compatible server: mlx_lm.server --model "npario/Qwen3.8-27B-Abliterated-MLX-6bit"