Qwen3.6-27B-Architect-DS9-1M-qx86-hi-mlx

BeamMeUp

'Beam me up': Zeiss ZF-100-T/Nikon D300

This model is a NuSLERP merge using Qwen3.6-27B as a base:

  • nightmedia/Qwen3.5-27B-Engineer-Deckard-Claude-TNG-C
    • nightmedia/Qwen3.5-27B-Engineer-Deckard-Claude
      • DavidAU/Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking
      • DavidAU/Qwen3.5-27B-Claude-4.6-OS-INSTRUCT
    • DavidAU/Qwen3.5-27B-Star-Trek-TNG-DS9-Heretic-Uncensored-Thinking
  • DavidAU/Qwen3.5-27B-Claude-4.6-OS-INSTRUCT

Brainwaves

         arc   arc/e boolq hswag obkqa piqa  wino
bf16     0.678,0.852,0.911
mxfp8    0.690,0.867,0.909
qx86-hi  0.663,0.832,0.911
qx64-hi  0.685,0.855,0.903
mxfp4    0.679,0.858,0.911

Quant    Perplexity      Peak Memory   Tokens/sec
bf16     4.017 ± 0.026   60.75 GB      262
mxfp8    4.026 ± 0.026   34.74 GB      178
qx86-hi  3.917 ± 0.025   32.36 GB      180
qx64-hi  4.036 ± 0.026   25.64 GB      218
mxfp4    4.102 ± 0.027   21.30 GB      221

Component metrics

Qwen3.6-27B-Claude-4.6-OS

         arc   arc/e boolq hswag obkqa piqa  wino
bf16     0.683,0.858,0.910,0.797,0.494,0.820,0.755
mxfp8    0.695,0.869,0.910,0.791,0.504,0.824,0.760
qx64-hi  0.688,0.859,0.903

Quant    Perplexity      Peak Memory   Tokens/sec
mxfp8    4.006 ± 0.026   34.74 GB      187  
qx64-hi  4.098 ± 0.027   25.64 GB      208

Qwen3.6-27B-Deckard-Claude-DS9

         arc   arc/e boolq hswag obkqa piqa  wino
mxfp8    0.672,0.845,0.909
qx64-hi  0.685,0.851,0.903

Baseline model

         arc   arc/e boolq hswag obkqa piqa  wino
Qwen3.6-27B-Instruct
mxfp8    0.647,0.803,0.910,0.773,0.450,0.806,0.742
qx86-hi  0.637,0.798,0.911,0.775,0.442,0.807,0.737

This model is using the fixed jinja template from froggeric/Qwen-Fixed-Chat-Templates

Thinking toggle

Drop <|think_on|> or <|think_off|> anywhere in your system or user prompt. The template intercepts the tag, removes it from context so the model never sees it, and flips the mode.

Fast answer, no reasoning:

System: You are a coding assistant. <|think_off|>
User: What's 2+2?

Deep reasoning:

System: You are a coding assistant. <|think_on|>
User: Implement a red-black tree in Rust.

The tag syntax (<|think_on|>, <|think_off|>) uses Qwen's control-token delimiters, so it will never collide with real text. Earlier community templates used /think, which broke legitimate paths like cd /mnt/project/think.


I added a similar set of tags for handling the preserve_thinking flag:

  • Drop <|think_forget|> or <|think_remember|> anywhere in your system or user prompt to flip the flag.
  • The template intercepts the tag, removes it from context so the model never sees it, and flips the mode.

-G

Model recipe

models:
  - model: Qwen/Qwen3.6-27B
    parameters:
      weight: 1.4
  - model: nightmedia/Qwen3.5-27B-Engineer-Deckard-Claude-TNG-C
    parameters:
      weight: 0.6
merge_method: nuslerp
dtype: bfloat16
name: Qwen3.6-27B-Deckard-Claude-DS9

models:
  - model: nightmedia/Qwen3.6-27B-Claude-4.6-OS
    parameters:
      weight: 1.4
  - model: nightmedia/Qwen3.6-27B-Deckard-Claude-DS9
    parameters:
      weight: 0.6
merge_method: nuslerp
dtype: bfloat16
name: Qwen3.6-27B-Architect-DS9

Use with mlx

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Qwen3.6-27B-DS9-1M-qx86-hi-mlx")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)
Downloads last month
34
Safetensors
Model size
7B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nightmedia/Qwen3.6-27B-Architect-DS9-1M-qx86-hi-mlx

Base model

Qwen/Qwen3.5-27B
Adapter
(41)
this model

Collections including nightmedia/Qwen3.6-27B-Architect-DS9-1M-qx86-hi-mlx