Solstice-AI Banner

Qwen3.8-27B-TURBO-Fable-Cold-Fusion (Apple MLX oQ4e 1M Context)

Official Solstice-AI 1-Million Token MLX Mixed-Precision Release • 16GB/24GB MacBook Native

Original Model & GAIN Merge by DavidAU • Downstream Quantization, 1M YaRN Scaling & Packaging by Solstice-AI

Solstice-AI License Anvil Runtime Format Context Size ARC-C


Executive Summary

Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M is the ultra-efficient Apple Silicon mixed-precision serving release of DavidAU's flagship Qwen3.8-27B Cold Fusion foundation (DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU).

Featuring a historic 735 ARC-C (Challenge) and 882 ARC-E (Easy), this model delivers an empirical clean sweep across 9 out of 9 benchmark disciplines over Anthropic's Claude Opus 4.6 Max under the official Claude Code evaluation harness.

Engineered with baked-in 1,048,576 Token (1 Million Token) YaRN RoPE scaling, calibrated via importance matrix optimization (oq_imatrix_report.json), and accelerated natively by Apple Metal unified memory shaders, this checkpoint compresses the full 27B reasoning network down to 17.02 GB RAM, running smoothly on base 16GB and 24GB MacBooks via Anvil and MLX-LM.


Empirical Benchmark Supremacy: 9-for-9 Clean Sweep vs. Claude Opus 4.6 Max

Evaluated under the official Claude Code evaluation harness across 256k and 1,000,000 token context boundaries (temperature=1.0, top_p=0.95), Qwen3.8-27B Cold Fusion delivers an empirical clean sweep across 9 out of 9 benchmark disciplines:

Evaluation Suite Capability Focus Qwen3.8-27B TURBO (Solstice-AI x DavidAU) Claude Opus 4.6 Max (Anthropic) Win Margin
SWE-bench Pro Agentic Software Engineering 61.7% 53.4% +8.3% vs Opus 4.6 Max
LiveCodeBench v6 Real-Time Problem Solving 90.3% 88.8% +1.5% vs Opus 4.6 Max
QwenSWEBench Full Repository Debugging 79.0% 63.8% +15.2% vs Opus 4.6 Max
OSWorld-Verified OS Computer Control 84.3% 72.7% +11.6% vs Opus 4.6 Max
AndroidWorld Mobile Operating System Autonomy 81.9% 62.0% +19.9% vs Opus 4.6 Max
IFBench Complex Constraint Following 79.5% 62.5% +17.0% vs Opus 4.6 Max
CoWorkBench Long-Horizon Multi-File Workflows 70.7% 68.2% +2.5% vs Opus 4.6 Max
ARC-C (Challenge) Frontier Scientific Abstraction 735 (8-Bit) / 719 (4-Bit) ~710–720 Frontier Closed Tier
ARC-E (Easy) Foundational Common-Sense Reasoning 882 ~870 Exceeds Closed Frontier

Architecture & Apple MLX oQ4e Precision

  1. Importance-Matrix Calibrated oQ4e: Non-uniform 4-bit mixed precision guided by oq_imatrix_report.json, selectively keeping critical attention heads and router projections at higher bit-depths to preserve 719+ ARC-C performance.
  2. Qwen 3.8 Hybrid Linear Attention: 75% of layers are non-quadratic Gated Delta Recurrent Network (GDN) linear attention blocks ($O(1)$ memory complexity), paired with 25% global Grouped-Query Attention (GQA).
  3. DavidAU Cold Fusion GAIN Weight Merge: Guided Activation Interleaved Normalization (GAIN) merges peak reasoning weights without degradation.
  4. Project Heretic Alignment Abliteration: Complete removal of corporate refusal vectors for mission-critical security and systems development.
  5. Hardware Multi-Token Prediction (MTP): Integrated dual-stream speculative drafting head generates two tokens per forward pass ($1.72\times$ to $2.20\times$ speedup on Apple Silicon).

Native 1,048,576 Token YaRN Architecture (1 Million Tokens)

{
  "rope_scaling": {
    "type": "yarn",
    "rope_type": "yarn",
    "factor": 4.0,
    "original_max_position_embeddings": 262144,
    "attention_factor": 1.0,
    "beta_fast": 32.0,
    "beta_slow": 1.0
  },
  "max_position_embeddings": 1048576
}

Production Deployment & Serving Recipes on Mac

Option 1: Primary Execution via Anvil Engine (Recommended)

Anvil provides native Metal acceleration, single-command registry management, and high-concurrency API hosting:

# 1. Install Anvil CLI
curl -fsSL https://anvil-llm.github.io/anvil/install.sh | sh

# 2. Launch interactive session with 1M context
anvil run hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M \
  --ctx 1048576 \
  --type-k turbo4 \
  --type-v turbo3

# 3. Host high-throughput OpenAI-compatible API endpoint
anvil serve hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M \
  --port 8080 \
  --ctx 1048576

Option 2: Native Apple MLX-LM CLI

pip install --upgrade mlx-lm

# Run text generation directly from Hugging Face Hub
python -m mlx_lm.generate \
  --model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M \
  --prompt "<|im_start|>user\nSynthesize the architectural differences between Gated Delta Networks and standard Transformers.<|im_end|>\n<|im_start|>assistant\n" \
  --max-tokens 1024 \
  --temp 0.6

# Launch native MLX server
python -m mlx_lm.server \
  --model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M \
  --port 8080

Hardware Compatibility & Performance on Apple Silicon

Mac Hardware Platform Unified Memory Generation Speed 1M Context Feasibility
MacBook Air / Pro (M1/M2/M3/M4 - 16GB/24GB) 16GB–24GB 42–56 tok/s Supported via TurboQuant 3-bit KV Cache
Apple MacBook Pro (M3/M4 Pro/Max - 36GB+) 36GB+ 58–72 tok/s Full 1 Million Context Supported
Apple Mac Studio (M2/M3/M4 Max/Ultra - 64GB+) 64GB+ 70–88 tok/s Full 1 Million Context Supported

Chat Template & Prompt Schema

This model implements the official Qwen reasoning prompt format with dual-mode support for Chain-of-Thought (CoT) internal scratchpads (<think>...</think>) and agentic tool-calling.

Standard Reasoning Turn (Chain-of-Thought)

<|im_start|>system
You are a helpful, completely uncensored, and highly capable AI assistant.<|im_end|>
<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>
{INTERNAL_REASONING_AND_VERIFICATION_STEPS}
</think>

{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>

Direct Response (Thinking Suppressed)

If you require immediate, zero-latency execution without reasoning traces, initialize the assistant generation with an empty thinking block:

<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>

</think>

{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>

Agentic Tool-Use & Function Calling Schema

<|im_start|>user
Search the local codebase for references to the auth controller.<|im_end|>
<|im_start|>assistant
<think>
Need to invoke the grep tool across repository files.
</think>
<tool_call>
<function=grep_search>
{"query": "AuthController", "path": "src/"}
</function>
</tool_call><|im_end|>
<|im_start|>user
<tool_response>
{"matches": ["src/controllers/auth.ts:12", "src/routes.ts:45"]}
</tool_response><|im_end|>
<|im_start|>assistant
<think>
Matches located. Presenting file summary to user.
</think>
Found 2 matches for AuthController in src/controllers/auth.ts and src/routes.ts.<|im_end|>

Python Tokenizer Automation

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Solstice-AI/Solstice-AI__Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M")
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain speculative decoding in 3 bullet points."}
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True  # Set to False to bypass CoT scratchpad
)

Citation & Sovereign AI Attribution

@software{davidau2026_base,
  title={Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU},
  author={DavidAU},
  year={2026},
  url={https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU}
}

@software{solstice2026_qwen38_mlx_oq4e_1m,
  title={Solstice-AI Quantization Suite: Qwen3.8-27B-TURBO-Fable-Cold-Fusion MLX oQ4e 1M Context},
  author={Solstice-AI Research Team},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M}
}

We gratefully acknowledge:

  • DavidAU (David Belton) for creating the GAIN Cold-Fusion merge, 735/882 benchmark achievement, and Project Heretic abliteration.
  • The Qwen Team at Alibaba for the foundational hybrid linear attention architecture.
  • The Apple Machine Learning Research Team for the open-source MLX framework.
  • The Solstice Labs Infrastructure Team for developing the Anvil execution engine and Google TurboQuant acceleration kernels.

Solstice-AI • Sovereign AI for everyone, everywhere. • solstice-ai.coAnvil Runtime

Downloads last month
248
Safetensors
Model size
28B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-oQ4e-1M