REQUIRES ik_llama.cpp or its derivatives

INCOMPATIBLE with mainline llama.cpp and its derivatives (as of August 16th, 2026)

What's that?

A ik_llama.cpp compatible quantization of Gryphe/WorldSim-Opus-3.6-35B-A3B, utilizing:

  • Q8_0 for SSM tensors
  • IQ6_K for embeddings, output, shared experts, attention tensors in global self-attn layers
  • IQ4_KT for sparse experts and attention in local attn layers.

The intention behind such a recipe is squeezing the most sparse experts brainpower in the least amount of space, with absolutely no compromise on long-context attention.

Deliberately stepping away from mixed math/article/story/rp soup datasets, imatrix dataset is a random conversation ~250000-token prune of Squish42/bluemoon-fandom-1-1-rp-cleaned with some turn-based conversation data added from bartowski's calibration data. The imatrix is provided.

For RP-specific finetunes not aimed at being a "general assistant", the weights engaged in creative writing must be preserved first.
At least that's how the theory goes!

Disclosure

WYSIWYG. My only contribution is compute. This is not my merge. Assume WTFPL licensing where no other license implicitly applies. Have fun.

Model card incomplete. Tests and comparisons may be uploaded at a later date.

Downloads last month
266
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-IQ4_KT-GGUF

Quantized
(7)
this model

Dataset used to train Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-IQ4_KT-GGUF