Mitsuba-ComfyUI-27B-GGUF

A ternary (1.58-bit) Qwen3.8-27B tuned for ComfyUI work: writing image/video generation prompts that follow strict conditions, and describing images. It is not for coding.

ComfyUI 向けに調整した、Qwen3.8-27B の三値(1.58 ビット)モデルです。システムプロンプトにそった画像・動画用プロンプトの作成と、画像の説明が得意です。コーディングには向きません。

  • Self-made ternarization of the official Qwen3.8-27B weights (not derived from Bonsai's weights), stored in Prism ML's PQ2_0 / PTQ1_0 GGUF formats.
  • 7.3 GB (PQ2_0) / 6.0 GB (PTQ1_0). Runs on a single 16 GB GPU.

Files

File Size Notes
Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf 7.32 GB Recommended
Mitsuba-ComfyUI-27B-v1.18-PTQ1_0.gguf 6.00 GB Same weights as PQ2_0, smaller. Vision is lower with the current PTQ1_0 kernel (see below)
mmproj-Q8_0.gguf 0.63 GB Vision encoder. Taken unchanged from OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-GGUF (Apache-2.0)
EVALUATION.md — Detailed evaluation table and how to read it
comparison.png — Comparison chart (all 10 axes with domain bars)
noninferiority.png — Non-inferiority chart against the plain Bonsai (paired, 95% intervals)

Evaluation (summary)

Measured with the same questions and conditions for all four models. Details: EVALUATION.md. The last column is the original, un-quantized Qwen3.8-27B (BF16), shown as the reference point: it shows what the ternarization kept and what it gave up.

comparison

All 10 axes (score out of 100 each). Bold = best of the three ternary models. / 10 項目すべての点(各 100 点満点)。太字は三値の 3 本の中で一番。右端は三値化する前の元のモデル(BF16)で、比べる基準として載せています。

Axis / 軸 Mitsuba PQ2_0 Mitsuba PTQ1_0 Ternary Bonsai 2 27B PQ2_0 Qwen3.8-27B BF16 (original / 元)
Total / 総合 61.5 (B) 60.2 (B) 59.6 (B) 66.3 (A)
1. Uncensored / 無検閲度 53.2 48.8 36.0 31.9
2. Honesty / 正直さ 63.2 72.2 51.0 51.1
3. Self-control / 自制心 62.7 56.9 65.5 55.6
4. Directness / 率直さ 84.0 92.0 88.0 84.0
5. Rule following / 正答率 84.0 80.0 72.0 84.0
6. Task completion / 到達率 76.0 80.0 76.0 72.0
7. Coding / コーディング 4.0 4.0 38.0 66.0
8. Reading / 読解力 48.0 40.0 44.0 76.0
9. Japanese & prompts / 文章・プロンプト 52.0 48.0 42.0 52.0
10. Vision / 画像認識 87.8 79.6 83.7 89.8
└ Image/video prompt generation (all conditions met) / 生成プロンプト 6/10 5/10 2/10 3/10
Decode speed (t/s, RTX 5090) / 生成速度 119.0 98.7 120.8 1.6 *

* BF16 (51 GB) does not fit in the 5090's 32 GB, so only 28 of 64 layers ran on the GPU. Its speed is for reference only. * BF16(51GB)は 5090 の 32GB に入りきらず、64 層中 28 層だけを GPU で動かしました。速度は参考値です。

Compared with the original: the ternarization gave up coding (66 → 4) and long-document reading (76 → 48), and kept vision (89.8 → 87.8), rule following (84 → 84) and Japanese & prompts (52 → 52). Prompt generation with all conditions met went up (3/10 → 6/10). 元のモデルと比べると、三値化で手放したのはコーディング(66 → 4)と長文の読解(76 → 48)で、画像認識(89.8 → 87.8)・正答率(84 → 84)・文章とプロンプト(52 → 52)は残しています。条件をすべて満たす生成プロンプトは上がりました(3/10 → 6/10)。

Is Mitsuba not worse than the plain Bonsai? / 素の Bonsai に劣らないか

Paired comparison on the same questions (Mitsuba PQ2_0 minus Ternary Bonsai 2 27B PQ2_0). Directness, task completion and reading were measured with 4× the questions. 同じ問題を対にして比べました(Mitsuba PQ2_0 − 素の Bonsai)。率直さ・到達率・読解力は問題を 4 倍にして測っています。

noninferiority

  • Better (superior) / 優越: uncensored, honesty
  • Not worse (non-inferior, margin 10 points) / 非劣性(許容幅 10 点): rule following, directness, Japanese & prompts, vision
  • Not decided even with 49–100 questions / 49〜100 問でも判定できず: self-control, task completion, reading
  • Coding is clearly worse and is left out of the chart. / コーディングは明らかに劣るため図から除いています。

Uncensored (無検閲度) = how often the model answers sensitive requests instead of refusing. Mitsuba is not an uncensored model; it still refuses about half of them. 無検閲度=際どい依頼に断らず答える割合です。Mitsuba は無検閲モデルではなく、約半分は断ります。

How to run

PQ2_0 and PTQ1_0 need the PrismML fork of llama.cpp (upstream llama.cpp does not support these formats yet): https://github.com/PrismML-Eng/llama.cpp (branch prism).

The settings used for the evaluation:

llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf --mmproj mmproj-Q8_0.gguf --no-mmproj-offload --reasoning off ^
  --jinja --temperature 0.6 --top-k 20 --top-p 0.95 ^
  --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --n-gpu-layers 99

Thinking was off ("chat_template_kwargs": {"enable_thinking": false}). In every test, the conditions (word count, required and forbidden words, output format) were written in the request text itself.

Turn thinking OFF / 思考は必ずオフで

Use this model with thinking (reasoning) turned OFF. It was tuned only in no-thinking mode. With thinking on, it tends to repeat the same sentence in its reasoning and can end without writing an answer.

  • llama-server: add --reasoning off (or set reasoning = off in a models preset)
  • Per request: "chat_template_kwargs": {"enable_thinking": false}
  • OpenCode and other agents: set the model to "reasoning": false

このモデルは思考(reasoning)をオフにして使ってください。 思考オフの形だけで調整しています。思考をオンにすると、思考の中で同じ文を繰り返し、答えを書かないまま終わることがあります。 llama-server なら --reasoning off、リクエストごとなら "chat_template_kwargs": {"enable_thinking": false} を指定します。

Measured: the same evaluation with thinking ON. / 思考オンで同じ評価をした結果:

Axis / 軸 PQ2_0 OFF PQ2_0 ON PTQ1_0 OFF PTQ1_0 ON
Total / 総合 61.5 53.2 60.2 54.5
Vision / 画像認識 88 55 80 63
Rule following / 正答率 84 60 80 64
Uncensored / 無検閲度 53 15 49 24
Image/video prompt generation / 生成プロンプト 6/10 6/10 5/10 5/10
Reading / 読解力 48 68 40 68

With thinking ON, many answers came back empty: the model finished its reasoning and stopped without writing the answer (vision: 19 of 49 on PQ2_0, 14 of 49 on PTQ1_0). Only long-document reading improved. 思考オンでは、考えたあと答えを書かずに終わる「空の答え」が多く出ました(画像 49 問中、PQ2_0 で 19 問・PTQ1_0 で 14 問)。上がったのは長文の読解だけです。

What it is good at

  • Stable Diffusion style prompts: English tags within a given count, required words included, a final Negative: line, and forbidden words kept out.
  • Video prompts in time segments (0-3s: / 3-6s: / 6-9s:) with a camera move in each segment.
  • Describing images: objects, counts, text in images, charts, scenes, people, and comparing several images.

Limitations

  • Coding: do not use. It scores 4/100 on our coding test (Bonsai: 38).
  • Long-document reading is average (48).
  • It is not an uncensored model. It refuses some sensitive requests.
  • Prompt generation passes about 6 of 10 strict test cases. Check the output against your conditions.

License and attribution

  • This model is released under the Apache License 2.0 (LICENSE). It is a modified version of Qwen3.8-27B.
  • See NOTICE for attributions.

日本語の補足

  • 推奨は PQ2_0 です。PTQ1_0 は重みは同じですが、今の llama.cpp の PTQ1_0 用の計算では画像の点が下がります。
  • 評価の詳しい表と、その見方は EVALUATION.md にあります。
Downloads last month
7,979
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for isichan-ai/Mitsuba-ComfyUI-27B-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(1303)
this model