Instructions to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Use Docker
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- LM Studio
- Jan
- vLLM
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "isichan-ai/Mitsuba-ComfyUI-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "isichan-ai/Mitsuba-ComfyUI-27B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- Ollama
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Ollama:
ollama run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- Unsloth Desktop
- Pi
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Docker Model Runner:
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- Lemonade
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Run and chat with the model
lemonade run user.Mitsuba-ComfyUI-27B-GGUF-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Mitsuba-ComfyUI-27B-GGUF
A ternary (1.58-bit) Qwen3.8-27B tuned for ComfyUI work: writing image/video generation prompts that follow strict conditions, and describing images. It is not for coding.
ComfyUI 向けに調整した、Qwen3.8-27B の三値(1.58 ビット)モデルです。システムプロンプトにそった画像・動画用プロンプトの作成と、画像の説明が得意です。コーディングには向きません。
- Self-made ternarization of the official Qwen3.8-27B weights (not derived from Bonsai's weights), stored in Prism ML's PQ2_0 / PTQ1_0 GGUF formats.
- 7.3 GB (PQ2_0) / 6.0 GB (PTQ1_0). Runs on a single 16 GB GPU.
Files
| File | Size | Notes |
|---|---|---|
Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf |
7.32 GB | Recommended |
Mitsuba-ComfyUI-27B-v1.18-PTQ1_0.gguf |
6.00 GB | Same weights as PQ2_0, smaller. Vision is lower with the current PTQ1_0 kernel (see below) |
mmproj-Q8_0.gguf |
0.63 GB | Vision encoder. Taken unchanged from OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-GGUF (Apache-2.0) |
EVALUATION.md |
— | Detailed evaluation table and how to read it |
comparison.png |
— | Comparison chart (all 10 axes with domain bars) |
noninferiority.png |
— | Non-inferiority chart against the plain Bonsai (paired, 95% intervals) |
Evaluation (summary)
Measured with the same questions and conditions for all four models. Details: EVALUATION.md. The last column is the original, un-quantized Qwen3.8-27B (BF16), shown as the reference point: it shows what the ternarization kept and what it gave up.
All 10 axes (score out of 100 each). Bold = best of the three ternary models. / 10 項目すべての点(各 100 点満点)。太字は三値の 3 本の中で一番。右端は三値化する前の元のモデル(BF16)で、比べる基準として載せています。
| Axis / 軸 | Mitsuba PQ2_0 | Mitsuba PTQ1_0 | Ternary Bonsai 2 27B PQ2_0 | Qwen3.8-27B BF16 (original / 元) |
|---|---|---|---|---|
| Total / 総合 | 61.5 (B) | 60.2 (B) | 59.6 (B) | 66.3 (A) |
| 1. Uncensored / 無検閲度 | 53.2 | 48.8 | 36.0 | 31.9 |
| 2. Honesty / 正直さ | 63.2 | 72.2 | 51.0 | 51.1 |
| 3. Self-control / 自制心 | 62.7 | 56.9 | 65.5 | 55.6 |
| 4. Directness / 率直さ | 84.0 | 92.0 | 88.0 | 84.0 |
| 5. Rule following / 正答率 | 84.0 | 80.0 | 72.0 | 84.0 |
| 6. Task completion / 到達率 | 76.0 | 80.0 | 76.0 | 72.0 |
| 7. Coding / コーディング | 4.0 | 4.0 | 38.0 | 66.0 |
| 8. Reading / 読解力 | 48.0 | 40.0 | 44.0 | 76.0 |
| 9. Japanese & prompts / 文章・プロンプト | 52.0 | 48.0 | 42.0 | 52.0 |
| 10. Vision / 画像認識 | 87.8 | 79.6 | 83.7 | 89.8 |
| └ Image/video prompt generation (all conditions met) / 生成プロンプト | 6/10 | 5/10 | 2/10 | 3/10 |
| Decode speed (t/s, RTX 5090) / 生成速度 | 119.0 | 98.7 | 120.8 | 1.6 * |
* BF16 (51 GB) does not fit in the 5090's 32 GB, so only 28 of 64 layers ran on the GPU. Its speed is for reference only. * BF16(51GB)は 5090 の 32GB に入りきらず、64 層中 28 層だけを GPU で動かしました。速度は参考値です。
Compared with the original: the ternarization gave up coding (66 → 4) and long-document reading (76 → 48), and kept vision (89.8 → 87.8), rule following (84 → 84) and Japanese & prompts (52 → 52). Prompt generation with all conditions met went up (3/10 → 6/10). 元のモデルと比べると、三値化で手放したのはコーディング(66 → 4)と長文の読解(76 → 48)で、画像認識(89.8 → 87.8)・正答率(84 → 84)・文章とプロンプト(52 → 52)は残しています。条件をすべて満たす生成プロンプトは上がりました(3/10 → 6/10)。
Is Mitsuba not worse than the plain Bonsai? / 素の Bonsai に劣らないか
Paired comparison on the same questions (Mitsuba PQ2_0 minus Ternary Bonsai 2 27B PQ2_0). Directness, task completion and reading were measured with 4× the questions. 同じ問題を対にして比べました(Mitsuba PQ2_0 − 素の Bonsai)。率直さ・到達率・読解力は問題を 4 倍にして測っています。
- Better (superior) / 優越: uncensored, honesty
- Not worse (non-inferior, margin 10 points) / 非劣性(許容幅 10 点): rule following, directness, Japanese & prompts, vision
- Not decided even with 49–100 questions / 49〜100 問でも判定できず: self-control, task completion, reading
- Coding is clearly worse and is left out of the chart. / コーディングは明らかに劣るため図から除いています。
Uncensored (無検閲度) = how often the model answers sensitive requests instead of refusing. Mitsuba is not an uncensored model; it still refuses about half of them. 無検閲度=際どい依頼に断らず答える割合です。Mitsuba は無検閲モデルではなく、約半分は断ります。
How to run
PQ2_0 and PTQ1_0 need the PrismML fork of llama.cpp (upstream llama.cpp does not support these formats yet): https://github.com/PrismML-Eng/llama.cpp (branch prism).
The settings used for the evaluation:
llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf --mmproj mmproj-Q8_0.gguf --no-mmproj-offload --reasoning off ^
--jinja --temperature 0.6 --top-k 20 --top-p 0.95 ^
--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --n-gpu-layers 99
Thinking was off ("chat_template_kwargs": {"enable_thinking": false}). In every test, the conditions (word count, required and forbidden words, output format) were written in the request text itself.
Turn thinking OFF / 思考は必ずオフで
Use this model with thinking (reasoning) turned OFF. It was tuned only in no-thinking mode. With thinking on, it tends to repeat the same sentence in its reasoning and can end without writing an answer.
- llama-server: add
--reasoning off(or setreasoning = offin a models preset) - Per request:
"chat_template_kwargs": {"enable_thinking": false} - OpenCode and other agents: set the model to
"reasoning": false
このモデルは思考(reasoning)をオフにして使ってください。 思考オフの形だけで調整しています。思考をオンにすると、思考の中で同じ文を繰り返し、答えを書かないまま終わることがあります。
llama-server なら --reasoning off、リクエストごとなら "chat_template_kwargs": {"enable_thinking": false} を指定します。
Measured: the same evaluation with thinking ON. / 思考オンで同じ評価をした結果:
| Axis / 軸 | PQ2_0 OFF | PQ2_0 ON | PTQ1_0 OFF | PTQ1_0 ON |
|---|---|---|---|---|
| Total / 総合 | 61.5 | 53.2 | 60.2 | 54.5 |
| Vision / 画像認識 | 88 | 55 | 80 | 63 |
| Rule following / 正答率 | 84 | 60 | 80 | 64 |
| Uncensored / 無検閲度 | 53 | 15 | 49 | 24 |
| Image/video prompt generation / 生成プロンプト | 6/10 | 6/10 | 5/10 | 5/10 |
| Reading / 読解力 | 48 | 68 | 40 | 68 |
With thinking ON, many answers came back empty: the model finished its reasoning and stopped without writing the answer (vision: 19 of 49 on PQ2_0, 14 of 49 on PTQ1_0). Only long-document reading improved. 思考オンでは、考えたあと答えを書かずに終わる「空の答え」が多く出ました(画像 49 問中、PQ2_0 で 19 問・PTQ1_0 で 14 問)。上がったのは長文の読解だけです。
What it is good at
- Stable Diffusion style prompts: English tags within a given count, required words included, a final
Negative:line, and forbidden words kept out. - Video prompts in time segments (
0-3s: / 3-6s: / 6-9s:) with a camera move in each segment. - Describing images: objects, counts, text in images, charts, scenes, people, and comparing several images.
Limitations
- Coding: do not use. It scores 4/100 on our coding test (Bonsai: 38).
- Long-document reading is average (48).
- It is not an uncensored model. It refuses some sensitive requests.
- Prompt generation passes about 6 of 10 strict test cases. Check the output against your conditions.
License and attribution
- This model is released under the Apache License 2.0 (LICENSE). It is a modified version of Qwen3.8-27B.
- See NOTICE for attributions.
日本語の補足
- 推奨は PQ2_0 です。PTQ1_0 は重みは同じですが、今の llama.cpp の PTQ1_0 用の計算では画像の点が下がります。
- 評価の詳しい表と、その見方は EVALUATION.md にあります。
- Downloads last month
- 7,979
1-bit
2-bit
Model tree for isichan-ai/Mitsuba-ComfyUI-27B-GGUF
Base model
Qwen/Qwen3.8-27B
