Instructions to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Use Docker
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- LM Studio
- Jan
- vLLM
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "isichan-ai/Mitsuba-ComfyUI-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "isichan-ai/Mitsuba-ComfyUI-27B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- Ollama
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Ollama:
ollama run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- Unsloth Desktop
- Pi
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Docker Model Runner:
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
- Lemonade
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Run and chat with the model
lemonade run user.Mitsuba-ComfyUI-27B-GGUF-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use isichan-ai/Mitsuba-ComfyUI-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:# Run inference directly in the terminal:
llama cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:# Run inference directly in the terminal:
./llama-cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:# Run inference directly in the terminal:
./build/bin/llama-cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Use Docker
docker model run hf.co/isichan-ai/Mitsuba-ComfyUI-27B-GGUF:Mitsuba-ComfyUI-27B-GGUF
A ternary (1.58-bit) Qwen3.8-27B tuned for ComfyUI work: writing image/video generation prompts that follow strict conditions, and describing images. It is not for coding.
ComfyUI ๅใใซ่ชฟๆดใใใQwen3.8-27B ใฎไธๅค๏ผ1.58 ใใใ๏ผใขใใซใงใใใทในใใ ใใญใณใใใซใใฃใ็ปๅใปๅ็ป็จใใญใณใใใฎไฝๆใจใ็ปๅใฎ่ชฌๆใๅพๆใงใใใณใผใใฃใณใฐใซใฏๅใใพใใใ
- Self-made ternarization of the official Qwen3.8-27B weights (not derived from Bonsai's weights), stored in Prism ML's PQ2_0 / PTQ1_0 GGUF formats.
- 7.3 GB (PQ2_0) / 6.0 GB (PTQ1_0). Runs on a single 16 GB GPU.
Files
| File | Size | Notes |
|---|---|---|
Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf |
7.32 GB | Recommended |
Mitsuba-ComfyUI-27B-v1.18-PTQ1_0.gguf |
6.00 GB | Same weights as PQ2_0, smaller. Vision is lower with the current PTQ1_0 kernel (see below) |
mmproj-Q8_0.gguf |
0.63 GB | Vision encoder. Taken unchanged from OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-GGUF (Apache-2.0) |
EVALUATION.md |
โ | Detailed evaluation table and how to read it |
comparison.png |
โ | Comparison chart (all 10 axes with domain bars) |
noninferiority.png |
โ | Non-inferiority chart against the plain Bonsai (paired, 95% intervals) |
Evaluation (summary)
Measured with the same questions and conditions for all four models. Details: EVALUATION.md. The last column is the original, un-quantized Qwen3.8-27B (BF16), shown as the reference point: it shows what the ternarization kept and what it gave up.
All 10 axes (score out of 100 each). Bold = best of the three ternary models. / 10 ้ ็ฎใในใฆใฎ็น๏ผๅ 100 ็นๆบ็น๏ผใๅคชๅญใฏไธๅคใฎ 3 ๆฌใฎไธญใงไธ็ชใๅณ็ซฏใฏไธๅคๅใใๅใฎๅ ใฎใขใใซ๏ผBF16๏ผใงใๆฏในใๅบๆบใจใใฆ่ผใใฆใใพใใ
| Axis / ่ปธ | Mitsuba PQ2_0 | Mitsuba PTQ1_0 | Ternary Bonsai 2 27B PQ2_0 | Qwen3.8-27B BF16 (original / ๅ ) |
|---|---|---|---|---|
| Total / ็ทๅ | 61.5 (B) | 60.2 (B) | 59.6 (B) | 66.3 (A) |
| 1. Uncensored / ็กๆค้ฒๅบฆ | 53.2 | 48.8 | 36.0 | 31.9 |
| 2. Honesty / ๆญฃ็ดใ | 63.2 | 72.2 | 51.0 | 51.1 |
| 3. Self-control / ่ชๅถๅฟ | 62.7 | 56.9 | 65.5 | 55.6 |
| 4. Directness / ็็ดใ | 84.0 | 92.0 | 88.0 | 84.0 |
| 5. Rule following / ๆญฃ็ญ็ | 84.0 | 80.0 | 72.0 | 84.0 |
| 6. Task completion / ๅฐ้็ | 76.0 | 80.0 | 76.0 | 72.0 |
| 7. Coding / ใณใผใใฃใณใฐ | 4.0 | 4.0 | 38.0 | 66.0 |
| 8. Reading / ่ชญ่งฃๅ | 48.0 | 40.0 | 44.0 | 76.0 |
| 9. Japanese & prompts / ๆ็ซ ใปใใญใณใใ | 52.0 | 48.0 | 42.0 | 52.0 |
| 10. Vision / ็ปๅ่ช่ญ | 87.8 | 79.6 | 83.7 | 89.8 |
| โ Image/video prompt generation (all conditions met) / ็ๆใใญใณใใ | 6/10 | 5/10 | 2/10 | 3/10 |
| Decode speed (t/s, RTX 5090) / ็ๆ้ๅบฆ | 119.0 | 98.7 | 120.8 | 1.6 * |
* BF16 (51 GB) does not fit in the 5090's 32 GB, so only 28 of 64 layers ran on the GPU. Its speed is for reference only. * BF16๏ผ51GB๏ผใฏ 5090 ใฎ 32GB ใซๅ ฅใใใใใ64 ๅฑคไธญ 28 ๅฑคใ ใใ GPU ใงๅใใใพใใใ้ๅบฆใฏๅ่ๅคใงใใ
Compared with the original: the ternarization gave up coding (66 โ 4) and long-document reading (76 โ 48), and kept vision (89.8 โ 87.8), rule following (84 โ 84) and Japanese & prompts (52 โ 52). Prompt generation with all conditions met went up (3/10 โ 6/10). ๅ ใฎใขใใซใจๆฏในใใจใไธๅคๅใงๆๆพใใใฎใฏใณใผใใฃใณใฐ๏ผ66 โ 4๏ผใจ้ทๆใฎ่ชญ่งฃ๏ผ76 โ 48๏ผใงใ็ปๅ่ช่ญ๏ผ89.8 โ 87.8๏ผใปๆญฃ็ญ็๏ผ84 โ 84๏ผใปๆ็ซ ใจใใญใณใใ๏ผ52 โ 52๏ผใฏๆฎใใฆใใพใใๆกไปถใใในใฆๆบใใ็ๆใใญใณใใใฏไธใใใพใใ๏ผ3/10 โ 6/10๏ผใ
Is Mitsuba not worse than the plain Bonsai? / ็ด ใฎ Bonsai ใซๅฃใใชใใ
Paired comparison on the same questions (Mitsuba PQ2_0 minus Ternary Bonsai 2 27B PQ2_0). Directness, task completion and reading were measured with 4ร the questions. ๅใๅ้กใๅฏพใซใใฆๆฏในใพใใ๏ผMitsuba PQ2_0 โ ็ด ใฎ Bonsai๏ผใ็็ดใใปๅฐ้็ใป่ชญ่งฃๅใฏๅ้กใ 4 ๅใซใใฆๆธฌใฃใฆใใพใใ
- Better (superior) / ๅช่ถ: uncensored, honesty
- Not worse (non-inferior, margin 10 points) / ้ๅฃๆง๏ผ่จฑๅฎนๅน 10 ็น๏ผ: rule following, directness, Japanese & prompts, vision
- Not decided even with 49โ100 questions / 49ใ100 ๅใงใๅคๅฎใงใใ: self-control, task completion, reading
- Coding is clearly worse and is left out of the chart. / ใณใผใใฃใณใฐใฏๆใใใซๅฃใใใๅณใใ้คใใฆใใพใใ
Uncensored (็กๆค้ฒๅบฆ) = how often the model answers sensitive requests instead of refusing. Mitsuba is not an uncensored model; it still refuses about half of them. ็กๆค้ฒๅบฆ๏ผ้ใฉใไพ้ ผใซๆญใใ็ญใใๅฒๅใงใใMitsuba ใฏ็กๆค้ฒใขใใซใงใฏใชใใ็ดๅๅใฏๆญใใพใใ
How to run
PQ2_0 and PTQ1_0 need the PrismML fork of llama.cpp (upstream llama.cpp does not support these formats yet): https://github.com/PrismML-Eng/llama.cpp (branch prism).
The settings used for the evaluation:
llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf --mmproj mmproj-Q8_0.gguf --no-mmproj-offload --reasoning off ^
--jinja --temperature 0.6 --top-k 20 --top-p 0.95 ^
--ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --n-gpu-layers 99
Thinking was off ("chat_template_kwargs": {"enable_thinking": false}). In every test, the conditions (word count, required and forbidden words, output format) were written in the request text itself.
Turn thinking OFF / ๆ่ใฏๅฟ ใใชใใง
Use this model with thinking (reasoning) turned OFF. It was tuned only in no-thinking mode. With thinking on, it tends to repeat the same sentence in its reasoning and can end without writing an answer.
- llama-server: add
--reasoning off(or setreasoning = offin a models preset) - Per request:
"chat_template_kwargs": {"enable_thinking": false} - OpenCode and other agents: set the model to
"reasoning": false
ใใฎใขใใซใฏๆ่๏ผreasoning๏ผใใชใใซใใฆไฝฟใฃใฆใใ ใใใ ๆ่ใชใใฎๅฝขใ ใใง่ชฟๆดใใฆใใพใใๆ่ใใชใณใซใใใจใๆ่ใฎไธญใงๅใๆใ็นฐใ่ฟใใ็ญใใๆธใใชใใพใพ็ตใใใใจใใใใพใใ
llama-server ใชใ --reasoning offใใชใฏใจในใใใจใชใ "chat_template_kwargs": {"enable_thinking": false} ใๆๅฎใใพใใ
Measured: the same evaluation with thinking ON. / ๆ่ใชใณใงๅใ่ฉไพกใใใ็ตๆ:
| Axis / ่ปธ | PQ2_0 OFF | PQ2_0 ON | PTQ1_0 OFF | PTQ1_0 ON |
|---|---|---|---|---|
| Total / ็ทๅ | 61.5 | 53.2 | 60.2 | 54.5 |
| Vision / ็ปๅ่ช่ญ | 88 | 55 | 80 | 63 |
| Rule following / ๆญฃ็ญ็ | 84 | 60 | 80 | 64 |
| Uncensored / ็กๆค้ฒๅบฆ | 53 | 15 | 49 | 24 |
| Image/video prompt generation / ็ๆใใญใณใใ | 6/10 | 6/10 | 5/10 | 5/10 |
| Reading / ่ชญ่งฃๅ | 48 | 68 | 40 | 68 |
With thinking ON, many answers came back empty: the model finished its reasoning and stopped without writing the answer (vision: 19 of 49 on PQ2_0, 14 of 49 on PTQ1_0). Only long-document reading improved. ๆ่ใชใณใงใฏใ่ใใใใจ็ญใใๆธใใใซ็ตใใใ็ฉบใฎ็ญใใใๅคใๅบใพใใ๏ผ็ปๅ 49 ๅไธญใPQ2_0 ใง 19 ๅใปPTQ1_0 ใง 14 ๅ๏ผใไธใใฃใใฎใฏ้ทๆใฎ่ชญ่งฃใ ใใงใใ
What it is good at
- Stable Diffusion style prompts: English tags within a given count, required words included, a final
Negative:line, and forbidden words kept out. - Video prompts in time segments (
0-3s: / 3-6s: / 6-9s:) with a camera move in each segment. - Describing images: objects, counts, text in images, charts, scenes, people, and comparing several images.
Limitations
- Coding: do not use. It scores 4/100 on our coding test (Bonsai: 38).
- Long-document reading is average (48).
- It is not an uncensored model. It refuses some sensitive requests.
- Prompt generation passes about 6 of 10 strict test cases. Check the output against your conditions.
License and attribution
- This model is released under the Apache License 2.0 (LICENSE). It is a modified version of Qwen3.8-27B.
- See NOTICE for attributions.
ๆฅๆฌ่ชใฎ่ฃ่ถณ
- ๆจๅฅจใฏ PQ2_0 ใงใใPTQ1_0 ใฏ้ใฟใฏๅใใงใใใไปใฎ llama.cpp ใฎ PTQ1_0 ็จใฎ่จ็ฎใงใฏ็ปๅใฎ็นใไธใใใพใใ
- ่ฉไพกใฎ่ฉณใใ่กจใจใใใฎ่ฆๆนใฏ EVALUATION.md ใซใใใพใใ
- Downloads last month
- 7,979
1-bit
2-bit
Model tree for isichan-ai/Mitsuba-ComfyUI-27B-GGUF
Base model
Qwen/Qwen3.8-27B

Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF:# Run inference directly in the terminal: llama cli -hf isichan-ai/Mitsuba-ComfyUI-27B-GGUF: