Instructions to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
- Ollama
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with Ollama:
ollama run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with Docker Model Runner:
docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
- Lemonade
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.ThinkingCap-Qwen3.8-27B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF
GGUF / llama.cpp quantizations of bottlecapai/ThinkingCap-Qwen3.8-27B — the ThinkingCap finetune of Qwen3.8-27B that cuts reasoning tokens by 37% on average while holding 85.8% average accuracy against the base model's 86.6%.
➡️ Full model description, evaluation results (multi-seed, statistically tested), the measured thinking-token reduction, recommended sampling params, and citation: see the main model card at bottlecapai/ThinkingCap-Qwen3.8-27B.
About GGUF and quantization
GGUF is a single-file model format for running LLMs locally with llama.cpp and compatible runtimes (Ollama, LM Studio, …). The quantized variants below store weights at reduced precision — e.g. ≈5.1 bits per weight for Q4_K_M instead of the 16-bit f16 source — cutting download size and memory severalfold; the quality cost is measured under Expected performance.
The low-bit files are built with an importance matrix (activation statistics from a chat-templated calibration corpus) and a per-tensor precision layout: the attention projections of the full-attention layers and the linear-attention output projections stay at 6–8 bit while the feed-forward weights take the 4-bit hit.
Files
| File | Quant | Size |
|---|---|---|
ThinkingCap-Qwen3.8-27B-IQ4_XS.gguf |
IQ4_XS | 15.5 GB |
ThinkingCap-Qwen3.8-27B-Q4_K_M.gguf |
Q4_K_M | 17.4 GB |
ThinkingCap-Qwen3.8-27B-Q6_K.gguf |
Q6_K | 23.9 GB |
ThinkingCap-Qwen3.8-27B-Q8_0.gguf |
Q8_0 | 29.0 GB |
ThinkingCap-Qwen3.8-27B-f16.gguf |
f16 | 54.7 GB |
mmproj-ThinkingCap-Qwen3.8-27B-f16.gguf |
mmproj (vision) | 931 MB |
f16 is the unquantized GGUF conversion, used as the llama.cpp comparison in the evaluation below. Q4_K_M is a common default for local setups and IQ4_XS the smallest file. Against bf16 on the five evaluated benchmarks, IQ4_XS and Q6_K scored 2.8 pp lower on GPQA-Diamond, and Q6_K 1.7 pp higher on MMLU-Pro and 7.0 pp lower on AA-LCR; these are observed differences, and no file is shown to be lossless (see Expected performance).
Usage (llama.cpp)
# pull a specific quant straight from the Hub and chat
llama-cli -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M -p "Hi"
# or download one file and run it
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF ThinkingCap-Qwen3.8-27B-Q4_K_M.gguf --local-dir .
llama-cli -m ThinkingCap-Qwen3.8-27B-Q4_K_M.gguf -p "Hi"
Use the sampling settings from the main model card (the base model's recommended thinking-mode settings). Greedy decoding can loop; keep temperature at the recommended value.
Speculative decoding (MTP)
These GGUFs carry the model's MTP (multi-token-prediction) head, so llama.cpp can run self-speculative decoding for a decode speed-up — no separate draft model needed. Add --spec-type draft-mtp when serving:
llama-server -hf bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF:Q4_K_M --spec-type draft-mtp --spec-draft-n-max 3
Requires a llama.cpp build with MTP support for this architecture (v0.4.1 or newer). It speeds up decoding at 4 parallel slots (see Decode speed and MTP under Expected performance); larger batches are untested. Runtimes that predate MTP support for this architecture may refuse to load the file (missing tensor blk.64…) — update the runtime.
Vision (image input)
ThinkingCap is a vision-language model. Image input needs the multimodal projector
mmproj-ThinkingCap-Qwen3.8-27B-f16.gguf (in this repo) loaded alongside a text GGUF — the
single f16 mmproj pairs with any of the quants above.
- LM Studio / Jan / Ollama, …: download the
mmproj-*.gguffrom this repo; LM Studio auto-detects it and enables the image (🖼️) button. - llama.cpp CLI:
huggingface-cli download bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF \
ThinkingCap-Qwen3.8-27B-Q4_K_M.gguf mmproj-ThinkingCap-Qwen3.8-27B-f16.gguf --local-dir .
llama-mtmd-cli -m ThinkingCap-Qwen3.8-27B-Q4_K_M.gguf \
--mmproj mmproj-ThinkingCap-Qwen3.8-27B-f16.gguf --image photo.jpg -p "Describe this image."
- llama-server: add
--mmproj mmproj-ThinkingCap-Qwen3.8-27B-f16.ggufto expose an OpenAI-compatible vision endpoint.
Expected performance
Paired comparison of each file with the bf16 weights served by vLLM 0.29.0 on an H200. The GGUF files are served by llama.cpp with the mmproj loaded for the RealWorldQA images, one GPU per benchmark (GPUs below). Both sides think at the chat template's default reasoning effort (xhigh) with sampled decoding (temperature 1.0, top_p 0.95, top_k 20, min_p 0.0) and a 65,536-token generation cap; AA-LCR answers are graded by Gemma-4-26B-A4B-it with thinking off.
Full plan — RealWorldQA 765 questions × 2 seeds, GPQA-Diamond 198 × 4, MMLU-Pro 1,500 × 1 (a fixed slice), IFBench 300 × 2, AA-LCR 100 × 1. IQ4_XS, Q4_K_M, Q6_K and Q8_0 ran on llama.cpp: GPQA-Diamond, MMLU-Pro, IFBench and AA-LCR on an RTX PRO 6000 Max-Q (8 parallel slots, 2 for AA-LCR), RealWorldQA on an RTX PRO 5000 (4 slots; 3 for Q6_K and 2 for Q8_0, which is all the 48 GB card holds beside those files). f16 ran on llama.cpp on the same H200 as the bf16 reference (16 slots, 4 for AA-LCR), so f16 against bf16 keeps the GPU and changes the engine and the 16-bit weight format, with no low-bit quantization. llama.cpp commit bfd73a8, served with --jinja on the file's embedded chat template.
Each cell: Δ accuracy in pp (this file minus bf16 on matched (seed, question) cells) with an approximate 95% interval over questions (seeds averaged per question first) and the exact McNemar p over cells (which counts seeds as independent); below it, the change in mean completion tokens (reasoning plus answer) with its 95% interval, and the change in median tokens. Bold: that interval excludes zero; nothing is adjusted for multiple comparisons.
| file | RealWorldQA | GPQA-Diamond | MMLU-Pro | IFBench | AA-LCR |
|---|---|---|---|---|---|
| bf16 accuracy % (reference) | 83.1 | 88.0 | 84.1 | 79.7 | 81.0 |
| f16 (unquantized) | −0.4 [−2.1, +1.3] (p 0.693) tok +15.5% [−0.6, +34.5], med +2.7% |
+1.1 [−1.1, +3.3] (p 0.336) tok −5.6% [−12.8, +1.3], med −7.6% |
−0.1 [−1.4, +1.2] (p 1.000) tok −2.2% [−13.9, +11.1], med +3.6% |
+1.0 [−2.2, +4.2] (p 0.610) tok −6.6% [−13.7, +0.5], med −8.5% |
0.0 [−7.8, +7.8] (p 1.000) tok −9.2% [−29.3, +17.1], med +12.7% |
| Q8_0 | +0.3 [−1.3, +1.9] (p 0.805) tok +12.8% [−2.4, +30.4], med +4.5% |
−0.8 [−3.0, +1.5] (p 0.556) tok −5.0% [−12.9, +2.6], med 0.0% |
+1.0 [−0.3, +2.3] (p 0.159) tok −0.7% [−12.6, +13.5], med +1.2% |
+0.5 [−2.4, +3.4] (p 0.822) tok −6.2% [−13.7, +1.7], med −6.3% |
−2.0 [−7.5, +3.5] (p 0.727) tok −15.1% [−26.8, −0.8], med +7.1% |
| Q6_K | −0.7 [−2.3, +0.9] (p 0.439) tok +9.3% [−6.2, +27.1], med +4.0% |
−2.8 [−5.2, −0.4] (p 0.017) tok −6.7% [−15.6, +2.7], med −8.3% |
+1.7 [+0.4, +3.1] (p 0.014) tok −13.6% [−25.3, −0.4], med −0.6% |
−0.2 [−3.0, +2.7] (p 1.000) tok −7.1% [−14.9, +1.7], med −6.4% |
−7.0 [−13.4, −0.6] (p 0.065) tok −16.1% [−33.7, +5.8], med +23.0% |
| Q4_K_M | −0.2 [−1.8, +1.4] (p 0.869) tok +18.2% [+1.4, +38.0], med +6.2% |
−0.9 [−2.9, +1.1] (p 0.477) tok −1.2% [−8.3, +6.6], med +3.1% |
+0.8 [−0.6, +2.2] (p 0.290) tok −9.6% [−20.1, +2.1], med +4.2% |
+0.7 [−2.2, +3.5] (p 0.741) tok −8.6% [−15.3, −1.7], med −7.7% |
−3.0 [−9.5, +3.5] (p 0.549) tok −24.4% [−37.9, −7.1], med +21.4% |
| IQ4_XS | −1.2 [−3.0, +0.5] (p 0.151) tok +14.1% [−2.4, +34.4], med +9.8% |
−2.8 [−5.2, −0.4] (p 0.023) tok +1.4% [−6.1, +9.4], med +14.2% |
+0.8 [−0.5, +2.1] (p 0.281) tok −0.1% [−12.2, +13.6], med −1.2% |
−0.8 [−4.0, +2.3] (p 0.675) tok −4.9% [−11.7, +2.8], med −3.0% |
−1.0 [−7.5, +5.5] (p 1.000) tok −15.4% [−31.0, +3.6], med +8.9% |
- Accuracy: three of the 25 comparisons have McNemar p below 0.05: IQ4_XS and Q6_K score 2.8 pp lower on GPQA-Diamond, and Q6_K 1.7 pp higher on MMLU-Pro. They are the only three among the 45 accuracy comparisons of the nine builds we evaluated on this plan, and none survives a multiple-comparison correction. Q6_K on AA-LCR scored 74 against 81 of 100 (−7.0 pp): its approximate interval excludes zero, but the exact McNemar p is 0.065. The estimates are not ordered by bit width, and each comparison changes engine and GPU along with the weights, so the cause of the GPQA-Diamond and AA-LCR deficits is unresolved. Q4_K_M and Q8_0 show no detected difference, but their intervals still allow losses of up to 9.5 and 7.5 pp on AA-LCR: evaluated, not proven lossless.
- Unquantized comparison: f16 on llama.cpp against bf16 on vLLM, on the same H200, shows no detected accuracy difference; its intervals allow both losses and gains (−7.8 to +7.8 pp on AA-LCR).
- Tokens: every file has higher RealWorldQA mean tokens (+9.3% to +18.2%; only Q4_K_M's interval excludes zero), unquantized f16 included, which points to the engine contributing; quantization and GPU are not separated (the quantized files ran RealWorldQA on an RTX PRO 5000). Means and typical answers can move apart: every AA-LCR mean falls (−9% to −24%) while every AA-LCR median rises (+7% to +23%).
- Earlier screens asked nested subsets of these questions, so they are not independent confirmations, and they disagreed with this plan: IQ4_XS on GPQA-Diamond was +6.7 [+0.8, +12.5] on the smallest screen (60 questions × 2 seeds), +0.5 [−4.2, +5.2] on a larger one (100 × 2) and is −2.8 [−5.2, −0.4] here (198 × 4).
Decode speed and MTP self-speculative decoding (MMLU-Pro) — 24 questions × 1 seed, llama.cpp on one H200, 4 parallel slots
| config | median tokens | tok/s | s / task | MTP speedup | accept_len (max 4) |
|---|---|---|---|---|---|
| f16 · standard | 195 | 29.9 | 5.5 | 1.00× | — |
| f16 · MTP | 213 | 52.7 | 3.7 | 1.76× | 2.44 |
| Q8_0 · standard | 330 | 43.6 | 8.9 | 1.00× | — |
| Q8_0 · MTP | 217 | 46.6 | 5.5 | 1.07× | 2.52 |
| Q6_K · standard | 162 | 36.5 | 4.1 | 1.00× | — |
| Q6_K · MTP | 224 | 48.9 | 6.3 | 1.34× | 2.36 |
| Q4_K_M · standard | 241 | 37.5 | 6.4 | 1.00× | — |
| Q4_K_M · MTP | 180 | 48.9 | 4.3 | 1.30× | 2.33 |
| IQ4_XS · standard | 190 | 44.8 | 4.6 | 1.00× | — |
| IQ4_XS · MTP | 243 | 54.1 | 4.0 | 1.21× | 2.41 |
Where to find us
Need even more efficiency? The open release is production-ready. Our enterprise versions go further — fewer thinking tokens still, tuned to your workload, at matched accuracy on your own tasks. Built for AI labs, inference providers and enterprises running models at scale. Deployed on your infrastructure, or in the cloud and region you choose. Talk to our team
License
ThinkingCap: PolyForm Small Business 1.0.0 + BottleCap personal-use grant (see LICENSE).
Upstream Qwen materials: Apache-2.0 (see NOTICE).
Commercial license: contact BottleCap AI.
Citation
If you use this model, please cite:
@misc{ThinkingCap-Qwen3.8-27B,
title = {bottlecapai/ThinkingCap-Qwen3.8-27B},
author = {Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and Bartek, Vojtech and Jirak, Jiri and Kubista, Daniel and Krus, Frantisek and Mikolov, Tomas},
year = {2026},
}
- Downloads last month
- 196,432
4-bit
6-bit
8-bit
16-bit

docker model run hf.co/bottlecapai/ThinkingCap-Qwen3.8-27B-GGUF: