Instructions to use bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ") config = load_config("bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ: direct link, hf CLI and curl.
- Browser
- Download file 7.58 kB
-
https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ/resolve/main/README.md
- Command line
-
hf download hf://bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ/README.md
-
curl -L -o README.md https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ/resolve/main/README.md
7.58 kB
| base_model: bottlecapai/ThinkingCap-Qwen3.8-27B | |
| base_model_relation: quantized | |
| library_name: mlx | |
| pipeline_tag: image-text-to-text | |
| license: other | |
| license_name: polyform-small-business-1.0.0 | |
| license_link: LICENSE | |
| extra_gated_heading: "Request access to the BottleCap AI model" | |
| extra_gated_prompt: | | |
| Request access below. Add your company if you're evaluating this for work — we have enterprise versions that go further, and we'll make sure you hear about them first. We'll also send you new ThinkingCap releases and early access before they're public. | |
| Tell us how you plan to use the model and we can help you get the most out of it — there's [a short form](https://docs.google.com/forms/d/e/1FAIpQLSdU8MyVP_mVx0_y55d6QCMXVyCKsQ6yg68KEqWm_EIptKB0Nw/viewform) for that too. | |
| extra_gated_fields: | |
| Name: text | |
| "Company name (enter “N/A” if none)": text | |
| "Work email (enter your personal email if none)": text | |
| extra_gated_button_content: "Agree and request access" | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - quantized | |
| - 4-bit | |
| # ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ | |
| MLX build for Apple Silicon, in a mixed 4/8-bit affine layout (group 64) whose quantization | |
| parameters are fitted to the bf16 model's own outputs by distillation (DWQ) rather than set by | |
| rounding alone. **22.5 GB** on disk (21.0 GiB), against ≈56 GB for bf16 — it runs on a 32 GB Mac. | |
| Built from [bottlecapai/ThinkingCap-Qwen3.8-27B](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.8-27B) (bf16). | |
| The MLP projections of all but the last 8 layers are 4-bit; self-attention, the wide | |
| Gated-DeltaNet projections, the last 8 layers' MLPs, `lm_head` and the embeddings are 8-bit. The | |
| vision tower, the MTP head's `fc`, the narrow Gated-DeltaNet input gates and all norms stay | |
| bf16. Vision input and MTP self-speculative decoding both work from this repo — the drafter is in | |
| `mtp-drafter/` and is also embedded in the main shards. | |
| Decode on Apple Silicon is bound by memory bandwidth, so bits per weight set the speed; the 8-bit | |
| share is placed where a flat 4-bit layout costs the most accuracy. | |
| ## Serving | |
| Use **mlx-vlm** or **oMLX**. mlx-lm drops the MTP head and the vision tower when it loads a | |
| checkpoint, so it will serve this one as text-only without speculative decoding. | |
| ```bash | |
| pip install mlx-vlm | |
| python -m mlx_vlm generate \ | |
| --model bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ \ | |
| --prompt "Explain the Cauchy–Schwarz inequality." \ | |
| --enable-thinking --max-tokens 4096 | |
| ``` | |
| ### MTP self-speculative decoding | |
| The model's own next-token head drafts for it; there is no draft model to download. | |
| ```bash | |
| python -m mlx_vlm generate \ | |
| --model bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ \ | |
| --draft-model bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ/mtp-drafter \ | |
| --draft-kind mtp --draft-block-size 3 \ | |
| --prompt "..." --enable-thinking --max-tokens 4096 | |
| ``` | |
| Acceptance is **2.41 tokens per round**, against 2.43 for the bf16 weights on the same prompts. | |
| ### Faster prefill on M5 | |
| M5 has INT8 matrix units, so you can turn on oMLX's INT8-activation prefill with these weights — | |
| no separate build needed, the 4-bit projections are already in the affine group-64 format those | |
| kernels take. Turn it on in oMLX under Model Settings → Experimental Features; it can't be | |
| combined with ANE prefill. oMLX measured 34% faster prefill at 32K on an M5 Max, 615 → 827 | |
| tok/s, with generation slightly slower. It's off by default and can change outputs; the numbers | |
| below were measured with it off. | |
| ### Sampling | |
| Thinking mode, as the base model: `temperature 1.0, top_p 0.95, top_k 20, min_p 0.0`. | |
| ## Expected performance | |
| Accuracy and completion length against the bf16 weights, on the same questions at the same | |
| sampling seeds — GPQA-Diamond 198, MMLU-Pro 1,500, RealWorldQA 765 (images), IFBench 300, AA-LCR | |
| 100 (long-context prompts) — thinking at the chat template's default reasoning effort (`xhigh`), | |
| with sampled decoding (temperature 1.0, top_p 0.95, top_k 20, min_p 0.0). The MMLU-Pro set is a | |
| uniform sample of the 12,032-question test split; AA-LCR answers are graded by | |
| Gemma-4-26B-A4B-it with thinking off, the same judge the comparison's bf16 row used. `median | |
| tokens` / `mean tokens` = completion length (reasoning plus answer) over the questions. | |
| Every benchmark lands within about two points of bf16, in both directions, and completion length | |
| is unchanged. Compared question by question, wins and losses are evenly split on all five. | |
| Evaluated, not proven lossless. | |
| **GPQA-Diamond (graduate-level science)** — 198 questions × 4 seeds | |
| | config | acc | median tokens | mean tokens | | |
| |---|---|---|---| | |
| | ThinkingCap-Qwen3.8-27B bf16 | 0.880 | 1033 | 7115 | | |
| | MLX-4bit-DWQ | 0.862 | 1142 | 6916 | | |
| **MMLU-Pro (knowledge)** — 1,500 questions × 1 seed | |
| | config | acc | median tokens | mean tokens | | |
| |---|---|---|---| | |
| | ThinkingCap-Qwen3.8-27B bf16 | 0.841 | 166 | 1436 | | |
| | MLX-4bit-DWQ | 0.849 | 167 | 1546 | | |
| **RealWorldQA (vision)** — 765 questions × 2 seeds | |
| | config | acc | median tokens | mean tokens | | |
| |---|---|---|---| | |
| | ThinkingCap-Qwen3.8-27B bf16 | 0.831 | 112 | 488 | | |
| | MLX-4bit-DWQ | 0.818 | 114 | 484 | | |
| **IFBench (instruction following)** — 300 questions × 2 seeds | |
| | config | acc | median tokens | mean tokens | | |
| |---|---|---|---| | |
| | ThinkingCap-Qwen3.8-27B bf16 | 0.797 | 1819 | 4531 | | |
| | MLX-4bit-DWQ | 0.785 | 1911 | 4529 | | |
| **AA-LCR (long-context reasoning)** — 100 questions × 1 seed | |
| | config | acc | median tokens | mean tokens | | |
| |---|---|---|---| | |
| | ThinkingCap-Qwen3.8-27B bf16 | 0.810 | 844 | 1718 | | |
| | MLX-4bit-DWQ | 0.830 | 899 | 1843 | | |
| ## Where to find us | |
| <table style="border-collapse:collapse;border:0;margin:0"><tbody><tr> | |
| <td style="border:0;padding:0 18px 0 0"><a href="https://www.bottlecapai.com/"><img src="social-web.png" alt="Website" width="34" height="34"></a></td> | |
| <td style="border:0;padding:0 18px 0 0"><a href="https://www.linkedin.com/company/bottlecap-ai/"><img src="social-linkedin.png" alt="LinkedIn" width="34" height="34"></a></td> | |
| <td style="border:0;padding:0 18px 0 0"><a href="https://www.instagram.com/bottlecapai/"><img src="social-instagram.png" alt="Instagram" width="34" height="34"></a></td> | |
| <td style="border:0;padding:0"><a href="https://x.com/BottleCapAI"><img src="social-x.png" alt="X" width="34" height="34"></a></td> | |
| </tr></tbody></table> | |
| Need even more efficiency? The open release is production-ready. Our enterprise versions go further — fewer thinking tokens still, tuned to your workload, at matched accuracy on your own tasks. Built for AI labs, inference providers and enterprises running models at scale. Deployed on your infrastructure, or in the cloud and region you choose. | |
| [Talk to our team](mailto:enterprise@bottlecapai.com) | |
| ## License | |
| ThinkingCap: PolyForm Small Business 1.0.0 + BottleCap personal-use grant (see [LICENSE](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ/blob/main/LICENSE)). | |
| Upstream Qwen materials: Apache-2.0 (see [NOTICE](https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.8-27B-MLX-4bit-DWQ/blob/main/NOTICE)). | |
| Commercial license: [contact BottleCap AI](mailto:enterprise@bottlecapai.com). | |
| ## Citation | |
| If you use this model, please cite: | |
| ```bibtex | |
| @misc{ThinkingCap-Qwen3.8-27B, | |
| title = {bottlecapai/ThinkingCap-Qwen3.8-27B}, | |
| author = {Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and Bartek, Vojtech and Jirak, Jiri and Kubista, Daniel and Krus, Frantisek and Mikolov, Tomas}, | |
| year = {2026}, | |
| } | |
| ``` | |