Instructions to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Use Docker
docker model run hf.co/islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
- Ollama
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with Ollama:
ollama run hf.co/islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
- Unsloth Desktop
- Pi
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with Docker Model Runner:
docker model run hf.co/islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
- Lemonade
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Run and chat with the model
lemonade run user.Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF-IQ4_XS
List all available models
lemonade list
- Hermes Agent
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "islamsidratul/Qwen3.8-27B-ByteShape-IQ4_XS-ASCII-GGUF:IQ4_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-27B ByteShape IQ4_XS-3.84bpw — ASCII-P1M vocab
ByteShape's Qwen3.8-27B IQ4_XS-3.84bpw (ShapeLearn per-tensor quantization) with its vocabulary pruned to ASCII plus math/typography symbols, using bsaleh03/ASCII-Condensed-prune-tools (--policy P1M).
English and code only. See Limitations.
| File | Qwen3.8-27B-ByteShape-IQ4_XS-3.84bpw-ASCII-P1M.gguf |
| Size | 12.25 GB (source 13.08 GB) |
| Vocab | 129,272 tokens (source 248,320) |
| sha256 | d533c568ba028fb7bc24d646d9f0259341a388fb2535e559fd0bc62275a55a4d |
| Source sha256 | 89434f23dc89c5f990894e3fe9fdad19d88c370f0d3638a176f29933f218b78b |
| Architecture | qwen35 (hybrid Gated DeltaNet + attention), 65 blocks, MTP head kept |
| Parameters | Qwen3.8-27B (27.3B). The Hub shows 26.1B because it counts tensor shapes, and the pruned vocab removes 119,048 rows from both token_embd and output (~1.22B parameters). Every transformer weight is unchanged. |
What changed
Only token_embd.weight (IQ4_XS, 0.675 → 0.352 GB) and output.weight (Q6_K, 1.043 → 0.543 GB) were touched. Rows are gathered in quantized space: no dequantize/requantize, so every kept row is bit-identical to the source. The tokenizer arrays and merges were rewritten and the special-token ids remapped. The 256 byte-fallback tokens, all specials and all partial-UTF-8 fragments are kept, so any text can still be tokenized.
verify_prune.py --policy P1M passed every check:
- all 864 non-vocab tensors byte-identical to the source
- sampled vocab rows identical to their source rows
- metadata preserved
- kept set exactly what P1M specifies
Tool commit: 376a426d5c6b30f28e03c6efaa9edcf85048773f.
Measurements
Measured on an RTX 5060 Ti 16 GB with llama.cpp v0.4.1.
Perplexity: ctx 4096, 30 chunks, f16 KV. Lower is better, and the figures are deterministic, so they compare exactly across rows.
| model | code (mixed TS/JS project source) | code, ASCII-only lines | wikitext-2 |
|---|---|---|---|
| ByteShape IQ4_XS-3.84bpw (source) | 1.6671 | 1.6490 | 5.9480 |
| this file | 1.7156 | 1.6488 | 5.9612 |
| Unsloth UD-IQ4_XS (reference) | 1.6583 | — | 5.8322 |
On ASCII-only text the prune costs nothing (1.6488 vs 1.6490). The +2.9% on the mixed code corpus comes entirely from its 1.27% non-ASCII characters, mostly Bangla string literals.
Context and speed: whole model on GPU, q4_0 KV cache, -ub 512, MTP off.
| ctx | VRAM allocated | prefill | decode |
|---|---|---|---|
| 96k | 13,952 MiB | 939 t/s | 28.7 t/s |
| 128k | 14,391 MiB | 939 t/s | 28.8 t/s |
| 160k | 14,873 MiB | 895 t/s | 27.7 t/s |
A 16 GB card with a light desktop fits about 128k of context; 160k needs a nearly idle desktop.
Agentic coding eval (10 executable tasks, thinking on, temp 1.0 / top_p 0.95 / top_k 20): 8/10, 18,381 tokens total. One of the two failures is a grading artifact: the answer ended with a usage-example code block, which the grader took for the solution.
Tool calling and mid-conversation system messages work with a patched chat template (see below).
Usage
llama-server -m Qwen3.8-27B-ByteShape-IQ4_XS-3.84bpw-ASCII-P1M.gguf \
-c 131072 -ngl 99 -fa on --cache-type-k q4_0 --cache-type-v q4_0 \
--jinja --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
- MTP: the head is still in the file; add
--spec-type draft-mtp --spec-draft-n-max 2if you have the VRAM for the draft context. - Claude Code and similar agents: the stock Qwen3.8 template raises
System message must be at the beginning.on mid-conversation system messages. Patch that branch to render them as normal system blocks.
Limitations
- Non-Latin text breaks. Scripts outside ASCII and the P1M symbol set (Bangla, CJK, Arabic, Cyrillic, …) fall back to one token per UTF-8 byte. The model was never trained on those byte sequences, so it reads them as garbled text. A Bangla prompt got a romanized reply saying the input looked garbled; the full-vocab model answers it correctly. For multilingual use, take the source file instead.
- Accented Latin (é, ü, ñ) is not in P1M either, so it also degrades.
- Not affiliated with ByteShape or the Qwen team. All numbers above come from one machine.
Credits
- Base model: Qwen/Qwen3.8-27B (Apache-2.0)
- Quantization: ByteShape (ShapeLearn)
- Vocab pruning: bsaleh03/ASCII-Condensed-prune-tools
- Downloads last month
- 997
4-bit