Text Generation
GGUF
Mixture of Experts
apex
quantized
granite
mamba
hybrid
llama.cpp
imatrix
conversational
Instructions to use Myric/granite-4.0-h-tiny-APEX-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/granite-4.0-h-tiny-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/granite-4.0-h-tiny-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Myric/granite-4.0-h-tiny-APEX-GGUF # Run inference directly in the terminal: ./llama-cli -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Myric/granite-4.0-h-tiny-APEX-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Use Docker
docker model run hf.co/Myric/granite-4.0-h-tiny-APEX-GGUF
- LM Studio
- Jan
- vLLM
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Myric/granite-4.0-h-tiny-APEX-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Myric/granite-4.0-h-tiny-APEX-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Myric/granite-4.0-h-tiny-APEX-GGUF
- Ollama
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with Ollama:
ollama run hf.co/Myric/granite-4.0-h-tiny-APEX-GGUF
- Unsloth Desktop
- Pi
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Myric/granite-4.0-h-tiny-APEX-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with Docker Model Runner:
docker model run hf.co/Myric/granite-4.0-h-tiny-APEX-GGUF
- Lemonade
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Myric/granite-4.0-h-tiny-APEX-GGUF
Run and chat with the model
lemonade run user.granite-4.0-h-tiny-APEX-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Myric/granite-4.0-h-tiny-APEX-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Myric/granite-4.0-h-tiny-APEX-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/granite-4.0-h-tiny-APEX-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Myric/granite-4.0-h-tiny-APEX-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -35,8 +35,15 @@ Perplexity on wikitext-2-raw (test, 200×512-token windows), `llama-perplexity`.
|
|
| 35 |
|------|------|-----|-----|-----------|
|
| 36 |
| bf16 (reference) | 13 GB | 16.0 | 8.868 | — |
|
| 37 |
| **APEX i-quality** (ssm@Q6) | **4.4 GB** | 5.40 | **8.901** | **+0.38%** |
|
|
|
|
| 38 |
| APEX hand-roll (ssm@Q8) | 4.5 GB | 5.54 | 8.913 | +0.51% |
|
| 39 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
Within **0.38% of full-precision perplexity at ~3× smaller**, and it runs comfortably
|
| 41 |
on modest hardware (~14 GB bf16 → 4.4 GB). Coherent, ~110 tok/s on a single GPU.
|
| 42 |
Built with a diverse imatrix (Bartowski `calibration_datav3`), full expert coverage.
|
|
@@ -71,6 +78,48 @@ the linear-recurrence state helps: it doesn't — it's **larger and slightly wor
|
|
| 71 |
hybrid architectures the SSM/recurrence tensors just aren't precision-sensitive. The
|
| 72 |
hand-roll is included to document the experiment; `i-quality` is the recommended tier.
|
| 73 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
## Attribution & licenses
|
| 75 |
See [`LICENSE`](LICENSE) (Apache-2.0) and [`NOTICE`](NOTICE).
|
| 76 |
- Base: **IBM** ([@ibm-granite](https://huggingface.co/ibm-granite)) — [granite-4.0-h-tiny](https://huggingface.co/ibm-granite/granite-4.0-h-tiny) (Apache-2.0)
|
|
|
|
| 35 |
|------|------|-----|-----|-----------|
|
| 36 |
| bf16 (reference) | 13 GB | 16.0 | 8.868 | — |
|
| 37 |
| **APEX i-quality** (ssm@Q6) | **4.4 GB** | 5.40 | **8.901** | **+0.38%** |
|
| 38 |
+
| **APEX i-quality — torch-imatrix** | **4.4 GB** | 5.40 | **8.864** | **−0.04%** |
|
| 39 |
| APEX hand-roll (ssm@Q8) | 4.5 GB | 5.54 | 8.913 | +0.51% |
|
| 40 |
|
| 41 |
+
The **torch-imatrix** row is the *same* APEX i-quality recipe, quantized with an
|
| 42 |
+
importance matrix computed independently in PyTorch (see [*The torch-imatrix
|
| 43 |
+
variant*](#the-torch-imatrix-variant)) instead of `llama-imatrix`. Its PPL (8.864)
|
| 44 |
+
is inside the ±0.11 error bar of both bf16 and the standard i-quality — i.e. the
|
| 45 |
+
two imatrices are interchangeable in quality.
|
| 46 |
+
|
| 47 |
Within **0.38% of full-precision perplexity at ~3× smaller**, and it runs comfortably
|
| 48 |
on modest hardware (~14 GB bf16 → 4.4 GB). Coherent, ~110 tok/s on a single GPU.
|
| 49 |
Built with a diverse imatrix (Bartowski `calibration_datav3`), full expert coverage.
|
|
|
|
| 78 |
hybrid architectures the SSM/recurrence tensors just aren't precision-sensitive. The
|
| 79 |
hand-roll is included to document the experiment; `i-quality` is the recommended tier.
|
| 80 |
|
| 81 |
+
## The torch-imatrix variant
|
| 82 |
+
|
| 83 |
+
`granite-4.0-h-tiny-APEX-i-quality-torch.gguf` (+ `granite-4.0-h-tiny-torch.imatrix`)
|
| 84 |
+
is a companion build that answers one question: **can the calibration imatrix be
|
| 85 |
+
produced without llama.cpp, and does that change the result?**
|
| 86 |
+
|
| 87 |
+
**What it is.** An importance matrix (imatrix) records, per weight tensor, the
|
| 88 |
+
per-input-channel sum of squared activations over calibration text — it tells
|
| 89 |
+
`llama-quantize` where to spend bits. Normally you get it from `llama-imatrix`,
|
| 90 |
+
which needs the whole model resident in RAM+VRAM. Here it was instead computed
|
| 91 |
+
**directly from the Hugging Face model in PyTorch**, by registering forward hooks
|
| 92 |
+
on every matmul (attention, Mamba-2 in/out projections, router, shared experts,
|
| 93 |
+
and per-expert routed FFNs) and accumulating Σx² as calibration text streams
|
| 94 |
+
through. The generator is **band-serialized**: it loads a few decoder layers at a
|
| 95 |
+
time, runs all calibration chunks through them, caches activations, frees them,
|
| 96 |
+
and moves on — so peak VRAM is a few layers, **not the whole model** (this run
|
| 97 |
+
peaked at ~6.4 GB on a 16 GB card, vs the ~14 GB a full-resident forward needs).
|
| 98 |
+
That decouples imatrix generation from llama.cpp's memory model: you can calibrate
|
| 99 |
+
a model far larger than your GPU by streaming it in bands.
|
| 100 |
+
|
| 101 |
+
**Everything else is identical.** Same bf16 base, same APEX `--tensor-type-file`
|
| 102 |
+
recipe, same `calibration_datav3`, same 126×512 calibration windows. The *only*
|
| 103 |
+
difference from the standard `i-quality` file is which tool produced the imatrix.
|
| 104 |
+
The two quants therefore differ **only** in the tensors whose quantization
|
| 105 |
+
consumes an imatrix (the K-/I-quant expert, attention and Mamba weights); the
|
| 106 |
+
Q8_0 shared experts and F32 tensors are bit-identical.
|
| 107 |
+
|
| 108 |
+
**Validation.** The PyTorch imatrix matches `llama-imatrix`'s output tensor-for-
|
| 109 |
+
tensor (368/368 names) with **median per-tensor correlation 0.995**. The handful
|
| 110 |
+
of lower-correlation tensors are the post-nonlinearity inputs (`ffn_down_shexp`,
|
| 111 |
+
`attn_output`) where transformers' naive Mamba scan and SiLU-gated intermediates
|
| 112 |
+
differ numerically from llama.cpp's kernels — a known wrinkle that does not affect
|
| 113 |
+
quality, since the quantizer only needs *relative* per-channel importance. The
|
| 114 |
+
proof is in the table above: the torch-imatrix quant scores **PPL 8.864**, inside
|
| 115 |
+
the ±0.11 error bar of both bf16 and the standard i-quality. The two imatrices are
|
| 116 |
+
interchangeable.
|
| 117 |
+
|
| 118 |
+
**Pick whichever you like** — they are equivalent in quality. The standard
|
| 119 |
+
`i-quality` is the reference; the `torch` build is provided for anyone who wants a
|
| 120 |
+
llama.cpp-free, VRAM-bounded path to the same result (e.g. calibrating very large
|
| 121 |
+
MoEs on modest hardware).
|
| 122 |
+
|
| 123 |
## Attribution & licenses
|
| 124 |
See [`LICENSE`](LICENSE) (Apache-2.0) and [`NOTICE`](NOTICE).
|
| 125 |
- Base: **IBM** ([@ibm-granite](https://huggingface.co/ibm-granite)) — [granite-4.0-h-tiny](https://huggingface.co/ibm-granite/granite-4.0-h-tiny) (Apache-2.0)
|