Instructions to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Ollama
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Ollama:
ollama run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Docker Model Runner:
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Lemonade
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.VeriLoop-E2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download QUALITY_CARD.md from tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 7.46 kB
-
https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF/resolve/main/QUALITY_CARD.md
- Command line
-
hf download hf://tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF/QUALITY_CARD.md
-
curl -L -o QUALITY_CARD.md https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF/resolve/main/QUALITY_CARD.md
VeriLoop E2 — Q8_0 Quality Card
Release type: high-fidelity Q8_0 GGUF
Baseline: canonical VeriLoop E2 BF16 GGUF
Purpose: quantify BF16 → Q8_0 distortion under a fixed llama.cpp protocol
Formal parent-model benchmark re-run: no
Release status: PASS
Positioning: High-fidelity / near-reference Q8_0 tier. This card evaluates quantization fidelity directly against BF16; it does not re-label the parent model's downstream benchmark scores as Q8-specific results.
1. Artifact Identity
| Property | Q8_0 main | Optional MTP draft |
|---|---|---|
| Filename | VeriLoop-E2-Q8_0.gguf |
mtp-VeriLoop-E2-Q8_0.gguf |
| Exact bytes | 28595765600 | 3287070240 |
| Binary size | 26.632 GiB | 3.061 GiB |
| Decimal size | 28.596 GB | 3.287 GB |
| SHA256 | see Q8_RELEASE_MANIFEST.json`` |
see Q8_RELEASE_MANIFEST.json`` |
| Status | Verified | Verified |
BF16 reference SHA256:
11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d
Release finalized:
see `Q8_RELEASE_MANIFEST.json`
2. Quantization Build
| Item | Value |
|---|---|
| Parent model | VeriLoop E2 |
| Base family | Qwen3.8-27B |
| HF architecture | Qwen3_5ForConditionalGeneration |
| Quant type | Q8_0 |
| Main tensors | 851 |
| Q8_0 tensors | 498 |
| Retained F32 tensors | 353 |
| Source model size reported by quantizer | 51,305.09 MiB |
| Quantized model size reported by quantizer | 27,260.56 MiB |
| Effective density reported by quantizer | 8.50 BPW |
| Size reduction vs BF16 GGUF | 46.86% |
| Quantization threads | 25 |
| Quantization time | 38.315 s |
| Quantization return code | 0 |
| Main build gate | PASS |
| llama.cpp quantization/QC revision | 42916d83f4a225e56709f873aa8050ac11f5b6a4 |
The 353 F32 tensors are intentionally retained by the llama.cpp quantized layout; the release does not claim every tensor is stored as Q8_0.
3. Fixed BF16 → Q8_0 Comparison Protocol
| Item | Value |
|---|---|
| Reference model | VeriLoop E2 BF16 GGUF |
| Candidate | VeriLoop E2 Q8_0 GGUF |
| Corpus | WikiText-2 raw test |
| Context length | 2,048 |
| Chunks | 8 |
| Seed | 42 |
| GPU layers | 40 |
| Flash Attention | Off |
| KV cache | F16 / F16 |
| Batch size | 512 |
| Micro-batch size | 512 |
| Evaluation executable | llama-perplexity |
| Reference-logit method | BF16 logits persisted with --kl-divergence-base |
| Candidate comparison | Q8_0 evaluated with --kl-divergence against the frozen BF16 logit file |
This protocol is a paired quantization-fidelity test. It is not a new leaderboard campaign.
4. Quantization Fidelity
Perplexity
| Metric | Value |
|---|---|
| BF16 mean PPL | 4.840423 ± 0.119931 |
| Q8_0 mean PPL | 4.843536 ± 0.120062 |
| Absolute ΔPPL | +0.003113 ± 0.003651 |
| PPL ratio | 1.000643 ± 0.000754 |
| Relative PPL increase | +0.0643% |
| Cor(ln PPL(Q8), ln PPL(BF16)) | 99.95% |
KL divergence
| Metric | Value |
|---|---|
| Mean KLD | 0.002176 ± 0.000668 |
| Median KLD | 0.000276 |
| 90th percentile | 0.001548 |
| 95th percentile | 0.002856 |
| 99th percentile | 0.009561 |
| 99.9th percentile | 0.201704 |
| Maximum | 4.066900 |
Token-probability stability
| Metric | Value |
|---|---|
| Mean Δp | −0.007 ± 0.014% |
| RMS Δp | 1.251 ± 0.137% |
| Same top-p | 98.815 ± 0.120% |
Interpretation
The measured Q8_0 distortion is small under the fixed protocol:
- predictive loss rises by 0.0643% relative to BF16;
- average distributional divergence is 0.002176 KLD;
- the top-probability token remains the same 98.815% of the time;
- mean probability shift is close to zero, while RMS Δp quantifies the remaining quantization noise.
Important: the +0.0643% figure is a relative perplexity increase, not a universal downstream benchmark or capability-loss percentage.
5. Release Gates
| Gate | Predeclared release threshold | Measured | Result |
|---|---|---|---|
| PPL ratio | ≤ 1.005 | 1.000643 | PASS |
| Mean KLD | ≤ 0.005 | 0.002176 | PASS |
| Same top-p | ≥ 97.0% | 98.815% | PASS |
| Main tensor count | 851 | 851 | PASS |
| Non-zero / structural audit | required | PASS | PASS |
| llama.cpp model load | required | PASS | PASS |
| OpenAI-compatible generation | required | PASS | PASS |
| Optional MTP loader compatibility | required for MTP artifact | PASS | PASS |
| Optional MTP speculative draft smoke | required for MTP artifact | PASS | PASS |
These thresholds are VeriLoop internal release-QA criteria and are not presented as universal quantization standards.
6. Runtime Compatibility
Validated llama.cpp revision:
42916d83f4a225e56709f873aa8050ac11f5b6a4
Validated MTP compatibility smoke:
| Item | Value |
|---|---|
| Main model | VeriLoop-E2-Q8_0.gguf |
| Draft model | mtp-VeriLoop-E2-Q8_0.gguf |
| Context | 8,192 |
| Main GPU offload | enabled |
| Draft GPU layers | 0 |
| Model load | PASS |
| Generation | PASS |
| Draft tokens proposed | 3 |
| Draft tokens accepted | 3 |
An earlier 32K all-GPU main+draft attempt exhausted available CUDA memory while allocating the draft runtime buffer. It did not produce a tensor-layout or GGUF-loader error. The resource-conservative 8K main-GPU/draft-CPU configuration loaded and generated successfully.
This card therefore distinguishes format/runtime compatibility from hardware-specific memory fit.
7. Size Efficiency
| Artifact | Size |
|---|---|
| BF16 reference GGUF | 50.113 GiB |
| Q8_0 GGUF | 26.632 GiB |
| Reduction | 46.86% |
The final exact Q8_0 byte count is 28595765600 bytes.
8. Throughput Note
A standalone Q8_0 llama-bench run produced:
| Workload | Q8_0 |
|---|---|
Prompt processing (pp512) |
3102.722776 tok/s |
Token generation (tg128) |
41.999677 tok/s |
The paired BF16 speed run did not complete under the original full-offload benchmark configuration, so no BF16→Q8 speedup factor is claimed. These Q8_0 throughput numbers are descriptive only and are not part of the quality gate.
9. Parent-Model Scores
The public VeriLoop E2 benchmark record belongs to the parent release. Q8_0 was not re-scored across the full benchmark suite for this quantization card.
That separation is intentional:
Parent-model evaluation
→ establishes model capability
BF16 vs Q8_0 paired PPL/KLD/token analysis
→ establishes quantization fidelity
This prevents format conversion and model capability from being conflated.
10. Final Decision
MAIN_Q8_BUILD=PASS
Q8_STRUCTURE_GATE=PASS
Q8_QUANTITATIVE_QUALITY_GATE=PASS
Q8_STOCK_LLAMA_CPP_COMPATIBILITY=PASS
MTP_RUNTIME_GATE=PASS
Q8_RELEASE_QUALITY_GATE=PASS
Release conclusion: Q8_0 is accepted as the high-fidelity local-deployment quantization of VeriLoop E2 under the frozen release protocol.
The strongest numerical statement supported by the evidence is:
Compared with the canonical BF16 GGUF under the identical fixed protocol, Q8_0 reduces file size by 46.86%, increases mean perplexity by 0.0643%, yields mean KLD 0.002176, and preserves the same top-probability token 98.815% of the time.
No broader capability-loss percentage is inferred from these statistics.