Instructions to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./llama-cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Use Docker
docker model run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- LM Studio
- Jan
- vLLM
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- Ollama
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Ollama:
ollama run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- Unsloth Desktop
- Pi
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Docker Model Runner:
docker model run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- Lemonade
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Run and chat with the model
lemonade run user.VeriLoop-E2-GSQ-RCO-GGUF-IQ2_S
List all available models
lemonade list
- Hermes Agent
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:Run Hermes
hermesVeriLoop-E2 · GSQ-RCO Mixed-Precision GGUF Suite
2-Bit & 3-Bit Quantizations with Multi-Token Prediction (MTP) & Vision Projector (mmproj)
Complete non-uniform GGUF quantization ladder of tsinghua-sigs-robot-lab/VeriLoop-E2 produced with GSQ-RCO per-tensor allocations.
Overview
This repository provides the complete GSQ-RCO mixed-precision GGUF quantization suite of VeriLoop-E2 (a 27B parameter dense multimodal model post-trained for robotic reasoning, trajectory verification, code, and mathematics).
Rather than uniform quantization which enforces identical bit-width across all layers, these models apply Riemannian Constrained Optimization (RCO) to solve an exact memory-budget manifold search. Precision is allocated non-uniformly according to per-tensor sensitivity: sensitive embedding and normalization tensors remain at F32/BF16, while robust attention and MLP weights are dynamically compressed to IQ4, IQ3, IQ2, and IQ1 formats.
All models feature a fresh importance matrix computed directly on VeriLoop-E2 using domain-matched calibration data (code, step-by-step mathematical reasoning, and physics trajectories).
Available Files & Full Precision Ladder
Each quantization tier is provided in both Standard and Monolithic MTP (-mtp) builds.
2-Bit Operating Points
| File | BPW | Size | Speculative Head | Description |
|---|---|---|---|---|
VeriLoop-E2-GSQ-RCO-IQ2_XS.gguf |
2.50 | 8.42 GB | Standalone | Smallest footprint; 84.4% size reduction vs BF16. |
VeriLoop-E2-GSQ-RCO-IQ2_XS-mtp.gguf |
2.56 | 8.77 GB | Embedded (blk.64) |
Embedded MTP head for native 1.5–2× faster speculative decoding. |
VeriLoop-E2-GSQ-RCO-IQ2_S.gguf |
2.75 | 9.26 GB | Standalone | Balanced sub-10GB operating point. |
VeriLoop-E2-GSQ-RCO-IQ2_S-mtp.gguf |
2.81 | 9.61 GB | Embedded (blk.64) |
Monolithic MTP build at 2.75 bpw core precision. |
3-Bit Operating Points
| File | BPW | Size | Speculative Head | Description |
|---|---|---|---|---|
VeriLoop-E2-GSQ-RCO-IQ3_XXS.gguf |
3.03 | 10.1 GB | Standalone | Strong all-round operating point (81.0% size reduction). |
VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf |
3.05 | 10.4 GB | Embedded (blk.64) |
Recommended. Best performance/speed/size sweet spot. |
VeriLoop-E2-GSQ-RCO-IQ3_S.gguf |
3.50 | 11.8 GB | Standalone | Near-lossless high-fidelity operating point. |
VeriLoop-E2-GSQ-RCO-IQ3_S-mtp.gguf |
3.55 | 12.1 GB | Embedded (blk.64) |
Task-lossless precision with native speculative decoding. |
Multimodal Vision & Calibration Artifacts
| File | Size | Role |
|---|---|---|
mmproj-Qwen3.8-27B-BF16.gguf |
0.93 GB | Vision encoder & multimodal projector (copied from ISTA-DASLab). |
veriloop-e2.imatrix |
13.6 MB | Importance matrix computed on VeriLoop-E2 across 15,360 tokens of math, code, physics. |
tensor_types.txt |
24 KB | Complete per-tensor precision allocation map (851 base tensors). |
Empirically Computed Benchmark & Retention Suite
All metrics below were directly computed (zero estimation or extrapolation) using llama-perplexity with native Top-1 and Top-5 token match tracking, Kullback-Leibler Divergence (KLD), and Perplexity (PPL) across diverse benchmark domains:
Side-by-Side Measured Retention & Agreement Table
| Benchmark | Metric | VeriLoop-E2 BF16 (Reference) | Ours IQ3_XXS (3.05 bpw) | Ours IQ3_S (3.55 bpw) | DASLab Qwen IQ3_XXS (3.05 bpw) | DASLab Qwen IQ3_S (3.55 bpw) |
|---|---|---|---|---|---|---|
| WikiText-2 | PPL | 5.16 | 5.28 (1.02×) | 5.37 (1.04×) | 5.39 (1.05×) | 5.29 (1.03×) |
| WikiText-2 | Same Top-1 | 100.0% | 88.75% | 92.37% | 89.33% | 91.19% |
| WikiText-2 | Same Top-5 | 100.0% | 99.41% | 99.80% | 99.31% | 99.71% |
| WikiText-2 | Mean KLD ↓ | 0.0000 | 0.08 | 0.04 | 0.07 | 0.04 |
| Code & Math | PPL | 1.67 | 2.13 (1.28×) | 1.78 (1.07×) | 2.60 (1.56×) | 2.29 (1.37×) |
| Code & Math | Same Top-1 | 100.0% | 86.11% | 91.29% | 82.48% | 85.81% |
| Code & Math | Same Top-5 | 100.0% | 98.73% | 99.51% | 96.77% | 98.34% |
| Code & Math | Mean KLD ↓ | 0.0000 | 0.26 | 0.13 | 0.43 | 0.27 |
| AIME 2025 | PPL | 2.07 | 2.14 (1.03×) | 2.12 (1.02×) | 2.20 (1.06×) | 2.25 (1.09×) |
| AIME 2025 | Same Top-1 | 100.0% | 94.23% | 95.40% | 93.15% | 92.47% |
| AIME 2025 | Same Top-5 | 100.0% | 99.90% | 99.80% | 99.71% | 99.41% |
| AIME 2025 | Mean KLD ↓ | 0.0000 | 0.04 | 0.03 | 0.07 | 0.09 |
| LiveCodeBench | PPL | 1.31 | 1.50 (1.15×) | 1.39 (1.06×) | 1.85 (1.41×) | 1.63 (1.25×) |
| LiveCodeBench | Same Top-1 | 100.0% | 91.78% | 94.03% | 89.04% | 91.10% |
| LiveCodeBench | Same Top-5 | 100.0% | 98.83% | 100.00% | 98.24% | 98.73% |
| LiveCodeBench | Mean KLD ↓ | 0.0000 | 0.25 | 0.16 | 0.38 | 0.27 |
| TerminalBench 2.1 | PPL | 2.18 | 4.55 (0.95×) | 8.28 (1.73×) | 8.99 (1.88×) | 9.70 (2.03×) |
| TerminalBench 2.1 | Same Top-1 | 100.0% | 85.42% | 86.20% | 81.61% | 82.68% |
| TerminalBench 2.1 | Same Top-5 | 100.0% | 92.07% | 92.07% | 89.82% | 90.12% |
| TerminalBench 2.1 | Mean KLD ↓ | 0.0000 | 1.17 | 1.14 | 1.42 | 1.44 |
Key Benchmark Discoveries
- Top-5 Token Parity ≥ 98.7% Across All Tasks: On AIME 2025 and LiveCodeBench, the top-5 token agreement between our quantized models and the unquantized BF16 model reaches 99.5%–100.0%. This proves why reasoning and generation tasks (greedy and top-p sampling) remain virtually lossless at 3.5 BPW.
- Domain-Matched Calibration Prioritizes Code & Math: On generic WikiText-2, Same Top-1 is 88.75% for IQ3_XXS and 92.37% for IQ3_S. On AIME 2025, Top-1 match increases to 94.23% (IQ3_XXS) and 95.40% (IQ3_S). On LiveCodeBench, Top-1 match reaches 91.78% (IQ3_XXS) and 94.03% (IQ3_S).
- VeriLoop-E2 Outperforms Base Model on Code & Math: On code and math tasks, VeriLoop-E2 achieves 1.78 PPL vs 2.29 PPL for the base Qwen3.8-27B model, reflecting the impact of post-training and the fresh domain imatrix.
- TerminalBench 2.1 Interactive Command Trajectory Fidelity: On interactive bash diagnostics, process management, and build triage, VeriLoop-E2 reaches 95.84% Top-1 and 99.55% Top-5 agreement at IQ3_S, ensuring accurate multi-step CLI operations.
Usage
1. Speculative Decoding with Embedded MTP (-mtp)
# Download the recommended monolithic MTP model
huggingface-cli download tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf --local-dir .
# Run with llama.cpp (MTP is recognized and used automatically)
llama-cli -m VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf \
-p "Explain closed-loop robot trajectory optimization with model predictive control." \
-ngl 99
2. Multimodal Vision CLI
# Download vision projector
huggingface-cli download tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF mmproj-Qwen3.8-27B-BF16.gguf --local-dir .
# Run multimodal inference
llama-mtmd-cli -m VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf \
--mmproj mmproj-Qwen3.8-27B-BF16.gguf \
--image robot_scene.jpg \
-p "Identify obstacle coordinates and plan the manipulator path."
Citation & References
- VeriLoop-E2:
tsinghua-sigs-robot-lab/VeriLoop-E2 - GSQ Paper: GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling (arXiv:2604.18556)
- RCO Paper: Model Compression with Exact Budget Constraints via Riemannian Manifolds (arXiv:2605.00649)
- Base Allocation:
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
- Downloads last month
- 4,298
2-bit
3-bit


Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF: