How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

VeriLoop-E2 · GSQ-RCO Mixed-Precision GGUF Suite

2-Bit & 3-Bit Quantizations with Multi-Token Prediction (MTP) & Vision Projector (mmproj)

Complete non-uniform GGUF quantization ladder of tsinghua-sigs-robot-lab/VeriLoop-E2 produced with GSQ-RCO per-tensor allocations.

Base Model: VeriLoop-E2 arXiv: GSQ arXiv: RCO License: Apache 2.0


VeriLoop-E2 GSQ-RCO Suite Benchmark Comparison


Overview

This repository provides the complete GSQ-RCO mixed-precision GGUF quantization suite of VeriLoop-E2 (a 27B parameter dense multimodal model post-trained for robotic reasoning, trajectory verification, code, and mathematics).

Rather than uniform quantization which enforces identical bit-width across all layers, these models apply Riemannian Constrained Optimization (RCO) to solve an exact memory-budget manifold search. Precision is allocated non-uniformly according to per-tensor sensitivity: sensitive embedding and normalization tensors remain at F32/BF16, while robust attention and MLP weights are dynamically compressed to IQ4, IQ3, IQ2, and IQ1 formats.

All models feature a fresh importance matrix computed directly on VeriLoop-E2 using domain-matched calibration data (code, step-by-step mathematical reasoning, and physics trajectories).


Available Files & Full Precision Ladder

Each quantization tier is provided in both Standard and Monolithic MTP (-mtp) builds.

2-Bit Operating Points

File BPW Size Speculative Head Description
VeriLoop-E2-GSQ-RCO-IQ2_XS.gguf 2.50 8.42 GB Standalone Smallest footprint; 84.4% size reduction vs BF16.
VeriLoop-E2-GSQ-RCO-IQ2_XS-mtp.gguf 2.56 8.77 GB Embedded (blk.64) Embedded MTP head for native 1.5–2× faster speculative decoding.
VeriLoop-E2-GSQ-RCO-IQ2_S.gguf 2.75 9.26 GB Standalone Balanced sub-10GB operating point.
VeriLoop-E2-GSQ-RCO-IQ2_S-mtp.gguf 2.81 9.61 GB Embedded (blk.64) Monolithic MTP build at 2.75 bpw core precision.

3-Bit Operating Points

File BPW Size Speculative Head Description
VeriLoop-E2-GSQ-RCO-IQ3_XXS.gguf 3.03 10.1 GB Standalone Strong all-round operating point (81.0% size reduction).
VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf 3.05 10.4 GB Embedded (blk.64) Recommended. Best performance/speed/size sweet spot.
VeriLoop-E2-GSQ-RCO-IQ3_S.gguf 3.50 11.8 GB Standalone Near-lossless high-fidelity operating point.
VeriLoop-E2-GSQ-RCO-IQ3_S-mtp.gguf 3.55 12.1 GB Embedded (blk.64) Task-lossless precision with native speculative decoding.

Multimodal Vision & Calibration Artifacts

File Size Role
mmproj-Qwen3.8-27B-BF16.gguf 0.93 GB Vision encoder & multimodal projector (copied from ISTA-DASLab).
veriloop-e2.imatrix 13.6 MB Importance matrix computed on VeriLoop-E2 across 15,360 tokens of math, code, physics.
tensor_types.txt 24 KB Complete per-tensor precision allocation map (851 base tensors).

Empirically Computed Benchmark & Retention Suite

All metrics below were directly computed (zero estimation or extrapolation) using llama-perplexity with native Top-1 and Top-5 token match tracking, Kullback-Leibler Divergence (KLD), and Perplexity (PPL) across diverse benchmark domains:

Empirical Benchmark Deep Dive

Side-by-Side Measured Retention & Agreement Table

Benchmark Metric VeriLoop-E2 BF16 (Reference) Ours IQ3_XXS (3.05 bpw) Ours IQ3_S (3.55 bpw) DASLab Qwen IQ3_XXS (3.05 bpw) DASLab Qwen IQ3_S (3.55 bpw)
WikiText-2 PPL 5.16 5.28 (1.02×) 5.37 (1.04×) 5.39 (1.05×) 5.29 (1.03×)
WikiText-2 Same Top-1 100.0% 88.75% 92.37% 89.33% 91.19%
WikiText-2 Same Top-5 100.0% 99.41% 99.80% 99.31% 99.71%
WikiText-2 Mean KLD ↓ 0.0000 0.08 0.04 0.07 0.04
Code & Math PPL 1.67 2.13 (1.28×) 1.78 (1.07×) 2.60 (1.56×) 2.29 (1.37×)
Code & Math Same Top-1 100.0% 86.11% 91.29% 82.48% 85.81%
Code & Math Same Top-5 100.0% 98.73% 99.51% 96.77% 98.34%
Code & Math Mean KLD ↓ 0.0000 0.26 0.13 0.43 0.27
AIME 2025 PPL 2.07 2.14 (1.03×) 2.12 (1.02×) 2.20 (1.06×) 2.25 (1.09×)
AIME 2025 Same Top-1 100.0% 94.23% 95.40% 93.15% 92.47%
AIME 2025 Same Top-5 100.0% 99.90% 99.80% 99.71% 99.41%
AIME 2025 Mean KLD ↓ 0.0000 0.04 0.03 0.07 0.09
LiveCodeBench PPL 1.31 1.50 (1.15×) 1.39 (1.06×) 1.85 (1.41×) 1.63 (1.25×)
LiveCodeBench Same Top-1 100.0% 91.78% 94.03% 89.04% 91.10%
LiveCodeBench Same Top-5 100.0% 98.83% 100.00% 98.24% 98.73%
LiveCodeBench Mean KLD ↓ 0.0000 0.25 0.16 0.38 0.27
TerminalBench 2.1 PPL 2.18 4.55 (0.95×) 8.28 (1.73×) 8.99 (1.88×) 9.70 (2.03×)
TerminalBench 2.1 Same Top-1 100.0% 85.42% 86.20% 81.61% 82.68%
TerminalBench 2.1 Same Top-5 100.0% 92.07% 92.07% 89.82% 90.12%
TerminalBench 2.1 Mean KLD ↓ 0.0000 1.17 1.14 1.42 1.44

Key Benchmark Discoveries

  1. Top-5 Token Parity ≥ 98.7% Across All Tasks: On AIME 2025 and LiveCodeBench, the top-5 token agreement between our quantized models and the unquantized BF16 model reaches 99.5%–100.0%. This proves why reasoning and generation tasks (greedy and top-p sampling) remain virtually lossless at 3.5 BPW.
  2. Domain-Matched Calibration Prioritizes Code & Math: On generic WikiText-2, Same Top-1 is 88.75% for IQ3_XXS and 92.37% for IQ3_S. On AIME 2025, Top-1 match increases to 94.23% (IQ3_XXS) and 95.40% (IQ3_S). On LiveCodeBench, Top-1 match reaches 91.78% (IQ3_XXS) and 94.03% (IQ3_S).
  3. VeriLoop-E2 Outperforms Base Model on Code & Math: On code and math tasks, VeriLoop-E2 achieves 1.78 PPL vs 2.29 PPL for the base Qwen3.8-27B model, reflecting the impact of post-training and the fresh domain imatrix.
  4. TerminalBench 2.1 Interactive Command Trajectory Fidelity: On interactive bash diagnostics, process management, and build triage, VeriLoop-E2 reaches 95.84% Top-1 and 99.55% Top-5 agreement at IQ3_S, ensuring accurate multi-step CLI operations.

Usage

1. Speculative Decoding with Embedded MTP (-mtp)

# Download the recommended monolithic MTP model
huggingface-cli download tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf --local-dir .

# Run with llama.cpp (MTP is recognized and used automatically)
llama-cli -m VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf \
          -p "Explain closed-loop robot trajectory optimization with model predictive control." \
          -ngl 99

2. Multimodal Vision CLI

# Download vision projector
huggingface-cli download tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF mmproj-Qwen3.8-27B-BF16.gguf --local-dir .

# Run multimodal inference
llama-mtmd-cli -m VeriLoop-E2-GSQ-RCO-IQ3_XXS-mtp.gguf \
               --mmproj mmproj-Qwen3.8-27B-BF16.gguf \
               --image robot_scene.jpg \
               -p "Identify obstacle coordinates and plan the manipulator path."

Citation & References

  • VeriLoop-E2: tsinghua-sigs-robot-lab/VeriLoop-E2
  • GSQ Paper: GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling (arXiv:2604.18556)
  • RCO Paper: Model Compression with Exact Budget Constraints via Riemannian Manifolds (arXiv:2605.00649)
  • Base Allocation: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
Downloads last month
4,185
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(9)
this model

Space using tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF 1

Papers for tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF