Text Generation
GGUF
English
Chinese
llama.cpp
bfloat16
q8_0
q6_k
q5_k_m
q4_k_m
q3_k
iq2_s
iq1_m
mixed-precision
imatrix
mtp
speculative-decoding
veriloop
post-training
coding-agent
software-engineering
mathematical-reasoning
scientific-reasoning
long-context
apache-2.0
conversational
Instructions to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Ollama
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Ollama:
ollama run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Docker Model Runner:
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Lemonade
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.VeriLoop-E2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md
Browse files
README.md
CHANGED
|
@@ -70,7 +70,7 @@ tags:
|
|
| 70 |
| **Default** | **`VeriLoop-E2-Q6_K.gguf`** | **20.566 GiB** | BF16-paired PPL parity within uncertainty; KLD **0.004409**; Same top-p **98.204%** | Strongest overall quality / footprint balance |
|
| 71 |
| **Tighter memory** | **`VeriLoop-E2-Q5_K_M.gguf`** | **18.965 GiB** | PPL **+0.4450%**; KLD **0.006919**; Same top-p **97.251%** | Smaller deployment with a wider fidelity margin |
|
| 72 |
| **Released low-footprint sweet spot** | **`VeriLoop-E2-IQ2_S.gguf`** | **16.799 GiB** | PPL **+0.3457%**; KLD **0.014023**; Same top-p **95.516%** | Smallest fully release-verified low-footprint tier |
|
| 73 |
-
| **Ultra-low-footprint frontier candidate** | **`VeriLoop-E2-IQ1_M.gguf`** | **16.790 GiB** | PPL **+0.3191%**; KLD **0.014357**; Same top-p **95.870%** | Smallest measured candidate;
|
| 74 |
| **Low-footprint alternative** | `VeriLoop-E2-Q3_K_M.gguf` | **16.826 GiB** | PPL **+0.4090%**; KLD **0.014349**; Same top-p **95.919%** | Adjacent fully validated option |
|
| 75 |
| **High fidelity** | `VeriLoop-E2-Q8_0.gguf` | **26.632 GiB** | PPL **+0.0643%**; KLD **0.002176**; Same top-p **98.815%** | Distributional fidelity matters more than footprint |
|
| 76 |
| **Reference** | `VeriLoop-E2-BF16.gguf` | **50.113 GiB** | Canonical reference | Reproducibility, attribution, quantization research |
|
|
@@ -83,9 +83,9 @@ tags:
|
|
| 83 |
|
| 84 |
**IQ2_S remains the current released low-footprint sweet spot.** It is fully verified through structure, BF16-paired fidelity, engineering reserve, stock llama.cpp runtime, MTP engagement, immutable SHA256, and remote readability.
|
| 85 |
|
| 86 |
-
**IQ1_M is the current ultra-low-footprint frontier candidate
|
| 87 |
|
| 88 |
-
**
|
| 89 |
|
| 90 |
> **Naming note:** `VeriLoop-E2-IQ1_M.gguf` is not a uniform 1-bit model. Its measured policy is **353 F32 + 1 IQ1_M + 2 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K = 851 tensors**, with **5.36 effective BPW**. `VeriLoop-E2-IQ2_S.gguf` is likewise a mixed-precision release rather than a uniform 2-bit model.
|
| 91 |
|
|
@@ -97,10 +97,10 @@ tags:
|
|
| 97 |
| **Q8_0** | High-fidelity local | **28.596 GB / 26.632 GiB** | **46.86%** | **8.50 BPW** | Verified |
|
| 98 |
| **Q6_K** | **Overall sweet spot** | **22.083 GB / 20.566 GiB** | **58.96%** | **6.57 BPW** | Verified / Release Quality PASS |
|
| 99 |
| **Q5_K_M** | **Memory-quality sweet spot** | **20.364 GB / 18.965 GiB** | **62.16%** | **6.05 BPW** | Verified / Release Quality PASS |
|
| 100 |
-
| **Q4_K_M** | Intermediate mixed-precision candidate | **19.651 GB / 18.301 GiB** | **63.48%** | **5.84 BPW** |
|
| 101 |
| **Q3_K_M** | Adjacent low-footprint alternative | **18.067 GB / 16.826 GiB** | **66.42%** | **5.37 BPW** | Full release PASS |
|
| 102 |
| **IQ2_S** | **Released low-footprint sweet spot** | **18.037 GB / 16.799 GiB** | **66.48%** | **5.36 BPW** | **Full release PASS** |
|
| 103 |
-
| **IQ1_M** | **Ultra-low-footprint frontier candidate** | **18.028 GB / 16.790 GiB** | **66.50%** | **5.36 BPW** | **
|
| 104 |
|
| 105 |
## Quantization-retention benchmark
|
| 106 |
|
|
@@ -290,7 +290,7 @@ hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-Q6_K.gguf --loc
|
|
| 290 |
hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-IQ2_S.gguf --local-dir .
|
| 291 |
```
|
| 292 |
|
| 293 |
-
> **IQ1_M is
|
| 294 |
|
| 295 |
Run IQ2_S:
|
| 296 |
|
|
@@ -317,11 +317,11 @@ llama-server -m ./VeriLoop-E2-IQ2_S.gguf --model-draft ./mtp-VeriLoop-E2-Q5_
|
|
| 317 |
| Q4_K_M | **18.301 GiB** | Intermediate quantitative candidate |
|
| 318 |
| Q3_K_M | **16.826 GiB** | Adjacent low-footprint release |
|
| 319 |
| **IQ2_S** | **16.799 GiB** | **Released low-footprint sweet spot** |
|
| 320 |
-
| **IQ1_M** | **16.790 GiB** | **Ultra-low-footprint frontier candidate
|
| 321 |
|
| 322 |
A **16.79 GiB model file does not imply full offload on a 16 GiB GPU**. Runtime memory also includes KV cache, compute buffers, allocator overhead, and optional MTP weights.
|
| 323 |
|
| 324 |
-
A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128 41.999677 tok/s**), but there is no frozen paired BF16 throughput campaign. No universal speedup percentage is claimed for Q6/Q5/Q4/Q3/IQ2_S/IQ1_M.
|
| 325 |
|
| 326 |
## Release positioning
|
| 327 |
|
|
@@ -334,7 +334,7 @@ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128
|
|
| 334 |
| Q4_K_M | Intermediate mixed-precision candidate |
|
| 335 |
| Q3_K_M | Adjacent low-footprint release |
|
| 336 |
| **IQ2_S** | **Current released low-footprint sweet spot / lower-KLD frontier point** |
|
| 337 |
-
| **IQ1_M** | **Ultra-low-footprint frontier candidate
|
| 338 |
|
| 339 |
## Measurement boundaries
|
| 340 |
|
|
@@ -342,7 +342,7 @@ A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128
|
|
| 342 |
- IQ1_M's **+0.3191% PPL** does not mean a 0.3191% drop on SWE-bench, Terminal-Bench, AIME, GPQA, or another downstream benchmark.
|
| 343 |
- IQ2_S's **+0.3457% PPL** likewise does not imply a 0.3457% downstream-score loss.
|
| 344 |
- IQ1_M and IQ2_S are both mixed-precision artifacts; the nominal tier name is not the model-wide effective bit width.
|
| 345 |
-
- IQ1_M
|
| 346 |
- The IQ1_M vs IQ2_S point-estimate differences are small relative to the reported uncertainty; this repository does not claim categorical fidelity superiority for either frontier point.
|
| 347 |
- Native 262K context comes from the parent configuration; practical context depends on runtime memory.
|
| 348 |
- MTP acceptance is prompt/workload-dependent.
|
|
|
|
| 70 |
| **Default** | **`VeriLoop-E2-Q6_K.gguf`** | **20.566 GiB** | BF16-paired PPL parity within uncertainty; KLD **0.004409**; Same top-p **98.204%** | Strongest overall quality / footprint balance |
|
| 71 |
| **Tighter memory** | **`VeriLoop-E2-Q5_K_M.gguf`** | **18.965 GiB** | PPL **+0.4450%**; KLD **0.006919**; Same top-p **97.251%** | Smaller deployment with a wider fidelity margin |
|
| 72 |
| **Released low-footprint sweet spot** | **`VeriLoop-E2-IQ2_S.gguf`** | **16.799 GiB** | PPL **+0.3457%**; KLD **0.014023**; Same top-p **95.516%** | Smallest fully release-verified low-footprint tier |
|
| 73 |
+
| **Ultra-low-footprint frontier candidate** | **`VeriLoop-E2-IQ1_M.gguf`** | **16.790 GiB** | PPL **+0.3191%**; KLD **0.014357**; Same top-p **95.870%** | Smallest measured candidate; quantitative fidelity and local identity verified |
|
| 74 |
| **Low-footprint alternative** | `VeriLoop-E2-Q3_K_M.gguf` | **16.826 GiB** | PPL **+0.4090%**; KLD **0.014349**; Same top-p **95.919%** | Adjacent fully validated option |
|
| 75 |
| **High fidelity** | `VeriLoop-E2-Q8_0.gguf` | **26.632 GiB** | PPL **+0.0643%**; KLD **0.002176**; Same top-p **98.815%** | Distributional fidelity matters more than footprint |
|
| 76 |
| **Reference** | `VeriLoop-E2-BF16.gguf` | **50.113 GiB** | Canonical reference | Reproducibility, attribution, quantization research |
|
|
|
|
| 83 |
|
| 84 |
**IQ2_S remains the current released low-footprint sweet spot.** It is fully verified through structure, BF16-paired fidelity, engineering reserve, stock llama.cpp runtime, MTP engagement, immutable SHA256, and remote readability.
|
| 85 |
|
| 86 |
+
**IQ1_M is the current ultra-low-footprint frontier candidate.** Its measured V2 artifact is **16.790078 GiB**, **66.4955% smaller than BF16**, **8.633 MiB / 0.0502% smaller than IQ2_S**, and **36.523 MiB / 0.2120% smaller than Q3_K_M**. Against IQ2_S, IQ1_M has a slightly better PPL ratio (**1.003191 vs 1.003457**) and higher Same top-p (**95.870% vs 95.516%**), but a slightly higher Mean KLD (**0.014357 vs 0.014023**). The uncertainty bands substantially overlap, so this release does **not** claim that either low-footprint tier is universally superior in fidelity.
|
| 87 |
|
| 88 |
+
**Positioning.** IQ1_M defines the current minimum-footprint frontier, while IQ2_S remains the released low-footprint balance point / lower-KLD alternative.
|
| 89 |
|
| 90 |
> **Naming note:** `VeriLoop-E2-IQ1_M.gguf` is not a uniform 1-bit model. Its measured policy is **353 F32 + 1 IQ1_M + 2 IQ2_S + 64 Q4_K + 429 Q5_K + 2 Q6_K = 851 tensors**, with **5.36 effective BPW**. `VeriLoop-E2-IQ2_S.gguf` is likewise a mixed-precision release rather than a uniform 2-bit model.
|
| 91 |
|
|
|
|
| 97 |
| **Q8_0** | High-fidelity local | **28.596 GB / 26.632 GiB** | **46.86%** | **8.50 BPW** | Verified |
|
| 98 |
| **Q6_K** | **Overall sweet spot** | **22.083 GB / 20.566 GiB** | **58.96%** | **6.57 BPW** | Verified / Release Quality PASS |
|
| 99 |
| **Q5_K_M** | **Memory-quality sweet spot** | **20.364 GB / 18.965 GiB** | **62.16%** | **6.05 BPW** | Verified / Release Quality PASS |
|
| 100 |
+
| **Q4_K_M** | Intermediate mixed-precision candidate | **19.651 GB / 18.301 GiB** | **63.48%** | **5.84 BPW** | Quantitatively validated candidate |
|
| 101 |
| **Q3_K_M** | Adjacent low-footprint alternative | **18.067 GB / 16.826 GiB** | **66.42%** | **5.37 BPW** | Full release PASS |
|
| 102 |
| **IQ2_S** | **Released low-footprint sweet spot** | **18.037 GB / 16.799 GiB** | **66.48%** | **5.36 BPW** | **Full release PASS** |
|
| 103 |
+
| **IQ1_M** | **Ultra-low-footprint frontier candidate** | **18.028 GB / 16.790 GiB** | **66.50%** | **5.36 BPW** | **Quantitative fidelity PASS; local identity verified** |
|
| 104 |
|
| 105 |
## Quantization-retention benchmark
|
| 106 |
|
|
|
|
| 290 |
hf download tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF VeriLoop-E2-IQ2_S.gguf --local-dir .
|
| 291 |
```
|
| 292 |
|
| 293 |
+
> **IQ1_M is staged as the ultra-low-footprint frontier artifact.** Its quantitative evidence is PASS and its local SHA256 is frozen as `e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`.
|
| 294 |
|
| 295 |
Run IQ2_S:
|
| 296 |
|
|
|
|
| 317 |
| Q4_K_M | **18.301 GiB** | Intermediate quantitative candidate |
|
| 318 |
| Q3_K_M | **16.826 GiB** | Adjacent low-footprint release |
|
| 319 |
| **IQ2_S** | **16.799 GiB** | **Released low-footprint sweet spot** |
|
| 320 |
+
| **IQ1_M** | **16.790 GiB** | **Ultra-low-footprint frontier candidate** |
|
| 321 |
|
| 322 |
A **16.79 GiB model file does not imply full offload on a 16 GiB GPU**. Runtime memory also includes KV cache, compute buffers, allocator overhead, and optional MTP weights.
|
| 323 |
|
| 324 |
+
A standalone Q8_0 `llama-bench` record exists (**pp512 3102.722776 tok/s; tg128 41.999677 tok/s**), but there is no frozen paired BF16 throughput campaign. No universal speedup percentage is claimed for Q6/Q5/Q4/Q3/IQ2_S/IQ1_M.
|
| 325 |
|
| 326 |
## Release positioning
|
| 327 |
|
|
|
|
| 334 |
| Q4_K_M | Intermediate mixed-precision candidate |
|
| 335 |
| Q3_K_M | Adjacent low-footprint release |
|
| 336 |
| **IQ2_S** | **Current released low-footprint sweet spot / lower-KLD frontier point** |
|
| 337 |
+
| **IQ1_M** | **Ultra-low-footprint frontier candidate** |
|
| 338 |
|
| 339 |
## Measurement boundaries
|
| 340 |
|
|
|
|
| 342 |
- IQ1_M's **+0.3191% PPL** does not mean a 0.3191% drop on SWE-bench, Terminal-Bench, AIME, GPQA, or another downstream benchmark.
|
| 343 |
- IQ2_S's **+0.3457% PPL** likewise does not imply a 0.3457% downstream-score loss.
|
| 344 |
- IQ1_M and IQ2_S are both mixed-precision artifacts; the nominal tier name is not the model-wide effective bit width.
|
| 345 |
+
- IQ1_M has a frozen local SHA256 (`e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b`) and quantitative PASS evidence.
|
| 346 |
- The IQ1_M vs IQ2_S point-estimate differences are small relative to the reported uncertainty; this repository does not claim categorical fidelity superiority for either frontier point.
|
| 347 |
- Native 262K context comes from the parent configuration; practical context depends on runtime memory.
|
| 348 |
- MTP acceptance is prompt/workload-dependent.
|