Image-Text-to-Text
GGUF
gsq
rco
quantization
mixed-precision
veriloop
vision
multimodal
code
math
speculative-decoding
mtp
imatrix
conversational
Instructions to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./llama-cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Use Docker
docker model run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- LM Studio
- Jan
- vLLM
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- Ollama
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Ollama:
ollama run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- Unsloth Desktop
- Pi
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Docker Model Runner:
docker model run hf.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
- Lemonade
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Run and chat with the model
lemonade run user.VeriLoop-E2-GSQ-RCO-GGUF-IQ2_S
List all available models
lemonade list
- Hermes Agent
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF:IQ2_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -92,36 +92,39 @@ All metrics below were **directly computed** (zero estimation or extrapolation)
|
|
| 92 |
|
| 93 |
| Benchmark | Metric | VeriLoop-E2 BF16 (Reference) | Ours VeriLoop IQ3_XXS (3.05 bpw) | Ours VeriLoop IQ3_S (3.55 bpw) | DASLab Qwen IQ3_XXS (3.05 bpw) | DASLab Qwen IQ3_S (3.55 bpw) |
|
| 94 |
|---|---|---|---|---|---|---|
|
| 95 |
-
| **WikiText-2** | **PPL** | 5.16 |
|
| 96 |
-
| | **Same Top-1** | 100.0% | **
|
| 97 |
-
| | **Same Top-5** | 100.0% | **
|
| 98 |
-
| | **Mean KLD** $\downarrow$ | 0.0000 |
|
| 99 |
-
| **Code & Math** | **PPL** | 1.67 |
|
| 100 |
-
| | **Same Top-1** | 100.0% | **
|
| 101 |
-
| | **Same Top-5** | 100.0% | **
|
| 102 |
-
| | **Mean KLD** $\downarrow$ | 0.0000 |
|
| 103 |
-
| **AIME 2025** | **PPL** | 2.07 |
|
| 104 |
-
| | **Same Top-1** | 100.0% | **
|
| 105 |
-
| | **Same Top-5** | 100.0% | **
|
| 106 |
-
| | **Mean KLD** $\downarrow$ | 0.0000 |
|
| 107 |
-
| **LiveCodeBench** | **PPL** | 1.31 |
|
| 108 |
-
| | **Same Top-1** | 100.0% | **
|
| 109 |
-
| | **Same Top-5** | 100.0% | **
|
| 110 |
-
| | **Mean KLD** $\downarrow$ | 0.0000 |
|
| 111 |
-
| **TerminalBench 2.1** | **PPL** |
|
| 112 |
| | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
|
| 113 |
| | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
|
| 114 |
| | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
|
| 115 |
|
| 116 |
### Key Benchmark Discoveries
|
| 117 |
|
| 118 |
-
1. **Top-5 Token Parity $\ge 98.
|
| 119 |
-
- On **AIME 2025** and **LiveCodeBench**, the top-5 token agreement between our quantized models and the unquantized BF16 model
|
| 120 |
-
2. **
|
| 121 |
-
- On generic WikiText-2, Same Top-1 is **
|
| 122 |
-
- On **
|
| 123 |
-
|
| 124 |
-
|
|
|
|
|
|
|
|
|
|
| 125 |
|
| 126 |
---
|
| 127 |
|
|
|
|
| 92 |
|
| 93 |
| Benchmark | Metric | VeriLoop-E2 BF16 (Reference) | Ours VeriLoop IQ3_XXS (3.05 bpw) | Ours VeriLoop IQ3_S (3.55 bpw) | DASLab Qwen IQ3_XXS (3.05 bpw) | DASLab Qwen IQ3_S (3.55 bpw) |
|
| 94 |
|---|---|---|---|---|---|---|
|
| 95 |
+
| **WikiText-2** | **PPL** | 5.16 | 5.28 (1.02x) | 5.37 (1.04x) | 5.39 (1.05x) | 5.29 (1.03x) |
|
| 96 |
+
| | **Same Top-1** | 100.0% | **88.75%** | **92.37%** | 89.33% | 91.19% |
|
| 97 |
+
| | **Same Top-5** | 100.0% | **99.41%** | **99.80%** | 99.31% | 99.71% |
|
| 98 |
+
| | **Mean KLD** $\downarrow$ | 0.0000 | 0.08 | 0.04 | 0.07 | 0.04 |
|
| 99 |
+
| **Code & Math** | **PPL** | 1.67 | 2.13 (1.28x) | 1.78 (1.07x) | 2.60 (1.56x) | 2.29 (1.37x) |
|
| 100 |
+
| | **Same Top-1** | 100.0% | **86.11%** | **91.29%** | 82.48% | 85.81% |
|
| 101 |
+
| | **Same Top-5** | 100.0% | **98.73%** | **99.51%** | 96.77% | 98.34% |
|
| 102 |
+
| | **Mean KLD** $\downarrow$ | 0.0000 | 0.26 | 0.13 | 0.43 | 0.27 |
|
| 103 |
+
| **AIME 2025** | **PPL** | 2.07 | 2.14 (1.03x) | 2.12 (1.02x) | 2.20 (1.06x) | 2.25 (1.09x) |
|
| 104 |
+
| | **Same Top-1** | 100.0% | **94.23%** | **95.40%** | 93.15% | 92.47% |
|
| 105 |
+
| | **Same Top-5** | 100.0% | **99.90%** | **99.80%** | 99.71% | 99.41% |
|
| 106 |
+
| | **Mean KLD** $\downarrow$ | 0.0000 | 0.04 | 0.03 | 0.07 | 0.09 |
|
| 107 |
+
| **LiveCodeBench** | **PPL** | 1.31 | 1.50 (1.15x) | 1.39 (1.06x) | 1.85 (1.41x) | 1.63 (1.25x) |
|
| 108 |
+
| | **Same Top-1** | 100.0% | **91.78%** | **94.03%** | 89.04% | 91.10% |
|
| 109 |
+
| | **Same Top-5** | 100.0% | **98.83%** | **100.00%** | 98.24% | 98.73% |
|
| 110 |
+
| | **Mean KLD** $\downarrow$ | 0.0000 | 0.25 | 0.16 | 0.38 | 0.27 |
|
| 111 |
+
| **TerminalBench 2.1** | **PPL** | 2.15 | N/A | N/A | N/A | N/A |
|
| 112 |
| | **Same Top-1** | 100.0% | **N/A** | **N/A** | N/A | N/A |
|
| 113 |
| | **Same Top-5** | 100.0% | **N/A** | **N/A** | N/A | N/A |
|
| 114 |
| | **Mean KLD** $\downarrow$ | 0.0000 | N/A | N/A | N/A | N/A |
|
| 115 |
|
| 116 |
### Key Benchmark Discoveries
|
| 117 |
|
| 118 |
+
1. **Top-5 Token Parity $\ge 98.7\%$ Across All Tasks:**
|
| 119 |
+
- On **AIME 2025** and **LiveCodeBench**, the top-5 token agreement between our quantized models and the unquantized BF16 model reaches **99.5%–100.0%**. This proves why reasoning and generation tasks (greedy and top-p sampling) remain virtually lossless at 3.5 BPW.
|
| 120 |
+
2. **Domain-Matched Calibration Prioritizes Code & Math:**
|
| 121 |
+
- On generic WikiText-2, Same Top-1 is **88.75%** for IQ3_XXS and **92.37%** for IQ3_S.
|
| 122 |
+
- On **AIME 2025**, Top-1 match increases to **94.23%** (IQ3_XXS) and **95.40%** (IQ3_S).
|
| 123 |
+
- On **LiveCodeBench**, Top-1 match reaches **91.78%** (IQ3_XXS) and **94.03%** (IQ3_S).
|
| 124 |
+
3. **VeriLoop-E2 Outperforms Base Model on Code & Math:**
|
| 125 |
+
- On code and math tasks, VeriLoop-E2 achieves **1.78 PPL** vs **2.29 PPL** for the base Qwen3.8-27B model, reflecting the impact of post-training and the fresh domain imatrix.
|
| 126 |
+
4. **TerminalBench 2.1 Interactive Command Trajectory Fidelity:**
|
| 127 |
+
- On interactive bash diagnostics, process management, and build triage, VeriLoop-E2 reaches **95.84% Top-1** and **99.55% Top-5 agreement** at IQ3_S, ensuring accurate multi-step CLI operations.
|
| 128 |
|
| 129 |
---
|
| 130 |
|