Text Generation
GGUF
English
Chinese
llama.cpp
bfloat16
q8_0
q6_k
q5_k_m
q4_k_m
q3_k
iq2_s
iq1_m
mixed-precision
imatrix
mtp
speculative-decoding
veriloop
post-training
coding-agent
software-engineering
mathematical-reasoning
scientific-reasoning
long-context
apache-2.0
conversational
Instructions to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Ollama
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Ollama:
ollama run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Docker Model Runner:
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
- Lemonade
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.VeriLoop-E2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download RELEASE_MANIFEST.json from tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 10.3 kB
-
https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF/resolve/main/RELEASE_MANIFEST.json
- Command line
-
hf download hf://tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF/RELEASE_MANIFEST.json
-
curl -L -o RELEASE_MANIFEST.json https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF/resolve/main/RELEASE_MANIFEST.json
10.3 kB
| { | |
| "schema": "veriloop.e2.gguf.release_manifest.v9", | |
| "repository": "tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF", | |
| "model": { | |
| "name": "VeriLoop E2", | |
| "parent_model": "tsinghua-sigs-robot-lab/VeriLoop-E2", | |
| "base_model": "Qwen3.8-27B", | |
| "architecture": "Qwen3_5ForConditionalGeneration", | |
| "parameter_class": "27B", | |
| "native_context_tokens": 262144, | |
| "license": "Apache-2.0" | |
| }, | |
| "release_policy": { | |
| "reference_precision": "BF16", | |
| "quantization_reference": "VeriLoop-E2-BF16.gguf", | |
| "quantization_methodology": "direct_from_BF16_no_low_bit_requantization", | |
| "quality_method": "paired_BF16_vs_quant_fixed_protocol", | |
| "parent_benchmark_rerun_per_quant": false, | |
| "performance_claim_boundary": "PPL_KLD_and_token_drift_are_quantization_fidelity_metrics_not_universal_task_score_loss", | |
| "q3_filename_semantics": "mixed_precision_tier_minimum_active_precision_Q3_K", | |
| "q2_filename_semantics": "mixed_precision_tier_minimum_active_precision_IQ2_S", | |
| "q1_filename_semantics": "mixed_precision_tier_minimum_active_precision_IQ1_M", | |
| "internal_release_readiness": "PASS" | |
| }, | |
| "recommended_variants": { | |
| "overall_sweet_spot": "Q6_K", | |
| "memory_quality_sweet_spot": "Q5_K_M", | |
| "minimum_footprint_sweet_spot": "IQ1_M", | |
| "lower_kld_low_footprint_alternative": "IQ2_S", | |
| "adjacent_low_footprint_alternative": "Q3_K_M" | |
| }, | |
| "quality_protocol": { | |
| "name": "paired_BF16_vs_quant", | |
| "corpus": "WikiText-2 raw test", | |
| "context_tokens": 2048, | |
| "chunks": 8, | |
| "seed": 42, | |
| "gpu_layers": 40, | |
| "flash_attention": false, | |
| "kv_cache_key": "F16", | |
| "kv_cache_value": "F16", | |
| "batch_size": 512, | |
| "micro_batch_size": 512, | |
| "evaluation_executable": "llama-perplexity", | |
| "reference_logit_method": "--kl-divergence-base", | |
| "quantized_comparison_method": "--kl-divergence", | |
| "corpus_sha256": "173c87a53759e0201f33e0ccf978e510c2042d7f2cb78229d9a50d79b9e7dd08", | |
| "bf16_logits_reused_across_quant_tiers": true, | |
| "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4" | |
| }, | |
| "frozen_artifacts": { | |
| "VeriLoop-E2-BF16.gguf": { | |
| "tier": "BF16", | |
| "bytes": 53808284000, | |
| "binary_gib": 50.11286959052086, | |
| "sha256": "11bf5defde1a256b7582bc34fd2c4a85a61615ed88e7a422dcd24d814ea6d35d", | |
| "tensor_count": 851, | |
| "identity_gate": "PASS" | |
| }, | |
| "VeriLoop-E2-Q8_0.gguf": { | |
| "tier": "Q8_0", | |
| "bytes": 28595765600, | |
| "binary_gib": 26.631882041692734, | |
| "sha256": "6204a47274cfbc0c69c39877fb06615ce842bbab264eea77e2e0a5e3ae2fb8e8", | |
| "effective_bpw": 8.5, | |
| "identity_gate": "PASS" | |
| }, | |
| "VeriLoop-E2-Q6_K.gguf": { | |
| "tier": "Q6_K", | |
| "bytes": 22082532096, | |
| "binary_gib": 20.56596064567566, | |
| "sha256": "15d8f856471c4853f6bf0036b2a517426c6cb30a0313cef58a7f9577fd26fe9e", | |
| "effective_bpw": 6.57, | |
| "identity_gate": "PASS" | |
| }, | |
| "VeriLoop-E2-Q5_K_M.gguf": { | |
| "tier": "Q5_K_M", | |
| "bytes": 20363666176, | |
| "binary_gib": 18.965142011642456, | |
| "sha256": "f90ec14d7ec8f084a292413ae0e06483c87ab5e1422ddce51f2e3409d8853b61", | |
| "effective_bpw": 6.05, | |
| "identity_gate": "PASS" | |
| }, | |
| "VeriLoop-E2-Q3_K_M.gguf": { | |
| "tier": "Q3_K_M", | |
| "bytes": 18066506496, | |
| "binary_gib": 16.825745344161987, | |
| "sha256": "c2e9539cfb85d87a99605b6914ed602c0fa750164c153ce5574710123aba06fe", | |
| "effective_bpw": 5.37, | |
| "identity_gate": "PASS" | |
| }, | |
| "VeriLoop-E2-IQ2_S.gguf": { | |
| "tier": "IQ2_S", | |
| "bytes": 18037261056, | |
| "decimal_gb": 18.037261056, | |
| "binary_gib": 16.798508405685425, | |
| "sha256": "0dd44d41319efa9f383b9a867b5fd2114335ccd0a653d1c3e7605d535edc601b", | |
| "effective_bpw": 5.36, | |
| "tensor_count": 851, | |
| "tensor_type_counts": { | |
| "F32": 353, | |
| "IQ2_S": 3, | |
| "Q4_K": 64, | |
| "Q5_K": 429, | |
| "Q6_K": 2 | |
| }, | |
| "identity_gate": "PASS", | |
| "quality_gate": "PASS", | |
| "engineering_reserve_gate": "PASS", | |
| "runtime_gate": "PASS", | |
| "mtp_gate": "PASS" | |
| }, | |
| "VeriLoop-E2-IQ1_M.gguf": { | |
| "tier": "IQ1_M", | |
| "bytes": 18028208896, | |
| "decimal_gb": 18.028208896, | |
| "binary_gib": 16.790077924728394, | |
| "sha256": "e4d395806994cdbcfecf7711e22456af7a663b3ea37c89ef15d5bab4571cf68b", | |
| "effective_bpw": 5.36, | |
| "tensor_count": 851, | |
| "tensor_type_counts": { | |
| "F32": 353, | |
| "IQ1_M": 1, | |
| "IQ2_S": 2, | |
| "Q4_K": 64, | |
| "Q5_K": 429, | |
| "Q6_K": 2 | |
| }, | |
| "quantized_size_mib_reported_by_quantizer": 17182.55, | |
| "quantization_time_seconds": 167.77645, | |
| "end_to_end_quantization_seconds": 168.2609, | |
| "identity_gate": "PASS", | |
| "structure_gate": "PASS", | |
| "quality_gate": "PASS", | |
| "engineering_reserve_gate": "PASS", | |
| "runtime_gate": "PASS", | |
| "mtp_gate": "PASS", | |
| "release_readiness_gate": "PASS" | |
| } | |
| }, | |
| "measured_tiers": { | |
| "Q4_K_M": { | |
| "bytes": 19650634496, | |
| "binary_gib": 18.301079511642456, | |
| "effective_bpw": 5.84, | |
| "quantitative_quality_gate": "PASS" | |
| } | |
| }, | |
| "quality_results": { | |
| "BF16": { | |
| "mean_ppl": 4.840423, | |
| "ppl_uncertainty": 0.119931 | |
| }, | |
| "Q8_0": { | |
| "mean_ppl": 4.843536, | |
| "ppl_uncertainty": 0.120062, | |
| "ppl_ratio": 1.000643, | |
| "ppl_ratio_uncertainty": 0.000754, | |
| "mean_kld": 0.002176, | |
| "mean_kld_uncertainty": 0.000668, | |
| "same_top_p_percent": 98.815, | |
| "same_top_p_uncertainty_percent": 0.12, | |
| "rms_delta_p_percent": 1.251, | |
| "log_ppl_correlation_percent": 99.95, | |
| "quantitative_quality_gate": "PASS" | |
| }, | |
| "Q6_K": { | |
| "mean_ppl": 4.838514, | |
| "ppl_uncertainty": 0.119694, | |
| "ppl_ratio": 0.999605, | |
| "ppl_ratio_uncertainty": 0.001222, | |
| "mean_kld": 0.004409, | |
| "mean_kld_uncertainty": 0.000953, | |
| "same_top_p_percent": 98.204, | |
| "same_top_p_uncertainty_percent": 0.147, | |
| "rms_delta_p_percent": 1.929, | |
| "log_ppl_correlation_percent": 99.88, | |
| "quantitative_quality_gate": "PASS" | |
| }, | |
| "Q5_K_M": { | |
| "mean_ppl": 4.861965, | |
| "ppl_uncertainty": 0.12063, | |
| "ppl_ratio": 1.00445, | |
| "ppl_ratio_uncertainty": 0.001361, | |
| "mean_kld": 0.006919, | |
| "mean_kld_uncertainty": 0.000945, | |
| "same_top_p_percent": 97.251, | |
| "same_top_p_uncertainty_percent": 0.181, | |
| "rms_delta_p_percent": 2.49, | |
| "log_ppl_correlation_percent": 99.85, | |
| "quantitative_quality_gate": "PASS" | |
| }, | |
| "Q4_K_M": { | |
| "mean_ppl": 4.86376, | |
| "ppl_uncertainty": 0.12072, | |
| "ppl_ratio": 1.004821, | |
| "ppl_ratio_uncertainty": 0.001722, | |
| "mean_kld": 0.0097, | |
| "mean_kld_uncertainty": 0.00111, | |
| "same_top_p_percent": 96.786, | |
| "same_top_p_uncertainty_percent": 0.195, | |
| "rms_delta_p_percent": 2.843, | |
| "log_ppl_correlation_percent": 99.76, | |
| "quantitative_quality_gate": "PASS" | |
| }, | |
| "Q3_K_M": { | |
| "mean_ppl": 4.860222, | |
| "ppl_uncertainty": 0.120606, | |
| "ppl_ratio": 1.00409, | |
| "ppl_ratio_uncertainty": 0.002181, | |
| "mean_kld": 0.014349, | |
| "mean_kld_uncertainty": 0.001985, | |
| "same_top_p_percent": 95.919, | |
| "same_top_p_uncertainty_percent": 0.219, | |
| "rms_delta_p_percent": 3.515, | |
| "log_ppl_correlation_percent": 99.62, | |
| "quantitative_quality_gate": "PASS" | |
| }, | |
| "IQ2_S": { | |
| "mean_ppl": 4.857159, | |
| "ppl_uncertainty": 0.120482, | |
| "ppl_ratio": 1.003457, | |
| "ppl_ratio_uncertainty": 0.0021, | |
| "mean_kld": 0.014023, | |
| "mean_kld_uncertainty": 0.001408, | |
| "same_top_p_percent": 95.516, | |
| "same_top_p_uncertainty_percent": 0.229, | |
| "rms_delta_p_percent": 3.679, | |
| "log_ppl_correlation_percent": 99.64, | |
| "quantitative_quality_gate": "PASS", | |
| "engineering_reserve_gate": "PASS" | |
| }, | |
| "IQ1_M": { | |
| "mean_ppl": 4.85587, | |
| "ppl_uncertainty": 0.12048, | |
| "absolute_ppl_delta": 0.015446, | |
| "absolute_ppl_delta_uncertainty": 0.010522, | |
| "ppl_ratio": 1.003191, | |
| "ppl_ratio_uncertainty": 0.002175, | |
| "mean_kld": 0.014357, | |
| "mean_kld_uncertainty": 0.001317, | |
| "median_kld": 0.003778, | |
| "kld_p90": 0.02409, | |
| "kld_p95": 0.043233, | |
| "kld_p99": 0.161452, | |
| "kld_p999": 0.697101, | |
| "max_kld": 6.935023, | |
| "mean_delta_p_percent": -0.042, | |
| "mean_delta_p_uncertainty_percent": 0.041, | |
| "rms_delta_p_percent": 3.691, | |
| "rms_delta_p_uncertainty_percent": 0.205, | |
| "same_top_p_percent": 95.87, | |
| "same_top_p_uncertainty_percent": 0.22, | |
| "log_ppl_correlation_percent": 99.62, | |
| "size_reduction_vs_bf16_percent": 66.49547698640603, | |
| "size_reduction_vs_iq2_s_percent": 0.05018589004115371, | |
| "size_reduction_vs_q3_k_m_percent": 0.21198121512025167, | |
| "quantitative_quality_gate": "PASS", | |
| "engineering_reserve_gate": "PASS" | |
| } | |
| }, | |
| "runtime_validation": { | |
| "IQ2_S": { | |
| "stock_llama_cpp_only": true, | |
| "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4", | |
| "main_runtime": "PASS", | |
| "mtp_runtime": "PASS", | |
| "mtp_engagement": "PASS", | |
| "draft_tokens_generated": 331, | |
| "draft_tokens_accepted": 116, | |
| "draft_acceptance_rate": 0.35045 | |
| }, | |
| "IQ1_M": { | |
| "stock_llama_cpp_only": true, | |
| "llama_cpp_revision": "42916d83f4a225e56709f873aa8050ac11f5b6a4", | |
| "main_http_status": 200, | |
| "main_runtime": "PASS", | |
| "mtp_http_status": 200, | |
| "mtp_runtime": "PASS", | |
| "mtp_engagement": "PASS", | |
| "draft_tokens_generated": 104, | |
| "draft_tokens_accepted": 76, | |
| "draft_acceptance_rate": 0.73076923, | |
| "acceptance_rate_consistency_gate": "PASS", | |
| "output_identity_audit": "IDENTICAL", | |
| "runtime_mtp_gate": "PASS" | |
| } | |
| }, | |
| "release_validation": { | |
| "IQ1_M": { | |
| "identity_gate": "PASS", | |
| "structure_gate": "PASS", | |
| "hard_quality_gate": "PASS", | |
| "engineering_reserve_gate": "PASS", | |
| "stock_runtime_gate": "PASS", | |
| "mtp_runtime_gate": "PASS", | |
| "mtp_engagement_gate": "PASS", | |
| "release_readiness_gate": "PASS" | |
| } | |
| }, | |
| "throughput_claims": { | |
| "paired_bf16_speedup_claimed": false, | |
| "note": "MTP acceptance is prompt-dependent; no universal throughput speedup is claimed without a frozen paired throughput protocol." | |
| } | |
| } | |