rodrigoramosrs commited on
Commit
b76868f
·
verified ·
1 Parent(s): ac4c7ba

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +2 -7
README.md CHANGED
@@ -56,14 +56,9 @@ Quantized by [Rodrigo Ramos](https://github.com/rodrigoramosrs).
56
 
57
  ## Quantization Approach
58
 
59
- All quants were produced with [llama.cpp](https://github.com/ggml-org/llama.cpp) using a **code-specialized importance matrix (imatrix)** (`Qwen3.8-27B-imatrix.gguf`, same Qwen3.8-27B family — 100% of quantizable tensors covered, verified per-tensor). Unlike generic imatrix datasets, this one preserves fidelity on the distributions that matter most for coding tasks: repository-level code, patches, test suites, and agentic coding traces.
60
 
61
- The result is a set of GGUF files that retain the original model's strong software-engineering capabilities while being deployable via `llama.cpp`, `llama-cpp-python`, `Ollama`, `LM Studio`, and other GGUF-compatible runtimes (use a recent build with QWEN35 + MTP support, e.g. `ik_llama.cpp`).
62
-
63
- Notes:
64
-
65
- - The 34 `mtp.veriloop_engram_registry.*` tensors from the source checkpoint are **not model weights** (embedded data blobs: zips, JSONs, semantic keys — 92 MB total) and are excluded from these GGUFs. The source model's own serving contract ignores that namespace as well.
66
- - `IQ3_XS` keeps the MTP draft block (`blk.64`) at `Q8_0` for speculative-decoding quality.
67
 
68
  ## Available Quants
69
 
 
56
 
57
  ## Quantization Approach
58
 
59
+ All quants were produced with [llama.cpp](https://github.com/ggml-org/llama.cpp) using a **code-specialized importance matrix (imatrix)**. Unlike generic imatrix datasets, this one was curated from software engineering corpora — repository-level code, patches, test suites, and agentic coding traces — ensuring that quantization preserves fidelity on the distributions that matter most for coding tasks.
60
 
61
+ The result is a set of GGUF files that retain the original model's strong software-engineering capabilities while being deployable via `llama.cpp`, `llama-cpp-python`, `Ollama`, `LM Studio`, and other GGUF-compatible runtimes.
 
 
 
 
 
62
 
63
  ## Available Quants
64