Nairod785 commited on
Commit
008a6fe
·
verified ·
1 Parent(s): fe96519

Upload QUANTIZATION.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. QUANTIZATION.md +4 -3
QUANTIZATION.md CHANGED
@@ -1,11 +1,12 @@
1
  # Confucius4-R2T2 — Quantization Technical Report & Research Log
2
 
3
  > This report documents the empirical quantization research conducted for the Confucius4-R2T2 model in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp), leading to the creation of the `r2t2-q4_k_m.gguf` quantization.
 
4
 
5
  # Confucius4-R2T2 — quantization
6
 
7
- Companion to `docs/porting/families/confucius4_r2t2.md` (which covers porting and
8
- parity, not quantization) and to `docs/tools/quantization-arms.md` (which covers
9
  the method generally). This file records what was actually measured on R2T2.
10
 
11
  **Result: the floor is 1.187 GB — 52% below the 2.478 GB Q8_0 reference.**
@@ -39,7 +40,7 @@ the method generally). This file records what was actually measured on R2T2.
39
 
40
  `transcribe-quantize` could not build these arms: it is preset-only and offers no
41
  per-tensor control, and its preset menu has no Q2_K or Q3_K. See
42
- `docs/tools/quantization-arms.md` for the `--keep-type` mechanics, the name
43
  matching rules, and the traps.
44
 
45
  ## The arm inventory
 
1
  # Confucius4-R2T2 — Quantization Technical Report & Research Log
2
 
3
  > This report documents the empirical quantization research conducted for the Confucius4-R2T2 model in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp), leading to the creation of the `r2t2-q4_k_m.gguf` quantization.
4
+ > Companion to [`PORTING.md`](PORTING.md) (which covers porting and parity, not quantization) and to [`QUANTIZATION_ARMS.md`](QUANTIZATION_ARMS.md) (which covers the method generally).
5
 
6
  # Confucius4-R2T2 — quantization
7
 
8
+ Companion to [`PORTING.md`](PORTING.md) (which covers porting and
9
+ parity, not quantization) and to [`QUANTIZATION_ARMS.md`](QUANTIZATION_ARMS.md) (which covers
10
  the method generally). This file records what was actually measured on R2T2.
11
 
12
  **Result: the floor is 1.187 GB — 52% below the 2.478 GB Q8_0 reference.**
 
40
 
41
  `transcribe-quantize` could not build these arms: it is preset-only and offers no
42
  per-tensor control, and its preset menu has no Q2_K or Q3_K. See
43
+ [`QUANTIZATION_ARMS.md`](QUANTIZATION_ARMS.md) for the `--keep-type` mechanics, the name
44
  matching rules, and the traps.
45
 
46
  ## The arm inventory