cstr commited on
Commit
1c61dc4
·
verified ·
1 Parent(s): 41952e2

Document the F16 GGUF (#437)

Browse files
Files changed (1) hide show
  1. README.md +8 -4
README.md CHANGED
@@ -34,10 +34,14 @@ Converted and tested with [CrispASR](https://github.com/CrispStrobe/CrispASR), a
34
 
35
  | File | Quant | Size | Description |
36
  |------|-------|------|-------------|
37
- | `voxtral-mini-4b-realtime.gguf` | F16 | 8.3 GB | Full precision (reference) |
38
- | `voxtral-mini-4b-realtime-q8_0.gguf` | Q8_0 | 4.5 GB | 8-bit quantized |
39
  | `voxtral-mini-4b-realtime-q4_k.gguf` | Q4_K | 2.4 GB | 4-bit K-quant (recommended) |
40
 
 
 
 
 
41
  ## Performance (CPU, 4 threads, AVX2, jfk.wav 11s)
42
 
43
  | Quant | Encoder | Prefill | Decode (ms/tok) | Total | RTFx |
@@ -109,10 +113,10 @@ huggingface-cli download cstr/canary-ctc-aligner-GGUF \
109
  ```bash
110
  python models/convert-voxtral4b-to-gguf.py \
111
  --input /path/to/Voxtral-Mini-4B-Realtime-2602 \
112
- --output voxtral-mini-4b-realtime.gguf
113
 
114
  # Then quantize
115
- ./build/bin/cohere-quantize voxtral-mini-4b-realtime.gguf \
116
  voxtral-mini-4b-realtime-q4_k.gguf q4_k
117
  ```
118
 
 
34
 
35
  | File | Quant | Size | Description |
36
  |------|-------|------|-------------|
37
+ | `voxtral-mini-4b-realtime-f16.gguf` | F16 | 8.3 GB | Full precision (reference) — what the quants below are cut from |
38
+ | `voxtral-mini-4b-realtime-q8_0.gguf` | Q8_0 | 4.4 GB | 8-bit quantized |
39
  | `voxtral-mini-4b-realtime-q4_k.gguf` | Q4_K | 2.4 GB | 4-bit K-quant (recommended) |
40
 
41
+ ⚠ This table used to name the F16 `voxtral-mini-4b-realtime.gguf`, and no such
42
+ file was ever uploaded. The published name is `…-f16.gguf`, matching the quant
43
+ suffix CrispASR's `-m auto:f16` resolver expects.
44
+
45
  ## Performance (CPU, 4 threads, AVX2, jfk.wav 11s)
46
 
47
  | Quant | Encoder | Prefill | Decode (ms/tok) | Total | RTFx |
 
113
  ```bash
114
  python models/convert-voxtral4b-to-gguf.py \
115
  --input /path/to/Voxtral-Mini-4B-Realtime-2602 \
116
+ --output voxtral-mini-4b-realtime-f16.gguf
117
 
118
  # Then quantize
119
+ ./build/bin/crispasr-quantize voxtral-mini-4b-realtime-f16.gguf \
120
  voxtral-mini-4b-realtime-q4_k.gguf q4_k
121
  ```
122