aufklarer commited on
Commit
767be62
·
verified ·
1 Parent(s): 315f9f2

Clarify the runtime bit-width requirement and what is dropped from the export

Browse files
Files changed (1) hide show
  1. README.md +8 -6
README.md CHANGED
@@ -78,8 +78,9 @@ quantized; they stay float. Quantized tensors are stored as MLX's
78
  | `int8/tokenizer_config.json` | 16.7 kB | Chat template and special tokens |
79
 
80
  Each variant is self-contained: one directory is everything needed to run it.
81
- The multi-token-prediction draft head (`mtp.*`) and the vision tower are not
82
- included — no runtime here reads them, and they were 14 MB of dead download.
 
83
 
84
  ## Measured quality
85
 
@@ -140,10 +141,11 @@ let response = try model.generate(
140
  ```
141
 
142
  > **INT5 and INT8 need a runtime that reads the bit width from `config.json`.**
143
- > Older speech-swift versions build every `QuantizedLinear` with `bits = 4`
144
- > hardcoded and will misread an INT5 or INT8 file — it loads without error and
145
- > produces garbage. Use a speech-swift version that takes `quantization_bits`
146
- > and `quantization_group_size` from `config.json`, or stay on INT4.
 
147
 
148
  ### Python
149
 
 
78
  | `int8/tokenizer_config.json` | 16.7 kB | Chat template and special tokens |
79
 
80
  Each variant is self-contained: one directory is everything needed to run it.
81
+ The vision tower and the multi-token-prediction draft head (`mtp.*`) are not
82
+ included — no runtime here reads them. Dropping the draft head alone took 14 MB
83
+ off the INT4 download compared with the previous revision.
84
 
85
  ## Measured quality
86
 
 
141
  ```
142
 
143
  > **INT5 and INT8 need a runtime that reads the bit width from `config.json`.**
144
+ > Older speech-swift versions build every `QuantizedLinear` and
145
+ > `PreQuantizedEmbedding` with `bits = 4` hardcoded, which fits the INT4 file
146
+ > only; they will not read an INT5 or INT8 file correctly. Use a speech-swift
147
+ > version that takes `quantization_bits` and `quantization_group_size` from
148
+ > `config.json`, or stay on INT4.
149
 
150
  ### Python
151