Akahsizrr commited on
Commit
b5e26d0
Β·
verified Β·
1 Parent(s): 23e351b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +56 -4
README.md CHANGED
@@ -184,10 +184,12 @@ response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_spe
184
  print(response)
185
  ```
186
 
187
- ### 4-bit Quantization (bitsandbytes)
 
 
188
 
189
  ```python
190
- from transformers import AutoModelForCausalLM, BitsAndBytesConfig
191
  import torch
192
 
193
  bnb_config = BitsAndBytesConfig(
@@ -203,9 +205,10 @@ model = AutoModelForCausalLM.from_pretrained(
203
  device_map="auto",
204
  trust_remote_code=True,
205
  )
 
206
  ```
207
 
208
- ### 8-bit Quantization (bitsandbytes)
209
 
210
  ```python
211
  from transformers import AutoModelForCausalLM, BitsAndBytesConfig
@@ -231,6 +234,54 @@ model.set_coding_enabled(False)
231
  model.set_coding_enabled(True)
232
  ```
233
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
234
  ## Performance
235
 
236
  ### VRAM Requirements
@@ -262,10 +313,11 @@ The model produces complete implementations with docstrings, type hints, complex
262
 
263
  1. **Custom architecture**: Requires `trust_remote_code=True` β€” the model includes custom `Fuse3ForCausalLM` code
264
  2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine
265
- 3. **No GGUF conversion**: The custom architecture cannot be directly converted to GGUF format
266
  4. **Training data was small**: Only 55 examples were used for router training β€” the router may not generalize perfectly to all coding tasks
267
  5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β€” some expert knowledge may be lost in translation
268
  6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache
 
269
 
270
  ## Citation
271
 
 
184
  print(response)
185
  ```
186
 
187
+ ## Quantization & Deployment
188
+
189
+ ### bitsandbytes 4-bit (NF4) β€” ~4.5 GB VRAM
190
 
191
  ```python
192
+ from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
193
  import torch
194
 
195
  bnb_config = BitsAndBytesConfig(
 
205
  device_map="auto",
206
  trust_remote_code=True,
207
  )
208
+ tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite")
209
  ```
210
 
211
+ ### bitsandbytes 8-bit β€” ~7 GB VRAM
212
 
213
  ```python
214
  from transformers import AutoModelForCausalLM, BitsAndBytesConfig
 
234
  model.set_coding_enabled(True)
235
  ```
236
 
237
+ ### vLLM
238
+
239
+ > **Note:** vLLM does not currently support the custom `Fuse3ForCausalLM` architecture.
240
+ > The model uses LFM2's hybrid conv+attention backbone with augmented MoE layers,
241
+ > which requires a custom vLLM model implementation. Use transformers for inference.
242
+
243
+ To use with vLLM, you would need to:
244
+ 1. Write a custom vLLM model definition for `Fuse3ForCausalLM`
245
+ 2. Register it with vLLM's model registry
246
+ 3. Handle the hybrid conv+attention layers and expert routing
247
+
248
+ Contributions welcome β€” see the model code in `fuse3_model.py` for the full architecture.
249
+
250
+ ### MLX (Apple Silicon)
251
+
252
+ > **Note:** MLX does not currently support the custom `Fuse3ForCausalLM` architecture.
253
+ > The model's hybrid conv+attention layers and MoE expert routing require a custom
254
+ > MLX model implementation. Use transformers with MPS backend on Apple Silicon.
255
+
256
+ ```python
257
+ # Run on Apple Silicon with MPS backend
258
+ import torch
259
+ from transformers import AutoModelForCausalLM, AutoTokenizer
260
+
261
+ model = AutoModelForCausalLM.from_pretrained(
262
+ "Akahsizrr/fuse-1-Lite",
263
+ torch_dtype=torch.float16,
264
+ device_map="mps",
265
+ trust_remote_code=True,
266
+ )
267
+ tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite")
268
+ ```
269
+
270
+ ### GGUF / llama.cpp
271
+
272
+ > **Note:** GGUF conversion is not currently supported. The custom `Fuse3ForCausalLM`
273
+ > architecture is not recognized by llama.cpp's `convert_hf_to_gguf.py`. LFM2 itself
274
+ > IS supported by llama.cpp (see `lfm2.cpp`), but the expert augmentation layers
275
+ > require a custom llama.cpp model definition.
276
+
277
+ ### Transformers (Recommended)
278
+
279
+ The recommended way to run fuse-1 Lite is with transformers:
280
+
281
+ ```bash
282
+ pip install transformers torch bitsandbytes accelerate
283
+ ```
284
+
285
  ## Performance
286
 
287
  ### VRAM Requirements
 
313
 
314
  1. **Custom architecture**: Requires `trust_remote_code=True` β€” the model includes custom `Fuse3ForCausalLM` code
315
  2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine
316
+ 3. **No GGUF/MLX conversion**: The custom architecture cannot be directly converted to GGUF or MLX format
317
  4. **Training data was small**: Only 55 examples were used for router training β€” the router may not generalize perfectly to all coding tasks
318
  5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β€” some expert knowledge may be lost in translation
319
  6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache
320
+ 7. **bitsandbytes quantization**: 4-bit and 8-bit quantization work at runtime via `BitsAndBytesConfig` β€” pre-quantized saved versions are not available as separate repos
321
 
322
  ## Citation
323