Text Generation
Transformers
Safetensors
English
fuse3
mixture-of-experts
MoE
coding
python
code-generation
LFM2
Qwen
LiquidAI
small-language-model
SLM
agentic
fusion
expert-routing
5B
efficient-inference
conversational
custom_code
Eval Results (legacy)
Instructions to use Akahsizrr/fuse-1-Lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Akahsizrr/fuse-1-Lite with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Akahsizrr/fuse-1-Lite", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Akahsizrr/fuse-1-Lite", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Akahsizrr/fuse-1-Lite with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Akahsizrr/fuse-1-Lite" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akahsizrr/fuse-1-Lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Akahsizrr/fuse-1-Lite
- SGLang
How to use Akahsizrr/fuse-1-Lite with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Akahsizrr/fuse-1-Lite" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akahsizrr/fuse-1-Lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Akahsizrr/fuse-1-Lite" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akahsizrr/fuse-1-Lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Akahsizrr/fuse-1-Lite with Docker Model Runner:
docker model run hf.co/Akahsizrr/fuse-1-Lite
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -184,10 +184,12 @@ response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_spe
|
|
| 184 |
print(response)
|
| 185 |
```
|
| 186 |
|
| 187 |
-
##
|
|
|
|
|
|
|
| 188 |
|
| 189 |
```python
|
| 190 |
-
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
|
| 191 |
import torch
|
| 192 |
|
| 193 |
bnb_config = BitsAndBytesConfig(
|
|
@@ -203,9 +205,10 @@ model = AutoModelForCausalLM.from_pretrained(
|
|
| 203 |
device_map="auto",
|
| 204 |
trust_remote_code=True,
|
| 205 |
)
|
|
|
|
| 206 |
```
|
| 207 |
|
| 208 |
-
### 8-bit
|
| 209 |
|
| 210 |
```python
|
| 211 |
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
|
|
@@ -231,6 +234,54 @@ model.set_coding_enabled(False)
|
|
| 231 |
model.set_coding_enabled(True)
|
| 232 |
```
|
| 233 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 234 |
## Performance
|
| 235 |
|
| 236 |
### VRAM Requirements
|
|
@@ -262,10 +313,11 @@ The model produces complete implementations with docstrings, type hints, complex
|
|
| 262 |
|
| 263 |
1. **Custom architecture**: Requires `trust_remote_code=True` β the model includes custom `Fuse3ForCausalLM` code
|
| 264 |
2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine
|
| 265 |
-
3. **No GGUF conversion**: The custom architecture cannot be directly converted to GGUF format
|
| 266 |
4. **Training data was small**: Only 55 examples were used for router training β the router may not generalize perfectly to all coding tasks
|
| 267 |
5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β some expert knowledge may be lost in translation
|
| 268 |
6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache
|
|
|
|
| 269 |
|
| 270 |
## Citation
|
| 271 |
|
|
|
|
| 184 |
print(response)
|
| 185 |
```
|
| 186 |
|
| 187 |
+
## Quantization & Deployment
|
| 188 |
+
|
| 189 |
+
### bitsandbytes 4-bit (NF4) β ~4.5 GB VRAM
|
| 190 |
|
| 191 |
```python
|
| 192 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
|
| 193 |
import torch
|
| 194 |
|
| 195 |
bnb_config = BitsAndBytesConfig(
|
|
|
|
| 205 |
device_map="auto",
|
| 206 |
trust_remote_code=True,
|
| 207 |
)
|
| 208 |
+
tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite")
|
| 209 |
```
|
| 210 |
|
| 211 |
+
### bitsandbytes 8-bit β ~7 GB VRAM
|
| 212 |
|
| 213 |
```python
|
| 214 |
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
|
|
|
|
| 234 |
model.set_coding_enabled(True)
|
| 235 |
```
|
| 236 |
|
| 237 |
+
### vLLM
|
| 238 |
+
|
| 239 |
+
> **Note:** vLLM does not currently support the custom `Fuse3ForCausalLM` architecture.
|
| 240 |
+
> The model uses LFM2's hybrid conv+attention backbone with augmented MoE layers,
|
| 241 |
+
> which requires a custom vLLM model implementation. Use transformers for inference.
|
| 242 |
+
|
| 243 |
+
To use with vLLM, you would need to:
|
| 244 |
+
1. Write a custom vLLM model definition for `Fuse3ForCausalLM`
|
| 245 |
+
2. Register it with vLLM's model registry
|
| 246 |
+
3. Handle the hybrid conv+attention layers and expert routing
|
| 247 |
+
|
| 248 |
+
Contributions welcome β see the model code in `fuse3_model.py` for the full architecture.
|
| 249 |
+
|
| 250 |
+
### MLX (Apple Silicon)
|
| 251 |
+
|
| 252 |
+
> **Note:** MLX does not currently support the custom `Fuse3ForCausalLM` architecture.
|
| 253 |
+
> The model's hybrid conv+attention layers and MoE expert routing require a custom
|
| 254 |
+
> MLX model implementation. Use transformers with MPS backend on Apple Silicon.
|
| 255 |
+
|
| 256 |
+
```python
|
| 257 |
+
# Run on Apple Silicon with MPS backend
|
| 258 |
+
import torch
|
| 259 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 260 |
+
|
| 261 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 262 |
+
"Akahsizrr/fuse-1-Lite",
|
| 263 |
+
torch_dtype=torch.float16,
|
| 264 |
+
device_map="mps",
|
| 265 |
+
trust_remote_code=True,
|
| 266 |
+
)
|
| 267 |
+
tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite")
|
| 268 |
+
```
|
| 269 |
+
|
| 270 |
+
### GGUF / llama.cpp
|
| 271 |
+
|
| 272 |
+
> **Note:** GGUF conversion is not currently supported. The custom `Fuse3ForCausalLM`
|
| 273 |
+
> architecture is not recognized by llama.cpp's `convert_hf_to_gguf.py`. LFM2 itself
|
| 274 |
+
> IS supported by llama.cpp (see `lfm2.cpp`), but the expert augmentation layers
|
| 275 |
+
> require a custom llama.cpp model definition.
|
| 276 |
+
|
| 277 |
+
### Transformers (Recommended)
|
| 278 |
+
|
| 279 |
+
The recommended way to run fuse-1 Lite is with transformers:
|
| 280 |
+
|
| 281 |
+
```bash
|
| 282 |
+
pip install transformers torch bitsandbytes accelerate
|
| 283 |
+
```
|
| 284 |
+
|
| 285 |
## Performance
|
| 286 |
|
| 287 |
### VRAM Requirements
|
|
|
|
| 313 |
|
| 314 |
1. **Custom architecture**: Requires `trust_remote_code=True` β the model includes custom `Fuse3ForCausalLM` code
|
| 315 |
2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine
|
| 316 |
+
3. **No GGUF/MLX conversion**: The custom architecture cannot be directly converted to GGUF or MLX format
|
| 317 |
4. **Training data was small**: Only 55 examples were used for router training β the router may not generalize perfectly to all coding tasks
|
| 318 |
5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β some expert knowledge may be lost in translation
|
| 319 |
6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache
|
| 320 |
+
7. **bitsandbytes quantization**: 4-bit and 8-bit quantization work at runtime via `BitsAndBytesConfig` β pre-quantized saved versions are not available as separate repos
|
| 321 |
|
| 322 |
## Citation
|
| 323 |
|