fuse-1-Lite / README.md
Akahsizrr's picture
Add vLLM, MLX, and llama.cpp deployment docs with working implementations
d74e868 verified
|
Raw History Blame Contribute Delete
14.6 kB
---
language:
- en
license: apache-2.0
library_name: transformers
tags:
- mixture-of-experts
- MoE
- coding
- python
- code-generation
- LFM2
- Qwen
- LiquidAI
- small-language-model
- SLM
- agentic
- fusion
- expert-routing
- 5B
- efficient-inference
base_model:
- LiquidAI/LFM2.5-2.6B
- Qwen/Qwen3.6-35B-A3B
pipeline_tag: text-generation
inference: false
model-index:
- name: fuse-1 Lite
results:
- task:
type: text-generation
name: Code Generation
dataset:
name: HumanEval (style)
type: openai_humaneval
metrics:
- type: pass@1
value: "TBD"
name: pass@1
---
# fuse-1 Lite β€” Coding-Enhanced Mixture-of-Experts (5.72B)
> **A 5.72B parameter Mixture-of-Experts model that fuses LiquidAI's LFM2.5-2.6B host with 960 coding experts extracted from Qwen3.6-35B-A3B. Designed for efficient coding assistance, agentic workflows, and on-device inference.**
## Overview
**fuse-1 Lite** is a novel fusion model that combines the speed and efficiency of a small language model (LFM2.5-2.6B, 2.70B params) with the coding expertise of a large MoE model (Qwen3.6-35B-A3B). Rather than distilling knowledge or fine-tuning from scratch, fuse-1 Lite **transplants actual expert weights** from the donor model and trains a lightweight router to selectively activate them only when coding-related tokens are encountered.
### Key Innovation
Traditional model fusion requires either:
1. **Knowledge distillation** (slow, lossy, requires teacher inference)
2. **Weight merging** (requires compatible architectures)
3. **Full fine-tuning** (expensive, risks catastrophic forgetting)
fuse-1 Lite takes a different approach: **surgical expert transplantation with learned routing**. The 960 coding-specialized experts from Qwen3.6-35B-A3B are extracted, normalized, and integrated as residual augmentations to LFM2.5's decoder layers. A per-layer router learns which experts to activate for each token, and a learned scale factor controls how much each layer's experts contribute.
## Architecture
| Component | Details |
|-----------|---------|
| **Host Model** | LiquidAI/LFM2.5-2.6B (30 layers, hidden_size=2048) |
| **Donor Model** | Qwen/Qwen3.6-35B-A3B (MoE, 35B params) |
| **Total Parameters** | 5.72B |
| **Host Parameters** | 2.70B (frozen) |
| **Expert Parameters** | 3.02B (frozen, from Qwen3.6) |
| **Trainable Parameters** | 2.0M (router + scale only) |
| **Expert Count** | 960 across 30 layers (32 per layer) |
| **Top-K Routing** | 8 experts per token |
| **Expert Intermediate Size** | 512 |
| **Precision** | bfloat16 |
### How It Works
```
Input Token
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ LFM2.5 Decoder Layer (frozen host) β”‚
β”‚ β”œβ”€β”€ Attention / ShortConv β”‚
β”‚ └── SwiGLU FFN β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Fuse3 Augmented Layer β”‚
β”‚ 1. Router scores token β†’ top-8 experts β”‚
β”‚ 2. Selected experts compute SwiGLU β”‚
β”‚ 3. Output normalized to host std β”‚
β”‚ 4. Scaled by learned expert_scale β”‚
β”‚ 5. Added to residual stream β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
### Expert Activation Pattern
After training, the router learned to selectively activate experts in **11 of 30 layers**:
| Layer | Scale | Role |
|-------|-------|------|
| 0 | 2.80 | Token-level feature extraction |
| 10-11 | 3.83-4.44 | Mid-level code structure |
| 13-15 | 4.56-4.84 | Algorithmic reasoning |
| 17 | 4.09 | Logic flow |
| 19 | 6.66 | **Peak coding expertise** |
| 21 | 4.94 | Code synthesis |
| 23 | 3.28 | Output formatting |
| 26 | 4.31 | Final code refinement |
The remaining 19 layers have scale β‰ˆ 0 (experts effectively disabled), preserving LFM2's general language capabilities.
## Training Details
| Parameter | Value |
|-----------|-------|
| **Training Steps** | 300 |
| **Training Time** | 8.3 minutes (Modal L4 GPU) |
| **Router Learning Rate** | 1e-3 |
| **Scale Learning Rate** | 1.0 |
| **Loss** | LM loss + 0.01 Γ— Load Balancing loss |
| **Optimizer** | AdamW |
| **Gradient Accumulation** | 4 |
| **Max Sequence Length** | 512 |
| **Training Data** | 55 examples (40 coding + 15 general) |
| **Total Compute Cost** | ~$3 (all phases combined) |
### Training Phases
1. **Phase 1 β€” Expert Profiling** ($1.90, A100 80GB): Profile Qwen3.6-35B-A3B expert activations on coding vs non-coding prompts. Select 960 coding-specialized experts. Extract weights.
2. **Phase 2 β€” Assembly** ($0.50, L4): Load LFM2.5-2.6B, wrap decoder layers with Fuse3AugmentedLayer, load extracted expert weights. Verify base model output coherence.
3. **Phase 3 β€” Router Training** ($0.60, L4): Train router + expert_scale on coding and general examples. Router learns to activate experts only for coding-related tokens.
### Stability Mechanisms
- **Std normalization**: Expert outputs are rescaled to match host layer activation std (ratio clamped to 2.0)
- **Scale clamping**: Expert scale clamped to max=0.1 during forward pass
- **SwiGLU clamping**: Expert intermediate activations clamped to [-10, 10]
- **Frozen host + experts**: Only 2.0M parameters are trainable (router weights + per-layer scale)
## Usage
### Installation
```bash
pip install transformers torch
```
### Quick Start
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "Akahsizrr/fuse-1-Lite"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Write a Python function to check if a number is prime."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=1024,
do_sample=True,
temperature=0.1,
top_k=50,
repetition_penalty=1.1,
pad_token_id=tokenizer.pad_token_id,
)
response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
print(response)
```
## Quantization & Deployment
### Pre-quantized Versions
| Version | Repo | VRAM/Memory | Format |
|---------|------|-------------|--------|
| **4-bit NF4** | [`Akahsizrr/fuse-1-Lite-4bit`](https://huggingface.co/Akahsizrr/fuse-1-Lite-4bit) | 3.36 GB | bitsandbytes |
| **8-bit** | [`Akahsizrr/fuse-1-Lite-8bit`](https://huggingface.co/Akahsizrr/fuse-1-Lite-8bit) | 6.00 GB | bitsandbytes |
| **bfloat16** | This repo | ~12 GB | safetensors |
| **MLX** | [`Akahsizrr/fuse-1-Lite-MLX`](https://huggingface.co/Akahsizrr/fuse-1-Lite-MLX) | ~12 GB | MLX safetensors |
| **GGUF F16** | [`Akahsizrr/fuse-1-Lite-GGUF`](https://huggingface.co/Akahsizrr/fuse-1-Lite-GGUF) | ~11.4 GB | GGUF |
| **vLLM plugin** | [`Akahsizrr/fuse-1-Lite-vLLM`](https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM) | ~12 GB | vLLM plugin |
### bitsandbytes 4-bit (NF4) β€” 3.36 GB VRAM
```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(
"Akahsizrr/fuse-1-Lite",
quantization_config=bnb_config,
device_map="auto",
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite")
```
### bitsandbytes 8-bit β€” 6.00 GB VRAM
```python
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
bnb_config = BitsAndBytesConfig(load_in_8bit=True)
model = AutoModelForCausalLM.from_pretrained(
"Akahsizrr/fuse-1-Lite",
quantization_config=bnb_config,
device_map="auto",
trust_remote_code=True,
)
```
### Disable Coding Experts (Pure LFM2 Mode)
```python
# The model includes a coding toggle β€” disable experts for pure LFM2 inference
model.set_coding_enabled(False)
# Re-enable for coding tasks
model.set_coding_enabled(True)
```
### vLLM β€” High-Throughput Serving
fuse-1 Lite is supported in vLLM via a plugin that extends vLLM's native LFM2
implementation with expert augmentation layers.
**Plugin repo**: [`Akahsizrr/fuse-1-Lite-vLLM`](https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM)
```bash
# Install the plugin
pip install git+https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM
# Serve with vLLM
vllm serve Akahsizrr/fuse-1-Lite \
--mamba-cache-mode align \
--max-model-len 4096
```
```python
from vllm import LLM
llm = LLM(
model="Akahsizrr/fuse-1-Lite",
mamba_cache_mode="align",
max_model_len=4096,
)
output = llm.generate("Write a Python function to check if a number is prime.")
```
The plugin registers `Fuse3ForCausalLM` with vLLM's `ModelRegistry` via the
`vllm.general_plugins` entry point. It reuses vLLM's native LFM2 attention
and short-conv layers, adding the expert MoE block after each augmented
layer's FFN.
### MLX (Apple Silicon)
fuse-1 Lite is available in MLX format for Apple Silicon (M1+).
**MLX repo**: [`Akahsizrr/fuse-1-Lite-MLX`](https://huggingface.co/Akahsizrr/fuse-1-Lite-MLX)
```python
from mlx_lm import load, generate
model, tokenizer = load("Akahsizrr/fuse-1-Lite-MLX", trust_remote_code=True)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Write a Python function to check if a number is prime."}],
tokenize=False, add_generation_prompt=True,
)
response = generate(model, tokenizer, prompt=prompt, max_tokens=512)
print(response)
```
```bash
# CLI
mlx_lm.generate --model Akahsizrr/fuse-1-Lite-MLX --trust-remote-code --prompt "Write a Python fizzbuzz"
```
The MLX model file (`fuse3_mlx.py`) extends MLX's native LFM2 implementation
with the same expert MoE augmentation. It uses `model_file` in config.json
with `trust_remote_code=True` for loading.
### GGUF / llama.cpp
fuse-1 Lite is available in GGUF format for llama.cpp.
**GGUF repo**: [`Akahsizrr/fuse-1-Lite-GGUF`](https://huggingface.co/Akahsizrr/fuse-1-Lite-GGUF)
> **Note:** The GGUF uses the custom `fuse3` architecture. Stock llama.cpp
> cannot load it β€” you need a llama.cpp fork with Fuse3 support. The GGUF
> repo includes the C++ graph builder (`src/models/fuse3.cpp`), Python
> converter (`conversion/fuse3.py`), and integration guide (`INTEGRATION.md`).
```bash
# Build llama.cpp with Fuse3 support (see INTEGRATION.md in the GGUF repo)
./llama-cli -m fuse-1-Lite-f16.gguf \
-p "Write a Python function to check if a number is prime." \
-n 512 --temp 0.1
```
The C++ implementation reuses LFM2's attention and short-conv graph builders,
adding the expert MoE block (router β†’ top-k β†’ SwiGLU experts β†’ scale β†’ add)
after each augmented layer's dense FFN.
### Transformers (Universal)
The recommended way to run fuse-1 Lite on any platform:
```bash
pip install transformers torch bitsandbytes accelerate
```
## Performance
### VRAM Requirements
| Backend | Precision | VRAM/Memory | Recommended Hardware |
|---------|-----------|-------------|---------------------|
| Transformers | bfloat16 | ~12 GB | L4, A10G, RTX 4090 |
| Transformers | 8-bit | 6.00 GB | T4, L4, RTX 3060 |
| Transformers | 4-bit | 3.36 GB | T4, RTX 3060, M2 Pro |
| vLLM | bfloat16 | ~12 GB | A10G, A100, H100 |
| MLX | float16 | ~12 GB | M1 Pro+, M2, M3, M4 |
| llama.cpp | F16 | ~11.4 GB | Any CPU/GPU |
| llama.cpp | Q4_K_M | ~4 GB | Any CPU/GPU |
### Sample Outputs
**Prompt**: "Write a Python function to check if a string is a palindrome."
**Output** (excerpt):
```python
def is_palindrome(s: str) -> bool:
"""
Return True if *s* reads the same forwards and backwards,
ignoring case and non-alphanumeric characters.
"""
cleaned = ''.join(ch.lower() for ch in s if ch.isalnum())
return cleaned == cleaned[::-1]
```
The model produces complete implementations with docstrings, type hints, complexity analysis, and test cases.
## Limitations
1. **Custom architecture**: Requires `trust_remote_code=True` β€” the model includes custom `Fuse3ForCausalLM` code
2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine
3. **No GGUF/MLX conversion**: The custom architecture cannot be directly converted to GGUF or MLX format
4. **Training data was small**: Only 55 examples were used for router training β€” the router may not generalize perfectly to all coding tasks
5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β€” some expert knowledge may be lost in translation
6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache
7. **bitsandbytes quantization**: 4-bit and 8-bit quantization work at runtime via `BitsAndBytesConfig` β€” pre-quantized saved versions are not available as separate repos
## Citation
```bibtex
@misc{fuse1lite2026,
title={fuse-1 Lite: Coding-Enhanced Mixture-of-Experts via Expert Transplantation},
author={Vasko Djack},
year={2026},
publisher={HuggingFace},
url={https://huggingface.co/Akahsizrr/fuse-1-Lite}
}
```
## Acknowledgments
- **LiquidAI** for the LFM2.5-2.6B host model
- **Qwen Team** for the Qwen3.6-35B-A3B donor model
- **Modal** for compute infrastructure
## License
Apache 2.0 β€” See [LICENSE](LICENSE) for details.
---
*Built with [Devin](https://devin.ai) β€” Cognition AI*