File size: 14,551 Bytes
23e351b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5e26d0
 
97cc08f
 
d74e868
 
97cc08f
 
 
d74e868
 
 
97cc08f
d74e868
23e351b
 
b5e26d0
23e351b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5e26d0
23e351b
 
d74e868
23e351b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d74e868
b5e26d0
d74e868
 
b5e26d0
d74e868
b5e26d0
d74e868
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5e26d0
 
 
d74e868
 
 
b5e26d0
 
d74e868
b5e26d0
d74e868
 
 
 
 
b5e26d0
d74e868
 
 
 
 
 
 
 
b5e26d0
 
d74e868
 
 
 
b5e26d0
 
d74e868
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5e26d0
d74e868
b5e26d0
d74e868
b5e26d0
 
 
 
 
23e351b
 
 
 
d74e868
 
 
 
 
 
 
 
 
23e351b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b5e26d0
23e351b
 
 
b5e26d0
23e351b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
---

language:
  - en
license: apache-2.0
library_name: transformers
tags:
  - mixture-of-experts
  - MoE
  - coding
  - python
  - code-generation
  - LFM2
  - Qwen
  - LiquidAI
  - small-language-model
  - SLM
  - agentic
  - fusion
  - expert-routing
  - 5B
  - efficient-inference
base_model:
  - LiquidAI/LFM2.5-2.6B
  - Qwen/Qwen3.6-35B-A3B
pipeline_tag: text-generation
inference: false
model-index:
  - name: fuse-1 Lite
    results:
      - task:
          type: text-generation
          name: Code Generation
        dataset:
          name: HumanEval (style)
          type: openai_humaneval
        metrics:
          - type: pass@1
            value: "TBD"
            name: pass@1
---


# fuse-1 Lite β€” Coding-Enhanced Mixture-of-Experts (5.72B)

> **A 5.72B parameter Mixture-of-Experts model that fuses LiquidAI's LFM2.5-2.6B host with 960 coding experts extracted from Qwen3.6-35B-A3B. Designed for efficient coding assistance, agentic workflows, and on-device inference.**

## Overview

**fuse-1 Lite** is a novel fusion model that combines the speed and efficiency of a small language model (LFM2.5-2.6B, 2.70B params) with the coding expertise of a large MoE model (Qwen3.6-35B-A3B). Rather than distilling knowledge or fine-tuning from scratch, fuse-1 Lite **transplants actual expert weights** from the donor model and trains a lightweight router to selectively activate them only when coding-related tokens are encountered.

### Key Innovation

Traditional model fusion requires either:
1. **Knowledge distillation** (slow, lossy, requires teacher inference)
2. **Weight merging** (requires compatible architectures)
3. **Full fine-tuning** (expensive, risks catastrophic forgetting)

fuse-1 Lite takes a different approach: **surgical expert transplantation with learned routing**. The 960 coding-specialized experts from Qwen3.6-35B-A3B are extracted, normalized, and integrated as residual augmentations to LFM2.5's decoder layers. A per-layer router learns which experts to activate for each token, and a learned scale factor controls how much each layer's experts contribute.

## Architecture

| Component | Details |
|-----------|---------|
| **Host Model** | LiquidAI/LFM2.5-2.6B (30 layers, hidden_size=2048) |

| **Donor Model** | Qwen/Qwen3.6-35B-A3B (MoE, 35B params) |

| **Total Parameters** | 5.72B |

| **Host Parameters** | 2.70B (frozen) |

| **Expert Parameters** | 3.02B (frozen, from Qwen3.6) |

| **Trainable Parameters** | 2.0M (router + scale only) |

| **Expert Count** | 960 across 30 layers (32 per layer) |

| **Top-K Routing** | 8 experts per token |

| **Expert Intermediate Size** | 512 |

| **Precision** | bfloat16 |



### How It Works



```

Input Token

    β”‚

    β–Ό

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”

β”‚  LFM2.5 Decoder Layer (frozen host)     β”‚

β”‚  β”œβ”€β”€ Attention / ShortConv              β”‚

β”‚  └── SwiGLU FFN                         β”‚

β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

              β”‚

              β–Ό

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”

β”‚  Fuse3 Augmented Layer                  β”‚

β”‚  1. Router scores token β†’ top-8 experts β”‚

β”‚  2. Selected experts compute SwiGLU     β”‚

β”‚  3. Output normalized to host std       β”‚

β”‚  4. Scaled by learned expert_scale      β”‚
β”‚  5. Added to residual stream            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```



### Expert Activation Pattern



After training, the router learned to selectively activate experts in **11 of 30 layers**:



| Layer | Scale | Role |

|-------|-------|------|

| 0 | 2.80 | Token-level feature extraction |

| 10-11 | 3.83-4.44 | Mid-level code structure |

| 13-15 | 4.56-4.84 | Algorithmic reasoning |

| 17 | 4.09 | Logic flow |

| 19 | 6.66 | **Peak coding expertise** |

| 21 | 4.94 | Code synthesis |

| 23 | 3.28 | Output formatting |

| 26 | 4.31 | Final code refinement |



The remaining 19 layers have scale β‰ˆ 0 (experts effectively disabled), preserving LFM2's general language capabilities.



## Training Details



| Parameter | Value |

|-----------|-------|

| **Training Steps** | 300 |

| **Training Time** | 8.3 minutes (Modal L4 GPU) |

| **Router Learning Rate** | 1e-3 |

| **Scale Learning Rate** | 1.0 |

| **Loss** | LM loss + 0.01 Γ— Load Balancing loss |

| **Optimizer** | AdamW |

| **Gradient Accumulation** | 4 |

| **Max Sequence Length** | 512 |

| **Training Data** | 55 examples (40 coding + 15 general) |

| **Total Compute Cost** | ~$3 (all phases combined) |



### Training Phases



1. **Phase 1 β€” Expert Profiling** ($1.90, A100 80GB): Profile Qwen3.6-35B-A3B expert activations on coding vs non-coding prompts. Select 960 coding-specialized experts. Extract weights.



2. **Phase 2 β€” Assembly** ($0.50, L4): Load LFM2.5-2.6B, wrap decoder layers with Fuse3AugmentedLayer, load extracted expert weights. Verify base model output coherence.



3. **Phase 3 β€” Router Training** ($0.60, L4): Train router + expert_scale on coding and general examples. Router learns to activate experts only for coding-related tokens.



### Stability Mechanisms



- **Std normalization**: Expert outputs are rescaled to match host layer activation std (ratio clamped to 2.0)

- **Scale clamping**: Expert scale clamped to max=0.1 during forward pass

- **SwiGLU clamping**: Expert intermediate activations clamped to [-10, 10]

- **Frozen host + experts**: Only 2.0M parameters are trainable (router weights + per-layer scale)



## Usage



### Installation



```bash

pip install transformers torch

```

### Quick Start

```python

from transformers import AutoTokenizer, AutoModelForCausalLM

import torch



model_id = "Akahsizrr/fuse-1-Lite"



tokenizer = AutoTokenizer.from_pretrained(model_id)

model = AutoModelForCausalLM.from_pretrained(

    model_id,

    torch_dtype=torch.bfloat16,

    device_map="auto",

    trust_remote_code=True,

)



messages = [{"role": "user", "content": "Write a Python function to check if a number is prime."}]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

inputs = tokenizer(text, return_tensors="pt").to(model.device)



with torch.no_grad():

    outputs = model.generate(

        **inputs,

        max_new_tokens=1024,

        do_sample=True,

        temperature=0.1,

        top_k=50,

        repetition_penalty=1.1,

        pad_token_id=tokenizer.pad_token_id,

    )



response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)

print(response)

```

## Quantization & Deployment

### Pre-quantized Versions

| Version | Repo | VRAM/Memory | Format |
|---------|------|-------------|--------|
| **4-bit NF4** | [`Akahsizrr/fuse-1-Lite-4bit`](https://huggingface.co/Akahsizrr/fuse-1-Lite-4bit) | 3.36 GB | bitsandbytes |
| **8-bit** | [`Akahsizrr/fuse-1-Lite-8bit`](https://huggingface.co/Akahsizrr/fuse-1-Lite-8bit) | 6.00 GB | bitsandbytes |
| **bfloat16** | This repo | ~12 GB | safetensors |
| **MLX** | [`Akahsizrr/fuse-1-Lite-MLX`](https://huggingface.co/Akahsizrr/fuse-1-Lite-MLX) | ~12 GB | MLX safetensors |
| **GGUF F16** | [`Akahsizrr/fuse-1-Lite-GGUF`](https://huggingface.co/Akahsizrr/fuse-1-Lite-GGUF) | ~11.4 GB | GGUF |
| **vLLM plugin** | [`Akahsizrr/fuse-1-Lite-vLLM`](https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM) | ~12 GB | vLLM plugin |

### bitsandbytes 4-bit (NF4) β€” 3.36 GB VRAM

```python

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

import torch



bnb_config = BitsAndBytesConfig(

    load_in_4bit=True,

    bnb_4bit_quant_type="nf4",

    bnb_4bit_compute_dtype=torch.bfloat16,

    bnb_4bit_use_double_quant=True,

)



model = AutoModelForCausalLM.from_pretrained(

    "Akahsizrr/fuse-1-Lite",

    quantization_config=bnb_config,

    device_map="auto",

    trust_remote_code=True,

)

tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite")

```

### bitsandbytes 8-bit β€” 6.00 GB VRAM

```python

from transformers import AutoModelForCausalLM, BitsAndBytesConfig

import torch



bnb_config = BitsAndBytesConfig(load_in_8bit=True)



model = AutoModelForCausalLM.from_pretrained(

    "Akahsizrr/fuse-1-Lite",

    quantization_config=bnb_config,

    device_map="auto",

    trust_remote_code=True,

)

```

### Disable Coding Experts (Pure LFM2 Mode)

```python

# The model includes a coding toggle β€” disable experts for pure LFM2 inference

model.set_coding_enabled(False)



# Re-enable for coding tasks

model.set_coding_enabled(True)

```

### vLLM β€” High-Throughput Serving

fuse-1 Lite is supported in vLLM via a plugin that extends vLLM's native LFM2
implementation with expert augmentation layers.

**Plugin repo**: [`Akahsizrr/fuse-1-Lite-vLLM`](https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM)

```bash

# Install the plugin

pip install git+https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM



# Serve with vLLM

vllm serve Akahsizrr/fuse-1-Lite \

  --mamba-cache-mode align \

  --max-model-len 4096

```

```python

from vllm import LLM



llm = LLM(

    model="Akahsizrr/fuse-1-Lite",

    mamba_cache_mode="align",

    max_model_len=4096,

)

output = llm.generate("Write a Python function to check if a number is prime.")

```

The plugin registers `Fuse3ForCausalLM` with vLLM's `ModelRegistry` via the
`vllm.general_plugins` entry point. It reuses vLLM's native LFM2 attention
and short-conv layers, adding the expert MoE block after each augmented
layer's FFN.

### MLX (Apple Silicon)

fuse-1 Lite is available in MLX format for Apple Silicon (M1+).

**MLX repo**: [`Akahsizrr/fuse-1-Lite-MLX`](https://huggingface.co/Akahsizrr/fuse-1-Lite-MLX)

```python

from mlx_lm import load, generate



model, tokenizer = load("Akahsizrr/fuse-1-Lite-MLX", trust_remote_code=True)



prompt = tokenizer.apply_chat_template(

    [{"role": "user", "content": "Write a Python function to check if a number is prime."}],

    tokenize=False, add_generation_prompt=True,

)



response = generate(model, tokenizer, prompt=prompt, max_tokens=512)

print(response)

```

```bash

# CLI

mlx_lm.generate --model Akahsizrr/fuse-1-Lite-MLX --trust-remote-code --prompt "Write a Python fizzbuzz"

```

The MLX model file (`fuse3_mlx.py`) extends MLX's native LFM2 implementation
with the same expert MoE augmentation. It uses `model_file` in config.json
with `trust_remote_code=True` for loading.

### GGUF / llama.cpp

fuse-1 Lite is available in GGUF format for llama.cpp.

**GGUF repo**: [`Akahsizrr/fuse-1-Lite-GGUF`](https://huggingface.co/Akahsizrr/fuse-1-Lite-GGUF)

> **Note:** The GGUF uses the custom `fuse3` architecture. Stock llama.cpp
> cannot load it β€” you need a llama.cpp fork with Fuse3 support. The GGUF
> repo includes the C++ graph builder (`src/models/fuse3.cpp`), Python
> converter (`conversion/fuse3.py`), and integration guide (`INTEGRATION.md`).

```bash

# Build llama.cpp with Fuse3 support (see INTEGRATION.md in the GGUF repo)

./llama-cli -m fuse-1-Lite-f16.gguf \

  -p "Write a Python function to check if a number is prime." \

  -n 512 --temp 0.1

```

The C++ implementation reuses LFM2's attention and short-conv graph builders,
adding the expert MoE block (router β†’ top-k β†’ SwiGLU experts β†’ scale β†’ add)
after each augmented layer's dense FFN.

### Transformers (Universal)

The recommended way to run fuse-1 Lite on any platform:

```bash

pip install transformers torch bitsandbytes accelerate

```

## Performance

### VRAM Requirements

| Backend | Precision | VRAM/Memory | Recommended Hardware |
|---------|-----------|-------------|---------------------|
| Transformers | bfloat16 | ~12 GB | L4, A10G, RTX 4090 |
| Transformers | 8-bit | 6.00 GB | T4, L4, RTX 3060 |
| Transformers | 4-bit | 3.36 GB | T4, RTX 3060, M2 Pro |
| vLLM | bfloat16 | ~12 GB | A10G, A100, H100 |
| MLX | float16 | ~12 GB | M1 Pro+, M2, M3, M4 |
| llama.cpp | F16 | ~11.4 GB | Any CPU/GPU |
| llama.cpp | Q4_K_M | ~4 GB | Any CPU/GPU |

### Sample Outputs

**Prompt**: "Write a Python function to check if a string is a palindrome."

**Output** (excerpt):
```python

def is_palindrome(s: str) -> bool:

    """

    Return True if *s* reads the same forwards and backwards,

    ignoring case and non-alphanumeric characters.

    """

    cleaned = ''.join(ch.lower() for ch in s if ch.isalnum())

    return cleaned == cleaned[::-1]

```

The model produces complete implementations with docstrings, type hints, complexity analysis, and test cases.

## Limitations

1. **Custom architecture**: Requires `trust_remote_code=True` β€” the model includes custom `Fuse3ForCausalLM` code
2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine
3. **No GGUF/MLX conversion**: The custom architecture cannot be directly converted to GGUF or MLX format
4. **Training data was small**: Only 55 examples were used for router training β€” the router may not generalize perfectly to all coding tasks
5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β€” some expert knowledge may be lost in translation
6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache

7. **bitsandbytes quantization**: 4-bit and 8-bit quantization work at runtime via `BitsAndBytesConfig` β€” pre-quantized saved versions are not available as separate repos



## Citation



```bibtex

@misc{fuse1lite2026,

  title={fuse-1 Lite: Coding-Enhanced Mixture-of-Experts via Expert Transplantation},

  author={Vasko Djack},

  year={2026},

  publisher={HuggingFace},

  url={https://huggingface.co/Akahsizrr/fuse-1-Lite}

}

```



## Acknowledgments



- **LiquidAI** for the LFM2.5-2.6B host model

- **Qwen Team** for the Qwen3.6-35B-A3B donor model

- **Modal** for compute infrastructure



## License



Apache 2.0 β€” See [LICENSE](LICENSE) for details.



---

*Built with [Devin](https://devin.ai) β€” Cognition AI*