Fuse Models
Collection
Models I release using my fusion ( experts merging ) architecture. • 9 items • Updated
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
vLLM plugin for serving Fuse3 (fuse-1 Lite) models — LFM2 host + Qwen3.6 coding experts.
pip install -e .
Once installed, vLLM automatically discovers the plugin via Python entry points.
# Serve fuse-1 Lite with vLLM
vllm serve Akahsizrr/fuse-1-Lite \
--mamba-cache-mode align \
--max-model-len 4096
# Or with the Python API
from vllm import LLM
llm = LLM(
model="Akahsizrr/fuse-1-Lite",
mamba_cache_mode="align",
max_model_len=4096,
)
The plugin registers Fuse3ForCausalLM with vLLM's ModelRegistry. The model
extends vLLM's native LFM2 implementation:
Lfm2AttentionDecoderLayer and Lfm2ShortConvDecoderLayerAutoWeightsLoader with a WeightsMapper that handles
the host_layer. prefix and conv weight renamingThe plugin does NOT require --trust-remote-code — the model is registered natively.