Text Generation
Transformers
Safetensors
English
fuse3
mixture-of-experts
MoE
coding
python
code-generation
LFM2
Qwen
LiquidAI
small-language-model
SLM
agentic
fusion
expert-routing
5B
efficient-inference
conversational
custom_code
Eval Results (legacy)
Instructions to use Akahsizrr/fuse-1-Lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Akahsizrr/fuse-1-Lite with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Akahsizrr/fuse-1-Lite", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Akahsizrr/fuse-1-Lite", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Akahsizrr/fuse-1-Lite with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Akahsizrr/fuse-1-Lite" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akahsizrr/fuse-1-Lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Akahsizrr/fuse-1-Lite
- SGLang
How to use Akahsizrr/fuse-1-Lite with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Akahsizrr/fuse-1-Lite" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akahsizrr/fuse-1-Lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Akahsizrr/fuse-1-Lite" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akahsizrr/fuse-1-Lite", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Akahsizrr/fuse-1-Lite with Docker Model Runner:
docker model run hf.co/Akahsizrr/fuse-1-Lite
|
Download README.md from Akahsizrr/fuse-1-Lite: direct link, hf CLI and curl.
- Browser
- Download file 14.6 kB
-
https://huggingface.co/Akahsizrr/fuse-1-Lite/resolve/main/README.md
- Command line
-
hf download hf://Akahsizrr/fuse-1-Lite/README.md
-
curl -L -o README.md https://huggingface.co/Akahsizrr/fuse-1-Lite/resolve/main/README.md
14.6 kB
| language: | |
| - en | |
| license: apache-2.0 | |
| library_name: transformers | |
| tags: | |
| - mixture-of-experts | |
| - MoE | |
| - coding | |
| - python | |
| - code-generation | |
| - LFM2 | |
| - Qwen | |
| - LiquidAI | |
| - small-language-model | |
| - SLM | |
| - agentic | |
| - fusion | |
| - expert-routing | |
| - 5B | |
| - efficient-inference | |
| base_model: | |
| - LiquidAI/LFM2.5-2.6B | |
| - Qwen/Qwen3.6-35B-A3B | |
| pipeline_tag: text-generation | |
| inference: false | |
| model-index: | |
| - name: fuse-1 Lite | |
| results: | |
| - task: | |
| type: text-generation | |
| name: Code Generation | |
| dataset: | |
| name: HumanEval (style) | |
| type: openai_humaneval | |
| metrics: | |
| - type: pass@1 | |
| value: "TBD" | |
| name: pass@1 | |
| # fuse-1 Lite β Coding-Enhanced Mixture-of-Experts (5.72B) | |
| > **A 5.72B parameter Mixture-of-Experts model that fuses LiquidAI's LFM2.5-2.6B host with 960 coding experts extracted from Qwen3.6-35B-A3B. Designed for efficient coding assistance, agentic workflows, and on-device inference.** | |
| ## Overview | |
| **fuse-1 Lite** is a novel fusion model that combines the speed and efficiency of a small language model (LFM2.5-2.6B, 2.70B params) with the coding expertise of a large MoE model (Qwen3.6-35B-A3B). Rather than distilling knowledge or fine-tuning from scratch, fuse-1 Lite **transplants actual expert weights** from the donor model and trains a lightweight router to selectively activate them only when coding-related tokens are encountered. | |
| ### Key Innovation | |
| Traditional model fusion requires either: | |
| 1. **Knowledge distillation** (slow, lossy, requires teacher inference) | |
| 2. **Weight merging** (requires compatible architectures) | |
| 3. **Full fine-tuning** (expensive, risks catastrophic forgetting) | |
| fuse-1 Lite takes a different approach: **surgical expert transplantation with learned routing**. The 960 coding-specialized experts from Qwen3.6-35B-A3B are extracted, normalized, and integrated as residual augmentations to LFM2.5's decoder layers. A per-layer router learns which experts to activate for each token, and a learned scale factor controls how much each layer's experts contribute. | |
| ## Architecture | |
| | Component | Details | | |
| |-----------|---------| | |
| | **Host Model** | LiquidAI/LFM2.5-2.6B (30 layers, hidden_size=2048) | | |
| | **Donor Model** | Qwen/Qwen3.6-35B-A3B (MoE, 35B params) | | |
| | **Total Parameters** | 5.72B | | |
| | **Host Parameters** | 2.70B (frozen) | | |
| | **Expert Parameters** | 3.02B (frozen, from Qwen3.6) | | |
| | **Trainable Parameters** | 2.0M (router + scale only) | | |
| | **Expert Count** | 960 across 30 layers (32 per layer) | | |
| | **Top-K Routing** | 8 experts per token | | |
| | **Expert Intermediate Size** | 512 | | |
| | **Precision** | bfloat16 | | |
| ### How It Works | |
| ``` | |
| Input Token | |
| β | |
| βΌ | |
| βββββββββββββββββββββββββββββββββββββββββββ | |
| β LFM2.5 Decoder Layer (frozen host) β | |
| β βββ Attention / ShortConv β | |
| β βββ SwiGLU FFN β | |
| βββββββββββββββ¬ββββββββββββββββββββββββββββ | |
| β | |
| βΌ | |
| βββββββββββββββββββββββββββββββββββββββββββ | |
| β Fuse3 Augmented Layer β | |
| β 1. Router scores token β top-8 experts β | |
| β 2. Selected experts compute SwiGLU β | |
| β 3. Output normalized to host std β | |
| β 4. Scaled by learned expert_scale β | |
| β 5. Added to residual stream β | |
| βββββββββββββββββββββββββββββββββββββββββββ | |
| ``` | |
| ### Expert Activation Pattern | |
| After training, the router learned to selectively activate experts in **11 of 30 layers**: | |
| | Layer | Scale | Role | | |
| |-------|-------|------| | |
| | 0 | 2.80 | Token-level feature extraction | | |
| | 10-11 | 3.83-4.44 | Mid-level code structure | | |
| | 13-15 | 4.56-4.84 | Algorithmic reasoning | | |
| | 17 | 4.09 | Logic flow | | |
| | 19 | 6.66 | **Peak coding expertise** | | |
| | 21 | 4.94 | Code synthesis | | |
| | 23 | 3.28 | Output formatting | | |
| | 26 | 4.31 | Final code refinement | | |
| The remaining 19 layers have scale β 0 (experts effectively disabled), preserving LFM2's general language capabilities. | |
| ## Training Details | |
| | Parameter | Value | | |
| |-----------|-------| | |
| | **Training Steps** | 300 | | |
| | **Training Time** | 8.3 minutes (Modal L4 GPU) | | |
| | **Router Learning Rate** | 1e-3 | | |
| | **Scale Learning Rate** | 1.0 | | |
| | **Loss** | LM loss + 0.01 Γ Load Balancing loss | | |
| | **Optimizer** | AdamW | | |
| | **Gradient Accumulation** | 4 | | |
| | **Max Sequence Length** | 512 | | |
| | **Training Data** | 55 examples (40 coding + 15 general) | | |
| | **Total Compute Cost** | ~$3 (all phases combined) | | |
| ### Training Phases | |
| 1. **Phase 1 β Expert Profiling** ($1.90, A100 80GB): Profile Qwen3.6-35B-A3B expert activations on coding vs non-coding prompts. Select 960 coding-specialized experts. Extract weights. | |
| 2. **Phase 2 β Assembly** ($0.50, L4): Load LFM2.5-2.6B, wrap decoder layers with Fuse3AugmentedLayer, load extracted expert weights. Verify base model output coherence. | |
| 3. **Phase 3 β Router Training** ($0.60, L4): Train router + expert_scale on coding and general examples. Router learns to activate experts only for coding-related tokens. | |
| ### Stability Mechanisms | |
| - **Std normalization**: Expert outputs are rescaled to match host layer activation std (ratio clamped to 2.0) | |
| - **Scale clamping**: Expert scale clamped to max=0.1 during forward pass | |
| - **SwiGLU clamping**: Expert intermediate activations clamped to [-10, 10] | |
| - **Frozen host + experts**: Only 2.0M parameters are trainable (router weights + per-layer scale) | |
| ## Usage | |
| ### Installation | |
| ```bash | |
| pip install transformers torch | |
| ``` | |
| ### Quick Start | |
| ```python | |
| from transformers import AutoTokenizer, AutoModelForCausalLM | |
| import torch | |
| model_id = "Akahsizrr/fuse-1-Lite" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto", | |
| trust_remote_code=True, | |
| ) | |
| messages = [{"role": "user", "content": "Write a Python function to check if a number is prime."}] | |
| text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tokenizer(text, return_tensors="pt").to(model.device) | |
| with torch.no_grad(): | |
| outputs = model.generate( | |
| **inputs, | |
| max_new_tokens=1024, | |
| do_sample=True, | |
| temperature=0.1, | |
| top_k=50, | |
| repetition_penalty=1.1, | |
| pad_token_id=tokenizer.pad_token_id, | |
| ) | |
| response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True) | |
| print(response) | |
| ``` | |
| ## Quantization & Deployment | |
| ### Pre-quantized Versions | |
| | Version | Repo | VRAM/Memory | Format | | |
| |---------|------|-------------|--------| | |
| | **4-bit NF4** | [`Akahsizrr/fuse-1-Lite-4bit`](https://huggingface.co/Akahsizrr/fuse-1-Lite-4bit) | 3.36 GB | bitsandbytes | | |
| | **8-bit** | [`Akahsizrr/fuse-1-Lite-8bit`](https://huggingface.co/Akahsizrr/fuse-1-Lite-8bit) | 6.00 GB | bitsandbytes | | |
| | **bfloat16** | This repo | ~12 GB | safetensors | | |
| | **MLX** | [`Akahsizrr/fuse-1-Lite-MLX`](https://huggingface.co/Akahsizrr/fuse-1-Lite-MLX) | ~12 GB | MLX safetensors | | |
| | **GGUF F16** | [`Akahsizrr/fuse-1-Lite-GGUF`](https://huggingface.co/Akahsizrr/fuse-1-Lite-GGUF) | ~11.4 GB | GGUF | | |
| | **vLLM plugin** | [`Akahsizrr/fuse-1-Lite-vLLM`](https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM) | ~12 GB | vLLM plugin | | |
| ### bitsandbytes 4-bit (NF4) β 3.36 GB VRAM | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig | |
| import torch | |
| bnb_config = BitsAndBytesConfig( | |
| load_in_4bit=True, | |
| bnb_4bit_quant_type="nf4", | |
| bnb_4bit_compute_dtype=torch.bfloat16, | |
| bnb_4bit_use_double_quant=True, | |
| ) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "Akahsizrr/fuse-1-Lite", | |
| quantization_config=bnb_config, | |
| device_map="auto", | |
| trust_remote_code=True, | |
| ) | |
| tokenizer = AutoTokenizer.from_pretrained("Akahsizrr/fuse-1-Lite") | |
| ``` | |
| ### bitsandbytes 8-bit β 6.00 GB VRAM | |
| ```python | |
| from transformers import AutoModelForCausalLM, BitsAndBytesConfig | |
| import torch | |
| bnb_config = BitsAndBytesConfig(load_in_8bit=True) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "Akahsizrr/fuse-1-Lite", | |
| quantization_config=bnb_config, | |
| device_map="auto", | |
| trust_remote_code=True, | |
| ) | |
| ``` | |
| ### Disable Coding Experts (Pure LFM2 Mode) | |
| ```python | |
| # The model includes a coding toggle β disable experts for pure LFM2 inference | |
| model.set_coding_enabled(False) | |
| # Re-enable for coding tasks | |
| model.set_coding_enabled(True) | |
| ``` | |
| ### vLLM β High-Throughput Serving | |
| fuse-1 Lite is supported in vLLM via a plugin that extends vLLM's native LFM2 | |
| implementation with expert augmentation layers. | |
| **Plugin repo**: [`Akahsizrr/fuse-1-Lite-vLLM`](https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM) | |
| ```bash | |
| # Install the plugin | |
| pip install git+https://huggingface.co/Akahsizrr/fuse-1-Lite-vLLM | |
| # Serve with vLLM | |
| vllm serve Akahsizrr/fuse-1-Lite \ | |
| --mamba-cache-mode align \ | |
| --max-model-len 4096 | |
| ``` | |
| ```python | |
| from vllm import LLM | |
| llm = LLM( | |
| model="Akahsizrr/fuse-1-Lite", | |
| mamba_cache_mode="align", | |
| max_model_len=4096, | |
| ) | |
| output = llm.generate("Write a Python function to check if a number is prime.") | |
| ``` | |
| The plugin registers `Fuse3ForCausalLM` with vLLM's `ModelRegistry` via the | |
| `vllm.general_plugins` entry point. It reuses vLLM's native LFM2 attention | |
| and short-conv layers, adding the expert MoE block after each augmented | |
| layer's FFN. | |
| ### MLX (Apple Silicon) | |
| fuse-1 Lite is available in MLX format for Apple Silicon (M1+). | |
| **MLX repo**: [`Akahsizrr/fuse-1-Lite-MLX`](https://huggingface.co/Akahsizrr/fuse-1-Lite-MLX) | |
| ```python | |
| from mlx_lm import load, generate | |
| model, tokenizer = load("Akahsizrr/fuse-1-Lite-MLX", trust_remote_code=True) | |
| prompt = tokenizer.apply_chat_template( | |
| [{"role": "user", "content": "Write a Python function to check if a number is prime."}], | |
| tokenize=False, add_generation_prompt=True, | |
| ) | |
| response = generate(model, tokenizer, prompt=prompt, max_tokens=512) | |
| print(response) | |
| ``` | |
| ```bash | |
| # CLI | |
| mlx_lm.generate --model Akahsizrr/fuse-1-Lite-MLX --trust-remote-code --prompt "Write a Python fizzbuzz" | |
| ``` | |
| The MLX model file (`fuse3_mlx.py`) extends MLX's native LFM2 implementation | |
| with the same expert MoE augmentation. It uses `model_file` in config.json | |
| with `trust_remote_code=True` for loading. | |
| ### GGUF / llama.cpp | |
| fuse-1 Lite is available in GGUF format for llama.cpp. | |
| **GGUF repo**: [`Akahsizrr/fuse-1-Lite-GGUF`](https://huggingface.co/Akahsizrr/fuse-1-Lite-GGUF) | |
| > **Note:** The GGUF uses the custom `fuse3` architecture. Stock llama.cpp | |
| > cannot load it β you need a llama.cpp fork with Fuse3 support. The GGUF | |
| > repo includes the C++ graph builder (`src/models/fuse3.cpp`), Python | |
| > converter (`conversion/fuse3.py`), and integration guide (`INTEGRATION.md`). | |
| ```bash | |
| # Build llama.cpp with Fuse3 support (see INTEGRATION.md in the GGUF repo) | |
| ./llama-cli -m fuse-1-Lite-f16.gguf \ | |
| -p "Write a Python function to check if a number is prime." \ | |
| -n 512 --temp 0.1 | |
| ``` | |
| The C++ implementation reuses LFM2's attention and short-conv graph builders, | |
| adding the expert MoE block (router β top-k β SwiGLU experts β scale β add) | |
| after each augmented layer's dense FFN. | |
| ### Transformers (Universal) | |
| The recommended way to run fuse-1 Lite on any platform: | |
| ```bash | |
| pip install transformers torch bitsandbytes accelerate | |
| ``` | |
| ## Performance | |
| ### VRAM Requirements | |
| | Backend | Precision | VRAM/Memory | Recommended Hardware | | |
| |---------|-----------|-------------|---------------------| | |
| | Transformers | bfloat16 | ~12 GB | L4, A10G, RTX 4090 | | |
| | Transformers | 8-bit | 6.00 GB | T4, L4, RTX 3060 | | |
| | Transformers | 4-bit | 3.36 GB | T4, RTX 3060, M2 Pro | | |
| | vLLM | bfloat16 | ~12 GB | A10G, A100, H100 | | |
| | MLX | float16 | ~12 GB | M1 Pro+, M2, M3, M4 | | |
| | llama.cpp | F16 | ~11.4 GB | Any CPU/GPU | | |
| | llama.cpp | Q4_K_M | ~4 GB | Any CPU/GPU | | |
| ### Sample Outputs | |
| **Prompt**: "Write a Python function to check if a string is a palindrome." | |
| **Output** (excerpt): | |
| ```python | |
| def is_palindrome(s: str) -> bool: | |
| """ | |
| Return True if *s* reads the same forwards and backwards, | |
| ignoring case and non-alphanumeric characters. | |
| """ | |
| cleaned = ''.join(ch.lower() for ch in s if ch.isalnum()) | |
| return cleaned == cleaned[::-1] | |
| ``` | |
| The model produces complete implementations with docstrings, type hints, complexity analysis, and test cases. | |
| ## Limitations | |
| 1. **Custom architecture**: Requires `trust_remote_code=True` β the model includes custom `Fuse3ForCausalLM` code | |
| 2. **No vLLM support**: The custom MoE augmentation is not yet supported by vLLM's optimized inference engine | |
| 3. **No GGUF/MLX conversion**: The custom architecture cannot be directly converted to GGUF or MLX format | |
| 4. **Training data was small**: Only 55 examples were used for router training β the router may not generalize perfectly to all coding tasks | |
| 5. **Expert compatibility**: Qwen3.6 experts operate on LFM2's activation space with std normalization β some expert knowledge may be lost in translation | |
| 6. **`use_cache=False` during training**: Augmented layers don't propagate KV cache correctly during training; generation uses the standard cache | |
| 7. **bitsandbytes quantization**: 4-bit and 8-bit quantization work at runtime via `BitsAndBytesConfig` β pre-quantized saved versions are not available as separate repos | |
| ## Citation | |
| ```bibtex | |
| @misc{fuse1lite2026, | |
| title={fuse-1 Lite: Coding-Enhanced Mixture-of-Experts via Expert Transplantation}, | |
| author={Vasko Djack}, | |
| year={2026}, | |
| publisher={HuggingFace}, | |
| url={https://huggingface.co/Akahsizrr/fuse-1-Lite} | |
| } | |
| ``` | |
| ## Acknowledgments | |
| - **LiquidAI** for the LFM2.5-2.6B host model | |
| - **Qwen Team** for the Qwen3.6-35B-A3B donor model | |
| - **Modal** for compute infrastructure | |
| ## License | |
| Apache 2.0 β See [LICENSE](LICENSE) for details. | |
| --- | |
| *Built with [Devin](https://devin.ai) β Cognition AI* | |