` CoT from the answer (see above). |
| `--kv-cache-dtype fp8` | optional, ~2Γ the KV capacity for even more parallel sessions. |
| `--tensor-parallel-size` | **leave at 1** β the model fits one GPU; sharding an 8B-A1B rarely pays. |
### Offline
```python
from vllm import LLM, SamplingParams
llm = LLM("sakamakismile/LFM2.5-8B-A1B-NVFP4", quantization="modelopt",
max_model_len=32768, gpu_memory_utilization=0.90, max_num_seqs=8)
tok = llm.get_tokenizer()
chat = tok.apply_chat_template(
[{"role": "user", "content": "ζ₯ζ¬θͺγ§θͺε·±η΄Ήδ»γγ¦γ"}],
tokenize=False, add_generation_prompt=True)
print(llm.generate([chat], SamplingParams(
temperature=0.2, top_k=80, repetition_penalty=1.05, max_tokens=512))[0].outputs[0].text)
```
### Container
A ready `Dockerfile` + `compose.yaml` + `entrypoint.sh` + `run.sh` are bundled β see
**`USAGE.md`** for `./run.sh up | test | bench | logs | down` and every env knob.
**Sampling (Liquid's recommendation):** `temperature=0.2`, `top_k=80`,
`repetition_penalty=1.05`. It thinks first, so give it `max_tokens β₯ 512`.
---
## β οΈ Usage notes & caveats
- **Needs Blackwell (SM120) + a recent vLLM** (β₯0.21 with NVFP4/modelopt) and
`flashinfer` β the FP4 GEMM and MoE run on FlashInfer-CUTLASS kernels.
- `ModuleNotFoundError: No module named 'trinity_turbo'` in the logs is **harmless**
(optional plugin auto-probe); the engine continues.
- If a MoE backend objects to the FP4 scales, force Marlin: `VLLM_USE_FLASHINFER_MOE_FP4=0`.
- Use `--quantization modelopt` only β **not** `fp8`/`awq`/`gptq`.
- A handful of rarely-routed ("cold") experts are calibrated from limited activation
coverage; for the overwhelming majority of tokens, output tracks the BF16 source
closely. As with the base model, heavy programming / knowledge-heavy QA without
retrieval isn't its strong suit.
- Straight quantization of the base **instruct** model β no refusal-reduction or other
behavioral changes.
## π¬ What's quantized
NVFP4 = e2m1 weights, 16-wide blocks, FP8-e4m3 block scales + FP32 global scale, static
per-tensor `input_scale`.
- **β NVFP4:** all 32 MoE experts (per layer) + the 2 dense MLP layers.
- **kept BF16:** attention (q/k/v/out), short-conv projections, the MoE router
(`feed_forward.gate`), token embeddings, and `lm_head`.
Full recipe + scripts: Lna-Lab `lnarizer/recipes/lfm2_moe/` (includes the one modelopt
calibration patch needed for `lfm2_moe`, and the expert key remap for vLLM).
## License
Inherits the base model's license (LFM Open License v1.0, `license_name: lfm1.0`) β see
the bundled `LICENSE`. Base model: `LiquidAI/LFM2.5-8B-A1B`.
## Credits
- Base model: **Liquid AI** β LFM2.5-8B-A1B.
- NVFP4 quantization & packaging: **Lna-Lab**.
- Tooling: NVIDIA TensorRT Model-Optimizer, vLLM, FlashInfer.
---
**π¬ Lna-Lab** Β· NVFP4 for Blackwell Β· *LLMs without colored glasses, in 4-bit, for the edge*
Quantized & verified on 7Γ RTX PRO 2000 Blackwell Β· 2026