# USAGE — LFM2.5-8B-A1B-NVFP4 (Lna-Lab ops cheat-sheet) NVFP4 (modelopt W4A4) build of LFM2.5-8B-A1B. Fits **one 16 GB Blackwell card**. Tested: vLLM 0.21, CUDA 12.8/13.x, SM120. ~117 tok/s single, ~326 tok/s @4 concurrent. --- ## 0. TL;DR (single GPU, no container) ```bash CUDA_VISIBLE_DEVICES=0 vllm serve /path/to/LFM2.5-8B-A1B-NVFP4 \ --served-model-name lfm25-8b-a1b \ --quantization modelopt \ --kv-cache-dtype fp8 \ --max-model-len 128000 \ --max-num-seqs 16 \ --gpu-memory-utilization 0.90 \ --reasoning-parser deepseek_r1 \ --port 8000 ``` Smoke test: ```bash curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{ "model":"lfm25-8b-a1b", "messages":[{"role":"user","content":"東京の名所を3つ。"}], "temperature":0.2,"top_k":80,"repetition_penalty":1.05,"max_tokens":512}' | jq . ``` --- ## 1. The flags, and why | flag | value | why | |---|---|---| | `--quantization modelopt` | **required** | tells vLLM the checkpoint is NVFP4 (`hf_quant_config.json`). Without it vLLM treats weights as raw uint8 → garbage. | | `--max-model-len` | `32768` (up to `131072` native) | context window. KV is cheap here (only 6 of 24 layers carry KV), so you can go long or run many sessions. | | `--max-num-seqs` | `16` | concurrent sequences. Measured on a 16 GB card: with `--kv-cache-dtype fp8` the KV pool is **1.38 M tokens** → **~10 sessions each at a FULL 128 K context** (bf16 KV → ~5). 16 leaves headroom for shorter sessions via paged KV; pin to `10` if you want every slot guaranteed at full 128 K. | | `--gpu-memory-utilization` | `0.90` | fraction of VRAM vLLM may use. 0.90 is safe on a clean 16 GB card; lower if the card is shared. | | `--reasoning-parser deepseek_r1` | recommended | LFM2.5 is a **reasoning model** emitting `…` then the answer. This parser puts the CoT in `reasoning_content` and the answer in `content`. Omit it to get the raw text (think tags included). | | `--enable-auto-tool-choice --tool-call-parser pythonic` | optional (agentic) | LFM2.5 emits **Pythonic** tool calls between `<\|tool_call_start\|>` / `<\|tool_call_end\|>`. Pass tools via `apply_chat_template(tools=[...])`. Use these flags if your vLLM Pythonic parser handles the special-token wrapper; else parse the block yourself. | | `--kv-cache-dtype fp8` | optional | ~doubles KV capacity (→ ~1.3 M tokens) for even more concurrency; tiny quality cost. Default (omit) keeps bf16 KV. | | `--tensor-parallel-size` | `1` | **this model fits one GPU — keep TP=1.** Only raise if you deliberately shard (rarely worth it for an 8B-A1B). | | `--dtype` | `auto` | leave as is (bf16 compute around the FP4 GEMMs). | **Sampling (Liquid's recommendation):** `temperature=0.2`, `top_k=80`, `repetition_penalty=1.05`. It's a CoT model, so give it room: `max_tokens≥512`. --- ## 2. Offline (no server) ```python from vllm import LLM, SamplingParams llm = LLM("/path/to/LFM2.5-8B-A1B-NVFP4", quantization="modelopt", max_model_len=32768, gpu_memory_utilization=0.90, max_num_seqs=8) tok = llm.get_tokenizer() chat = tok.apply_chat_template([{"role":"user","content":"Reverse a linked list in Python."}], tokenize=False, add_generation_prompt=True) sp = SamplingParams(temperature=0.2, top_k=80, repetition_penalty=1.05, max_tokens=512) print(llm.generate([chat], sp)[0].outputs[0].text) ``` --- ## 3. Containers (docker compose) Bundled here: `Dockerfile`, `compose.yaml`, `entrypoint.sh`, `run.sh`. ```bash # edit compose.yaml -> volumes: point /model at this folder (or it defaults to the repo dir) ./run.sh up # build + start detached on GPU 0, port 8000 ./run.sh test # poll /v1/models until ready ./run.sh bench # one-shot chat completion ./run.sh logs # tail ./run.sh down # stop ``` All serve flags are env-overridable (see `entrypoint.sh`): `PORT`, `MAX_MODEL_LEN`, `MAX_NUM_SEQS`, `GPU_MEM_UTIL`, `KV_CACHE_DTYPE`, `REASONING_PARSER`, `TP_SIZE`, `CUDA_VISIBLE_DEVICES`. Example: run on GPU 3 with fp8 KV and 64 K context: ```bash CUDA_VISIBLE_DEVICES=3 MAX_MODEL_LEN=65536 KV_CACHE_DTYPE=fp8 MAX_NUM_SEQS=16 ./run.sh up ``` The image installs vLLM on a CUDA base; `runtime: nvidia` + `NVIDIA_VISIBLE_DEVICES` pass the GPU through. One GPU is enough — `compose.yaml` reserves `count: 1`. --- ## 4. Requirements / gotchas - **Blackwell (SM120) + recent vLLM** (≥0.21 with NVFP4 / modelopt support). The FP4 GEMM + MoE run on FlashInfer-CUTLASS kernels; needs `flashinfer`. - `ModuleNotFoundError: No module named 'trinity_turbo'` in the logs is **harmless** — it's an optional vLLM plugin auto-probe; the engine continues. - If a MoE backend complains about FP4 scales, force the Marlin path: `VLLM_USE_FLASHINFER_MOE_FP4=0`. - This is a straight NVFP4 quant of the base instruct model — no behavioral changes. - Don't pass `--quantization fp8`/`awq`/etc. — it's `modelopt` (NVFP4) only. --- ## 5. What's inside - Quantized → NVFP4: all 32 MoE experts + the 2 dense MLP layers. - Kept BF16: attention (q/k/v/out), short-conv, MoE router (`feed_forward.gate`), embeddings, `lm_head`. - Recipe + scripts: Lna-Lab `lnarizer/recipes/lfm2_moe/`.