--- license: other base_model: mistralai/Mistral-7B-v0.1 base_model_relation: quantized library_name: transformers pipeline_tag: text-generation tags: - fp8 - compressed-tensors - vllm - quantized quantized_by: liodon-ai --- # Mistral-7B-v0.1 — FP8 (dynamic) FP8 quantization of [mistralai/Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1), published by [Liodon AI](https://huggingface.co/liodon-ai). Quantized with [llm-compressor](https://github.com/vllm-project/llm-compressor) using the `FP8_DYNAMIC` scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this scheme, so the quantized weights are numerically just a direct cast of the original — no calibration-set bias to worry about. `lm_head` is left unquantized (standard practice — negligible size, disproportionate quality impact if quantized). Original size: 14.5 GB → Quantized: 7.5 GB. ## Quick Start **vLLM** ```bash vllm serve liodon-ai/Mistral-7B-v0.1-FP8 ``` **Text Generation Inference (TGI)** ```bash docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \ --model-id liodon-ai/Mistral-7B-v0.1-FP8 ``` **SGLang** ```bash python -m sglang.launch_server --model-path liodon-ai/Mistral-7B-v0.1-FP8 ``` FP8 execution requires an NVIDIA GPU with compute capability ≥ 8.9 (Ada/Hopper/Blackwell — RTX 40-series, L4/L40S, H100/H200, B100/B200/GB10). On older GPUs, vLLM/TGI will dequantize to run, which loses the speed/memory benefit. ## Source - **Model**: [mistralai/Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1) - **License**: other ## Citation ```bibtex @misc{liodonai_mistral_7b_v0_1_fp8, title = {Mistral-7B-v0.1 — FP8}, author = {{Liodon AI}}, year = {2026}, howpublished = {\url{https://huggingface.co/liodon-ai/Mistral-7B-v0.1-FP8}}, note = {FP8 (dynamic) quantization of mistralai/Mistral-7B-v0.1} } ``` --- *Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*