Mistral-7B-v0.1-FP8 / README.md
PY-AI-Dev's picture
Add FP8 (dynamic) quantization for Mistral-7B-v0.1
7356f46 verified
|
Raw
History Blame Contribute Delete
2.1 kB
---
license: other
base_model: mistralai/Mistral-7B-v0.1
base_model_relation: quantized
library_name: transformers
pipeline_tag: text-generation
tags:
- fp8
- compressed-tensors
- vllm
- quantized
quantized_by: liodon-ai
---
# Mistral-7B-v0.1 β€” FP8 (dynamic)
FP8 quantization of [mistralai/Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1), published by [Liodon AI](https://huggingface.co/liodon-ai).
Quantized with [llm-compressor](https://github.com/vllm-project/llm-compressor) using the
`FP8_DYNAMIC` scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are
quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this
scheme, so the quantized weights are numerically just a direct cast of the original β€” no calibration-set
bias to worry about. `lm_head` is left unquantized (standard practice β€” negligible size, disproportionate
quality impact if quantized).
Original size: 14.5 GB β†’ Quantized: 7.5 GB.
## Quick Start
**vLLM**
```bash
vllm serve liodon-ai/Mistral-7B-v0.1-FP8
```
**Text Generation Inference (TGI)**
```bash
docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \
--model-id liodon-ai/Mistral-7B-v0.1-FP8
```
**SGLang**
```bash
python -m sglang.launch_server --model-path liodon-ai/Mistral-7B-v0.1-FP8
```
FP8 execution requires an NVIDIA GPU with compute capability β‰₯ 8.9 (Ada/Hopper/Blackwell β€” RTX 40-series,
L4/L40S, H100/H200, B100/B200/GB10). On older GPUs, vLLM/TGI will dequantize to run, which loses the
speed/memory benefit.
## Source
- **Model**: [mistralai/Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1)
- **License**: other
## Citation
```bibtex
@misc{liodonai_mistral_7b_v0_1_fp8,
title = {Mistral-7B-v0.1 β€” FP8},
author = {{Liodon AI}},
year = {2026},
howpublished = {\url{https://huggingface.co/liodon-ai/Mistral-7B-v0.1-FP8}},
note = {FP8 (dynamic) quantization of mistralai/Mistral-7B-v0.1}
}
```
---
*Quantized by [Liodon AI](https://huggingface.co/liodon-ai)*