--- license: other license_name: fish-audio-research-license license_link: https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE base_model: fishaudio/s2-pro pipeline_tag: text-to-speech library_name: transformers tags: - text-to-speech - tts - voice-cloning - bitsandbytes - nf4 - 4-bit - quantized language: - en - zh - ja - ko - id - de - fr - ar --- # Fish Speech S2-Pro — NF4 (bitsandbytes 4-bit) A **bitsandbytes NF4 4-bit** quantization of [`fishaudio/s2-pro`](https://huggingface.co/fishaudio/s2-pro), the 4B-parameter flagship Fish Audio S2 text-to-speech model. Roughly **4.9 GB** on disk against 8.5 GB of bf16 safetensors, and it loads directly as 4-bit — no quantization step at startup. Everything the base model does is preserved: voice cloning from a short reference clip, ~50 languages, and inline expressiveness tags such as `[laugh]`, `[whispers]`, `[sighs]`. ## Intended use This checkpoint is built for [scarxity/fish-speech-int4-patch](https://github.com/scarxity/fish-speech-int4-patch), a fork whose `server-4gb` / `webui-4gb` Docker profiles run S2-Pro on a **4 GB** laptop GPU (measured 3294 MiB peak on a simulated 4096 MiB budget, faster than realtime on an RTX 4060 Ti). > **The 4 GB figure comes from the fork's serving code, not from these weights > alone.** The quantization is what fits ~4.9 GB on disk; getting into 4 GB of > VRAM additionally needs the fork's decode-only codec, host-resident embedding > table, sliced `lm_head`, and chunked codec decode. Stock upstream code with > this checkpoint will need more VRAM than that. ```bash git clone https://github.com/scarxity/fish-speech-int4-patch cd fish-speech-int4-patch huggingface-cli download scarxity/fish-speech-s2-pro-nf4 --local-dir checkpoints/s2-pro-nf4 docker compose --profile server-4gb build docker compose --profile server-4gb up -d ``` Then POST to `http://localhost:8880/v1/tts`: ```json { "text": "Hello, world.", "reference_id": "my-voice" } ``` An OpenAI-compatible `/v1/audio/speech` endpoint and a Gradio WebUI (port 7860) are included. See the fork's [4 GB guide](https://github.com/scarxity/fish-speech-int4-patch/blob/main/docs/en/4gb-laptop.md) for reference-voice setup and long-text handling. ## Requirements - NVIDIA GPU (bitsandbytes NF4 is CUDA-only) - `bitsandbytes >= 0.43` - Loading flags `--bnb4 --half` This is **NF4**, not GPTQ, AWQ, or int4 — those formats are not interchangeable, and loaders for them will not read this checkpoint. ## Quantization details | | | |---|---| | Scheme | bitsandbytes NF4 (normal-float 4-bit) | | Double quantization | enabled (`compress_statistics=True`) | | Compute dtype | fp16 | | Quantized | the 36-layer slow backbone and the 4-layer depth transformer | | Effective | ~4.13 bits/weight | ## Architecture S2-Pro is a Dual-AR model over an RVQ audio codec (10 codebooks, 21.53 Hz): - **Slow AR** (36 layers, dim 2560) predicts the semantic codebook along time — words, pronunciation, language, and the inline emotion tags. - **Fast AR** (4 layers, dim 2560) runs nine times per frame for the residual codebooks — timbre and fine acoustic texture. - **Codec** upsamples the codes ×2048 to a 44.1 kHz waveform. ## Notes - `tokenizer_config.json` in this repo declares `PreTrainedTokenizerFast`. Some S2-Pro copies ship `TokenizersBackend`, a transformers v5 name that older loaders reject with `UnboundLocalError: tokenizer`. - Voice cloning is **in-context**: the reference clip's codes are prepended to every request rather than fine-tuned in. Keep reference clips near 10 seconds — a 31.7 s clip costs 683 prompt tokens on every request against 220 for a 10.2 s one. ## License Released under the **Fish Audio Research License**, inherited from the base model. Review it before any commercial or redistributive use. ## Credits - [Fish Audio](https://fish.audio/) — the S2 model, all core research and architecture - [groxaxo/fish-speech-int4-patch](https://github.com/groxaxo/fish-speech-int4-patch) — the original NF4 quantization path - [Technical report](https://arxiv.org/abs/2603.08823)