File size: 12,807 Bytes
bb055dd
a36d000
 
 
 
 
 
 
 
 
 
 
 
 
 
f5c0d85
 
a36d000
bb055dd
a36d000
f5c0d85
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6b18be5
f5c0d85
 
 
 
 
a36d000
 
 
f5c0d85
 
 
 
 
 
 
 
 
 
 
 
 
 
a36d000
 
 
f5c0d85
a36d000
 
 
 
 
 
f5c0d85
a36d000
 
 
 
 
 
 
f5c0d85
a36d000
f5c0d85
 
 
 
 
 
a36d000
ed21fbb
a36d000
ed21fbb
 
 
a36d000
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f5c0d85
ffc80f4
f5c0d85
ffc80f4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f5c0d85
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ffc80f4
 
f5c0d85
 
ffc80f4
f5c0d85
 
 
ffc80f4
a36d000
 
f5c0d85
a36d000
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
---
license: other
license_name: fish-audio-research-license
license_link: https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md
base_model: fishaudio/s2-pro
pipeline_tag: text-to-speech
library_name: sglang
tags:
- text-to-speech
- fish-audio
- s2-pro
- sglang
- realtime
- rtx-5090
- no-weights
- local-serving
- neko-legends
inference: false
---

<div style="font-family: Inter, ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; border: 1px solid #2f2118; border-radius: 18px; overflow: hidden; background: #0b0b0f; box-shadow: 0 20px 48px rgba(0,0,0,0.26); margin: 0 0 28px 0;">
  <div style="padding: 30px 28px 24px 28px; background: radial-gradient(circle at 6% 0%, rgba(255,122,26,0.32), transparent 34%), radial-gradient(circle at 95% 12%, rgba(255,184,107,0.16), transparent 28%), linear-gradient(135deg, #050507 0%, #101014 56%, #201208 100%); border-bottom: 1px solid rgba(255,122,26,0.35);">
    <div style="display: flex; flex-wrap: wrap; gap: 14px; align-items: center; justify-content: space-between;">
      <div>
        <div style="font-size: 11px; font-weight: 900; color: #ffb86b; letter-spacing: 1.8px; text-transform: uppercase;">Neko Legends local voice profile</div>
        <h1 style="margin: 8px 0 0 0; color: #fff7ed; font-size: 30px; line-height: 1.12; font-weight: 950; border: 0;">Fish Audio S2-Pro Realtime Optimized for RTX 5090</h1>
      </div>
      <div style="background: rgba(255,122,26,0.14); border: 1px solid rgba(255,122,26,0.72); color: #ffd7ad; font-size: 12px; font-weight: 900; padding: 8px 12px; border-radius: 999px;">model-card-only</div>
    </div>
    <p style="margin: 14px 0 0 0; max-width: 900px; color: #d6d3d1; font-size: 14px; line-height: 1.7;">
      A no-weights release documenting a local RTX 5090 serving profile for <a href="https://huggingface.co/fishaudio/s2-pro" target="_blank" style="color:#ffb86b; text-decoration:none; font-weight:900;">Fish Audio S2-Pro</a>: Docker-backed SGLang Omni, CUDA graphs, cached reference audio, and a graph-safe decoder fallback for SM120.
    </p>
  </div>

  <div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(150px, 1fr)); gap: 1px; background: #2f2118;">
    <div style="background:#111116; padding: 15px 16px;"><span style="display:block; color:#a8a29e; font-size:11px; font-weight:900; text-transform:uppercase;">Base model</span><b style="display:block; margin-top:5px; color:#fff7ed; font-size:18px;">fishaudio/s2-pro</b></div>
    <div style="background:#111116; padding: 15px 16px;"><span style="display:block; color:#a8a29e; font-size:11px; font-weight:900; text-transform:uppercase;">Target GPU</span><b style="display:block; margin-top:5px; color:#ffb86b; font-size:18px;">RTX 5090</b></div>
    <div style="background:#111116; padding: 15px 16px;"><span style="display:block; color:#a8a29e; font-size:11px; font-weight:900; text-transform:uppercase;">Server</span><b style="display:block; margin-top:5px; color:#fff7ed; font-size:18px;">SGLang Omni</b></div>
    <div style="background:#111116; padding: 15px 16px;"><span style="display:block; color:#a8a29e; font-size:11px; font-weight:900; text-transform:uppercase;">Weights</span><b style="display:block; margin-top:5px; color:#ffb86b; font-size:18px;">not hosted</b></div>
    <div style="background:#111116; padding: 15px 16px;"><span style="display:block; color:#a8a29e; font-size:11px; font-weight:900; text-transform:uppercase;">Goal</span><b style="display:block; margin-top:5px; color:#fff7ed; font-size:18px;">realtime TTS</b></div>
    <div style="background:#111116; padding: 15px 16px;"><span style="display:block; color:#a8a29e; font-size:11px; font-weight:900; text-transform:uppercase;">First audio</span><b style="display:block; margin-top:5px; color:#ffb86b; font-size:18px;">0.36s</b></div>
  </div>
</div>

> [!IMPORTANT]
> This repository does **not** redistribute Fish Audio S2-Pro weights. Users must accept the upstream Fish Audio license and download the official checkpoint from [`fishaudio/s2-pro`](https://huggingface.co/fishaudio/s2-pro).

## What This Is

<div style="font-family: Inter, ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; display: grid; grid-template-columns: repeat(auto-fit, minmax(245px, 1fr)); gap: 14px; margin: 18px 0 26px 0;">
  <div style="border:1px solid #3a2a1f; background:#111116; border-radius:14px; padding:16px;">
    <div style="color:#ffb86b; font-size:12px; font-weight:950; letter-spacing:0.8px; text-transform:uppercase;">Runtime profile</div>
    <p style="margin:8px 0 0 0; color:#e7e5e4; font-size:13px; line-height:1.65;">A local serving recipe for Fish Audio S2-Pro using Docker-backed SGLang Omni on a 32 GB RTX 5090.</p>
  </div>
  <div style="border:1px solid #3a2a1f; background:#111116; border-radius:14px; padding:16px;">
    <div style="color:#ffb86b; font-size:12px; font-weight:950; letter-spacing:0.8px; text-transform:uppercase;">No model fork</div>
    <p style="margin:8px 0 0 0; color:#e7e5e4; font-size:13px; line-height:1.65;">No fine-tuning, quantization, checkpoint conversion, or rehosted upstream model files are claimed here.</p>
  </div>
  <div style="border:1px solid #3a2a1f; background:#111116; border-radius:14px; padding:16px;">
    <div style="color:#ffb86b; font-size:12px; font-weight:950; letter-spacing:0.8px; text-transform:uppercase;">Live setup</div>
    <p style="margin:8px 0 0 0; color:#e7e5e4; font-size:13px; line-height:1.65;">Designed for single-user realtime local TTS. A separate GPU can handle realtime STT/ASR in a split workstation setup.</p>
  </div>
</div>

## Why There Are No Weights Here

The local checkpoint files match the official upstream `fishaudio/s2-pro` release. To avoid duplicating gated model files or confusing the license boundary, this repository does not upload:

- `model-00001-of-00002.safetensors`
- `model-00002-of-00002.safetensors`
- `codec.pth`
- tokenizer/config files from the upstream model

Download the official files from:

```text
https://huggingface.co/fishaudio/s2-pro
```

## Local Benchmark Summary

Measured locally on an RTX 5090. Speedup is the headline; raw before/after values are included for reproducibility.

<div style="font-family: Inter, ui-sans-serif, system-ui, -apple-system, BlinkMacSystemFont, 'Segoe UI', sans-serif; border: 1px solid #3a2a1f; background: #0f0f13; border-radius: 16px; overflow: hidden; margin: 18px 0 24px 0;">
  <div style="padding: 14px 18px; background: linear-gradient(90deg, #1a120d 0%, #2b1708 100%); border-bottom: 1px solid rgba(255,122,26,0.35); color: #ffd7ad; font-weight: 950;">RTX 5090 before/after serving benchmark</div>
  <a href="./assets/rtx5090-benchmark-bars.svg" target="_blank" style="display:block; background:#050507;">
    <img src="./assets/rtx5090-benchmark-bars.svg" alt="RTX 5090 benchmark bar chart for Fish Audio S2-Pro realtime serving" style="display:block; width:100%; border:0;" />
  </a>
</div>

| Metric | Before: Python Fish server | After: SGLang Omni + cached reference | Speedup |
|---|---:|---:|---:|
| First audio | 25.1s | 0.36s | 69.9x faster, about 6,890% faster |
| Total request time | 25.1s | 2.10s | 12.0x faster, about 1,100% faster |
| Estimated RTF | 5.51 | 0.48 | 11.5x faster, about 1,050% faster |

Benchmark notes:

- Baseline was the local Python Fish server path.
- Optimized path was the Docker-backed SGLang Omni path with a warm server and cached reference voice.
- Warm cached SGLang generation was below realtime for the measured short live TTS sample.
- The chart is a local serving benchmark, not an upstream Fish Audio benchmark claim.

## Optimization Profile

The measured realtime profile used:

- SGLang Omni instead of the eager Python Fish server path.
- The Docker image's pinned Torch/SGLang/FlashInfer stack.
- SGLang CUDA graph replay enabled.
- RTX 5090 / SM120-safe Fish audio-decoder path by disabling the incompatible `sgl-kernel` KV-cache flash-attention path.
- Graph-safe fixed-cache SDPA fallback for the Fish audio decoder.
- `flashinfer` text attention backend.
- Single-user live memory profile:
  - `mem_fraction_static=0.50`
  - `chunked_prefill_size=2048`
  - `max_running_requests=4`
- Preloaded/cached reference VQ codes for repeated voice-reference requests.
- Docker model/runtime volumes to avoid repeated slow checkpoint reads through Windows `/mnt/d` bind mounts.

Measured local VRAM after the tuned live restart was about 24.6 GB on the RTX 5090, down from an earlier near-full 32.2 GB SGLang container reservation.

## Recommended RTX 5090 Environment

```bash
FISH_S2PRO_DISABLE_FLASH_ATTN_KVCACHE=1
FISH_S2PRO_KVCACHE_FALLBACK_BACKEND=sdpa
FISH_S2PRO_SGLANG_ATTENTION_BACKEND=flashinfer
FISH_S2PRO_SGLANG_DISABLE_CUDA_GRAPH=0
FISH_S2PRO_SGLANG_MEM_FRACTION_STATIC=0.50
FISH_S2PRO_SGLANG_CHUNKED_PREFILL_SIZE=2048
FISH_S2PRO_SGLANG_MAX_RUNNING_REQUESTS=4
FISH_S2PRO_PRELOAD_REFERENCE_PATHS=/path/to/reference.wav
```

Recommended serve shape:

```bash
sgl-omni serve \
  --model-path /path/to/s2-pro \
  --config examples/configs/s2pro_tts.yaml \
  --host 0.0.0.0 \
  --port 8000
```

## For AI Agents

If a user asks you to apply the RTX 5090 realtime optimizations from this model card, treat this repository as documentation for a local runtime setup. Do not edit, quantize, rehost, or replace the Fish Audio S2-Pro weights.

Agent checklist:

1. Confirm the user has accepted the upstream Fish Audio license and has the official `fishaudio/s2-pro` checkpoint locally.
2. If weights are missing, instruct the user to download them from `https://huggingface.co/fishaudio/s2-pro`; do not download weights from this repository.
3. Locate or clone a writable `sgl-project/sglang-omni` checkout.
4. Preserve the serving image's pinned Torch/SGLang/FlashInfer stack. If you need an editable local checkout inside Docker, install it without replacing the pinned dependency stack.
5. Apply the runtime optimizations to the local SGLang Omni FishAudio S2-Pro implementation, not to the model checkpoint.

Patch targets and intent:

- In `sglang_omni/models/fishaudio_s2_pro/fish_speech/models/text2semantic/modeling.py`: add `FISH_S2PRO_DISABLE_FLASH_ATTN_KVCACHE` and `FISH_S2PRO_KVCACHE_FALLBACK_BACKEND`.
- In that same file, when `FISH_S2PRO_DISABLE_FLASH_ATTN_KVCACHE=1`, bypass the `sgl-kernel` / `flash_attn_with_kvcache` path for the Fish audio decoder and use a graph-safe fixed-cache SDPA fallback.
- The fallback should update the K/V cache at the current `cache_seqlens`, repeat KV heads to match query heads, mask positions beyond `cache_seqlens`, call `torch.nn.functional.scaled_dot_product_attention`, and return the normal attention output shape.
- Add non-persistent `batch_indices` and `key_positions` buffers to the Fish decoder `KVCache` helper so the fallback is CUDA-graph friendly.
- In `sglang_omni/models/fishaudio_s2_pro/pipeline/stages.py`: make SGLang `ServerArgs` configurable from env vars: `FISH_S2PRO_SGLANG_ATTENTION_BACKEND`, `FISH_S2PRO_SGLANG_DISABLE_CUDA_GRAPH`, `FISH_S2PRO_SGLANG_MEM_FRACTION_STATIC`, `FISH_S2PRO_SGLANG_CHUNKED_PREFILL_SIZE`, and `FISH_S2PRO_SGLANG_MAX_RUNNING_REQUESTS`.
- In that same preprocessing stage, add an in-process reference VQ cache keyed by reference audio path and local file signature. Preload any paths from `FISH_S2PRO_PRELOAD_REFERENCE_PATHS`, and use cached VQ codes for repeated `audio_path` references.

Docker/runtime notes for agents:

- Use NVIDIA GPU passthrough and select the RTX 5090 explicitly when multiple GPUs are visible.
- Prefer Docker volumes or Linux-native storage for the model and runtime cache; repeated checkpoint reads through a Windows `/mnt/d` bind mount can add many minutes to cold startup.
- Keep CUDA graphs enabled unless debugging compatibility.
- Do not switch to `trtllm_mha` on RTX 5090 / SM120 unless upstream support has changed; local tests rejected that backend for this GPU.
- Benchmark after every change with the same prompt and reference. Report first-audio latency, total request time, estimated audio duration, RTF, VRAM, and whether the run was cold, first-after-health, or warm cached.
- Expected warm target on the measured machine: first audio under 1 second and RTF below 1.0 for the short live TTS benchmark.

## License

The base model is governed by the [Fish Audio Research License](https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md). Research and non-commercial use are permitted by Fish Audio under that license. Commercial use requires a separate license from Fish Audio.

This repository does not grant additional rights to the Fish Audio model weights.

## Attribution

Built with Fish Audio S2-Pro. Fish Audio S2-Pro is developed by Fish Audio / 39 AI, INC.

Upstream model:

```text
fishaudio/s2-pro
```