neko-legends commited on
Commit
a36d000
·
1 Parent(s): bb055dd

Update model card for 5090 optimized Fish S2

Browse files
Files changed (3) hide show
  1. NOTICE.md +14 -0
  2. README.md +112 -1
  3. assets/rtx5090-benchmark-bars.svg +111 -0
NOTICE.md ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ This repository documents an RTX 5090 local serving profile for Fish Audio S2-Pro.
2
+
3
+ It does not redistribute Fish Audio model weights. Users must obtain the base model
4
+ from the upstream Hugging Face repository:
5
+
6
+ https://huggingface.co/fishaudio/s2-pro
7
+
8
+ Fish Audio attribution:
9
+
10
+ This model is licensed under the Fish Audio Research License, Copyright (c) 39 AI,
11
+ INC. All Rights Reserved.
12
+
13
+ Commercial use of Fish Audio S2-Pro or derivative works requires a separate license
14
+ from Fish Audio.
README.md CHANGED
@@ -1,3 +1,114 @@
1
  ---
2
- license: mit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: other
3
+ license_name: fish-audio-research-license
4
+ license_link: https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md
5
+ base_model: fishaudio/s2-pro
6
+ pipeline_tag: text-to-speech
7
+ library_name: sglang
8
+ tags:
9
+ - text-to-speech
10
+ - fish-audio
11
+ - s2-pro
12
+ - sglang
13
+ - realtime
14
+ - rtx-5090
15
+ - no-weights
16
+ inference: false
17
  ---
18
+
19
+ # Fish Audio S2-Pro Realtime Optimized for RTX 5090
20
+
21
+ This repository is a model-card-only release for an RTX 5090 local serving profile for
22
+ [Fish Audio S2-Pro](https://huggingface.co/fishaudio/s2-pro).
23
+
24
+ It does **not** redistribute Fish Audio S2-Pro weights. The underlying model files are
25
+ the official upstream `fishaudio/s2-pro` checkpoint files, so users should accept the
26
+ upstream license and download the weights directly from Fish Audio.
27
+
28
+ ## What This Is
29
+
30
+ This release documents a local realtime optimization profile:
31
+
32
+ - Base model: `fishaudio/s2-pro`
33
+ - Target GPU: NVIDIA GeForce RTX 5090, 32 GB VRAM
34
+ - Serving path: Docker-backed SGLang Omni
35
+ - Intended use: single-user realtime local TTS
36
+ - Example workstation split: RTX 5090 handles TTS; a separate RTX 3090 can handle realtime STT/ASR
37
+
38
+ No fine-tuning, quantization, or checkpoint conversion is claimed here. The optimization
39
+ work is in the serving/runtime profile around the official S2-Pro weights.
40
+
41
+ ## Why There Are No Weights Here
42
+
43
+ The local checkpoint files match the official upstream `fishaudio/s2-pro` release. To
44
+ avoid duplicating gated model files or confusing the license boundary, this repository
45
+ does not upload:
46
+
47
+ - `model-00001-of-00002.safetensors`
48
+ - `model-00002-of-00002.safetensors`
49
+ - `codec.pth`
50
+ - tokenizer/config files from the upstream model
51
+
52
+ Download those files from:
53
+
54
+ ```text
55
+ https://huggingface.co/fishaudio/s2-pro
56
+ ```
57
+
58
+ ## Local Benchmark Summary
59
+
60
+ Measured locally on an RTX 5090. Lower is better.
61
+
62
+ ![RTX 5090 benchmark bar chart](assets/rtx5090-benchmark-bars.svg)
63
+
64
+ | Metric | Before: Python Fish server | After: SGLang Omni + cached reference | Change |
65
+ |---|---:|---:|---:|
66
+ | First audio | 25.1s | 0.36s | 98.6% lower latency |
67
+ | Total request time | 25.1s | 2.10s | 91.6% lower latency |
68
+ | Estimated RTF | 5.51 | 0.48 | 91.3% lower RTF |
69
+
70
+ Benchmark notes:
71
+
72
+ - Baseline was the local Python Fish server path.
73
+ - Optimized path was the Docker-backed SGLang Omni path with a warm server and cached reference voice.
74
+ - Warm cached SGLang generation was below realtime for the measured short live TTS sample.
75
+ - The chart is a local serving benchmark, not an upstream Fish Audio benchmark claim.
76
+
77
+ ## Optimization Profile
78
+
79
+ The measured realtime profile used:
80
+
81
+ - SGLang Omni instead of the eager Python Fish server path.
82
+ - The Docker image's pinned Torch/SGLang/FlashInfer stack.
83
+ - SGLang CUDA graph replay enabled.
84
+ - RTX 5090 / SM120-safe Fish audio-decoder path by disabling the incompatible `sgl-kernel` KV-cache flash-attention path.
85
+ - Graph-safe fixed-cache SDPA fallback for the Fish audio decoder.
86
+ - `flashinfer` text attention backend.
87
+ - Single-user live memory profile:
88
+ - `mem_fraction_static=0.50`
89
+ - `chunked_prefill_size=2048`
90
+ - `max_running_requests=4`
91
+ - Preloaded/cached reference VQ codes for repeated voice-reference requests.
92
+ - Docker model/runtime volumes to avoid repeated slow checkpoint reads through Windows `/mnt/d` bind mounts.
93
+
94
+ Measured local VRAM after the tuned live restart was about 24.6 GB on the RTX 5090,
95
+ down from an earlier near-full 32.2 GB SGLang container reservation.
96
+
97
+ ## License
98
+
99
+ The base model is governed by the
100
+ [Fish Audio Research License](https://huggingface.co/fishaudio/s2-pro/blob/main/LICENSE.md).
101
+ Research and non-commercial use are permitted by Fish Audio under that license.
102
+ Commercial use requires a separate license from Fish Audio.
103
+
104
+ This repository does not grant additional rights to the Fish Audio model weights.
105
+
106
+ ## Attribution
107
+
108
+ Built with Fish Audio S2-Pro. Fish Audio S2-Pro is developed by Fish Audio / 39 AI, INC.
109
+
110
+ Upstream model:
111
+
112
+ ```text
113
+ fishaudio/s2-pro
114
+ ```
assets/rtx5090-benchmark-bars.svg ADDED