Instructions to use AlicanKiraz0/Kizagan-TTS-v1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VoxCPM
How to use AlicanKiraz0/Kizagan-TTS-v1.0 with VoxCPM:
import soundfile as sf from voxcpm import VoxCPM model = VoxCPM.from_pretrained("AlicanKiraz0/Kizagan-TTS-v1.0") wav = model.generate( text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.", prompt_wav_path=None, # optional: path to a prompt speech for voice cloning prompt_text=None, # optional: reference text cfg_value=2.0, # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse inference_timesteps=10, # LocDiT inference timesteps, higher for better result, lower for fast speed normalize=True, # enable external TN tool denoise=True, # enable external Denoise tool retry_badcase=True, # enable retrying mode for some bad cases (unstoppable) retry_badcase_max_times=3, # maximum retrying times retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech ) sf.write("output.wav", wav, 16000) print("saved: output.wav") - Notebooks
- Google Colab
- Kaggle
Document voice cloning and advanced VoxCPM2 inference
Browse files- README.md +225 -0
- SHA256SUMS +2 -2
- release_manifest.json +10 -3
README.md
CHANGED
|
@@ -79,6 +79,231 @@ Each command writes `audio.wav`, individual `sentence_*.wav` files and `metrics.
|
|
| 79 |
|
| 80 |
`--text` is one call: it does not automatically split a paragraph. For long answers, use a sentences file. The script does not automatically resolve Turkish abbreviations, numbers or ambiguous sentence boundaries. Spell out numbers as they should be spoken and keep each line a natural utterance.
|
| 81 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
## Measured inference examples
|
| 83 |
|
| 84 |
Measurements below come from the fixed **5 September 2026** development run on an **NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB**, CUDA BF16, CFG 2.0, 16 diffusion steps, seed 42 and the same reference clip. They are observations from two texts, not a throughput SLA.
|
|
|
|
| 79 |
|
| 80 |
`--text` is one call: it does not automatically split a paragraph. For long answers, use a sentences file. The script does not automatically resolve Turkish abbreviations, numbers or ambiguous sentence boundaries. Spell out numbers as they should be spoken and keep each line a natural utterance.
|
| 81 |
|
| 82 |
+
## Voice cloning guide
|
| 83 |
+
|
| 84 |
+
**Yes: Kahya accepts a voice reference and generates new text conditioned on that recording.** This happens during inference; supplying a new recording does not train another adapter or change the model weights. A transcript is optional: the packaged `inference.py` uses reference audio alone.
|
| 85 |
+
|
| 86 |
+
Kahya was fine-tuned on a Turkish single-speaker corpus. Its recorded listening evaluation used that speaker, so it does **not** establish cloning similarity for unfamiliar speakers, accents or languages. The API can accept another speaker's recording, but speaker similarity, style adherence and intelligibility need listening checks with that recording.
|
| 87 |
+
|
| 88 |
+
| Capability | How to use it | Evidence for this release |
|
| 89 |
+
|---|---|---|
|
| 90 |
+
| Reference-only voice cloning | `inference.py --reference` or Python `reference_wav_path` | Packaged command tested; limited same-speaker listening results below |
|
| 91 |
+
| Long text with the same reference | `--sentences-file`, one utterance per line | Preferred in the two recorded long-text listening trials |
|
| 92 |
+
| Reference plus its transcript | Python `reference_wav_path` + `prompt_wav_path` + `prompt_text` | Inherited VoxCPM2 API; separate functional examples below |
|
| 93 |
+
| Emotion/pace guidance with a reference | A parenthesized instruction at the start of Python `text` | Inherited control format; style accuracy not evaluated |
|
| 94 |
+
| Voice design without a reference | A voice description at the start of Python `text` | Inherited control format; designed-voice quality not evaluated |
|
| 95 |
+
| TTS without a reference | Python `generate(text=...)` | Inherited API; default-voice quality not evaluated |
|
| 96 |
+
| Stream audio chunks | Python `generate_streaming(...)` | Python yields chunks; packaged CLI saves completed WAV files |
|
| 97 |
+
|
| 98 |
+
### 1. Prepare a reference recording
|
| 99 |
+
|
| 100 |
+
Use your own voice or a recording you have permission to use. Choose one speaker, a quiet room, clear pronunciation and a natural speaking style. Avoid background music, overlapping speakers, clipped peaks and long stretches of silence. Cut at word boundaries and keep the beginning and end of the spoken phrase intact.
|
| 101 |
+
|
| 102 |
+
Start with a short, representative utterance. Our measured reference was **6.175 seconds**; this is an evaluated example length, not a minimum or maximum accepted by the API. A longer recording is not automatically a better reference and consumes more model context. The style in the recording can influence the result as well as the voice identity.
|
| 103 |
+
|
| 104 |
+
The packaged CLI requires **non-empty, finite, mono audio**. A mono WAV is the simplest input. VoxCPM2 internally resamples the input to **16 kHz**; the output is **48 kHz**. The measured 24 kHz reference therefore did not require manual conversion to 48 kHz.
|
| 105 |
+
|
| 106 |
+
If FFmpeg is installed, this optional conversion turns a phone recording into a mono 16 kHz PCM WAV:
|
| 107 |
+
|
| 108 |
+
```bash
|
| 109 |
+
ffmpeg -i ./my-recording.m4a -ac 1 -ar 16000 -c:a pcm_s16le ./speaker.wav
|
| 110 |
+
```
|
| 111 |
+
|
| 112 |
+
Conversion does not remove noise or repair clipping. Listen to the converted file before using it. Neither the supplied CLI nor the Python examples below enable denoising.
|
| 113 |
+
|
| 114 |
+
### 2. Clone the voice without a transcript
|
| 115 |
+
|
| 116 |
+
Run this after the installation and model download above:
|
| 117 |
+
|
| 118 |
+
```bash
|
| 119 |
+
python ./Kahya-TTS-v1.0/inference.py \
|
| 120 |
+
--model ./Kahya-TTS-v1.0 \
|
| 121 |
+
--reference ./speaker.wav \
|
| 122 |
+
--text "Merhaba, yeni bir metni referans sesle okuyorum." \
|
| 123 |
+
--output-dir ./kahya-voice-clone
|
| 124 |
+
```
|
| 125 |
+
|
| 126 |
+
Listen to `./kahya-voice-clone/audio.wav`. The reference may contain different words from `--text`: it supplies voice conditioning, while `--text` supplies the new words to synthesize. No `speaker.txt`, speaker ID, speaker embedding export or additional fine-tuning is required for this mode.
|
| 127 |
+
|
| 128 |
+
To try another voice, pass its recording to `--reference` and choose a new output directory. Compare the outputs using the same short text before moving to a longer passage. This is voice-conditioned TTS, not a speech-to-speech converter: it generates speech from the supplied text rather than replacing the voice in an existing recording.
|
| 129 |
+
|
| 130 |
+
### 3. Keep the reference consistent across a long text
|
| 131 |
+
|
| 132 |
+
Create a UTF-8 file with one complete utterance on each line:
|
| 133 |
+
|
| 134 |
+
```bash
|
| 135 |
+
cat > ./my-sentences.txt <<'TXT'
|
| 136 |
+
Merhaba, bugün birlikte kısa bir yolculuğa çıkacağız.
|
| 137 |
+
Önce planımızı gözden geçirelim.
|
| 138 |
+
Ardından her adımı sakin ve anlaşılır biçimde anlatalım.
|
| 139 |
+
TXT
|
| 140 |
+
|
| 141 |
+
python ./Kahya-TTS-v1.0/inference.py \
|
| 142 |
+
--model ./Kahya-TTS-v1.0 \
|
| 143 |
+
--reference ./speaker.wav \
|
| 144 |
+
--sentences-file ./my-sentences.txt \
|
| 145 |
+
--output-dir ./kahya-long-clone
|
| 146 |
+
```
|
| 147 |
+
|
| 148 |
+
Every line starts with a fresh cache from the **same original reference**. Generated sentences are not fed back as the next sentence's voice prompt. This is the evaluated long-text recipe. It can still produce imperfect sentence joins; the script adds no pauses or crossfades. Listen to transitions as well as individual sentences.
|
| 149 |
+
|
| 150 |
+
## Advanced Python examples
|
| 151 |
+
|
| 152 |
+
These use the pinned [VoxCPM Python API](https://github.com/OpenBMB/VoxCPM/blob/f772e498a45fbb5fb8e13fbf9b9c48be9fe33e69/src/voxcpm/core.py), with **Kahya's merged weights**. They do not require reapplying the old LoRA. Parameters such as `reference_wav_path` and `prompt_text` are Python arguments, not additional flags accepted by the packaged `inference.py`.
|
| 153 |
+
|
| 154 |
+
Run this setup once, then run the examples you want in the **same Python session or script**, from the directory containing `Kahya-TTS-v1.0` and your reference files. Use a new output folder when starting another run.
|
| 155 |
+
|
| 156 |
+
<!-- kahya-python: setup -->
|
| 157 |
+
```python
|
| 158 |
+
from pathlib import Path
|
| 159 |
+
import soundfile as sf
|
| 160 |
+
from voxcpm import VoxCPM
|
| 161 |
+
|
| 162 |
+
out = Path("kahya-python-examples")
|
| 163 |
+
out.mkdir(exist_ok=False)
|
| 164 |
+
model = VoxCPM.from_pretrained(
|
| 165 |
+
"./Kahya-TTS-v1.0",
|
| 166 |
+
device="cuda",
|
| 167 |
+
load_denoiser=False,
|
| 168 |
+
optimize=False,
|
| 169 |
+
)
|
| 170 |
+
sample_rate = model.tts_model.sample_rate
|
| 171 |
+
settings = dict(
|
| 172 |
+
cfg_value=2.0,
|
| 173 |
+
inference_timesteps=16,
|
| 174 |
+
seed=42,
|
| 175 |
+
min_len=2,
|
| 176 |
+
max_len=1024,
|
| 177 |
+
normalize=False,
|
| 178 |
+
denoise=False,
|
| 179 |
+
retry_badcase=False,
|
| 180 |
+
)
|
| 181 |
+
```
|
| 182 |
+
|
| 183 |
+
The examples keep the release's CFG, diffusion-step count and seed. The high-level API examples are separate from the packaged command's measured streaming recipe: matching these settings does not imply identical PCM or timing. Keep targets short for initial checks. The Python API returns waveforms; it does not create the wrapper's `metrics.json` or provide its length-cap error report.
|
| 184 |
+
|
| 185 |
+
### Reference-only cloning in Python
|
| 186 |
+
|
| 187 |
+
<!-- kahya-python: reference_clone -->
|
| 188 |
+
```python
|
| 189 |
+
wav = model.generate(
|
| 190 |
+
text="Merhaba, yeni bir metni referans sesle okuyorum.",
|
| 191 |
+
reference_wav_path="./speaker.wav",
|
| 192 |
+
**settings,
|
| 193 |
+
)
|
| 194 |
+
sf.write(out / "reference-clone.wav", wav, sample_rate, subtype="FLOAT")
|
| 195 |
+
```
|
| 196 |
+
|
| 197 |
+
### Cloning with the reference's exact transcript
|
| 198 |
+
|
| 199 |
+
Create `speaker.txt` in UTF-8 containing the **words actually spoken in `speaker.wav`**, in their original order. This is the reference transcript, not the new target sentence. Do not paste example words unless they match your recording.
|
| 200 |
+
|
| 201 |
+
<!-- kahya-python: transcript_clone -->
|
| 202 |
+
```python
|
| 203 |
+
reference_transcript = " ".join(
|
| 204 |
+
Path("speaker.txt").read_text(encoding="utf-8").split()
|
| 205 |
+
)
|
| 206 |
+
if not reference_transcript:
|
| 207 |
+
raise ValueError("speaker.txt must contain the reference recording's transcript")
|
| 208 |
+
|
| 209 |
+
wav = model.generate(
|
| 210 |
+
text="Şimdi yeni bir cümleyle devam ediyorum.",
|
| 211 |
+
reference_wav_path="./speaker.wav",
|
| 212 |
+
prompt_wav_path="./speaker.wav",
|
| 213 |
+
prompt_text=reference_transcript + " ",
|
| 214 |
+
**settings,
|
| 215 |
+
)
|
| 216 |
+
sf.write(out / "transcript-clone.wav", wav, sample_rate, subtype="FLOAT")
|
| 217 |
+
```
|
| 218 |
+
|
| 219 |
+
This combines a voice reference with audio/text continuation conditioning. The upstream documentation calls this **Ultimate Cloning**. Both audio arguments deliberately point to the same recording. `prompt_wav_path` and `prompt_text` must always be supplied together; the generated WAV contains the new output, not a file concatenation of the reference recording and the output.
|
| 220 |
+
|
| 221 |
+
The trailing space after `reference_transcript` preserves the boundary before the new target text: the pinned implementation concatenates prompt text and target text directly.
|
| 222 |
+
|
| 223 |
+
The API also accepts the audio/transcript pair without `reference_wav_path` for continuation-only conditioning. That variation is not part of our example validation. Transcript-conditioned generation is a different mode from the evaluated long-text reset recipe; higher similarity is not guaranteed for Kahya. Keep style instructions out of the reference transcript and do not combine the transcript mode with the style-control recipe below.
|
| 224 |
+
|
| 225 |
+
### Guide emotion or pace while keeping a reference
|
| 226 |
+
|
| 227 |
+
<!-- kahya-python: controlled_clone -->
|
| 228 |
+
```python
|
| 229 |
+
wav = model.generate(
|
| 230 |
+
text="(Calm, warm tone, speaking slowly)Merhaba, birlikte sakin bir başlangıç yapalım.",
|
| 231 |
+
reference_wav_path="./speaker.wav",
|
| 232 |
+
**settings,
|
| 233 |
+
)
|
| 234 |
+
sf.write(out / "controlled-clone.wav", wav, sample_rate, subtype="FLOAT")
|
| 235 |
+
```
|
| 236 |
+
|
| 237 |
+
The parenthesized prefix is the upstream format for a natural-language style instruction. It is not a `speed` multiplier or a guaranteed emotion label. Here the instruction is in English and the target speech is Turkish. Whether the requested pace and emotion are expressed correctly must be checked by listening; Kahya has not been separately evaluated for instruction adherence.
|
| 238 |
+
|
| 239 |
+
### Design a voice without reference audio
|
| 240 |
+
|
| 241 |
+
<!-- kahya-python: voice_design -->
|
| 242 |
+
```python
|
| 243 |
+
wav = model.generate(
|
| 244 |
+
text="(An adult man with a warm, gentle voice)Merhaba, bugün sana nasıl yardımcı olabilirim?",
|
| 245 |
+
**settings,
|
| 246 |
+
)
|
| 247 |
+
sf.write(out / "voice-design.wav", wav, sample_rate, subtype="FLOAT")
|
| 248 |
+
```
|
| 249 |
+
|
| 250 |
+
This requests a voice through a description. It does not clone a particular person. Single-speaker fine-tuning may constrain the variety retained from the base model, so this example is an API usage demonstration rather than evidence of broad voice-design quality.
|
| 251 |
+
|
| 252 |
+
For ordinary TTS without a reference or voice description:
|
| 253 |
+
|
| 254 |
+
<!-- kahya-python: no_reference -->
|
| 255 |
+
```python
|
| 256 |
+
wav = model.generate(
|
| 257 |
+
text="Merhaba, bugün sana nasıl yardımcı olabilirim?",
|
| 258 |
+
**settings,
|
| 259 |
+
)
|
| 260 |
+
sf.write(out / "no-reference.wav", wav, sample_rate, subtype="FLOAT")
|
| 261 |
+
```
|
| 262 |
+
|
| 263 |
+
### Consume streaming audio chunks
|
| 264 |
+
|
| 265 |
+
<!-- kahya-python: streaming_clone -->
|
| 266 |
+
```python
|
| 267 |
+
import numpy as np
|
| 268 |
+
|
| 269 |
+
chunks = []
|
| 270 |
+
stream = model.generate_streaming(
|
| 271 |
+
text="Merhaba, bu ses küçük parçalar halinde üretiliyor.",
|
| 272 |
+
reference_wav_path="./speaker.wav",
|
| 273 |
+
**settings,
|
| 274 |
+
)
|
| 275 |
+
try:
|
| 276 |
+
for chunk in stream:
|
| 277 |
+
# A mono float32 NumPy array on CPU, at sample_rate Hz.
|
| 278 |
+
# An application can send each chunk to its audio playback queue here.
|
| 279 |
+
chunks.append(chunk)
|
| 280 |
+
finally:
|
| 281 |
+
stream.close()
|
| 282 |
+
|
| 283 |
+
if not chunks:
|
| 284 |
+
raise RuntimeError("The model returned no audio chunks")
|
| 285 |
+
sf.write(out / "streaming-clone.wav", np.concatenate(chunks), sample_rate, subtype="FLOAT")
|
| 286 |
+
```
|
| 287 |
+
|
| 288 |
+
This example collects chunks into a WAV. It demonstrates the generator interface; it does not implement a live player, HTTP endpoint or streaming server. The packaged `inference.py` also streams internally, but its user-facing outputs are files.
|
| 289 |
+
|
| 290 |
+
### Troubleshooting and validation boundaries
|
| 291 |
+
|
| 292 |
+
| Symptom | Check |
|
| 293 |
+
|---|---|
|
| 294 |
+
| `--reference must contain finite, non-empty mono audio` | Convert stereo to mono; check that the file contains readable samples. |
|
| 295 |
+
| `--output-dir already exists` | Use a new directory; the packaged command preserves existing outputs. |
|
| 296 |
+
| `prompt_wav_path and prompt_text must both be provided` | Supply both the recording and its transcript, or use reference-only mode. |
|
| 297 |
+
| Voice differs from the intended speaker | Check reference quality and compare short targets first; unfamiliar-speaker similarity has not been established for this fine-tune. |
|
| 298 |
+
| Output drifts toward the end of a paragraph | Use explicit sentence lines with the same original reference for every line. |
|
| 299 |
+
| A style instruction is ineffective or spoken aloud | Remove the prefix and compare reference-only generation; instruction adherence is unmeasured. |
|
| 300 |
+
| Transcript-conditioned output repeats or omits words | Check the reference transcript and phrase boundary; compare against reference-only mode. |
|
| 301 |
+
| A generation is cut short | Inspect the packaged command's `metrics.json` and shorten the utterance/reference context. API examples do not implement that cap check. |
|
| 302 |
+
|
| 303 |
+
On **7 September 2026**, all six generation examples above produced finite, non-empty, mono **48 kHz** audio on the recorded RTX PRO 6000 environment, using Kahya weights from release commit `90d34bee01d7ffaf3d7cffce3fbabdb8f3329dde`. For this functional check, the published synthetic quickstart audio and its source text supplied the reference and transcript. The revised transcript example also passed with a transcript lacking terminal punctuation, verifying the explicit word boundary. The streaming example yielded multiple chunks whose concatenation matched the saved waveform. These are execution checks, not listening tests of the advanced modes.
|
| 304 |
+
|
| 305 |
+
API availability, successful waveform generation and perceived cloning quality are different checks. The recorded human ratings below apply to the original same-speaker reference recipe, not to new voices, transcript conditioning, voice design, style control or other languages. The published quickstart WAV is generated speech, not a replacement for a clean recording of the voice you want to use.
|
| 306 |
+
|
| 307 |
## Measured inference examples
|
| 308 |
|
| 309 |
Measurements below come from the fixed **5 September 2026** development run on an **NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB**, CUDA BF16, CFG 2.0, 16 diffusion steps, seed 42 and the same reference clip. They are observations from two texts, not a throughput SLA.
|
SHA256SUMS
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
4f10acc209addacfad28293315c74c4cd648f771ee1263a748f1781d1e0265e4 LICENSE
|
| 2 |
7c070bb12657803796a479ec173f6596f35756c49e3362d479fbfbf32e62e143 NOTICE
|
| 3 |
-
|
| 4 |
94b5d51e107e0507d4acc976cfdadb64edd6fd06d1f751dadbf2fd1594274bf1 audiovae.pth
|
| 5 |
405f0dcd92f7feba6011ed4eac5c8d4f74cba9712f07fd5cfa3063bbdd95402c config.json
|
| 6 |
9f0f2b3647a7643cbc8efdad667251006991e530bd38cb774d3da792bf392aa1 evaluation/README.md
|
|
@@ -13,7 +13,7 @@ eec23ec02cdb969462031031d6d7f36972f90df4e8ca57ec2946933b34cc5c7b examples/quick
|
|
| 13 |
7c0b2acb55672c1dd866237e4535bf1bcd1b6a798c5be163f31b376802066ff3 inference.py
|
| 14 |
3b9bc1f55569d5495f6dccf788c5da3349e606c200328af5cdf180a9bf320e0d merge_manifest.json
|
| 15 |
780bea8f21515c6dc97e577b37aa350f3cbdb4b1f669ee5ebb5bbe969b6b4ba5 model.safetensors
|
| 16 |
-
|
| 17 |
068594063e37662c02b21acf42ebb334ef6a74fb810e68a2368f88f08351de76 special_tokens_map.json
|
| 18 |
84489ea32b6ee0cae22ed5480cacb6df85c46624c3119be9a2021c3649a12729 tokenization_voxcpm2.py
|
| 19 |
f8984687e4a92a3503d521396d454b7d68e9fdaab2a0288eb3536c7c1aa4bc20 tokenizer.json
|
|
|
|
| 1 |
4f10acc209addacfad28293315c74c4cd648f771ee1263a748f1781d1e0265e4 LICENSE
|
| 2 |
7c070bb12657803796a479ec173f6596f35756c49e3362d479fbfbf32e62e143 NOTICE
|
| 3 |
+
ca0bfd2f4d1aa6f0f143a4a185d4e8f404361c7a5b3d6e00e45a7a2258da1bd1 README.md
|
| 4 |
94b5d51e107e0507d4acc976cfdadb64edd6fd06d1f751dadbf2fd1594274bf1 audiovae.pth
|
| 5 |
405f0dcd92f7feba6011ed4eac5c8d4f74cba9712f07fd5cfa3063bbdd95402c config.json
|
| 6 |
9f0f2b3647a7643cbc8efdad667251006991e530bd38cb774d3da792bf392aa1 evaluation/README.md
|
|
|
|
| 13 |
7c0b2acb55672c1dd866237e4535bf1bcd1b6a798c5be163f31b376802066ff3 inference.py
|
| 14 |
3b9bc1f55569d5495f6dccf788c5da3349e606c200328af5cdf180a9bf320e0d merge_manifest.json
|
| 15 |
780bea8f21515c6dc97e577b37aa350f3cbdb4b1f669ee5ebb5bbe969b6b4ba5 model.safetensors
|
| 16 |
+
eb2603a066ed4b936dda75bd8aa3c5200eb0d7418e60e01b3bd222d216b35c84 release_manifest.json
|
| 17 |
068594063e37662c02b21acf42ebb334ef6a74fb810e68a2368f88f08351de76 special_tokens_map.json
|
| 18 |
84489ea32b6ee0cae22ed5480cacb6df85c46624c3119be9a2021c3649a12729 tokenization_voxcpm2.py
|
| 19 |
f8984687e4a92a3503d521396d454b7d68e9fdaab2a0288eb3536c7c1aa4bc20 tokenizer.json
|
release_manifest.json
CHANGED
|
@@ -25,8 +25,8 @@
|
|
| 25 |
},
|
| 26 |
{
|
| 27 |
"path": "README.md",
|
| 28 |
-
"bytes":
|
| 29 |
-
"sha256": "
|
| 30 |
},
|
| 31 |
{
|
| 32 |
"path": "audiovae.pth",
|
|
@@ -109,5 +109,12 @@
|
|
| 109 |
"sha256": "e78a3ebb48a0b9437efd1823b6b726c823da89e49dd8bcc90c02419d9baa772b"
|
| 110 |
}
|
| 111 |
],
|
| 112 |
-
"manifest_scope": "Files listed here exclude release_manifest.json and SHA256SUMS themselves."
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 113 |
}
|
|
|
|
| 25 |
},
|
| 26 |
{
|
| 27 |
"path": "README.md",
|
| 28 |
+
"bytes": 23261,
|
| 29 |
+
"sha256": "ca0bfd2f4d1aa6f0f143a4a185d4e8f404361c7a5b3d6e00e45a7a2258da1bd1"
|
| 30 |
},
|
| 31 |
{
|
| 32 |
"path": "audiovae.pth",
|
|
|
|
| 109 |
"sha256": "e78a3ebb48a0b9437efd1823b6b726c823da89e49dd8bcc90c02419d9baa772b"
|
| 110 |
}
|
| 111 |
],
|
| 112 |
+
"manifest_scope": "Files listed here exclude release_manifest.json and SHA256SUMS themselves.",
|
| 113 |
+
"documentation_update": {
|
| 114 |
+
"date": "2026-09-07",
|
| 115 |
+
"topic": "Voice cloning and advanced Python API usage",
|
| 116 |
+
"parent_commit": "90d34bee01d7ffaf3d7cffce3fbabdb8f3329dde",
|
| 117 |
+
"weights_changed": false,
|
| 118 |
+
"inference_script_changed": false
|
| 119 |
+
}
|
| 120 |
}
|