Fun-CosyVoice3-0.5B-2512-GGUF

Unofficial community GGUF release for CosyVoice3-2512.

This model pack is intended for use with cosyvoice.cpp.

Important Notice

  • This is not an official CosyVoice release.
  • The implementation is maintained by an independent community developer.
  • The C++ runtime project (cosyvoice.cpp) is MIT.
  • Model artifacts in this pack remain Apache-2.0.

Available Files

Naming Changes

The following quantizations have been renamed:

  • Q6_K_S โ†’ Q6_K
  • Q5_K_S โ†’ Q5_K_XXS
  • Q4_K_S โ†’ Q4_K_XXS
  • Q3_K_S โ†’ Q3_K_XXS
  • Q2_K_S โ†’ removed (mostly noisy and unusable)

Current Variants

File Quantization Size (approx) Notes
CosyVoice3-2512_F32.gguf F32 3.21 GiB Highest precision, largest size
CosyVoice3-2512_F16.gguf F16 1.61 GiB High-quality baseline
CosyVoice3-2512_Q8_0.gguf Q8_0 0.88 GiB Near-F16 quality, smaller
CosyVoice3-2512_Q6_K.gguf Q6_K 0.78 GiB Formerly Q6_K_S; good quality/size balance
CosyVoice3-2512_Q5_1.gguf Q5_1 0.64 GiB Better quality than Q5_0
CosyVoice3-2512_Q5_K_M.gguf Q5_K_M 0.64 GiB New; maintains decent quality, good results
CosyVoice3-2512_Q5_K_S.gguf Q5_K_S 0.61 GiB New; maintains decent quality, good results
CosyVoice3-2512_Q5_K_XXS.gguf Q5_K_XXS 0.61 GiB Formerly Q5_K_S; further compressed
CosyVoice3-2512_Q5_0.gguf Q5_0 0.59 GiB Compact mid-quality
CosyVoice3-2512_Q4_K_M.gguf Q4_K_M 0.59 GiB New; quality drops from Q5, occasional muffled audio, mostly acceptable
CosyVoice3-2512_Q4_K_S.gguf Q4_K_S 0.56 GiB New; lower quality than Q4_K_M
CosyVoice3-2512_Q4_1.gguf Q4_1 0.54 GiB Smaller size, audible pronunciation artifacts
CosyVoice3-2512_Q4_K_XXS.gguf Q4_K_XXS 0.54 GiB Formerly Q4_K_S
CosyVoice3-2512_Q4_0.gguf Q4_0 0.49 GiB Further quality degradation, often unclear speech
CosyVoice3-2512_Q3_K_M.gguf Q3_K_M 0.49 GiB New; muffled audio, slightly below but close to Q4_K_XXS
CosyVoice3-2512_Q3_K_S.gguf Q3_K_S 0.46 GiB New; lower quality than Q3_K_M
CosyVoice3-2512_Q3_K_XXS.gguf Q3_K_XXS 0.44 GiB Formerly Q3_K_S
CosyVoice3-2512_Q2_K_L.gguf Q2_K_L 0.43 GiB New; significant quality drop, quiet and muffled with noise
CosyVoice3-2512_Q2_K.gguf Q2_K 0.42 GiB New; severe artifacts and noise, but speech content can still be discerned

Quick Recommendations

K-quant quality relationship within each series (high to low): K_M > K_S > K_XXS

  • Default recommendation: Q8_0 (near-lossless listening quality in current tests, with much smaller size than F16)
  • High-quality alternatives: F16 and Q6_K
  • Usable mid-tier options: Q5_1, Q5_K_M, Q5_K_S, Q5_0, Q5_K_XXS
  • Quality drops but still acceptable: Q4_K_M, Q4_K_S
  • Audible artifacts: Q4_1, Q4_K_XXS, Q4_0
  • Further quality degradation: Q3_K_M, Q3_K_S, Q3_K_XXS
  • Extreme compression only: Q2_K_L, Q2_K

Subjective listening notes from current tests:

  • Q8_0: quality remains strong with little audible loss in typical samples.
  • Q6_K: still sounds good and is often a practical choice.
  • Q5_K_M / Q5_K_S (new): maintain decent quality, good results.
  • Q5_K_XXS (formerly Q5_K_S): slightly higher than but close to Q4_K_M.
  • Q5 family (Q5_1 / Q5_K_M / Q5_K_S / Q5_0 / Q5_K_XXS): generally usable with moderate degradation.
  • Q4_K_M / Q4_K_S (new): quality drops compared to Q5, occasional muffled audio, but mostly acceptable.
  • Q4 family (including Q4_1 / Q4_K_XXS / Q4_0): audible pronunciation artifacts, often muffled.
  • Q3_K_M / Q3_K_S (new): further quality degradation, audio is muffled, slightly below but close to Q4_K_XXS.
  • Q3_K_XXS (formerly Q3_K_S): aggressive quantization, limited quality.
  • Q2_K_L (new): significant quality drop, quieter and muffled audio with noise.
  • Q2_K (new): severe artifacts and noise, but speech content can still be discerned.
  • Q2_K_S (removed): mostly degrades into noise and is usually unusable.

Runtime Requirements

  • Built for cosyvoice.cpp GGUF inference pipeline.
  • Parts of the runtime currently require AVX2 support.
  • GPU backend behavior depends on GGML backend and driver stack.

Basic Usage

cosyvoice-cli \
  --model CosyVoice3-2512_Q8_0.gguf \
  --prompt-speech prompt_speech.gguf \
  --text "Hello from CosyVoice" \
  --output out.wav

--prompt-speech expects a prompt-speech file in GGUF format (for example prompt_speech.gguf).

For runtime usage and documentation, see the cosyvoice.cpp repository:

Provenance

  • This work follows the architecture and inference behavior of upstream CosyVoice releases, then re-implements the pipeline for C++/GGML deployment.
  • Tokenizer implementation is adapted from llama.cpp and modified for this project.

License

  • Model files in this repository: Apache-2.0
  • Runtime code (cosyvoice.cpp): MIT
Downloads last month
4,388
GGUF
Model size
0.9B params
Architecture
cosyvoice3-2512
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Lourdle/Fun-CosyVoice3-0.5B-2512-GGUF

Quantized
(22)
this model

Spaces using Lourdle/Fun-CosyVoice3-0.5B-2512-GGUF 2