# ============ Common ============ numpy scipy pydub pyyaml onnxruntime>=1.15.0 gradio[mcp]>=4.0.0 # ============ Audiosplitter only ============ # ASR: faster-whisper (CTranslate2) faster-whisper>=1.0.0 # Clustering for diarization scikit-learn>=1.0.0 # ============ Voice Extractor only ============ setuptools # PyTorch CPU-only --extra-index-url https://download.pytorch.org/whl/cpu torch==2.5.1+cpu torchaudio==2.5.1+cpu # Audio soundfile==0.13.1 librosa==0.10.2.post1 ffmpeg-python==0.2.0 # ML models pyannote.audio==3.4.0 openai-whisper==20250625 speechbrain==1.0.2 silero-vad==5.1.2 # WeSpeaker wespeaker @ git+https://github.com/wenet-e2e/wespeaker.git # Deps s3prl==0.4.17 peft>=0.11.0 transformers>=4.40.0 pytorch-lightning==2.4.0 onnx==1.17.0 matplotlib==3.9.3 rich==14.0.0 tqdm==4.67.1