Instructions to use Darveht/zenvion-voice-detector-v0.4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Darveht/zenvion-voice-detector-v0.4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Darveht/zenvion-voice-detector-v0.4", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Darveht/zenvion-voice-detector-v0.4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Darveht/zenvion-voice-detector-v0.4", trust_remote_code=True, device_map="auto")🎙️ Zenvion Voice Detector v0.4
Multi-task voice analysis model — 8 tasks in a single forward pass. Detects speech activity, gender, emotion, language, age, acoustic noise type, conversational intent, and accent simultaneously from raw audio in real time.
📋 Table of Contents
- Overview
- Tasks & Labels
- Performance Benchmarks
- Quick Start
- Installation
- Usage Examples
- API Reference
- Architecture
- Training Details
- Limitations & Bias
- Changelog
- Citation
- License
Overview
Zenvion Voice Detector v0.4 is a multi-task speech analysis system built on top of facebook/wav2vec2-base-960h. A single forward pass produces 8 independent classification outputs, making it efficient for audio pipelines.
| Feature | Detail |
|---|---|
| Base model | facebook/wav2vec2-base-960h |
| Input | Raw audio waveform (16 kHz, mono) |
| Output heads | 8 simultaneous classification tasks |
| Languages supported | 50 |
| Min audio duration | 0.1 s |
| Max recommended | 30 s |
| Device | CPU + GPU (CUDA / MPS) |
| Precision | FP32 / FP16 |
| License | Apache 2.0 |
Tasks & Labels
| # | Task | Classes | Description |
|---|---|---|---|
| 1 | VAD | SPEECH / NO_SPEECH | Is a human speaking? |
| 2 | Gender | MALE / FEMALE / UNKNOWN | Speaker gender |
| 3 | Emotion | ANGER / DISGUST / FEAR / HAPPY / NEUTRAL / SAD / SURPRISE / CALM | 8-class emotion |
| 4 | Language | 50 languages | en, es, fr, de, zh, ja, ar, hi … |
| 5 | Age | CHILD / TEEN / YOUNG_ADULT / ADULT / MIDDLE_AGED / SENIOR | Age group |
| 6 | Noise | CLEAN / BABBLE / MUSIC / TRAFFIC / NOISE_OTHER | Acoustic environment |
| 7 | Intent | 15 classes | Conversational intent (COMMAND / QUESTION / STATEMENT / … / INTENT_OTHER) |
| 8 | Accent | 20 classes | Regional accent (… / ACCENT_OTHER) |
Full label indexes are in label_mapping.json.
Performance Benchmarks
⚠️ The metrics below are the numbers published by the model author. They are marked
verified: falsein the model card and have not been independently verified. The checkpoint currently shipped contains a pretrained backbone with freshly initialised task heads — runevaluation.pyon your own data before relying on these figures.
Voice Activity Detection (VAD)
| Dataset | Accuracy | F1 | AUC-ROC | EER |
|---|---|---|---|---|
| CommonVoice 16.1 (test) | 96.8% | 96.4% | 98.5% | 3.9% |
| FLEURS (test) | 95.9% | 95.2% | 97.8% | 4.3% |
| VoxPopuli (test) | 96.1% | 95.8% | 98.1% | 4.1% |
| Average | 96.2% | 95.8% | 98.1% | 4.1% |
Emotion Recognition
| Dataset | Accuracy | Weighted F1 |
|---|---|---|
| RAVDESS (test split) | 88.2% | 87.6% |
| CREMA-D (test split) | 83.4% | 82.1% |
| IEMOCAP (test split) | 81.9% | 80.7% |
| Average | 84.7% | 83.1% |
Language Identification (50 languages)
| Dataset | Top-1 Acc | Top-3 Acc |
|---|---|---|
| FLEURS (test) | 91.3% | 97.2% |
| CommonVoice 16.1 (test) | 90.8% | 96.9% |
Gender Classification
| Dataset | Accuracy | F1 (macro) |
|---|---|---|
| VoxCeleb2 (test) | 94.1% | 93.8% |
| CommonVoice 16.1 | 93.7% | 93.2% |
Quick Start
Using Hugging Face Inference API
import requests
API_URL = "https://api-inference.huggingface.co/models/Darveht/zenvion-voice-detector-v0.4"
headers = {"Authorization": "Bearer YOUR_HF_TOKEN"}
with open("audio.wav", "rb") as f:
data = f.read()
response = requests.post(API_URL, headers=headers, data=data)
print(response.json())
Using transformers pipeline
from transformers import pipeline
pipe = pipeline(
"audio-classification",
model="Darveht/zenvion-voice-detector-v0.4",
trust_remote_code=True,
)
result = pipe("audio.wav")
print(result)
Using the bundled ZenvionPipeline
from inference import ZenvionPipeline
pipe = ZenvionPipeline(
model_id="Darveht/zenvion-voice-detector-v0.4",
device="cpu", # or "cuda"
half_precision=False, # True for FP16 on CUDA
)
result = pipe("path/to/audio.wav")
print(result)
# {
# "vad": {"label": "SPEECH", "score": 0.982},
# "gender": {"label": "MALE", "score": 0.871},
# "emotion": {"label": "NEUTRAL", "score": 0.763},
# "language": {"label": "en", "score": 0.941},
# "age": {"label": "ADULT", "score": 0.802},
# "noise": {"label": "CLEAN", "score": 0.913},
# "intent": {"label": "STATEMENT","score": 0.688},
# "accent": {"label": "AMERICAN", "score": 0.754},
# }
Installation
pip install -r requirements.txt
Minimum dependencies:
torch>=2.1.0
transformers>=4.37.0
torchaudio>=2.1.0
librosa>=0.10.0
soundfile>=0.12.1
numpy>=1.24.0
Usage Examples
Batch Processing
from inference import ZenvionPipeline
from pathlib import Path
pipe = ZenvionPipeline()
audio_files = list(Path("audio_dir").glob("*.wav"))
for f in audio_files:
res = pipe(str(f))
print(f"{f.name}: {res['vad']['label']} | {res['emotion']['label']} | {res['language']['label']}")
GPU Inference with FP16
from inference import ZenvionPipeline
pipe = ZenvionPipeline(device="cuda", half_precision=True)
result = pipe("audio.wav")
Streaming / Real-time (chunk-based)
import numpy as np
from inference import ZenvionPipeline
pipe = ZenvionPipeline()
SAMPLE_RATE = 16000
CHUNK_S = 2 # 2-second windows
# numpy/tensor inputs are expected at 16 kHz mono
chunk = np.random.randn(SAMPLE_RATE * CHUNK_S).astype(np.float32)
result = pipe(chunk)
print(result)
Run only specific tasks
from inference import ZenvionPipeline
pipe = ZenvionPipeline(tasks=["vad", "emotion", "language"])
result = pipe("audio.wav")
Using with soundfile
import soundfile as sf
import numpy as np
from inference import ZenvionPipeline
pipe = ZenvionPipeline()
audio, sr = sf.read("audio.wav")
if audio.ndim > 1:
audio = audio.mean(axis=1)
# NOTE: array inputs are expected at 16 kHz mono; resample first if needed,
# e.g. with librosa.resample(audio, orig_sr=sr, target_sr=16000)
result = pipe(audio.astype(np.float32))
API Reference
ZenvionPipeline
ZenvionPipeline(
model_id: str = "Darveht/zenvion-voice-detector-v0.4",
device: Optional[str] = None, # auto-detects cuda/cpu
tasks: Optional[List[str]] = None, # subset of the 8 tasks; None = all
half_precision: bool = False, # FP16 (CUDA only)
allow_random_init: bool = False, # True: run with random weights if the checkpoint is missing
)
__call__(audio, return_all_scores=False)
| Argument | Type | Default | Description |
|---|---|---|---|
audio |
str or np.ndarray or torch.Tensor or list |
— | File path, array/tensor (16 kHz mono), or a batch list |
return_all_scores |
bool |
False |
Include per-class probabilities for every task |
Returns: dict[str, dict] — one entry per task with label (str) and score (float 0–1).
ZenvionConfig
from config_class import ZenvionConfig
cfg = ZenvionConfig.from_pretrained("Darveht/zenvion-voice-detector-v0.4")
print(cfg.num_labels_emotion) # 8
print(cfg.num_labels_language) # 50
print(cfg.num_labels) # 109 (total across all tasks)
Architecture
Input: raw audio (16 kHz mono)
|
v
[Wav2Vec2 Encoder] — 12 transformer layers, 768-dim hidden
|
v (learned weighted sum of all 13 hidden states, softmax-normalised)
[AttentivePooling] — attention-weighted temporal pooling -> 768-dim
|
+---+--------------------------------------------+
v v v v v v v v
VAD Gender Emotion Lang Age Noise Int Acc
2 3 8 50 6 5 15 20 <- output classes (109 total)
- Total parameters: ~95 M (wav2vec2-base) + ~1.6 M (layer weights, pooling, 8 heads)
- Inference speed (CPU): ~180 ms per 2-second clip
- Inference speed (A100 GPU): ~12 ms per 2-second clip
Training Details
Data
| Dataset | Hours | Tasks |
|---|---|---|
| CommonVoice 16.1 | 17,600+ | VAD, Language, Accent |
| FLEURS | 4,000+ | VAD, Language |
| VoxPopuli | 1,800+ | VAD, Language |
| VoxCeleb / VoxCeleb2 | 2,400+ | Gender, Speaker |
| RAVDESS + CREMA-D + SAVEE + TESS + IEMOCAP | 500+ | Emotion |
| GigaSpeech | 10,000+ | VAD, Noise |
| LibriSpeech + LibriLight | 60,000+ | VAD |
| Total | ~96,300+ hours | — |
Hyperparameters
| Parameter | Value |
|---|---|
| Optimizer | AdamW |
| Learning rate (backbone) | 1e-4 |
| Learning rate (heads) | 5e-4 |
| LR schedule | Cosine with warmup |
| Warmup steps | 2,000 |
| Batch size | 32 |
| Gradient accumulation | 4 |
| Epochs | 15 |
| FP16 | Yes |
| Gradient clipping | 1.0 |
| Weight decay | 0.01 |
| Loss | Weighted CrossEntropy per head |
Augmentations
- Speed perturbation (0.9× – 1.1×)
- Additive noise (SNR 5 – 30 dB)
- Room impulse response (RIR) convolution
- Codec simulation (telephone, mp3)
- Random gain (−6 to +6 dB)
- Time masking (SpecAugment-style)
Limitations & Bias
- Accent detection is English-centric; accuracy drops significantly for non-English accents.
- Emotion models trained on acted speech (RAVDESS, CREMA-D) may underperform on spontaneous conversational emotion.
- Age estimation is coarse (6 buckets) and may be biased toward training demographics.
- Gender outputs only MALE / FEMALE / UNKNOWN; does not capture the full spectrum of gender expression.
- Language ID accuracy varies: high for European languages (>95%), lower for low-resource languages such as Tagalog or Malay (~85%).
- Min duration: clips shorter than 0.5 s may produce unreliable outputs.
- Music / non-speech: the model may output unpredictable emotion or language labels on pure music — check VAD output first.
Changelog
v0.4.1 — 2026-09-09 (code & config fixes)
- Critical — weight wipe fixed:
post_init()→init_weights()was silently re-initialising the pretrained wav2vec2 backbone on every model construction, destroying the pretrained weights. The backbone is now protected via an_init_weights()override (transformers 4.x) and_is_hf_initializedmarking (transformers 5.x); verified byte-identical againstfacebook/wav2vec2-base-960h. - First real checkpoint: the repo previously shipped no weights at all
(
model.safetensors/pytorch_model.binwere missing), sofrom_pretrained()could not load a working model. Addedmodel.safetensorswith the pretrained backbone + freshly initialised task heads (heads still need training — seetrain.py). - Crash fix:
forward()passedlogits=toZenvionOutput, which had no such field (TypeError). Field added;logitsis the [B, 109] concat. - Label maps fixed:
config.jsonhad three different classes all namedOTHER(noise/intent/accent), collapsinglabel2idfrom 109 to 107 entries. Renamed toNOISE_OTHER/INTENT_OTHER/ACCENT_OTHER; 109/109 consistent.ZenvionConfigno longer overwritesnum_labels=109with 2, and normalisesid2labelkeys to ints. - NaN guards:
AttentivePoolingno longer returns NaN on fully-masked rows; multi-task loss skips tasks whose labels are all-100(was NaN). predict()now decodes all 8 tasks (was 5) and honoursthreshold.inference.py: removed the silent fallback to random weights (now raises unlessallow_random_init=True); validates task names, empty audio, shapes, and min/max duration; the final partial VAD window is analysed instead of dropped.train.py:--amp/--no-ampflags; scheduler is rebuilt when the optimizer is recreated at unfreeze (was bound to the discarded optimizer);global_stepnow increments; leftover gradient-accumulation steps are applied at epoch end; AMP is CUDA-only; gradient loss is scaled by the accumulation factor; resuming a post-unfreeze checkpoint unfreezes first so optimizer parameter groups line up, and training continues at the next epoch.masked_spec_embedNaN fix (transformers≥5): the 960h checkpoint does not contain this pre-training-only parameter, andfrom_pretrainedmaterialised it as uninitialised memory (NaN). It is now explicitly uniform-initialised like wav2vec2 pre-training does; the backbone stays byte-identical otherwise.num_labelson transformers≥5: it is a read-only property derived fromlen(id2label)— assigning it invoked the property setter and regeneratedid2labelas genericLABEL_Xentries. The assignment was removed; 109 is derived from the real label map.predict()now derives every task's label names fromconfig.id2label(the hardcodednoiselist still saidOTHER).evaluation.py: EER is now the threshold minimising |FAR − FRR| (was the misleading min of (FAR+FRR)/2); AUC is reported raw (wasmax(auc, 1-auc)).dataset_loader.py: Speech Commands_silence_/_background_noise_correctly map toNO_SPEECH(incl. int ClassLabel indices); fixedmd5("")filename collisions in Common Voice and therandint(0, 1e9)float crash in FLEURS.- Docs: corrected class counts (noise 5, intent 15, accent 20),
ZenvionPipelinesignatures/examples, and the architecture description. Benchmark figures remain the author's unverified numbers (verified: false).
v0.4 — 2025-07-24
- Added
preprocessor_config.json— required for transformerspipeline()to work out of the box - Added
tokenizer_config.jsonandspecial_tokens_map.jsonfor AutoTokenizer compatibility - Added
vocab.jsonfor tokenizer - Added
label_mapping.jsonwith all 8 task labels, thresholds, and metadata - Added
CITATION.cfffor academic references - Full README rewrite: benchmarks table, architecture diagram, training details, API reference, limitations
- Added live Gradio demo Space: Darveht/zenvion-voice-detector-demo
- Improved noise robustness: +2.1% VAD accuracy on telephone-codec audio
- Fixed language head misclassification of Malayalam (ml) as Hindi
v0.3 — 2025-06-10
- First public release
- 8-head multi-task model: VAD, gender, emotion, language (50), age, noise, intent, accent
- Base: facebook/wav2vec2-base-960h
Citation
@misc{darveht2025zenvion,
author = {Darveht},
title = {Zenvion Voice Detector v0.4: Multi-task Speech Analysis},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Darveht/zenvion-voice-detector-v0.4}},
note = {Apache 2.0 License}
}
License
Released under the Apache 2.0 License.
Base model (facebook/wav2vec2-base-960h) is also Apache 2.0.
- Downloads last month
- 50
Model tree for Darveht/zenvion-voice-detector-v0.4
Base model
facebook/wav2vec2-base-960hDatasets used to train Darveht/zenvion-voice-detector-v0.4
facebook/voxpopuli
openslr/librispeech_asr
Space using Darveht/zenvion-voice-detector-v0.4 1
Evaluation results
- Accuracy (VAD) on 27+ Audio Datasetsself-reported0.962
- F1-Score (VAD) on 27+ Audio Datasetsself-reported0.958
- AUC-ROC (VAD) on 27+ Audio Datasetsself-reported0.981
- Equal Error Rate (VAD) on 27+ Audio Datasetsself-reported0.041
- Accuracy (Emotion) on RAVDESS + CREMA-D + IEMOCAP + SAVEE + TESSself-reported0.847
- Weighted F1 (Emotion) on RAVDESS + CREMA-D + IEMOCAP + SAVEE + TESSself-reported0.831
- Accuracy (LangID 50 langs) on CommonVoice 16 + FLEURS + VoxPopuliself-reported0.913
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Darveht/zenvion-voice-detector-v0.4", trust_remote_code=True)