Instructions to use Darveht/zenvion-voice-detector-v0.4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
How to use Darveht/zenvion-voice-detector-v0.4 with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("audio-classification", model="Darveht/zenvion-voice-detector-v0.4", trust_remote_code=True)
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Darveht/zenvion-voice-detector-v0.4", trust_remote_code=True, device_map="auto")
Multi-task voice analysis model — 8 tasks in a single forward pass.
Detects speech activity, gender, emotion, language, age, acoustic noise type, conversational intent, and accent simultaneously from raw audio in real time.
Zenvion Voice Detector v0.4 is a multi-task speech analysis system built on top of facebook/wav2vec2-base-960h. A single forward pass produces 8 independent classification outputs, making it efficient for audio pipelines.
Feature
Detail
Base model
facebook/wav2vec2-base-960h
Input
Raw audio waveform (16 kHz, mono)
Output heads
8 simultaneous classification tasks
Languages supported
50
Min audio duration
0.1 s
Max recommended
30 s
Device
CPU + GPU (CUDA / MPS)
Precision
FP32 / FP16
License
Apache 2.0
Tasks & Labels
#
Task
Classes
Description
1
VAD
SPEECH / NO_SPEECH
Is a human speaking?
2
Gender
MALE / FEMALE / UNKNOWN
Speaker gender
3
Emotion
ANGER / DISGUST / FEAR / HAPPY / NEUTRAL / SAD / SURPRISE / CALM
⚠️ The metrics below are the numbers published by the model author. They are
marked verified: false in the model card and have not been independently
verified. The checkpoint currently shipped contains a pretrained backbone
with freshly initialised task heads — run evaluation.py on your own data
before relying on these figures.
from inference import ZenvionPipeline
from pathlib import Path
pipe = ZenvionPipeline()
audio_files = list(Path("audio_dir").glob("*.wav"))
for f in audio_files:
res = pipe(str(f))
print(f"{f.name}: {res['vad']['label']} | {res['emotion']['label']} | {res['language']['label']}")
GPU Inference with FP16
from inference import ZenvionPipeline
pipe = ZenvionPipeline(device="cuda", half_precision=True)
result = pipe("audio.wav")
Streaming / Real-time (chunk-based)
import numpy as np
from inference import ZenvionPipeline
pipe = ZenvionPipeline()
SAMPLE_RATE = 16000
CHUNK_S = 2# 2-second windows# numpy/tensor inputs are expected at 16 kHz mono
chunk = np.random.randn(SAMPLE_RATE * CHUNK_S).astype(np.float32)
result = pipe(chunk)
print(result)
Run only specific tasks
from inference import ZenvionPipeline
pipe = ZenvionPipeline(tasks=["vad", "emotion", "language"])
result = pipe("audio.wav")
Using with soundfile
import soundfile as sf
import numpy as np
from inference import ZenvionPipeline
pipe = ZenvionPipeline()
audio, sr = sf.read("audio.wav")
if audio.ndim > 1:
audio = audio.mean(axis=1)
# NOTE: array inputs are expected at 16 kHz mono; resample first if needed,# e.g. with librosa.resample(audio, orig_sr=sr, target_sr=16000)
result = pipe(audio.astype(np.float32))
API Reference
ZenvionPipeline
ZenvionPipeline(
model_id: str = "Darveht/zenvion-voice-detector-v0.4",
device: Optional[str] = None, # auto-detects cuda/cpu
tasks: Optional[List[str]] = None, # subset of the 8 tasks; None = all
half_precision: bool = False, # FP16 (CUDA only)
allow_random_init: bool = False, # True: run with random weights if the checkpoint is missing
)
__call__(audio, return_all_scores=False)
Argument
Type
Default
Description
audio
str or np.ndarray or torch.Tensor or list
—
File path, array/tensor (16 kHz mono), or a batch list
return_all_scores
bool
False
Include per-class probabilities for every task
Returns:dict[str, dict] — one entry per task with label (str) and score (float 0–1).
ZenvionConfig
from config_class import ZenvionConfig
cfg = ZenvionConfig.from_pretrained("Darveht/zenvion-voice-detector-v0.4")
print(cfg.num_labels_emotion) # 8print(cfg.num_labels_language) # 50print(cfg.num_labels) # 109 (total across all tasks)
Architecture
Input: raw audio (16 kHz mono)
|
v
[Wav2Vec2 Encoder] — 12 transformer layers, 768-dim hidden
|
v (learned weighted sum of all 13 hidden states, softmax-normalised)
[AttentivePooling] — attention-weighted temporal pooling -> 768-dim
|
+---+--------------------------------------------+
v v v v v v v v
VAD Gender Emotion Lang Age Noise Int Acc
2 3 8 50 6 5 15 20 <- output classes (109 total)
Total parameters: ~95 M (wav2vec2-base) + ~1.6 M (layer weights, pooling, 8 heads)
Inference speed (CPU): ~180 ms per 2-second clip
Inference speed (A100 GPU): ~12 ms per 2-second clip
Training Details
Data
Dataset
Hours
Tasks
CommonVoice 16.1
17,600+
VAD, Language, Accent
FLEURS
4,000+
VAD, Language
VoxPopuli
1,800+
VAD, Language
VoxCeleb / VoxCeleb2
2,400+
Gender, Speaker
RAVDESS + CREMA-D + SAVEE + TESS + IEMOCAP
500+
Emotion
GigaSpeech
10,000+
VAD, Noise
LibriSpeech + LibriLight
60,000+
VAD
Total
~96,300+ hours
—
Hyperparameters
Parameter
Value
Optimizer
AdamW
Learning rate (backbone)
1e-4
Learning rate (heads)
5e-4
LR schedule
Cosine with warmup
Warmup steps
2,000
Batch size
32
Gradient accumulation
4
Epochs
15
FP16
Yes
Gradient clipping
1.0
Weight decay
0.01
Loss
Weighted CrossEntropy per head
Augmentations
Speed perturbation (0.9× – 1.1×)
Additive noise (SNR 5 – 30 dB)
Room impulse response (RIR) convolution
Codec simulation (telephone, mp3)
Random gain (−6 to +6 dB)
Time masking (SpecAugment-style)
Limitations & Bias
Accent detection is English-centric; accuracy drops significantly for non-English accents.
Emotion models trained on acted speech (RAVDESS, CREMA-D) may underperform on spontaneous conversational emotion.
Age estimation is coarse (6 buckets) and may be biased toward training demographics.
Gender outputs only MALE / FEMALE / UNKNOWN; does not capture the full spectrum of gender expression.
Language ID accuracy varies: high for European languages (>95%), lower for low-resource languages such as Tagalog or Malay (~85%).
Min duration: clips shorter than 0.5 s may produce unreliable outputs.
Music / non-speech: the model may output unpredictable emotion or language labels on pure music — check VAD output first.
Changelog
v0.4.1 — 2026-09-09 (code & config fixes)
Critical — weight wipe fixed:post_init() → init_weights() was silently
re-initialising the pretrained wav2vec2 backbone on every model construction,
destroying the pretrained weights. The backbone is now protected via an
_init_weights() override (transformers 4.x) and _is_hf_initialized
marking (transformers 5.x); verified byte-identical against
facebook/wav2vec2-base-960h.
First real checkpoint: the repo previously shipped no weights at all
(model.safetensors / pytorch_model.bin were missing), so from_pretrained()
could not load a working model. Added model.safetensors with the pretrained
backbone + freshly initialised task heads (heads still need training — see
train.py).
Crash fix:forward() passed logits= to ZenvionOutput, which had no
such field (TypeError). Field added; logits is the [B, 109] concat.
Label maps fixed:config.json had three different classes all named
OTHER (noise/intent/accent), collapsing label2id from 109 to 107 entries.
Renamed to NOISE_OTHER / INTENT_OTHER / ACCENT_OTHER; 109/109 consistent.
ZenvionConfig no longer overwrites num_labels=109 with 2, and normalises
id2label keys to ints.
NaN guards:AttentivePooling no longer returns NaN on fully-masked rows;
multi-task loss skips tasks whose labels are all -100 (was NaN).
predict() now decodes all 8 tasks (was 5) and honours threshold.
inference.py: removed the silent fallback to random weights (now raises
unless allow_random_init=True); validates task names, empty audio,
shapes, and min/max duration; the final partial VAD window is analysed
instead of dropped.
train.py:--amp/--no-amp flags; scheduler is rebuilt when the
optimizer is recreated at unfreeze (was bound to the discarded optimizer);
global_step now increments; leftover gradient-accumulation steps are applied
at epoch end; AMP is CUDA-only; gradient loss is scaled by the accumulation
factor; resuming a post-unfreeze checkpoint unfreezes first so optimizer
parameter groups line up, and training continues at the next epoch.
masked_spec_embed NaN fix (transformers≥5): the 960h checkpoint does not
contain this pre-training-only parameter, and from_pretrained materialised
it as uninitialised memory (NaN). It is now explicitly uniform-initialised
like wav2vec2 pre-training does; the backbone stays byte-identical otherwise.
num_labels on transformers≥5: it is a read-only property derived from
len(id2label) — assigning it invoked the property setter and regenerated
id2label as generic LABEL_X entries. The assignment was removed; 109 is
derived from the real label map. predict() now derives every task's label
names from config.id2label (the hardcoded noise list still said OTHER).
evaluation.py: EER is now the threshold minimising |FAR − FRR| (was the
misleading min of (FAR+FRR)/2); AUC is reported raw (was max(auc, 1-auc)).
dataset_loader.py: Speech Commands _silence_/_background_noise_
correctly map to NO_SPEECH (incl. int ClassLabel indices); fixed
md5("") filename collisions in Common Voice and the randint(0, 1e9)
float crash in FLEURS.
Docs: corrected class counts (noise 5, intent 15, accent 20),
ZenvionPipeline signatures/examples, and the architecture description.
Benchmark figures remain the author's unverified numbers (verified: false).
v0.4 — 2025-07-24
Added preprocessor_config.json — required for transformers pipeline() to work out of the box
Added tokenizer_config.json and special_tokens_map.json for AutoTokenizer compatibility
Added vocab.json for tokenizer
Added label_mapping.json with all 8 task labels, thresholds, and metadata
Added CITATION.cff for academic references
Full README rewrite: benchmarks table, architecture diagram, training details, API reference, limitations