SwiftF0 (ONNX)

Re-hosted ONNX graph from lars76/swift-f0 (PyPI: swift-f0, v0.1.2), a fast STFT + 2D-CNN fundamental frequency (F0) detector. In the pitch-benchmark suite it beats CREPE on both speed and accuracy.

Upstream ships the model already exported to ONNX โ€” swift_f0/model.onnx โ€” and the Python package is itself just a thin onnxruntime.InferenceSession wrapper (see swift_f0/core.py). This is that exact graph, unmodified, re-hosted for direct download (e.g. by superassp, an R package for acoustic phonetic analysis).

Files

File Precision Size
onnx/model.onnx fp32 398 KB

I/O spec

input:  "input_audio"  float32 [1, audio_length]   # raw 16kHz mono waveform, dynamic length
output: "pitch_hz"      float32 [1, num_frames]      # per-frame F0 estimate, Hz
output: "confidence"    float32 [1, num_frames]      # per-frame voicing confidence, 0-1

Framing (STFT: 1024-sample window, 256-sample hop, 384-sample symmetric padding) is done inside the graph via an ONNX STFT op (opset โ‰ฅ17) โ€” unlike CREPE, you feed raw audio directly, no manual windowing.

  • Native frame rate: fixed 62.5 Hz (16 ms hop), not configurable without re-exporting the graph.
  • Frame i timestamp: t[i] = (i * 256 + 127.5) / 16000 seconds.
  • Model strictly supports 46.875โ€“2093.75 Hz (G1โ€“C7); values outside this range are not meaningful.
  • Minimum input length: 256 samples (shorter audio must be zero-padded).
  • Voicing decision (applied outside the graph, in upstream's Python wrapper): frame is voiced iff confidence > threshold (default 0.9) and fmin <= pitch_hz <= fmax.

Quick start (Python, onnxruntime)

import numpy as np
import onnxruntime as ort

sess = ort.InferenceSession("onnx/model.onnx", providers=["CPUExecutionProvider"])
audio = np.zeros((1, 16000), dtype=np.float32)  # replace with real 16kHz mono waveform
pitch_hz, confidence = sess.run(["pitch_hz", "confidence"], {"input_audio": audio})
voiced = (confidence[0] > 0.9) & (pitch_hz[0] >= 46.875) & (pitch_hz[0] <= 2093.75)

Validation

Unmodified upstream artifact โ€” no conversion step performed, so no PyTorch-vs-ONNX drift check applies (there is no separate PyTorch checkpoint upstream; the ONNX graph is the shipped model). Confirmed to load and run under onnxruntime 1.24.x CPU EP with no unsupported ops.

License

MIT, inherited from upstream lars76/swift-f0 (ยฉ Lars Nieradzik). See LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support