TFRCA-TAE-IARA-H-P8x8-F65xT40-3.88M

TFRCA-TAE is a masked time–frequency reconstruction Transformer autoencoder for passive underwater acoustics. It learns only from IARA H background recordings and assigns a continuous anomaly score to sounds that are difficult to reconstruct under the learned background model.

The checkpoint has 3,880,128 parameters. It uses a 10 s, 16 kHz waveform; a 513×313 log-magnitude STFT padded to 520×320; 8×8 patches; a 65×40 patch grid; separate time and frequency encoders; bidirectional cross-attention; context fusion; and a global Transformer decoder.

Interactive dataset demo: TFRCA-TAE IARA Explorer

image

What the model does

TFRCA-TAE is an unsupervised reconstruction model, not a vessel classifier. During training, 40% of the spectrogram patches are hidden. The model must reconstruct the hidden content using the remaining time and frequency context.

At inference, each 10 s window is evaluated with five fixed masking patterns arranged so every patch is tested twice. The anomaly score is the mean masked Huber reconstruction error over valid spectrogram bins. For a full IARA recording, four evenly spaced 10 s windows are evaluated and the maximum window score becomes the recording score.

The output is:

  • one continuous anomaly score per 10 s window
  • a reference threshold of 0.4438585638999939
  • a threshold-exceeded flag for convenience.

The flag does not identify a vessel type.

image

Quick use

pip install "torch>=2.11,<2.14" "numpy>=2.4,<3" "soundfile>=0.13,<0.15" "scikit-learn>=1.8,<2" "huggingface_hub>=1.4,<2"
from pathlib import Path
import sys
from huggingface_hub import snapshot_download

repo = Path(snapshot_download(
    "Moon-Young-Choi/TFRCA-TAE-IARA-H-P8x8-F65xT40-3.88M"
))
sys.path.insert(0, str(repo / "src"))

from tfrca_tae.inference import IARAAnomalyDetector

detector = IARAAnomalyDetector(
    repo / "tfrca_tae_iara_h_p8x8_f65xt40_3_88m.pt",
    device="cpu",
)
result = detector.analyze_file("recording.wav")
print(result.to_dict())

IARA comparison

The retained comparison uses five frozen H recording-level 29/9/9 splits.

Method Runs AUROC pAUROC at FPR≤0.10
P — TFRCA-TAE 15 0.9111 ± 0.0445 0.7797 ± 0.1065
B0 — PCA reconstruction 5 0.8457 ± 0.0804 0.6681 ± 0.1710
B1 — Plain masked Transformer 15 0.8559 ± 0.0446 0.6608 ± 0.0525

Retained IARA comparison

The selected deployment checkpoint itself achieved AUROC 0.8935, pAUROC 0.7661, and average precision 0.9847 on its held-out H-versus-F/G evaluation.

Efficiency

  • training time: 396.10 s
  • mean model inference per 10 s window: 30.93 ms
  • mean four-window recording inference: 123.72 ms

Dataset provenance

IARA v2 was released by the Brazilian Navy Research Institute and collaborators as IARA an Acoustical Recordings Archive, Zenodo DOI 10.5281/zenodo.15777429. The dataset contains 1,825 underwater acoustic recordings collected in the Santos Basin Soundscape Monitoring Project. The official Zenodo metadata identifies the dataset license as CC BY-NC 4.0.

Downloads last month
68
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Moon-Young-Choi/TFRCA-TAE-IARA-H-P8x8-F65xT40-3.88M 1