TFRCA-TAE-IARA-H-P8x8-F65xT40-3.88M
TFRCA-TAE is a masked time–frequency reconstruction Transformer autoencoder for passive underwater acoustics. It learns only from IARA H background recordings and assigns a continuous anomaly score to sounds that are difficult to reconstruct under the learned background model.
The checkpoint has 3,880,128 parameters. It uses a 10 s, 16 kHz waveform; a 513×313 log-magnitude STFT padded to 520×320; 8×8 patches; a 65×40 patch grid; separate time and frequency encoders; bidirectional cross-attention; context fusion; and a global Transformer decoder.
Interactive dataset demo: TFRCA-TAE IARA Explorer
What the model does
TFRCA-TAE is an unsupervised reconstruction model, not a vessel classifier. During training, 40% of the spectrogram patches are hidden. The model must reconstruct the hidden content using the remaining time and frequency context.
At inference, each 10 s window is evaluated with five fixed masking patterns arranged so every patch is tested twice. The anomaly score is the mean masked Huber reconstruction error over valid spectrogram bins. For a full IARA recording, four evenly spaced 10 s windows are evaluated and the maximum window score becomes the recording score.
The output is:
- one continuous anomaly score per 10 s window
- a reference threshold of
0.4438585638999939 - a threshold-exceeded flag for convenience.
The flag does not identify a vessel type.
Quick use
pip install "torch>=2.11,<2.14" "numpy>=2.4,<3" "soundfile>=0.13,<0.15" "scikit-learn>=1.8,<2" "huggingface_hub>=1.4,<2"
from pathlib import Path
import sys
from huggingface_hub import snapshot_download
repo = Path(snapshot_download(
"Moon-Young-Choi/TFRCA-TAE-IARA-H-P8x8-F65xT40-3.88M"
))
sys.path.insert(0, str(repo / "src"))
from tfrca_tae.inference import IARAAnomalyDetector
detector = IARAAnomalyDetector(
repo / "tfrca_tae_iara_h_p8x8_f65xt40_3_88m.pt",
device="cpu",
)
result = detector.analyze_file("recording.wav")
print(result.to_dict())
IARA comparison
The retained comparison uses five frozen H recording-level 29/9/9 splits.
| Method | Runs | AUROC | pAUROC at FPR≤0.10 |
|---|---|---|---|
| P — TFRCA-TAE | 15 | 0.9111 ± 0.0445 | 0.7797 ± 0.1065 |
| B0 — PCA reconstruction | 5 | 0.8457 ± 0.0804 | 0.6681 ± 0.1710 |
| B1 — Plain masked Transformer | 15 | 0.8559 ± 0.0446 | 0.6608 ± 0.0525 |
The selected deployment checkpoint itself achieved AUROC 0.8935, pAUROC 0.7661, and average precision 0.9847 on its held-out H-versus-F/G evaluation.
Efficiency
- training time: 396.10 s
- mean model inference per 10 s window: 30.93 ms
- mean four-window recording inference: 123.72 ms
Dataset provenance
IARA v2 was released by the Brazilian Navy Research Institute and collaborators as IARA an Acoustical Recordings Archive, Zenodo DOI 10.5281/zenodo.15777429. The dataset contains 1,825 underwater acoustic recordings collected in the Santos Basin Soundscape Monitoring Project. The official Zenodo metadata identifies the dataset license as CC BY-NC 4.0.
- Downloads last month
- 68


