File size: 6,299 Bytes
05fd534
 
 
 
 
 
 
2f2db3e
05fd534
 
10ede24
 
 
 
945cdd4
 
8085015
10ede24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8085015
10ede24
 
 
8085015
10ede24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8085015
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
---
title: PI-AcousticNeRF
colorFrom: blue
colorTo: indigo
sdk: docker
pinned: false
license: mit
short_description: 3D acoustic field reconstruction with physics constraints
---

# Physics-Informed Acoustic NeRF

Reconstruct a continuous 3D sound pressure field from sparse microphone measurements, with the Helmholtz wave equation enforced as a hard constraint during training.

[![HuggingFace Spaces](https://img.shields.io/badge/🤗%20HuggingFace-Live%20Demo-blue)](https://huggingface.co/spaces/Pooya-sh1998/acoustic-nerf)

On **Apple M4 Pro (48GB VRam) · No cloud GPU**

---

## What It Does

A microphone hears one point. This model hears the whole room.

Given a handful of microphone measurements at multiple frequencies, the network learns a function `(x, y, z, ω) → complex pressure p̂` that predicts acoustic pressure at any point in 3D space — not just where the mics are. Physics enforcement means the predicted field actually obeys the wave equation, not just the training data.

---

## Architecture

```
Input: (x, y, z, ω)

   Positional encoding
   L=10 spatial, L=6 frequency → 76 dims

   SIREN MLP  ──── skip at layer 4 ────┐
   8 layers, hidden_dim=256            │
   sine activations (ω₀=1.0)          │
         │◄───────────────────────────┘
   Output: (Re(p), Im(p))   ~500k params

Training losses:
  L_data     = MSE at microphone positions
  L_physics  = |∇²p̂ + k²p̂|²  at collocation points  (Helmholtz)
  L_boundary = |∂p̂/∂n + jkβp̂|²  at room walls       (impedance BC)
  L_total    = λ_d · L_data + λ_p · L_physics + λ_b · L_boundary
```

Physics weight schedule: `λ_physics` ramps from `1e-7 → 1e-3` over 3000 epochs so the data loss stabilises before the constraint kicks in.

---

## Results

| Config | Val NMSE | Phys. Residual | Notes |
|---|---|---|---|
| Baseline (data only) | 1.783 | 1124 Pa²/m⁴ | Noisy, no wave structure |
| + Boundary BC only | 1.140 | 1146 Pa²/m⁴ | No improvement |
| + Helmholtz only | **0.969** | **1.71 Pa²/m⁴** | Wave rings appear |
| Full (Helmholtz + BC) | 0.983 | **1.56 Pa²/m⁴** | Lowest residual |

**Physics residual reduction: 99.9%** (baseline 1124 → physics-informed 1.56 Pa²/m⁴)

Room: 6 × 5 × 3 m shoebox · 2000 microphones · 4 frequencies (500/1k/2k/4kHz)  
Training: ~45 min on Apple M4 Pro (MPS backend)

---

## Interactive Viewer

The web viewer renders a horizontal slice at z = 1.5 m with a diverging RdBu colormap (blue = low pressure, red = high). Toggle between physics-informed (A) and baseline (B) models to see the difference.

```bash
# Generate field data from trained checkpoints
python -m src.visualization.export

# Start the FastAPI server + viewer
uvicorn src.api.main:app --host 0.0.0.0 --port 8000

# Open browser
open http://localhost:8000/viewer/index.html
```

Features: frequency toggle (500 Hz / 1 kHz / 2 kHz / 4 kHz), model A/B comparison, microphone overlay, Helmholtz residual overlay.

---

## Quick Start

```bash
# 1. Environment
conda create -n acoustic-nerf python=3.11
conda activate acoustic-nerf
pip install -r requirements.txt

# 2. Generate synthetic room data and train
python -m src.training.trainer --config experiments/configs/physics_informed.yaml

# 3. Run ablation studies
python -m src.training.trainer --config experiments/configs/baseline.yaml
python -m src.training.trainer --config experiments/configs/ablation_physics_only.yaml

# 4. Collect metrics
python -m src.training.collect_metrics

# 5. Export and serve
python -m src.visualization.export
uvicorn src.api.main:app
```

**Requirements:** PyTorch (MPS or CUDA), pyroomacoustics, FastAPI, Three.js (CDN)

---

## The Physics

The Helmholtz wave equation for steady-state acoustics:

```
∇²p(x,y,z) + k²p(x,y,z) = 0
```

where `k = ω/c` (wave number, c = 343 m/s). Enforced as a residual loss at random collocation points throughout the room volume each training batch. Uses finite-difference Laplacian (MPS-stable) rather than autograd second derivatives.

Boundary condition at room walls:

```
∂p/∂n + jkβp = 0
```

where β = 0.4 is the specific admittance (absorption coefficient).

---

## Key Implementation Decisions

**SIREN activations (ω₀=1.0, not 30):** Sine activations give smooth second derivatives, critical for the Laplacian computation. ω₀=30 was tested but caused the physics loss to blow up (second derivatives O(1e9) with 8 stacked layers). Lowered to 1.0 since positional encoding already handles high frequencies.

**Finite difference Laplacian:** PyTorch autograd `create_graph=True` is unstable on MPS for second derivatives. Central differences with h=5e-3 give residual error ~0.05 on plane waves (acceptable). Coordinate scaling: multiply each dimension's contribution by (2/L_dim)² to convert from normalized to physical metres.

**2000 microphones:** The model with 128 mics produced visually noisy fields despite good physics metrics. *99.9%* Helmholtz residual reduction but no visible wave structure. Root cause: ~800:1 param-to-data ratio. pyroomacoustics computes all mic RIRs in a single call regardless of count (128 mics → 2000 mics costs 1.2s extra). 2000 mics × 4 frequencies = 8000 samples drops the ratio to ~62:1 and produces clear wave rings.

---

## Project Structure

```
src/
  data/        generate_synthetic.py, dataset.py
  model/       network.py (SIREN MLP), encoding.py, losses.py
  physics/     helmholtz.py, boundary.py, collocation.py
  training/    trainer.py, scheduler.py, metrics.py
  visualization/ field_viz.py, export.py
  api/         main.py, routes.py

viewer/        Three.js web viewer (index.html, field_renderer.js, controls.js)
experiments/   configs/*.yaml, results/ (checkpoints, metrics.json)
docs/          raw_log.md, article_draft.md
```

---

## References

- [SIREN: Implicit Neural Representations with Periodic Activation Functions](https://arxiv.org/abs/2006.09661)
- [Physics-Informed Neural Networks (Raissi et al.)](https://arxiv.org/abs/1711.10561)
- [NeRF: Representing Scenes as Neural Radiance Fields](https://arxiv.org/abs/2003.08934)
- [pyroomacoustics](https://github.com/LCAV/pyroomacoustics)

---

*Pooya Esfahani - April 2026*