COMPUTATIONAL ACOUSTICS · INDEPENDENT STUDY

Sound, space,
and reconstruction.

An independent study following sound from an acoustic PDE to listening: finite elements, inverse problems, reduced models, interactive acoustics, modern music and historical recordings.

WaveFEMListeningInverseROMRecordings
INTERACTIVE LISTENING Sound Lab Eight experiments: timbre · beating · spectra · density · space
FEATURED CASE STUDY · A RECORDING WITH A LONG AFTERLIFE A Beautiful Error

Recorded in wartime Vienna in December 1944, Furtwängler's Eroica later resurfaced as a controversial Urania LP, became the subject of a court case, and travelled through different tape, disc and digital-transfer lineages. The page follows that tangled history before asking how later editions changed what listeners actually heard.

1944 Vienna → Urania → Green Door / Tahra / Melodiya → listen to the three editions →
Urania Eroica cover Tahra FURT 2008 cover Melodiya Eroica cover
INTERACTIVE LISTENING Sound Lab Eight small experiments on pitch, timbre, beating, density and space Open →
SECOND CORE OF THE PROJECT Modern Music Nebula

A large field guide to sound, perception, material, space and technology after 1911 — with composers, works, aesthetics, acoustic analysis and listening experiments.

Enter the Nebula ↗

NUMERICAL ROOM USED THROUGHOUT

The room used in every calculation

Unless a section says otherwise, every result below uses this same 6 m × 4 m two-dimensional room.

Room6 m × 4 m
Sound speed343 m/s
WallsRigid (Neumann)
Loss factorη = 0.03
SourceGaussian at (1.2, 2.0) m
Source widthσ = 0.12 m

This is intentionally a simple 2D model. It is useful for testing the numerical ideas below, but it is not a full model of a real concert hall or a vibrating structure.

WAVE EQUATION → HELMHOLTZ → FEM

Turn the wave equation into a linear system

We start with acoustic pressure in space and time. A Fourier transform lets us work one frequency at a time. Finite elements then replace the continuous pressure field by a finite list of nodal values.

Pressure in time\[\frac{1}{c^2}p_{tt}-\Delta p=s\]
One frequency\[-\Delta P-k^2P=S\]
Numbers on a mesh\[A(\omega)u=b\]
What the last equation means

The mesh has one unknown pressure value at each node. The matrix A(ω) tells us how those nodes are coupled at frequency ω; b is the source; solving the linear system gives the complex pressure field u.

The matrix used here is \(A(\omega)=K-k^2(1-i\eta)M\), with \(k=\omega/c\).

Show the mathematical stepswave equation → weak form → FEM matrix
1

Move from time to frequency

Use \(P(x,\omega)=\int p(x,t)e^{-i\omega t}\,dt\). A second time derivative becomes multiplication by \(-\omega^2\):

\[\mathcal F[p_{tt}]=-\omega^2P.\]

So the wave equation becomes

\[-\Delta P-k^2P=S,\qquad k=\omega/c.\]
2

Reduce the derivative requirement

Multiply by a test function \(\overline v\), integrate over the room, and integrate the Laplacian by parts:

\[\int_\Omega \nabla P\cdot\nabla\overline v-k^2(1-i\eta)\int_\Omega P\overline v=\int_\Omega S\overline v.\]

Rigid walls mean \(\partial P/\partial n=0\), so the boundary term vanishes.

3

Replace the field by basis functions on the mesh

Write \(P_h=\sum_j u_j\phi_j\) and test with \(\phi_i\). This gives the stiffness matrix \(K\), mass matrix \(M\), and source vector \(b\):

\[[K-k^2(1-i\eta)M]u=b.\]
Pressure magnitude at 100 Hz
At 100 Hz, the colour shows the pressure magnitude at every mesh node. Bright areas have larger \(|P|\), dark areas smaller \(|P|\). The source position is marked by the white dot.
Implementation check

Before using the 2D room, the basic FEM assembly was tested on a 1D Helmholtz problem with a known solution. The error decreased approximately like \(h^2\) as the mesh was refined. This checks the implementation of the simple FEM machinery; it does not prove that the 2D room is a complete physical model.

Common questions about this section

Why is the pressure complex?

At one frequency we need both amplitude and phase. A complex number stores both in one quantity. The physical time signal is real; the complex representation is a convenient frequency-domain description.

Why make the walls perfectly rigid?

It gives a clean first model: no normal velocity through the boundary. Real concert-hall walls absorb and scatter sound in frequency-dependent ways, so this assumption is deliberately simplified.

What is the loss factor η = 0.03?

A small phenomenological damping term. It keeps resonances finite. It is not a measured wall-loss model.

Why only two dimensions?

Because the goal is to make the numerical chain inspectable. A realistic hall would need 3D geometry, real boundary impedances and much more computation.

FIELD → LISTENING POSITION

How strong is each frequency at three listening positions?

For each frequency, we solve the pressure field in the whole room and then read the complex pressure at Seat A, Seat B and Seat C. Repeating this from 20 to 400 Hz gives one frequency-response curve for each seat.

A(2,1)
B(3.5,2)
C(5,3)
6 m × 4 m room
1

Choose one frequency, for example 100 Hz.

2

Solve the whole room to get the pressure field at that frequency.

3

Read the pressure at A, B and C. Repeat for every frequency.

Receiver frequency responses
Each line shows how strongly one seat responds at each frequency. Every curve is divided by its own maximum, so compare the locations of peaks and dips rather than absolute loudness.
What is H(f)?

For one seat, \(H(f)\) is simply a frequency-by-frequency record of what the room does to the sound there: how much each frequency is changed in amplitude and phase.

\[H(f)\xrightarrow{\mathcal F^{-1}}h(t),\qquad y(t)=h*x\quad\Longleftrightarrow\quad Y(f)=H(f)X(f).\]
Why does a frequency response become an impulse response?H(f) → h(t) → sound
1

At one frequency

For a linear time-invariant model, the output spectrum is the input spectrum multiplied by the room response:

\[Y(f)=H(f)X(f).\]
2

Return to time

The inverse Fourier transform of \(H\) is the impulse response:

\[h(t)=\mathcal F^{-1}[H(f)].\]

The same operation in time is convolution: \(y=h*x\).

3

What the first listening test really used

The computed response only covered 20–400 Hz. It was interpolated, tapered near the band edges and transformed back to time. Because the frequency spacing is 2 Hz, the corresponding periodic time window is \(1/\Delta f=0.5\) s.

Before trusting the audio, check whether the mesh is fine enough

Higher frequencies have shorter wavelengths. If the mesh spacing stays fixed, fewer mesh points describe each wavelength, so the numerical solution becomes less reliable.

Resolution error versus frequency
The coarse and medium meshes are compared with the finer h = 0.02 mesh at the same nodes. The error grows as frequency rises. The fine mesh is only a numerical reference, not exact truth.
THREE-HARMONIC TEST

The same C3 note at three seats

The note contains only its first three harmonics. This makes the change in harmonic balance easy to hear without presenting it as a full-band room simulation.

Dry

Seat A

Seat B

Seat C

In this three-harmonic test, A has the weakest upper harmonics and sounds darkest; B has the strongest second harmonic and sounds brightest; C lies between them. Each seat is normalised to its own fundamental, so this compares colour rather than loudness.

REAL CONCERT HALLS

Do real concert halls also sound different from seat to seat?

Yes. Our small 2D room cannot tell us where to buy a ticket, but the basic question is real: move the listener and the balance of direct sound and reflections changes. Real hall studies therefore measure many source and receiver positions rather than one single “room response”.

Why compare these two halls?

They are almost opposite answers to the same problem. Amsterdam's Concertgebouw is a late-19th-century shoebox: long, rectangular, stage at one end, and designed before modern room-acoustics science existed. Hamburg's Elbphilharmonie is a 2017 vineyard: the stage sits near the centre and the audience climbs around it in terraces. Both are famous, but they create and distribute reflections in very different ways.

stage
AMSTERDAM · OPENED 1888 · SHOEBOX

Royal Concertgebouw

Why it is unusual: its Main Hall became famous before acousticians had modern measurement tools or computer models. The Concertgebouw says the design borrowed from successful older halls, especially Leipzig's old Gewandhaus, and later renovations tried to preserve the original geometry and finishes.

What that means acoustically: the long parallel side walls of a shoebox send strong reflections back toward the audience. The Concertgebouw's own history describes a characteristically warm sound and an occupied reverberation time of about 2.2 s.

Seat position still matters: Bradley's 1991 measurements of the Concertgebouw, Vienna Musikverein and Boston Symphony Hall explicitly studied how acoustic quantities changed when the source or receiver moved.

stage
HAMBURG · OPENED 2017 · VINEYARD

Elbphilharmonie Grand Hall

Why it is unusual: the audience surrounds the stage and no seat is more than about 30 m from the conductor. Nagata Acoustics broke the audience into small terraces, designed reflecting surfaces around those terraces, and tested the hall with a 1:10 physical model. More than 10,000 individually milled “white skin” panels scatter and redirect sound.

Why people call it revealing: the hall's own acoustics guide describes the sound as very transparent and says that tiny details are easy to hear — including a fluffed note or audience noise. So the listener comment that the hall can feel “almost too revealing” has a real acoustic basis; the phrase “too good” is subjective, but the transparency is not invented.

Seat position is especially obvious: the official guide says seats beside or behind the stage can emphasise whichever instrument group is nearest, while higher seats tend to give a more balanced overall impression.

CONCERTGEBOUWLong shoebox · side-wall reflections · warm blend

A classic hall whose successful acoustics were refined largely by experience and later measurement.

ELBPHILHARMONIECentral stage · terraces · very high clarity

A modern hall whose geometry and surfaces were deliberately engineered to bring a large audience close to the stage.

If we really wanted to ask “which seat is best?”, what would we measure?
How loud?sound strength · G How clear or blended?clarity · C80 How long does the sound hang in the room?EDT / reverberation time How much sound reaches us from the sides?LF / binaural measures such as IACC

There is no single universal “best seat”. A close seat may give more detail; another position may give a smoother blend or stronger sense of space. The best answer depends on the hall, the music and what the listener values.

Common questions about this section

Does a high peak in the response mean a good seat?

No. A peak means that frequency is strongly amplified relative to the rest of that seat's response. A good seat is not defined by one large resonance.

Why normalise every receiver curve separately?

To compare spectral shape. It deliberately removes absolute level differences, so the plot cannot tell us which seat is loudest.

Why does moving the seat change the response?

At a given frequency, direct and reflected waves arrive with different phases. Changing position changes how they add or cancel, so peaks and dips move.

Why only 20–400 Hz?

This was a small numerical demonstration. At higher frequencies the wavelength becomes shorter and the mesh would need to become much finer, which is exactly what the resolution audit shows.

MEASUREMENTS → UNKNOWN RESPONSE / FIELD

Can a few measurements tell us what the room is doing?

A forward problem starts with a model and predicts what microphones would measure. An inverse problem starts with measurements and tries to infer something we did not measure directly.

FORWARD

Known: room + source

Compute: microphone signal

INVERSE

Known: microphone signal

Infer: response or whole field

First, a small inverse problem: estimate one transfer function

Suppose the input spectrum \(X\) is known and the measured output is \(Y=HX+\varepsilon\). The tempting estimate \(H=Y/X\) becomes unstable wherever \(|X|\) is very small, because the noise is divided by a tiny number.

Measurement\[Y=HX+\varepsilon\]
Unstable division\[H=Y/X\]
Regularised estimate\[\widehat H_\lambda=\frac{\overline XY}{|X|^2+\lambda}\]
Show the Tikhonov derivationwhy λ stabilises the division
1

Penalise both data mismatch and an excessively large estimate

\[J(H)=|XH-Y|^2+\lambda|H|^2.\]
2

Set the derivative to zero

\[(|X|^2+\lambda)H=\overline X\,Y.\]
3

Solve for H

\[\widehat H_\lambda=\frac{\overline X\,Y}{|X|^2+\lambda}.\]

Larger \(\lambda\) suppresses the instability more strongly, but also shrinks the estimate and introduces bias.

Inverse reconstruction
In this synthetic test, direct division becomes noisy where the input is weak. Regularisation gives a much closer estimate. The best λ shown here uses the known synthetic truth, so it is an oracle demonstration rather than a practical tuning rule.

First reduce the problem: learn a small set of recurring room patterns

Suppose the FEM field has 15,251 nodes. Eight microphones give us only eight numbers at one frequency. We clearly cannot treat all 15,251 pressures as unrelated unknowns.

The key observation

We have already solved this same room at many nearby frequencies. Those fields look different, but they are not random: the same kinds of spatial shapes keep reappearing. POD learns a small set of whole-room patterns that can be mixed together to describe those familiar fields.

1 · START WITH MANY SOLVED FIELDS 69 complete FEM maps

40, 45, 50, …, 380 Hz. Each map contains 15,251 complex nodal values.

2 · FIND THE RECURRING BUILDING BLOCKS POD keeps a small library of whole-room patterns

Think of them as spatial building blocks. Each block already covers the entire room.

3 · DESCRIBE A FIELD BY A RECIPE Mix the same patterns with different weights
0.8 × pattern 10.3 × pattern 2+0.5 × pattern 3+ …

In this project we keep 40 patterns. A familiar field is then described by 40 weights instead of 15,251 separate nodal values.

before15,251 nodal values
after POD40 weights
\[u\approx V_ra.\]

Plain English: the columns of \(V_r\) are the 40 whole-room patterns; the entries of \(a\) are the 40 mixing weights. Multiply each pattern by its weight and add them together to get an approximate field.

If you want the matrix version: where do the 40 POD patterns come from?

1. Put all solved fields into one large table

Write each full FEM field as one long column of 15,251 values. Put the 69 columns side by side. That table is the snapshot matrix \(U\).

2. SVD is the sorting tool

We factor the snapshot matrix as \(U=V\Sigma W^*\). You do not need to think of this as a new physical model. It is a linear-algebra tool that finds directions that repeatedly appear in the snapshot data and orders them by how much of the snapshot variation they capture.

3. The columns of V are the candidate whole-room patterns

Each column \(v_i\) has one value at every FEM node, so each \(v_i\) is itself a complete spatial pattern over the room.

4. POD basis = keep the first r of those patterns

With \(r=40\), \(V_r=[v_1,\ldots,v_{40}]\). This collection of 40 columns is the POD basis used later in both sparse reconstruction and the reduced-order model.

5. What does “capture” mean?

The singular values in \(\Sigma\) tell us how much of the snapshot dataset is represented by the leading patterns. In this website that is a Euclidean coefficient-norm compression measure; it is not physical acoustic energy.

Spread the microphones around the room

We want a simple repeatable layout, not a claim of “optimal microphone placement”. Start near the centre. Each time we add a microphone, put it at the candidate point that is currently farthest from all microphones already chosen. This avoids wasting many sensors in the same corner.

1 2 3 4 5 6 7 8
first 8 microphone positions
1

Pick a starting point.

2

Look for the place farthest from the microphones we already have.

3

Put the next microphone there.

4

Repeat, so the sensors gradually cover the room.

8 ⊂ 16 ⊂ 32 ⊂ 64

The formal name is farthest-point selection. The 8-, 16-, 32- and 64-microphone layouts are nested: the larger set keeps all microphones from the smaller set.

The microphones now estimate only the 40 mixing weights

This is the important simplification. The microphones are not trying to guess every pressure value in the room directly. They only help us choose the 40 weights that best reproduce the measured microphone values.

MICROPHONES8 / 16 / 32 / 64 measured values

What we actually know.

ESTIMATE THE RECIPEa₁, a₂, …, a₄₀

How much of each POD pattern?

MIX THE 40 PATTERNSu ≈ Vᵣa

Build an approximate whole-room field.

Show the matrix version of the microphone reconstruction

What does C do?

Nothing mysterious: \(C\) simply picks out the rows of a whole-room field that correspond to microphone positions. So \(CV_ra\) means “build the field from the POD patterns, then read that field only where the microphones are”.

What are we solving?

We choose the weights \(a\) so that the predicted microphone values are close to the measured values \(y\), while a small regularisation term keeps the fit from becoming unstable in noise.

\[y\approx CV_ra+\varepsilon,\qquad \widehat a=\arg\min_a\|CV_ra-y\|_2^2+\lambda\|a\|_2^2.\]
Sparse field reconstruction
At 252.5 Hz, 8 microphones reproduce only part of the room pattern; 32 are better but still visibly wrong; in this particular layout, 64 microphones give a field close to the full FEM result. The number 64 is not a magic threshold.
Why checking only the microphone positions can fool us

A reconstructed field can match the measured points very well and still be wrong in the spaces between them. We care about the whole field, so the final comparison must be made against the full FEM field, not only against the sensor readings.

Questions a first-time reader may have

If there are 40 weights, why can’t 8 microphones determine them?

Imagine 40 sliders but only 8 readouts. Many different slider settings can produce the same 8 numbers. In linear-algebra language, there is a nullspace.

Then why not exactly 40 microphones?

Forty readings for forty weights leaves no spare information. If two POD patterns happen to look similar at those microphone positions, or if the measurements contain noise, a small error can cause a large change in the recovered weights. This is what “poor conditioning” means in practice.

Why can more than 40 microphones help?

Extra microphones give redundant information. Instead of forcing 40 equations to determine 40 weights exactly, we can find the set of weights that best fits many measurements at once, which is usually more robust to noise.

Does 64 microphones always work?

No. It works well in this one example with this basis, layout, frequency and noise level. A bad sensor layout can still miss important information, and optimal sensor placement is a separate research problem.

THE SAME POD PATTERNS, USED EARLIER

Use the 40 patterns to make the solve itself smaller

Section 03 used microphone data to choose the 40 POD weights. Here we ask a different question: can the Helmholtz equation itself choose those 40 weights, so we never have to solve for all 15,251 nodal pressures online?

SECTION 03 · RECONSTRUCTIONmicrophone readings → 40 weights → whole field

The measurements choose the recipe.

SECTION 04 · ROMHelmholtz equation → 40 weights → whole field

The governing equation chooses the recipe.

FULL FEM15,251 unknown nodal pressures

The computer solves for every nodal value in \(u\).

→ keep only combinations of the 40 POD patterns →
RANK-40 ROM40 unknown weights

The computer solves only for \(z_1,\ldots,z_{40}\), then rebuilds the field.

\[u\approx V_rz.\]

Plain English: we assume the new field can be made from the same 40 spatial building blocks learned from the snapshot fields. The only unknowns are how much of each block to use.

Show the matrix version: how does Galerkin turn the large FEM system into a 40 × 40 system?

1. Start with the full FEM equation

\[A(\omega)u=b.\]

The full unknown vector \(u\) has 15,251 entries on this mesh.

2. Restrict the answer to the POD space

\[u\approx V_rz.\]

Now the field is controlled by only 40 numbers in \(z\).

3. Ask the reduced field to satisfy the original equation as well as it can inside that 40-pattern space

\[V_r^*A(\omega)V_rz=V_r^*b.\]

This is the Galerkin projection. A useful plain-language picture is: after choosing only 40 allowed building blocks, we make sure the remaining equation error has no component along those same 40 directions.

4. For this Helmholtz model, precompute the reduced matrices

\[K_r=V_r^*KV_r,\qquad M_r=V_r^*MV_r,\qquad b_r=V_r^*b.\]
\[[K_r-k^2(1-i\eta)M_r]z=b_r,\qquad u_{ROM}=V_rz.\]

The online system is now 40 × 40 rather than 15,251 × 15,251.

ROM error versus frequency
We test the ROM at frequencies halfway between the frequencies used to build the POD basis. Rank 40 stays closest to the full h = 0.04 FEM result. This tests interpolation inside 40–380 Hz; it does not show that the model will work at completely new frequency ranges.
Full FEM and ROM fields
At 252.5 Hz, the rank-40 ROM produces almost the same spatial pattern as the full h = 0.04 FEM solve. This comparison checks the reduced model against the larger numerical model, not against a physical room measurement.
What the ROM saves — and what it does not fix

It saves online algebra: 40 unknown weights instead of 15,251 nodal pressures. It does not improve the mesh, change the PDE, or repair any error already present in the full FEM model.

Questions a first-time reader may have about POD and ROM

Isn’t POD already a reduced model?

POD gives us a reduced coordinate system: the 40 building blocks. In Section 03 we still used full solved fields to learn those blocks and then used microphone data to infer their weights. A Galerkin ROM goes one step further and uses the blocks inside the governing equations before the full solve.

Why is the coefficient called a in Section 03 and z here?

They play the same kind of role — both are weights on the POD patterns — but they come from different information. \(a\) is fitted to microphone measurements; \(z\) is solved from the reduced Helmholtz equations.

Why can resonance still be difficult?

Near a resonance, a small change in frequency can cause a large change in the field. A fixed set of 40 patterns may then need extra patterns or more careful stability treatment to follow that rapid change accurately.

Does a very accurate ROM prove that the physical model is correct?

No. It only proves that the ROM reproduces the chosen full FEM model well. If the full model has modelling or mesh error, the ROM can faithfully reproduce that error too.

Scope

This is a small 2-D numerical room. It is not a model of a measured space.

When meshes are compared, h = 0.02 is only the finest mesh used here. It is not an exact answer.

The microphone test uses synthetic data. The reduced model is checked against the full FEM solve, not against a real room.

Structural vibration and a full 3-D room model are outside this study.

SOUND LAB · LIVE

Eight interactive listening experiments

Eight interactive experiments now connect harmonic structure, beating, temporal fusion, spectral deformation, nonlinearity, sound masses and spatial hearing to concrete listening questions.

Open Sound Lab ↗

Timbre · perception · modern music · space