Skip to main content
MLAIA / Audio & Acoustics / Beamforming

Audio & Acoustics · Microphone arrays

Microphone array beamforming.
Designed around physics, then tuned on the device.

Microphone array beamforming lets a product listen in one direction and suppress the rest, but how well it works is largely fixed by geometry, microphone tolerances and the room before any algorithm runs. We help hardware and R&D teams choose the array, the beamformer and the evaluation plan together, so the result holds up on real units in real rooms.

A specialist practice of MLAIA Data ScienceLed by Dr. Yochai Edlitz · Ph.D., Weizmann Institute

What buyers get wrong

Most array problems are decided
before the first line of code.

Teams often fix the microphone count and board layout first, then expect the beamformer to deliver the directivity. Geometry sets hard limits. These are the issues we see most often, and they are where our Audio & Acoustics practice usually starts.

01

Aperture vs low-frequency directivity

Directivity depends on array size relative to wavelength. At 300 Hz the wavelength is over a metre, so a 6 cm delay-and-sum array is close to omnidirectional across much of the speech band. Small arrays can only gain low-frequency directivity through superdirective processing, which brings its own costs.

02

Spacing vs spatial aliasing

To avoid grating lobes, element spacing should stay below half the shortest wavelength of interest: about 2.1 cm for content up to 8 kHz. Wider spacing improves low-frequency resolution but creates false lobes and ambiguous DOA at high frequencies. Many designs need nested or non-uniform layouts to satisfy both.

03

White-noise gain and robustness

Superdirective and differential beamformers look excellent in simulation and then amplify sensor self-noise, wind and small errors on hardware. White-noise gain (WNG) measures that sensitivity. We constrain WNG explicitly, typically with diagonal loading, and accept a little less directivity in exchange for a beamformer that survives production.

04

Microphone mismatch and calibration

MEMS microphones are commonly specified with sensitivity tolerances around ±1 dB, plus phase differences between parts and acoustic ports. Small-aperture arrays are very sensitive to this. Decide early whether you will buy matched parts, calibrate per unit on the line, or estimate mismatch adaptively in the field.

05

Enclosure, ports and the room

Free-field steering vectors ignore the device body, gaskets, port tubes and reverberation. In a reverberant room the target arrives from many directions, and adaptive beamformers can cancel part of the desired speech. Measured or estimated relative transfer functions usually outperform idealised geometry.

Small aperture4 mics · 2 cm spacing · 6 cm across0° look90°180°−90°0 dB-10 dB-20 dB4 kHz — a usable beam500 Hz — almost omnidirectionalSpacing vs. aliasing4 mics at 6 kHz (half-wavelength ≈ 2.9 cm)0° look90°180°−90°0 dB-10 dB-20 dB6 cm spacing — grating lobes at ±72°2 cm spacing — no grating lobes
Delay-and-sum response of a four-microphone linear array, steered to broadside and computed from the array geometry. Left: a 6 cm array forms a beam at 4 kHz but is nearly omnidirectional at 500 Hz. Right: at 6 kHz, 6 cm spacing exceeds half a wavelength and grating lobes appear as strong as the main beam. Every pattern is mirrored front-to-back, because a linear array cannot tell front from back.

Methods

From delay-and-sum to neural front-ends.
Each with a known failure mode.

We start with a fixed beamformer as a baseline and add adaptivity or learned components only where measurements show they help.

Delay-and-sum and filter-and-sum

Time-align channels toward a look direction and sum. Robust, cheap and predictable, with uncorrelated-noise gain that grows with microphone count. Directivity is limited by aperture at low frequencies, so it is a baseline and a fallback rather than the whole answer.

See steering and mic count in the lab →

Superdirective and differential arrays

Optimise for diffuse-noise suppression with closely spaced elements, the usual choice for earphones and compact devices. Differential designs give frequency-invariant patterns such as cardioids, but need equalisation of their low-frequency roll-off and careful WNG control.

Adaptive beamforming: MVDR and GSC

MVDR minimises output power while passing the look direction undistorted; the generalised sidelobe canceller (GSC) implements the same idea with a blocking matrix and adaptive filter. Both need good noise-covariance estimates and steering vectors, or they suppress the target. Voice-activity control and RTF steering are essential.

Direction of arrival (DOA) estimation

GCC-PHAT, SRP-PHAT and subspace methods such as MUSIC locate talkers or sound sources. Reverberation creates spurious peaks, and linear arrays cannot separate front from back. We pair DOA with tracking and confidence measures so steering does not jump to reflections or loudspeakers.

Mask-based neural beamforming

A network estimates time-frequency masks for speech and noise; the masks drive spatial covariance estimates for an MVDR or GEV beamformer. This keeps the linear, low-distortion spatial filter while letting a model decide what is target. It also reduces dependence on explicit geometry.

Integration with AEC and post-filtering

Ordering matters. Echo cancellation per microphone before the beamformer is expensive; a single canceller after an adaptive beamformer sees a changing echo path. Post-filters and dereverberation such as WPE sit downstream. We design the chain as one system, not separate blocks.

wavefrontplane wave from θθd·sinθextra pathmic 1mic 2mic 3mic 4d+3Δ+2Δ+Δ+0Σtarget adds in phaseother directions partly cancelΔ = d·sinθ / cc ≈ 343 m/s
Delay-and-sum in one picture. Sound from angle θ reaches each microphone later by d·sinθ/c. Delaying each channel to remove that offset lines the target up before the sum, while sound from other directions stays misaligned and partly cancels.
A · Cancel echo per microphoneMicsN chAEC × None per micN chBeamformer1 chPost-filteroutto codec / ASRloudspeaker referenceCost grows with N.Each canceller models a fixed loudspeaker-to-mic path.B · Cancel echo after the beamMicsN chBeamformer1 chAEC × 1single canceller1 chPost-filteroutto codec / ASRloudspeaker referenceOne canceller.The echo path it models changes whenever the beam steers.
Where echo cancellation sits changes both cost and difficulty. One canceller per microphone multiplies compute by N; a single canceller after the beam is cheaper but must track an echo path that moves whenever the beamformer steers.

Where arrays are used

Different products.
Different array problems.

The same beamformer behaves very differently on a smart speaker, an earbud and a drone. Product constraints decide which methods are viable.

Smart devices and conferencing

Far-field pickup at several metres, talkers anywhere in the room, a loudspeaker on the same device. Circular arrays support 360° steering; the hard parts are reverberation, double-talk with echo cancellation and fast, stable talker tracking for wake words and ASR.

Earphones and hearables

Two or three microphones a centimetre or two apart, tight power budgets and very low latency targets. Wind noise is largely uncorrelated between microphones, so a directional beamformer can make it worse; wind detection and fallback modes matter as much as the beam.

Robots and drones

Loud, harmonic ego-noise from motors and rotors in the near field, often much louder than the target. Array placement relative to noise sources, reference sensors and spatial nulls toward known noise directions usually do more than adding microphones.

Industrial acoustic monitoring

Localising leaks, faults or abnormal machines in noisy halls. Larger apertures and wider bandwidths are possible, but spatial aliasing and multipath dominate. Beamformed outputs often feed acoustic event detection models, or complement vibration analysis.

What to measure

Judge the array by what the product needs.
Not by the simulated beam pattern.

A clean simulated beam pattern says little about a product. We define measurements up front and run them on representative units, at the distances, angles and noise conditions the device will meet.

Spatial metrics come first: measured beam patterns and directivity index per band, white-noise gain, and sensitivity to steering error and mismatch across a batch of units, not one golden sample. For speech products we then measure what the user and downstream model experience: SI-SDR or SNR improvement, quality estimates, and word error rate or wake-word false-accept and false-reject rates when ASR consumes the output. For localisation we report DOA error and tracking stability in reverberant rooms. Every result is paired with latency, memory and compute on the target processor; see edge audio AI for that side of the budget.

Typical evaluation setBeam pattern & directivity indexWhite-noise gainMismatch sensitivitySI-SDR / SNR gainASR word error rateDOA error & trackingLatency & compute

How an array engagement runs

Simulate. Record. Tune on real units.

Array work goes best when hardware and algorithm decisions are made together, before the board layout is frozen.

01

Requirements & geometry review

Agree target distances, angles, bandwidth, noise scenarios and success metrics. Review mic count, spacing, ports and enclosure against aliasing and low-frequency directivity limits.

02

Simulation & trade study

Model candidate geometries with simulated rooms and tolerance spreads. Compare fixed, superdirective and adaptive beamformers on directivity, WNG and robustness.

03

Multichannel recording & calibration

Record on prototype units in representative rooms and noise. Characterise mismatch, transfer functions and enclosure effects; define a calibration strategy.

04

Implementation & on-device validation

Tune the chosen chain, including DOA, AEC ordering and post-filtering, on the target processor. Validate across units and hand over code, test sets and documentation.

Talk through your array →

Before we start

Good questions. Straight answers.

How many microphones does my product need?

Fewer than many teams assume, if geometry is right. Uncorrelated-noise suppression grows with microphone count, but directivity at low frequencies is driven by aperture, and DOA ambiguity by layout. Two well-placed microphones can outperform four poorly spaced ones. We decide the count from the target scenarios, the bandwidth, the enclosure and the compute budget, then confirm it with simulation and prototype recordings.

Should we use a linear or circular array?

A linear array is simple and works for endfire or broadside pickup, but it cannot distinguish front from back and its resolution varies with angle. A circular or planar array supports uniform 360° steering and unambiguous azimuth, at the cost of more channels and board area. The right choice depends on where talkers or sources can be relative to the device.

Is a neural network better than MVDR?

Often the best systems use both. Mask-based beamforming uses a network to estimate which time-frequency bins are target and noise, then applies a linear MVDR or GEV filter that introduces little distortion, which suits ASR. Fully neural multichannel models can suppress more noise but need more compute and representative training data. We compare options on your recordings and on the target hardware.

Can you work with our existing hardware?

Yes. Many projects start after the geometry is fixed. We characterise what the current array can and cannot achieve, calibrate for mismatch and enclosure effects, and choose algorithms that suit the geometry. If a hardware change would make a large difference, such as a microphone position or port design, we show the expected benefit so you can decide whether it is worth a board revision.

Do microphones need per-unit calibration?

It depends on aperture and beamformer type. Delay-and-sum on a wide array tolerates typical sensitivity spread well. Small-aperture superdirective or differential designs can lose much of their directivity with a fraction of a decibel of mismatch. Options include matched-part grades, end-of-line calibration and online gain estimation. We quantify the sensitivity first, so calibration cost is justified by data.

Can you guarantee a dB of noise reduction?

No. Achievable suppression depends on geometry, the noise field, reverberation and microphone tolerances, and a single dB figure hides most of that. We agree the metrics that matter for your product, such as word error rate or SI-SDR at given distances and noise conditions, and measure them on your hardware. Our general approach is described on the Audio & Acoustics page.

Audio & Acoustics · Beamforming

Tell us what
you’re working on.

Share the problem, the data you have and what success would look like. We’ll discuss a practical next step.

Please don’t include confidential datasets, credentials or patient information.

yochai@mlaia.com
+972 52 484 6282

Your inquiry is handled under our Privacy Policy.