Skip to main content
MLAIA / Audio & Acoustics / Edge Audio AI

Audio & Acoustics · Edge & embedded

Audio models that run on the device.
Within its memory, latency and power.

On-device audio AI moves keyword spotting, sound event detection and audio classification onto phones, microcontrollers, DSPs and dedicated controllers. We engineer the path from microphone and front-end features to a quantized streaming model, and evaluate it on the target hardware with real microphones.

A specialist practice of MLAIA Data ScienceLed by Dr. Yochai Edlitz · Ph.D., Weizmann Institute

Why edge audio is its own discipline

The model is only part of the budget.
The audio pipeline is the rest.

A model that scores well in a notebook can still fail on a device. On constrained hardware, feature extraction, buffering, runtime and power management share one small budget with the network, and each choice changes what the model sees.

Our project experience includes deploying machine-learning and deep-learning models on Android and iOS phones and on dedicated controllers for applications involving drones and earphones. Across platforms, the lesson is the same: keep training and on-device features identical, and test with the shipping microphones. The Audio & Acoustics practice covers the wider signal chain.

01

Keyword spotting & wake words

Always-on detection of a small vocabulary, typically as a cascade: a cheap first stage runs continuously and a larger verifier wakes only on candidates. The design trades false accepts per hour against false rejects and idle power.

02

Sound event detection on the edge

Recognizing alarms, impacts, machinery states or other target sounds locally, without streaming audio off the device. Class definitions, onset handling and post-processing such as median filtering or hysteresis matter as much as the classifier. See the acoustic event detection guide.

03

Learned enhancement in the loop

Small networks that estimate noise-suppression masks or gains run on every frame, so their cost and algorithmic delay add directly to the audio path. They usually complement classical processing; see noise reduction & ANC.

ALWAYS ON · runs every hop · low-power core / DSPWOKEN ON DEMAND · main coremicfront-endlog-mel featuresstage 1: tiny detectorcheap, tuned for recallcandidates onlystage 2: verifierlarger, tuned for precisionactmost audio stops herefalse candidates rejectedIdle power is set by everything inside the blue frame, because it runs on every hop.
A wake-word cascade splits the work by power domain. Only the front-end and a tiny first stage run on every hop; the larger verifier runs on the main core only when the first stage flags a candidate. Idle power, and most false accepts, are decided by how that split is tuned.

Engineering the pipeline

Features, model, runtime.
Engineered as one system.

Most on-device audio failures come from mismatches between stages, not from the network itself. These are the areas we work through with your team.

Front-end feature cost

Log-mel or MFCC extraction means windowing, an FFT, a mel filterbank and a log on every hop. On a microcontroller this can cost as much as a small model. We size the FFT, mel bands and hop, use fixed-point or vendor DSP kernels, and verify on-device features match training numerically.

Streaming inference & latency

Frame length and hop set time resolution and minimum algorithmic delay. Convolutional models can stream by caching internal state, so each hop computes only new frames instead of re-running a full window. Buffering, inference time and decision smoothing go into one latency budget.

Quantization & pruning

Post-training int8 quantization with a representative calibration set comes first; quantization-aware training follows if accuracy drops. We check per-channel scales, activation ranges after the log and operator support. On most small cores, structured pruning and architecture changes cut compute more reliably than unstructured sparsity.

Runtime & deployment target

TensorFlow Lite and TFLite Micro, Core ML, ONNX Runtime and vendor NPU toolchains support different operators and quantization schemes. We choose architectures that convert cleanly, compare converted outputs with the reference, and profile per-layer latency and peak memory on the device.

Always-on power budgets

Continuous listening is dominated by what runs when nothing happens. We look at duty cycling, voice-activity or energy gating, a low-power first stage that wakes a larger core, and sample rate. Power is measured on the board in each operating state.

Microphone & acoustic path

Data recorded on studio or laptop microphones rarely matches a MEMS microphone behind a port and grille. We characterize the device response and noise floor, then augment with measured responses and recorded device noise. Multi-microphone products can add a beamforming front-end.

Re-run the windowevery frame, every hop20 frames per hopStream with cached statereuse earlier work1 frame per hoptime → newest frame on the rightcomputed this hopreused from cache
Streaming inference caches intermediate results so each hop processes only the newest frame instead of the whole window. For convolutional and recurrent audio models this cuts compute per hop by roughly the number of frames in the window, at the cost of state memory and careful conversion.
Engineering considerationsLog-mel / MFCC front-endsint8 & quantization-aware trainingStreaming & stateful inferenceTFLite, Core ML, ONNX RuntimeDuty cycling & wake cascadesOn-target profiling

What we measure

Success is measured on the device.
Not only on the test set.

We agree on metrics before optimizing, and report them for the deployed build, with the real microphone and runtime.

Detection error that matches the use

For wake words, false rejects at a fixed false-accept rate per hour of realistic background audio, not accuracy on balanced clips. For sound events, event-based precision and recall with onset tolerance, and false alarms per day in the field. Results are split by noise type and distance.

Latency from sound to decision

Time from the acoustic event to the device acting on it: buffering, feature hop, inference, smoothing and any core wake-up. We measure it end to end with a reference signal and report the distribution, not an average.

Resource footprint

Flash for weights and code, peak RAM including tensor arena and audio buffers, cycles or milliseconds per hop, and the CPU or DSP headroom left for the rest of the application.

Power & robustness on target

Average current while idle-listening and when active, measured on the board. Robustness tests cover the shipped microphone and enclosure, playback through the device's own speaker, and drift between float, quantized and on-device outputs.

Where the milliseconds go — an example budget25 ms window · 10 ms hop · 3-hop smoothing. Real values come from profiling on your device.25 msbuffer the analysis windowfeatures · 2 ms8 msinference30 msdecision smoothing (3 hops)wake + act · 5 mssound onsetdevice acts ≈ 70 ms01020304050607080 ms
An example budget from sound onset to the device acting. Here the model is a small share: buffering the analysis window and smoothing decisions over several hops take most of the time. Profiling the full chain on the target shows where reductions are actually available.

Achievable numbers depend on your hardware, environment and data. We set targets with you against representative recordings; we do not promise benchmark figures.

How an engagement runs

From a trained model to a working device.
In four stages.

Engagements range from reviewing an existing on-device model to building the pipeline with your firmware or mobile team. A typical path:

01

Budget the device

Agree on target hardware, runtime, latency, memory and power envelopes, and the error measures that matter. Profile a reference model on the device early, so the budget rests on measurements.

02

Match the data to the microphone

Record audio through the product's microphone and enclosure, build a negative set of realistic background audio, and lock one front-end specification for training and firmware.

03

Design for the target

Select architectures that fit the operator set and memory, train with device-realistic augmentation, then quantize and compare float, quantized and on-device outputs.

04

Validate & hand over

Run the agreed tests on production-representative units, and document the pipeline, conversion steps and test harness so your team can retrain and redeploy.

Discuss your device and model

Before we start

Good questions. Straight answers.

Can a useful audio model fit on a microcontroller?

For well-scoped tasks such as keyword spotting or a small set of sound events, often yes. Published keyword-spotting models commonly fit in tens to a few hundred kilobytes after int8 quantization. Whether yours fits depends on the classes, acoustic variability and what the rest of the firmware needs, so we profile a candidate on your hardware early.

Should we run inference on the phone or in the cloud?

On-device inference avoids network latency, works offline and keeps raw audio on the device. The cloud allows larger models and easier updates. Many products split the work: a small on-device model detects or gates, and sends only relevant segments or results onward. We compare both against your latency, privacy, cost and update requirements.

How much accuracy do we lose with int8 quantization?

It varies. Well-calibrated post-training quantization often loses little, but audio models can be sensitive where activation ranges are wide, for example right after the log in a log-mel front-end. We measure float, quantized and on-device outputs on your test set and use quantization-aware training when the gap matters.

Why does our model work in testing but not on the device?

Common causes are feature mismatches between Python and firmware (window, FFT scaling, mel filterbank or log floor), a microphone response and noise floor unlike the training data, and quantization or conversion differences. We trace the same audio through each stage on both sides and compare outputs numerically to find the divergence.

Do we need to record new data with our own hardware?

Usually some. Public datasets help with pre-training and negatives, but a modest set of recordings through your microphone and enclosure, in realistic conditions, often closes the gap between lab and field. We help design the recording protocol and use measured device responses to make existing data more representative.

Audio & Acoustics · Edge Audio AI

Tell us what
you’re working on.

Share the problem, the data you have and what success would look like. We’ll discuss a practical next step.

Please don’t include confidential datasets, credentials or patient information.

yochai@mlaia.com
+972 52 484 6282

Your inquiry is handled under our Privacy Policy.