Most on-device audio failures come from mismatches between stages, not from the network itself. These are the areas we work through with your team.
Front-end feature cost
Log-mel or MFCC extraction means windowing, an FFT, a mel filterbank and a log on every hop. On a microcontroller this can cost as much as a small model. We size the FFT, mel bands and hop, use fixed-point or vendor DSP kernels, and verify on-device features match training numerically.
Streaming inference & latency
Frame length and hop set time resolution and minimum algorithmic delay. Convolutional models can stream by caching internal state, so each hop computes only new frames instead of re-running a full window. Buffering, inference time and decision smoothing go into one latency budget.
Quantization & pruning
Post-training int8 quantization with a representative calibration set comes first; quantization-aware training follows if accuracy drops. We check per-channel scales, activation ranges after the log and operator support. On most small cores, structured pruning and architecture changes cut compute more reliably than unstructured sparsity.
Runtime & deployment target
TensorFlow Lite and TFLite Micro, Core ML, ONNX Runtime and vendor NPU toolchains support different operators and quantization schemes. We choose architectures that convert cleanly, compare converted outputs with the reference, and profile per-layer latency and peak memory on the device.
Always-on power budgets
Continuous listening is dominated by what runs when nothing happens. We look at duty cycling, voice-activity or energy gating, a low-power first stage that wakes a larger core, and sample rate. Power is measured on the board in each operating state.
Microphone & acoustic path
Data recorded on studio or laptop microphones rarely matches a MEMS microphone behind a port and grille. We characterize the device response and noise floor, then augment with measured responses and recorded device noise. Multi-microphone products can add a beamforming front-end.