The bug that doesn't crash: on-device speaker embeddings on iOS

I needed to tell people apart by voice in a meeting recording, on the phone, without sending audio anywhere for that step. The interesting part of the work turned out not to be the model. It was that the whole pipeline can be wrong and nothing anywhere will tell you.

No exception, no warning, no NaN. You get a 192-dimension vector of perfectly reasonable-looking floats. It is simply not the vector the model was trained to produce, so matching degrades to a coin flip. If you ship that, your bug report reads "sometimes it mixes up speakers", which is exactly the kind of thing you'll blame on a noisy room for two weeks.

First: Apple doesn't do this

Worth stating plainly because I went looking. There is no speaker recognition API on iOS. The Speech framework transcribes; it has nothing for who is talking. Server transcription APIs often give you diarization within a single session, but that doesn't help when you want the same person recognised across recordings made weeks apart.

So the model comes with you in the bundle.

Picking a model, and why the licence decides it

WeSpeaker publishes several speaker embedding models trained on VoxCeleb. The obvious pick is ResNet34 — more widely used, well documented. It's licensed CC-BY-4.0, which carries an attribution obligation in a commercially sold app.

CAM++ from the same project is Apache-2.0, with accuracy near enough for this job. That decided it. About 28 MB in the bundle.

If you're evaluating ML models for a paid app, check the licence before you benchmark anything. I nearly did those in the wrong order.

ONNX Runtime, not CoreML

The default instinct is to convert to CoreML for the Neural Engine. I didn't, and I'd make the same call again.

Conversion needed a working PyTorch environment, and the only prize was the Neural Engine. But this work happens after the recording stops, on clips a few seconds long. It is not a real-time path. The CPU handles it comfortably, so the conversion step would have bought a faster version of something already fast enough, in exchange for a toolchain I'd have to maintain.

ONNX Runtime ships as a Swift package and the call site is unremarkable:

let outputs = try session.run(withInputs: ["feats": tensor],
                              outputNames: ["embs"],
                              runOptions: nil)

Note the tensor name: feats. Not audio. That's the whole story.

The model does not accept audio

Its input is 80-band Kaldi-compatible filterbank features. Getting from 16 kHz PCM to those features is a specific sequence, and every constant has to match what the model saw in training. Mismatch produces no error — just a different, useless answer.

Here's what has to be right.

1. Don't normalise the audio

This is the one that will get you. Every audio pipeline you've written normalises samples to ±1. Kaldi doesn't. It works in int16 scale, ±32768.

Feed normalised floats and the log energies shift by a constant across every band — and since the log of a scaled value is just an offset, the features look completely plausible. They're describing audio roughly 90 dB quieter than anything in the training set.

2. The frame parameters

Window25 ms — 400 samples
Hop10 ms — 160 samples
FFT512, zero-padded from 400
Mel bands80, low edge 20 Hz
Pre-emphasis0.97
Dithernone

3. Kaldi's pre-emphasis quirk

Kaldi duplicates the first sample rather than assuming zero, so the first coefficient is x[0] - 0.97 * x[0], not x[0]:

var previous = frame[0]          // not 0
for i in 0..<frameLength {
    let current = frame[i]
    frame[i] = current - preEmphasis * previous
    previous = current
}

One sample per frame. At 10 ms hops that's 100 wrong samples a second, quietly.

4. vDSP's real FFT packing

Two Accelerate-specific things that are easy to miss.

A real-input FFT packs DC and Nyquist into the real and imaginary parts of element zero. And vDSP's real transform returns results at twice the mathematical scale, so everything needs halving before you square it:

let half = fftSize / 2
power[0]    = (outRealPtr[0] / 2) * (outRealPtr[0] / 2)   // DC
power[half] = (outImagPtr[0] / 2) * (outImagPtr[0] / 2)   // Nyquist
for k in 1..<half {
    let re = outRealPtr[k] / 2
    let im = outImagPtr[k] / 2
    power[k] = re * re + im * im
}

Forget the halving and every energy is 4× too large — a constant offset after the log, which again looks entirely reasonable.

5. Mean normalisation, and only mean

After the log, subtract each mel band's mean across the time axis. CMN, not CMVN — do not divide by the variance:

for bin in 0..<melBins {
    var sum: Float = 0
    for f in 0..<frames { sum += features[f * melBins + bin] }
    let average = sum / Float(frames)
    for f in 0..<frames { features[f * melBins + bin] -= average }
}

How to actually know it works

Since the failure is silent, correctness can't be argued from reading the code. It has to be measured.

What worked: write the feature extraction in Python first, against the real model, and check the numbers make sense. Two sentences from the same person should score high cosine similarity; two different people should score low. My reference script gave 0.83 for same-speaker and 0.20 for different-speaker pairs.

Then port that script to Swift line by line — not from the paper, not from a blog post, from the thing you measured. Keep the script in the repo. Any time someone touches a constant in the filterbank, it runs again.

The comment above my FilterBank says this outright: every constant here must match training exactly; a mismatch doesn't throw, it silently produces meaningless vectors and matching degrades to random. Future me needed that written down.

Two smaller things that mattered

A minimum clip length. Embeddings from short audio are unreliable in a specific way: how a word sounds depends more on the word than on the speaker. Below about three seconds you're clustering vocabulary, not people. I refuse to produce an embedding under that threshold rather than produce a bad one.

L2 normalise the output. Then cosine similarity is a plain dot product, and every downstream comparison gets simpler and faster.

What I'd tell someone starting this

The model is the easy part — it's a file and two lines of ONNX Runtime. The work is the twenty lines of signal processing in front of it, where a constant that disagrees with the training setup produces output that is wrong but looks fine.

So build the measurement before the implementation. Two audio clips of yourself, one of someone else, a reference script, three cosine similarities. If those three numbers aren't what they should be, nothing downstream can be trusted — and without them, nothing downstream can be checked at all.

Context: this came out of building Ses Notu, an iPhone app that transcribes in-person meetings and separates who said what. I'm an independent developer; it's just me. The speaker identification runs on the device as described above, though audio does go to a transcription service for the text itself — worth being precise about that rather than implying everything is local.