Most detection vendors publish one big lab number. We publish what our deployed production model actually measures — on real meeting audio, on tools it has never seen, with the weak spots stated in the same table.
Deployed checkpoint sonave-xlsr-meet · benchmarked 2026-08-12 · evaluation code: src/eval_xlsr.py
| Test | Sonave | Commodity detector* | What it measures |
|---|---|---|---|
| 27 unseen commercial voice-clone tools MLAAD subset, all tools absent from training, played through meeting audio |
95.2% | 1.9% | Catch rate on the clone tools an attacker would actually use today, in call-quality audio. Same clips for both models. |
| Real voices through the Opus meeting codec Unseen real-speech corpus, Opus-24k pipeline |
94.0% | — | Accuracy on genuine speakers — how rarely we cry wolf on the audio meetings actually produce. |
| In-the-Wild (hard public deepfakes) The hardest public real-world benchmark |
58.7% | 4.0% | Catch rate at 93.3% real-voice accuracy. This is our honest ceiling — see below. |
| In-the-Wild through Opus-24k Same benchmark, through the meeting pipeline |
59.3% | — | Catch rate at 94.0% real-voice accuracy — the model holds through the codec it was built for. |
*A widely-used open-source deepfake detector, evaluated on the identical clips. It advertises strong lab numbers; on unseen commercial tools through meeting audio it collapses — which is why "99% accurate" marketing tells you nothing about your next call.
sonave-xlsr-meet, XLS-R + SLS head, fine-tuned on meeting-codec audio) — not a lab build.Fraud teams have been burned by accuracy marketing. If a vendor won't tell you their catch rate on tools outside their training set, through a meeting codec, at a stated false-alarm rate — they're telling you a lab story. We'd rather you know exactly what you're buying, because the teams that compare honestly are the teams we win.
Try it on your next call — free