Sonave All guides

How accurate are deepfake voice detectors, really?

Every vendor page says 98% or 99%. Meanwhile, deepfake fraud keeps working. Both things are true, because "accuracy" without three qualifiers — on what audio, against which tools, at what false-alarm rate — is a marketing number, not an engineering one.

Why lab numbers collapse in the real world

Three gaps separate a benchmark table from your Tuesday finance call:

  1. The codec gap. Academic benchmarks use clean studio-quality files. Meetings compress voice with the Opus codec, noise suppression and echo cancellation — processing that erases many of the subtle artifacts detectors learn to key on. A model that has never heard meeting audio is scoring a signal it wasn't trained for.
  2. The generator gap. Detectors ace the synthesis tools in their training set. Fraudsters use this year's commercial tools, which weren't. The only number that predicts field performance is catch rate on tools excluded from training.
  3. The threshold gap. Any detector can "catch everything" if it also flags real speakers constantly — and a system that cries wolf gets muted within a week. Catch rate is meaningless without the real-voice accuracy it was measured at.

The size of these gaps is not subtle. In our published testing, a widely-used open-source detector with strong advertised lab numbers catches 1.9% of clips from 27 unseen commercial voice-clone tools played through meeting audio. A model trained specifically for that environment catches 95.2% of the identical clips. Ninety-seven points of difference, on the same audio, from training choices alone.

Four questions that expose any vendor (including us)

  1. "What's your catch rate on synthesis tools outside your training set?" If the answer is the same 99% as the homepage, the tools weren't really unseen.
  2. "Measured through a meeting codec, or on clean files?" If your use case is calls, clean-file numbers are irrelevant.
  3. "At what false-alarm rate on real voices?" Ask for both numbers from the same run.
  4. "Will you publish that, with methodology and a date?" This one ends most conversations.

What honest numbers look like

Here are ours, from the exact model checkpoint serving production, with methodology published: 95.2% catch on 27 unseen commercial tools through meeting audio (commodity detector: 1.9% on the same clips); 94.0% real-voice accuracy through the Opus codec; and on In-the-Wild — the hardest public real-world deepfake benchmark — 58.7% catch at 93.3% real-voice accuracy.

That last number is the one a marketing department would hide, and it's the most important one on the page. Adversarial, real-world deepfakes are genuinely hard; roughly 6-in-10 at a sub-7% false-alarm rate is the current honest state of a specialized detector, and near-zero is the honest state of a generic one. Anyone quoting 99%+ against that class of audio is describing their lab, not your risk.

The design consequence: because no detector is certain, detection should never be the last line of defense. It belongs where a probabilistic signal is valuable — watching every speaker on every call, and converting sustained suspicion into automatic action: an on-screen verdict, a wire-hold webhook that pauses the payment, a forensic report for the follow-up. Verification (callbacks, dual approval) stays in the loop; detection decides when the loop must run.

The bottom line

"How accurate is deepfake detection?" has a real answer, but it's conditional: very accurate against current commercial clone tools when the model is trained for your audio environment; partially effective against the adversarial frontier; and near-useless when a lab model meets meeting audio. Buy accordingly — and ask every vendor the four questions above. We answer them at usesonave.com/benchmarks.

See it on your own voice — free for 5 hours/month