Every vendor page says 98% or 99%. Meanwhile, deepfake fraud keeps working. Both things are true, because "accuracy" without three qualifiers — on what audio, against which tools, at what false-alarm rate — is a marketing number, not an engineering one.
Three gaps separate a benchmark table from your Tuesday finance call:
The size of these gaps is not subtle. In our published testing, a widely-used open-source detector with strong advertised lab numbers catches 1.9% of clips from 27 unseen commercial voice-clone tools played through meeting audio. A model trained specifically for that environment catches 95.2% of the identical clips. Ninety-seven points of difference, on the same audio, from training choices alone.
Here are ours, from the exact model checkpoint serving production, with methodology published: 95.2% catch on 27 unseen commercial tools through meeting audio (commodity detector: 1.9% on the same clips); 94.0% real-voice accuracy through the Opus codec; and on In-the-Wild — the hardest public real-world deepfake benchmark — 58.7% catch at 93.3% real-voice accuracy.
That last number is the one a marketing department would hide, and it's the most important one on the page. Adversarial, real-world deepfakes are genuinely hard; roughly 6-in-10 at a sub-7% false-alarm rate is the current honest state of a specialized detector, and near-zero is the honest state of a generic one. Anyone quoting 99%+ against that class of audio is describing their lab, not your risk.
"How accurate is deepfake detection?" has a real answer, but it's conditional: very accurate against current commercial clone tools when the model is trained for your audio environment; partially effective against the adversarial frontier; and near-useless when a lab model meets meeting audio. Buy accordingly — and ask every vendor the four questions above. We answer them at usesonave.com/benchmarks.
See it on your own voice — free for 5 hours/month