Sonave All guides

How to detect a deepfake voice on a live call

A cloned voice on a Google Meet, Zoom or Teams call is the hardest version of the problem: there's no file to upload, no time to send anything to a lab, and the person being imitated is often your own boss. Here's what actually works — and what doesn't.

Why your ears are the wrong tool

Modern voice cloning needs seconds of reference audio and reproduces timbre, accent and cadence well enough that listeners perform close to chance on good clones — especially over a meeting codec that compresses away many of the artifacts people are told to listen for. The classic advice ("listen for robotic tone") describes 2019-era clones, not current ones.

That said, there are still human-detectable tells worth knowing, as long as you treat them as weak signals, not verdicts:

The tells that actually matter are in the process, not the audio

Nearly every reported voice-fraud incident shares the same shape, and it's visible before any acoustic analysis:

Any two of those together should trigger verification regardless of how convincing the voice is. The voice is the least trustworthy part of the interaction; the request pattern is the reliable signal.

Verification that survives a perfect clone

  1. Out-of-band callback. Hang up and call the person back on a number you already had — not one from the meeting invite or email thread. This single habit defeats the majority of voice-fraud attempts.
  2. Pre-agreed code phrases for payment approvals, rotated like passwords.
  3. Dual authorization with the second approver outside the meeting.
  4. A hold window on first-time or changed-beneficiary wires — even 30 minutes changes outcomes, because deepfake fraud depends on same-call urgency.

What real-time detection adds

Automated detection does what ears can't: it scores the spectral and temporal fingerprints of synthesis continuously, on every speaker, without fatigue. Two things determine whether it actually works on a live call:

Honesty checkpoint: no detector is perfect. On the hardest public real-world deepfakes, our own deployed model catches about 59% at a 93%+ real-voice accuracy — far ahead of commodity tools, far short of certainty. Detection is a second factor that triggers verification; it does not replace callbacks. Any vendor telling you otherwise is selling a lab number.

Putting it together

The layered answer: process rules catch the request pattern, callbacks defeat even perfect clones, and live detection watches every voice on every call so the humans only have to be paranoid when the screen turns red. Sonave does the third layer: a visible bot joins your meeting, every speaker gets a live REAL / SUSPECT / FAKE verdict, and a sustained red verdict fires a wire-hold webhook into your approval flow with an exportable forensic report.

Protect your next call — free for 5 hours/month