How to detect a deepfake voice on a live call
A cloned voice on a Google Meet, Zoom or Teams call is the hardest version of the problem: there's no file to upload, no time to send anything to a lab, and the person being imitated is often your own boss. Here's what actually works — and what doesn't.
Why your ears are the wrong tool
Modern voice cloning needs seconds of reference audio and reproduces timbre, accent and cadence well enough that listeners perform close to chance on good clones — especially over a meeting codec that compresses away many of the artifacts people are told to listen for. The classic advice ("listen for robotic tone") describes 2019-era clones, not current ones.
That said, there are still human-detectable tells worth knowing, as long as you treat them as weak signals, not verdicts:
- Turn-taking latency. Live clone pipelines add processing delay. If every reply starts a beat late — consistently, not just from a bad connection — pay attention.
- Interruption behavior. Interrupt mid-sentence. Real speakers stop, adjust, talk over you. Some cloning setups plow onward or pause unnaturally before recovering.
- Emotional range under pressure. Ask something unexpected and personal ("what did we argue about in last week's 1:1?"). Cloned voices handle scripts well and improvisation less well — though a well-prepared attacker survives this.
The tells that actually matter are in the process, not the audio
Nearly every reported voice-fraud incident shares the same shape, and it's visible before any acoustic analysis:
- Urgency — the transfer must happen today, the deal collapses otherwise.
- Secrecy — "keep this between us," an acquisition, a regulator, a surprise.
- Channel or account change — new beneficiary account, new payment channel, a call instead of the usual ticketed process.
Any two of those together should trigger verification regardless of how convincing the voice is. The voice is the least trustworthy part of the interaction; the request pattern is the reliable signal.
Verification that survives a perfect clone
- Out-of-band callback. Hang up and call the person back on a number you already had — not one from the meeting invite or email thread. This single habit defeats the majority of voice-fraud attempts.
- Pre-agreed code phrases for payment approvals, rotated like passwords.
- Dual authorization with the second approver outside the meeting.
- A hold window on first-time or changed-beneficiary wires — even 30 minutes changes outcomes, because deepfake fraud depends on same-call urgency.
What real-time detection adds
Automated detection does what ears can't: it scores the spectral and temporal fingerprints of synthesis continuously, on every speaker, without fatigue. Two things determine whether it actually works on a live call:
- It has to be trained on meeting audio. Meeting platforms compress voice with the Opus codec and processing that destroys many lab-visible artifacts. In our published testing, a widely-used open-source detector that advertises strong lab numbers catches 1.9% of clips from 27 unseen commercial voice-clone tools played through meeting audio; a model trained for that environment catches 95.2% of the same clips. Same audio, opposite outcome — full table and methodology here.
- It has to be live. A verdict after the call is a post-mortem. A verdict ~4 seconds after a speaker starts talking, refreshed continuously, arrives while the wire can still be held.
Honesty checkpoint: no detector is perfect. On the hardest public real-world deepfakes, our own deployed model catches about 59% at a 93%+ real-voice accuracy — far ahead of commodity tools, far short of certainty. Detection is a second factor that triggers verification; it does not replace callbacks. Any vendor telling you otherwise is selling a lab number.
Putting it together
The layered answer: process rules catch the request pattern, callbacks defeat even perfect clones, and live detection watches every voice on every call so the humans only have to be paranoid when the screen turns red. Sonave does the third layer: a visible bot joins your meeting, every speaker gets a live REAL / SUSPECT / FAKE verdict, and a sustained red verdict fires a wire-hold webhook into your approval flow with an exportable forensic report.
Protect your next call — free for 5 hours/month