CEO voice fraud: how deepfake wire transfers happen — and the process that stops them
Executive-impersonation fraud predates AI — but voice and video cloning removed its biggest weakness. The email from "the CEO" could be doubted; the CEO's actual voice, on a live call, asking you by name, is engineered to end doubt. Here's the anatomy, the real cases, and the control stack that works.
The anatomy of the attack
- Reconnaissance. Executive voice samples are everywhere — earnings calls, podcasts, conference talks, sales videos. Seconds of audio suffice for a usable clone. Org charts and approval workflows come from LinkedIn and phishing.
- The pretext. A confidential situation that explains both urgency and secrecy: an acquisition closing, a supplier emergency, a regulator deadline. The pretext is chosen so that checking with colleagues feels like betraying the executive's trust.
- The live call. Increasingly not just a phone call — a video meeting, sometimes with multiple cloned "colleagues" present. Live conversation defeats the old advice ("call them back" gets reframed as "I'm calling you, we're talking right now").
- The transfer. A wire to a new beneficiary, often broken into several payments, timed near end-of-day or end-of-week so recovery windows close before anyone reconciles.
This is not hypothetical
- In the widely reported Arup case (Hong Kong, 2024), a finance employee joined a video call where the CFO and several colleagues appeared and spoke — all of them deepfaked — and subsequently authorized transfers totaling about US$25 million.
- As far back as 2019, public reporting described a UK energy-company executive wiring roughly €220,000 after a phone call cloning his parent-company CEO's voice — accent and cadence included.
- Law-enforcement and industry advisories since have repeatedly warned that generative voice is now standard tooling in business-email-compromise playbooks, which already account for billions in annual reported losses.
The pattern to internalize: in every case, the human on the call was convinced. The control failure wasn't gullibility — it was that the process allowed a live conversation to authorize money movement.
The control stack
No single control survives contact with a good attacker. The stack below assumes the voice will be convincing:
- Out-of-band callback, always. Any payment instruction that arrives by call or meeting gets verified by calling the requester back on a directory number. No exceptions for seniority — especially not for seniority.
- Dual authorization for wires above a threshold, with the second approver reached outside the original meeting.
- Beneficiary-change quarantine. First payment to any new account gets a mandatory hold window and independent confirmation.
- Urgency-plus-secrecy rule. Train the specific pattern: urgent + confidential + payment = stop, verify, document. Give staff explicit social cover: "our process requires a callback, even for the CEO."
- Live voice verification on the call itself. Real-time detection scores every speaker continuously and turns "I had a feeling" into an on-screen verdict — and, on sustained red, a machine-readable event.
Where detection fits (and where it doesn't)
Detection's job in this stack is specific: it watches every meeting so humans don't have to be suspicious of every meeting. When Sonave's model scores a voice in the red band for three consecutive windows, it raises a wire-hold webhook your approval system can consume — pausing the payment pending re-authentication — and produces an exportable forensic report (scores, timestamps, audio evidence) for the investigation that follows.
It is a second factor, not an oracle. Our deployed model catches 95% of clips from 27 unseen commercial clone tools through meeting audio, and about 59% on the hardest public real-world deepfakes — numbers we publish, with methodology, at usesonave.com/benchmarks. The callback still matters. The point is that the callback now gets triggered by an alarm instead of depending on a junior employee's willingness to challenge a CFO.
Put a verdict on every voice — free for 5 hours/month