Deepfakes Are Coming for Your Inbox

For most of phishing's history, the attack surface has been text and, occasionally, a spoofed website. Generative AI has expanded that surface to voice and video, and the implications for email-adjacent social engineering are significant: a phishing email is no longer the entire attack, it can be the opening move in a multi-step scheme that includes a follow-up phone call in a cloned voice, or a video message that appears to show a real executive making a request on camera.
This piece looks at how deepfake technology is being incorporated into phishing and fraud campaigns, why it's particularly effective against verification methods organizations have relied on for years, and what defending against it actually requires.
Quick Summary
- The shift: attackers are combining traditional email phishing with synthetic voice and video to make social engineering more convincing and harder to verify.
- Why it's effective: it directly undermines "call to verify" as a safeguard, since the voice on the other end of that call can now be synthetic too.
- What's being targeted: high-authority, high-trust moments — a voicemail appearing to confirm an urgent wire transfer, a video message appearing to come from a known executive.
- What detection requires: cross-channel correlation and content analysis capable of detecting synthetic media artifacts, not just text or link analysis.
From Text-Only to Multi-Modal Social Engineering
The traditional advice for verifying a suspicious request — "if an email asks for something unusual, call to confirm" — assumed the phone call was a trustworthy secondary channel, harder for an attacker to fake than an email. Generative voice cloning technology has meaningfully eroded that assumption. With a relatively small sample of someone's voice — sometimes obtainable from public video content, a conference talk, or even a company's own marketing material — modern voice synthesis tools can produce convincing audio in that person's voice, saying whatever an attacker scripts.
This doesn't mean every phishing attempt now involves a deepfake — the overwhelming majority of attacks remain text-based, because they're cheaper and faster to execute at scale. But for high-value, targeted attacks — the same category of attack executive protection and BEC detection are built to catch — deepfake audio and video are becoming a meaningful part of the toolkit for the subset of attackers willing to invest the extra effort for a high-value target.
How Deepfakes Get Incorporated Into a Campaign
A multi-channel attack incorporating deepfake elements typically follows a layered sequence: an initial email establishes a pretext — an urgent, time-sensitive request, often financial — followed by a voicemail or live call in a cloned voice that appears to corroborate the email's urgency and legitimacy, specifically designed to defeat the "call to verify" instinct the email alone might not have overcome. In more elaborate cases, a brief video clip — sent as an attachment or link, appearing to show a known executive making the request directly — adds a further layer of apparent legitimacy.
The email itself in these campaigns often looks unremarkable on its own; its role is to set up the pretext that the synthetic audio or video then reinforces. This is precisely why evaluating an email in isolation, without considering it as one step in a potentially multi-channel scheme, can miss the broader pattern.
Why Deepfakes Undermine Traditional Verification
Out-of-band verification — confirming an unusual request through a separate channel — has long been one of the most effective defenses against BEC and executive impersonation, precisely because it doesn't depend on any single channel's authenticity. Deepfake audio directly targets this assumption: if the "separate channel" being used for verification is a phone call, and the voice on that call can be synthesized, the verification step itself becomes vulnerable to the same manipulation it was designed to catch.
This doesn't make out-of-band verification useless — a phone call to a previously-known, independently-verified number remains far harder to compromise than one to a number provided in the original suspicious message. But it does mean organizations relying on "just call to confirm" as their sole safeguard need to reconsider what a genuinely independent verification channel looks like in a world where voice can be synthesized convincingly.
Detecting Synthetic Media Artifacts
Defending against deepfake-enhanced social engineering requires a different kind of content analysis than text-based phishing detection. Voice and video deepfakes, even convincing ones, typically contain subtle artifacts — unnatural pauses, inconsistent background audio characteristics, slight mismatches between lip movement and audio in video, or statistical patterns in the audio waveform that differ from genuine human speech — that specialized detection models can identify even when a human listener or viewer might not notice them consciously.
This kind of analysis is still a developing field relative to text-based phishing detection, which has had years longer to mature, and detection accuracy varies depending on the sophistication of the synthesis technique used. This is why cross-channel correlation matters as much as content analysis on its own: a suspicious email combined with an unusual follow-up call, evaluated together rather than as two independent, unrelated events, provides a stronger signal than either channel analyzed in isolation.
Practical Steps Organizations Can Take Now
Organizations don't need to wait for deepfake detection technology to mature perfectly before reducing their exposure. Establishing genuinely independent verification channels — a pre-agreed callback number that isn't provided in the original request, a secondary approval requirement for high-value transactions that doesn't rely on voice confirmation alone — reduces reliance on any single channel's authenticity. Awareness training that specifically covers voice cloning, rather than only text-based phishing indicators, helps employees understand that a familiar voice, on its own, is no longer sufficient grounds for trust in a high-stakes request. And technical monitoring that treats email, voice, and video as connected parts of a potential single campaign, rather than isolated channels, closes the gap that purely email-focused detection leaves open.
FAQ
Are deepfakes common in phishing attacks today?
Not yet at the scale of text-based phishing — deepfakes remain more resource-intensive to produce, so they're primarily used in high-value, targeted campaigns rather than mass phishing attempts.
How much audio does an attacker need to clone someone's voice?
Modern voice synthesis tools can work from relatively small audio samples, sometimes obtainable from public sources like conference talks, interviews, or a company's own marketing content.
Does "call to verify" still work as a safeguard?
It's weaker than it used to be if the number called is one provided in the suspicious request itself, since the voice on that call could be synthetic. It remains effective when the number is independently, previously verified — not supplied by the potentially fraudulent message.
How can synthetic voice or video be detected technically?
Specialized detection models look for artifacts often present in synthesized media — unnatural pauses, inconsistent background audio, or subtle statistical patterns that differ from genuine human speech and video — though detection accuracy varies with the sophistication of the synthesis technique.
What's the most effective organizational safeguard against deepfake-enhanced attacks?
Establishing verification channels that are genuinely independent of the original request — a pre-agreed callback number, a secondary approval step — rather than relying on voice recognition or a callback number supplied by the suspicious message itself.
Should awareness training be updated to cover deepfakes specifically?
Yes — training focused only on text-based phishing indicators doesn't prepare employees for the specific dynamics of a voice or video-based social engineering attempt, which requires a different kind of skepticism.
Key Takeaways
- Attackers are increasingly combining email phishing with synthetic voice and video for high-value, targeted campaigns.
- Deepfake audio specifically undermines "call to verify" as a safeguard, since the voice on the call can now be synthesized.
- Genuinely independent verification — a pre-agreed callback number, not one from the suspicious request — remains effective against this technique.
- Synthetic media detection looks for subtle artifacts in audio and video, though the field is still maturing relative to text-based detection.
- Effective defense requires treating email, voice, and video as parts of a potential single campaign, not isolated, independently-evaluated channels.