Protecting Against Deepfake Scams: Deepfake Phishing Is a Campaign, Not just a Call
Protecting Against Deepfake Scams: Deepfake Phishing Is a Campaign, Not a Call
The call you notice is stage four of six. The sequence has been running for a fortnight by the time the phone rings, and it works because your verification controls were written for the call.
Most people picture a deepfake scam as a moment. A cloned voice on the line, ninety seconds of pressure, a decision that either goes well or does not. That picture is the vulnerability. It is accurate about the call and wrong about when the campaign started, and it produces a defence that arrives about eleven days late.
We build and run these sequences against enterprise organizations as a paid engagement. What that work makes obvious is that the synthetic media is the last component an adversary builds, not the first. It is expensive relative to the rest of the sequence and it is the only stage carrying real risk of detection, so it is held back until everything else has already made the request plausible.
What Is Social Engineering?
The term gets used loosely, and the looseness is what lets a deepfake look like a new category rather than a new component.
Social engineering is the manipulation of a person into performing an action or disclosing information, by exploiting trust, authority, urgency, or routine rather than a technical flaw. The target is a decision, not a system.
Notice what is missing from that definition. There is no medium in it. Nothing about email, voice, or video, because the medium was never the mechanism. The mechanism is that somebody inside the organization holds the authority to move money, reset a credential, or grant access, and the whole operation exists to borrow that authority for ninety seconds.
What Is Deepfake Phishing?
Deepfake phishing is social engineering in which synthetic voice, video, or imagery is used to make the impersonation credible. The deepfake is a component of the operation, not a category of its own.
It is what an adversary reaches for when a request has to appear to come from one specific named person and text alone will not carry it: a finance approval that has to sound like the CFO, a support call that has to sound like the help desk, a payment instruction that has to look like the vendor. Naming the medium describes the delivery. It does not describe the operation, and defending against the medium leaves the operation intact.
This is why deepfake phishing does not respond well to channel-specific hardening. Ban unverified voice requests and the same ask arrives as a video meeting. Harden video and it arrives as a voice note attached to an email thread that already exists. The adversary does not pick a medium and then find a use for it. They pick a consequential action and then choose whichever medium reaches its weakest approach path.
The Six Stages of a Deepfake Phishing Campaign
Here is the sequence as we observe and reproduce it, ordered by when each stage happens relative to the contact the organization eventually notices. The day markers are indicative. Campaigns compress or stretch, but the order is stable. Note where the featured card sits.
Collection, with no contact at all
Reporting lines, approval thresholds, vendors, help desk hours, and travel patterns are assembled from public profiles, job postings, filings, and support documentation. Voice samples come from webinars, recorded panels, and voicemail greetings. Nothing here touches your perimeter, so nothing here produces a log entry you could review afterwards.
Corroboration through channels nobody reports
Small touches establish that a person exists and is expected. A connection request that gets accepted. A declined calendar invite that still leaves a name in the target's calendar history. Each is unremarkable on its own, and none of them is reportable, which is exactly the property being selected for.
Fitting the fake to a process you genuinely run
Only now is the synthetic media produced, and it is built to fit one named person making one request the organization performs routinely. The adversary is not looking for the most convincing clone they can build. They are looking for the least remarkable request they can attach one to. A flawless clone asking for something unusual fails. An adequate clone asking for something ordinary does not.
The synthetic contact
The call, the voice note, or the face on the video meeting. This is the entire event as far as most organizations are concerned, and the only stage most simulation programmes ever test. It is also the shortest stage, and the one arriving with the most accumulated credibility behind it.
The consequential action, routed to the weakest path
The payment, the vendor banking detail change, the credential reset, the access grant. Adversaries do not pick the most valuable action available to them. They pick the one with the thinnest verification requirement in front of it. A hardened wire process sitting next to an unwritten vendor bank detail change has not raised the bar. It has published a map.
Persistence, and reconnaissance that survives
Session tokens, enrolled devices, and forwarding rules outlive the contact. So does the research. Nothing gathered in stage one expires because one attempt was refused, so a target who correctly declined in March is a warm, fully mapped target in June with a better pretext. Refusal ends an attempt. It does not end a campaign.
What Phishing Simulation Platforms Can and Cannot See
Look back at those six stages and count how many a conventional test would observe. Most phishing simulation platforms were built around a single artifact, a link inside an email, and the metric they were designed to produce is whether that link was clicked. That instruments one half of stage four and nothing else.
It matters more than it sounds, because a deepfake phishing campaign frequently contains no link anywhere in it. A voice call, a voicemail, and a callback that ends in a credential handover produce no click to count, so a link-based tool reports an empty result for a population that actually failed. Adding a voice channel to an email tool is real progress, but sending a call is not the same as answering the callback, and neither is the same as following through into stage five.
So when you are comparing phishing simulation platforms, the useful question is not how many channels the platform can send on. It is which stages of the campaign the architecture is able to observe, and what it silently scores as a pass. We put the options side by side in our platform comparison and vendor breakdown, and the managed version of the full sequence is described under deepfake phishing simulation.
Why Detection Training Does Not Hold
Two controls stand between a synthetic media campaign and a consequential action, and they have opposite trajectories. That divergence is the whole argument.
Detection is a decaying control. Its effectiveness is an inverse function of the adversary's generation quality. Generation quality improves continuously and cheaply. Human perception does not improve at all. Every tell that training teaches, the flat prosody, the odd pause, the breath that is not there, is a defect in the current generation of synthesis, and those defects are being engineered out. The control loses value from the day it is delivered, at a rate the adversary sets.
Procedural verification is an invariant control. A callback to a number held in the directory, or a dual-approval gate, returns the same result against a crude fake and against a perfect one. The control never reads the artifact, because evaluating the artifact was never part of its specification.
Two bounds on that claim, because the overstated version is the one that gets dismantled in front of a board. First, invariance is a property of the specification, not a guarantee of the outcome. Better synthetic media does not defeat a callback directly. It raises the pressure applied against it, which raises how often somebody consciously waives a step they knew applied. Track that as exception rate. Second, the human is still in the measurement. Procedures are executed by people. What changes is the behaviour you hold them to, from perception, which is untrainable and unobservable, to compliance under pressure, which is both.
Source: Breacher.ai OSES™ engagement outcomes, published in the OSES Risk Index. Detection figures are measured across engagements covering more than one thousand individual targets. Reporting is organizational and sector level only.
Read the middle figure on its own and it invites the wrong conclusion, which is better detection training. Read it beside the third and it says something else. Those organizations were not scored on whether individuals noticed anything. They were scored on whether a verification requirement existed and held, which is a structural property of the control environment and does not move when the adversary ships a better model.
The Controls That Hold
Everything below shares one property. It works whether or not the person on the other end is real.
- Attach verification to the action, not the channel. The rule should read "a vendor banking detail change requires a callback to the number of record," not "phone requests require a callback." Channel-scoped rules just move the traffic.
- Define the number of record. A number offered on a call, in a signature block, or in a voicemail is not verified. Only a number looked up independently is. Extend this to calls the employee placed themselves, which is the largest uncovered path we find.
- Dual approval on the actions that end campaigns. Payments above a threshold, vendor bank detail changes, privileged access grants. Two people, separately reachable, neither able to complete the action alone.
- A standing rule with no exceptions. IT will never ask for a password, never ask you to read back a code, and never ask you to approve a prompt you did not initiate. A rule with exceptions is a rule an adversary can talk their way into.
- Coverage before adherence. Training people to follow a procedure that does not exist for the action being requested achieves nothing. Somebody has to write it first.
- Reporting that costs less than complying. If reporting takes six minutes and a form and complying takes ninety seconds, the outcome is already chosen. Measure reporting on speed rather than accuracy.
Start by Measuring Coverage
Coverage can be assessed before a simulation is ever commissioned, which makes it the cheapest place to start. These are the questions worth answering this week.
The fourth question is the one worth building reporting around. A requirement that was never written and a requirement that was waived look identical in an incident summary and need completely different fixes. One needs a procedure author. The other needs a friction review.
The voice mechanics, including why the return call carries more consequence than the outbound one, are covered in the callback breakdown. The help desk variant of stage five is in IT support impersonation, and the sequencing method under OSES™.
What to Do This Week
Write the list. Outbound payment, vendor banking detail change, credential reset, multi-factor enrollment or reset, privileged access grant, software installation on request, data export, physical access grant. Against each one, record whether a requirement exists, whether it is channel-independent, and whether it is reachable under time pressure. This takes an afternoon, and the unmarked rows are your actual finding.
Rewrite one rule to cover the exempt direction. Most verification rules are written for inbound contact. Amend the largest one so it also covers a self-initiated call, a number taken from a voicemail, and a request arriving through a collaboration platform.
Then test the sequence, not the moment. Not a survey and not a training module. A run that starts with reconnaissance and follows through into an actual consequential action, reported on coverage and process hold rate rather than on who answered.
Frequently Asked Questions
What is social engineering?
Social engineering is the manipulation of a person into performing an action or disclosing information, by exploiting trust, authority, urgency, or routine rather than a technical flaw. The target is a decision, not a system. Phishing, pretexting, vishing, and deepfake fraud are delivery methods for the same operation: persuading an authorised person to use their authority on the adversary's behalf.
What is deepfake phishing?
Deepfake phishing is social engineering in which synthetic voice, video, or imagery makes the impersonation credible. The deepfake is a component, not a category of its own. It is added at the point where credibility is the binding constraint, usually when a request must appear to come from a named executive, a known vendor, or the help desk. Naming the medium describes the delivery, not the operation.
Why is deepfake phishing rarely a single event?
Because the synthetic contact is the last and most expensive stage, and it only works if earlier stages have done the credibility work. Reconnaissance, low-stakes contact through other channels, and the choice of a request the organization genuinely performs all happen first. By the time the call arrives, the deepfake is not persuading the target. It is confirming something they already have reasons to believe.
Do phishing simulation platforms test deepfake phishing?
Most do not, because most were built around a link in an email and therefore measure whether the link was clicked. A deepfake phishing campaign frequently contains no link anywhere in it, which a link-based tool records as an empty result. Testing it properly requires placing the call, leaving the voicemail, answering the callback, and following through into the consequential action. When comparing phishing simulation platforms, the useful question is not which channels a platform can send on. It is which stages of the campaign the architecture is able to observe.
What is the best way of protecting against deepfake scams?
Attach the verification requirement to the action rather than to the channel it arrived on. A callback to a number held in the directory, a second-channel confirmation, or a dual-approval gate returns the same result against a crude fake and a convincing one, because the procedure never evaluates the media. Start by listing your consequential actions and recording which ones have a written requirement.
Can employees be trained to spot a deepfake?
Not reliably enough to build a programme on. Across engagements covering more than one thousand individual targets, most people tested could not distinguish synthetic voice from a real person. Every perceptual tell that training teaches is a defect in the current generation of synthesis, and those defects are being engineered out, so the control weakens over time. People still matter. What changes is that the behaviour worth measuring becomes whether the procedure was followed, not whether the fake was noticed.
Find Out Which Stage Your Process Stops At
We run the full sequence against your own organization, from reconnaissance through the consequential action, and report coverage and process hold rate against your vertical.
Book Your Demo
