Stop Training People to Spot Deepfakes. Train the Procedure, Then Test It.
Stop Training People to Spot Deepfakes. Train the Procedure, Then Test It.
Detection training is not worthless. It is depreciating, and the depreciation schedule is set by the adversary rather than by you. A verification procedure returns the same answer against a crude fake and a perfect one. That difference is the whole argument, and it decides where the budget should go.
Here is the part that sounds backwards. We build synthetic voice and video for a living, we run it against enterprise organizations as an authorized engagement, and the better our output gets, the less we care whether anyone can tell.
That is not bravado. It is the only conclusion the economics support. Two controls sit between a synthetic media pretext and a consequential action. One is human detection: the target recognizes the voice, the face, or the message as fake. The other is procedural: the target executes a verification step before acting, whatever they happen to believe. Those two controls are moving in opposite directions, and they have been for some time.
No bank measures whether a teller can eyeball a counterfeit note. They measure whether the teller ran the check. The bill got better. The check did not have to.
Two Controls, Opposite Trajectories
Detection is a decaying control. Its effectiveness is an inverse function of the adversary's generation quality. Generation quality improves continuously and cheaply. Human perceptual capability does not improve at all. Every dollar spent on detection training buys a control whose value starts declining the day the training is delivered, and the rate of decline is set by the adversary, not by the buyer.
Procedural verification is an invariant control. A callback to a known-good number, a second-channel confirmation, or a dual-approval gate returns the same result against a crude fake and against a perfect one. The control does not read the artifact. It does not care how good the artifact is.
The strategic point is not that process is better than detection in some abstract sense. It is that the gap between the two widens every quarter without anyone doing anything. A buyer who invests in detection accuracy is buying a depreciating asset. A buyer who invests in procedural coverage is buying one that holds its value.
Control effectiveness against rising generation quality
Horizontal axis: adversary generation quality, increasing left to right. Vertical axis: control effectiveness, 0 at the baseline to 100 at the top. Schematic of the argument, not plotted measurement data. The shape is what matters: the procedural line holds because the control never inspects the artifact, and the shaded gap widens on its own as generation quality rises.Why the Tell Is the Wrong Thing to Train
Three named mechanisms, rather than a general objection to awareness programs.
The asset depreciates on a schedule you do not control
Every artifact you teach people to listen for, the flat prosody, the odd pause, the breath that is not there, the hand that blurs at the edge of frame, is a defect in the current generation of synthesis. Those defects are being engineered out by people whose entire incentive is to remove them. You are buying a control and handing the adversary the depreciation schedule.
It puts a forensic judgment at the worst point in the process
Detection training asks a person, under time pressure, on a call they did not initiate, to make a media authenticity judgment they are untrained and unequipped to make. Trained analysts get this wrong with tooling in front of them. Expecting it of a payroll clerk between meetings is not a control, it is a hope with a completion rate attached.
Certainty is the failure mode, not ignorance
A person who has completed detection training and now believes they can tell is a harder target to protect than one who knows they cannot. Confidence is precisely what a good pretext needs. The user who is sure the voice is genuine is the user who skips the callback, and skipping the callback is the only thing that had to happen.
What the Numbers Say, With Their Denominators Attached
Three figures worth putting next to each other, each reported with the population it came from, because a percentage without a denominator is a slogan.
The middle figure is the one that usually gets quoted as proof that detection training is needed. Read it the other way. A population that cannot reliably tell is not a population that needs a better module on tells. It is a population whose safety has to come from somewhere that does not depend on telling. We broke the channel split down separately in our read of the 2026 Gartner deepfake data.
Where This Argument Stops
Publishing the bounds is what separates a position from a pitch, and every one of these is a place an informed buyer could otherwise take the argument apart.
- The human is still in the measurement. Procedures are executed by people. We have changed which human behavior we measure, from perception to compliance under pressure. That is a better variable because it is trainable and observable. It is still a human variable, and anyone selling you the idea that users do not matter is selling you something.
- Invariance is not total. Better synthetic media does not defeat a callback directly. It raises the pressure applied against it. A more convincing executive makes granting the exception feel more reasonable. What degrades as generation quality improves is not the control, it is the exception rate, and that is a number you can watch.
- Coverage is the limiting factor, not adherence. A procedural control only holds where a procedure exists. Invariance protects the covered paths and says nothing at all about the uncovered ones, which is where our engagements find most of the failures.
- We have not run the head-to-head. Nobody has published a clean comparison of a detection-trained population against a procedurally-trained one, measured on whether the verification step held. We think we know how it lands. We have not proved it, we should test it, and we would publish the result in either direction.
The Coverage Gap Is the Actual Finding
Most organizations have a verification requirement for outbound wire transfers, because that is where the money visibly leaves and because a regulator asked about it. Almost nothing else on the list below gets the same treatment, and the list below is what a synthetic voice is actually used to obtain.
| Consequential action | Verification requirement usually exists | What the pretext asks for |
|---|---|---|
| Outbound payment or wire transfer | Usually | An urgent transfer, approved outside the normal sequence |
| Vendor banking detail change | Sometimes | New remittance details before the next invoice run |
| Credential reset | Rarely | A password reset while the caller stays on the line |
| MFA enrollment or reset | Rarely | A new device enrolled, or a code read back |
| Privileged access grant | Sometimes | Temporary elevation to finish something time-critical |
| Software installation on request | Rarely | A support session joined so a certificate can be pushed |
| Data export or disclosure | Sometimes | A file sent across for an audit or a deal |
| Physical access grant | Rarely | A contractor let in ahead of the paperwork |
The coverage column reflects what we typically find in engagements rather than a published rate, and it is stated in bands for that reason. Build your own version of this table before you buy anything. It takes an afternoon, it needs no simulation, and it usually ends the argument about where the budget goes.
What to Teach Instead
Everything below shares one property: it works whether or not the person on the other end is real.
- State the verification step as an action, not a warning sign. "Call back on a number you look up yourself" is executable under pressure. "Be alert for unusual urgency" is not, and it never was.
- Out-of-band callback, from a number the person sources independently. Not a number given during the call, not the caller ID, not a link in the follow-up text.
- A standing negative rule with no exceptions. IT never asks for a password, never asks you to read back a code, never asks you to approve a prompt you did not personally start. A rule with no exceptions is a rule a person can apply while being pressured. A rule with exceptions is a rule an adversary can talk their way into.
- A pre-existing record for privileged actions. No reset, enrollment, or remote session without a ticket the user can independently see. This takes the judgment call out of the critical path, which is the point.
- A named exception path, logged rather than improvised. Exceptions will happen. If the only way to grant one is to go off-procedure quietly, you will never see the number that matters most.
- Reporting that costs less than complying. If reporting takes six minutes and a form, and complying takes ninety seconds, the outcome is already decided.
- Training keyed to the failed control, not the failed scenario. The module a person receives should be the procedure they did not follow, which is also what makes it re-testable later.
None of this asks anyone to become a detector. It asks them to follow a procedure, which is a thing organizations already know how to write, deliver, and measure. That is the shape of deepfake awareness training we think survives the next two years of model releases.
Then Test Whether It Holds
Teaching a procedure and testing a procedure are different budget lines, and only the second one produces evidence. These are the questions a test should answer.
Note what is absent from that list: a risk score attached to an individual employee. Per-user results are a routing input for us, used to aim training at the people and the paths that failed. They are not the outcome metric, because a score on a person does not tell you whether the payment went out. The organization-level version of this is what the Social Engineering Risk Index reports.
What We Would Do First
If you have one week and no budget, three things, in this order.
Write the list. Every consequential action in your organization on one page. Use the eight above as a starting set and add whatever is specific to you. This is the artifact almost nobody has, and producing it changes most of the conversations that follow.
Mark the coverage. Against each action, note whether a written verification requirement exists, whether it holds regardless of which channel the request arrived through, and whether the person under pressure can actually reach it. Three columns. The empty cells are your program.
Then run the call. Not a survey and not a module. A live vishing simulation against a real population, scored on whether the procedure held. If you want a low-commitment version of the demonstration first, our deepfake quiz shows a room how unreliable detection is in about two minutes, which is usually enough to get the procedural conversation started.
Frequently Asked Questions
Can employees be trained to detect deepfakes?
Not durably. Every tell taught in a detection module is a defect in the current generation of synthetic media, and those defects are being engineered out continuously. Human perceptual capability does not improve to compensate. That makes detection accuracy a control whose value declines from the day the training is delivered, at a rate the adversary sets rather than the buyer. Verification procedures return the same result against a crude fake and a convincing one, which is why they are the control worth building a program on.
Is deepfake detection training useless?
No, and that overstatement is worth avoiding. Exposure to a realistic synthetic voice or video changes what people treat as possible, and that is genuinely useful. The error is treating detection accuracy as the outcome the program is accountable for. Exposure should be used to make the verification step reflexive and to surface the paths where no verification step exists, not to certify that a population can now tell real from synthetic.
What should deepfake awareness training teach instead?
The procedure a person must execute before a consequential action, stated as an action rather than a warning sign. Out-of-band callback to a number the person looks up independently. A standing negative rule with no exceptions, such as IT never asking for a password or an MFA code. A pre-existing ticket or record before any privileged action. A named exception path that is logged rather than improvised. And a reporting route that costs the user less effort than complying with the caller.
Does procedural verification stop working as deepfakes improve?
The control itself does not degrade, because it never reads the artifact. A callback to a known-good number returns the same answer regardless of how convincing the voice was. What does move with generation quality is the exception rate: a more convincing executive makes waiving the step feel more reasonable to the person being pressured. That is the honest bound on the argument, it is measurable, and it is the number to watch as fidelity improves.
How do you test whether verification procedures actually hold?
In two passes. First, coverage: for every consequential action in the organization, establish whether a defined verification requirement exists, whether it is channel-independent, and whether the person under time pressure can reach it. That assessment needs no simulation. Second, pressure: request those same actions through a synthetic or impersonated channel and record whether the verification step was executed, skipped, or consciously waived. The gap between what people say they would do and what they do on a live call is the entire measurement.
Find Out Whether Your Procedures Hold
We run the simulation across voice, video, and messaging, report the process hold rate and the exception rate, and benchmark you against your sector.
Book Your Demo
