What We Find
When the Simulation Runs

Findings from enterprise engagements: which orchestrated sequences work, which departments give way first, and why the click was never the number that mattered.

92%
Of organizations tested were vulnerable
78%
Were highly vulnerable, meaning the process failed
1,000+
Simulated targets behind the benchmark

The Sequences That Work Best

Ranked by what people did, not by what they opened. Every one of these is a handoff between channels rather than a single message.

01

Deepfake Video, then Agentic Email

Video → Agentic AI → Email
21.78%
Took Action
33.0%
Clicked
02

Deepfake Call, then Agentic SMS

Voice → Agentic AI → SMS
14.75%
Took Action
23.0%
Clicked
03

Deepfake Call, then Calendar Invite

Voice → Agentic AI → Calendar
9.54%
Took Action
13.8%
Clicked

Note the gap between the two columns. Across all three sequences, roughly a third of the people who clicked did not go on to do the thing that would have caused a loss. A click rate counts both groups the same way.

Click Rate Is the Wrong Measure

The most useful thing we can tell you from three years of engagements is that the industry's default metric does not describe the risk.

A click tells you a person touched something. It does not tell you whether money moved, whether a credential was reset, or whether anyone in the building noticed. In the engagements where something consequential happened, three things failed at once: the tooling did not catch it, the process did not hold, and a person took the wrong action. A click metric captures the third one, partially, and only when a link exists.

Across a sample of voice campaigns covering roughly 900 targets, the weighted mean action rate was 14.5%, ranging from 0% in the strongest organizations to about 35% in the weakest. There was no link anywhere in those campaigns. A phishing platform would have reported nothing at all.

" The strongest organizations we test are not the ones with the sharpest people. They are the ones where a request outside normal process hits a step the requester cannot talk past.
What a click rate reports

A moment, on one channel

Whether someone touched a link, once
Only in scenarios where a link exists
Scored per person, as if risk were individual
A number that falls when people get cautious, not when the organization gets safer
What decides the outcome

The step after the moment

Whether the wire required a second channel
Whether the reset had a check that was not the caller's word
Whether escalation was reachable mid-call
Whether the same request would fail every time, not just this time
People

Did they act, push back, or report?

Depending on the path, a person is the first line of defense, the last, or the only one. We measure what they did with the request and how fast they raised it, not whether they could spot a rendering artifact.

Process

Did the procedure hold under pressure?

Callback rules, approval thresholds, identity checks at the helpdesk. This is where consequential failures actually occur, and it is different in every organization, which is why generic testing and generic training both miss it.

Technology

Did anything in the stack see it?

These sequences land inside trusted communication platforms and travel through the seams between security tools. Each tool works as designed and nobody owns the gap. That gap is where the attack surface has moved.

The highest-yield control we have measured

Procedural Verification

It is unglamorous, it is cheap, and it is the single thing that separates the organizations that hold from the ones that do not. Four steps carry most of the value.

1 Out-of-band callback to a number already on file, never a number supplied inside the request.
2 A second approver for anything that moves money, credentials, or access.
3 A named escalation path that is faster to use than to ignore, with no penalty for using it.
4 An explicit rule that urgency from an executive is a reason to verify, not a reason to skip.

None of these require a single person to correctly identify a deepfake.

What the Benchmark Shows

Aggregate figures are context, not a benchmark. The number that matters to you is how your organization compares to your own peer vertical.

Department Exposure

Average click rate by function
Finance
22.90%
HR
16.65%
Company Avg
14.04%

Finance runs 63% above the company average. It is also the function where a single wrong action moves money, which is why department-level reporting beats a single organizational number.

Training Impact

Action rate, trained against untrained
13%
No Training
~ -1/3
8%
With Training

Reported honestly: training cuts susceptibility by roughly a third and then stops. The remaining 8% is not a training problem. It is a process problem, and no amount of additional awareness content closes it.

Detection Reality

Human performance against synthetic media
38%
Average Accuracy
63%
Could Not Distinguish
11%
Reliably Identified

The headline finding: average accuracy sits below a coin flip, and it gets worse as generation quality improves. Any control that depends on a person telling real from synthetic is a control with a declining success rate.

A Chain That Walks Around the Controls

Adversaries are not defeating link protection. They are arriving somewhere it does not apply.

Mobile trust bypass

Voicemail Drop, then SMS

This chain exploits voicemail transcription on mobile to sidestep link protection entirely. The voicemail manufactures prior contact, so the SMS that follows arrives inside a conversation the target believes they already started.

1 Ringless voicemail drop. Synthetic audio deposited directly. No ring, no missed call notification.
2 Wait for transcription. The handset transcribes the message automatically within a few minutes.
3 Send the SMS. The link lands in a thread the device already treats as established contact.

Why It Works

The voicemail does the credibility work before any request is made. By the time the message arrives, the target is not evaluating a stranger, they are continuing a conversation.

The handset treats the thread as established contact, so protections written for unsolicited inbound links do not apply the same way.

The lesson is not that people should scrutinize voicemail transcripts more carefully. It is that a request to move money or credentials should require verification regardless of how established the thread feels.

Find Out Where Your Process Breaks

Thirty minutes. We walk through a real OSES™ engagement, scenario design through findings, and you decide whether your process would have held.

No IT integration required
Runs fully external
Peer vertical benchmark
Book Your Demo