OSES Risk Index - Breacher.ai

The OSES Risk Index

What we find when the simulation runs, published.

A benchmark computed from real engagement outcomes rather than survey answers: which orchestrated sequences work, where the process gives way first, and why the click was never the number that mattered. Every figure below came out of an OSES™ engagement against a live organization, scored the same way, so the comparison across sectors actually holds.

Action Rate Process Failure Escalation Time Detection Coverage Re-test Delta Sector Median

Organizational and sector reporting only. No named individuals, no department leaderboards.

The default metric is not describing the risk

92%
Of organizations tested were vulnerable to an orchestrated sequence
Source: Breacher.ai OSES engagements
78%
Were highly vulnerable, meaning the process itself failed
Source: Breacher.ai OSES engagements
14.5%
Weighted mean action rate in voice campaigns containing no link at all
Source: Breacher.ai OSES engagements

Three things decide the outcome

Every consequential failure we have observed involved all three at once. The index scores each separately, because they fail for different reasons and get fixed by different budgets.

People

Did they act, push back, or raise it. Depending on the path, a person is the first line of defense, the last, or the only one. We score what they did with the request and how fast they escalated, not whether they could spot a rendering artifact.

Action RateReport RateTime to Escalate
Process

Did the procedure hold under pressure. Callback rules, approval thresholds, identity checks at the service desk. This is where consequential failures actually occur, and it is different in every organization, which is why generic testing and generic training both miss it.

Callback DisciplineSecond ApproverStep-Up Path
Technology

Did anything in the stack see it. These sequences land inside trusted communication platforms and travel through the seams between security tools. Each tool works as designed and nobody owns the gap. That gap is where the attack surface has moved.

Detection CoverageChannel SeamsAlert Ownership

The sequences that work best

Ranked by what people did, not by what they opened. Every one of these is a handoff between channels rather than a single message.

01
Deepfake video, then agentic email
Video → Agentic AI → Email
21.78%
Took action
33.0%
Clicked
02
Deepfake call, then agentic SMS
Voice → Agentic AI → SMS
14.75%
Took action
23.0%
Clicked
03
Deepfake call, then calendar invite
Voice → Agentic AI → Calendar
9.54%
Took action
13.8%
Clicked
Note the gap between the two columns. Across all three sequences, roughly a third of the people who clicked did not go on to do the thing that would have caused a loss. A click rate counts both groups the same way.

Click rate is the wrong measure

The most useful thing three years of engagements can tell you is that the industry's default metric does not describe the risk it claims to.

What a click rate reports
A moment, on one channel
Whether someone touched a link, once
Only in scenarios where a link exists at all
Scored per person, as if risk were individual
A number that falls when people get cautious, not when the organization gets safer
In our voice campaigns there was no link anywhere. A phishing platform would have reported nothing at all, against a weighted mean action rate of 14.5 percent.
What the index reports
The step after the moment
Whether the wire required a second channel
Whether the credential reset had a check that was not the caller's word
Whether escalation was reachable mid-call
Whether the same request would fail every time, not just this time
Action rates ranged from zero percent in the strongest organizations to about 35 percent in the weakest, on comparable scenarios.

The strongest organizations we test are not the ones with the sharpest people. They are the ones where a request outside normal process hits a step the requester cannot talk past.

Breacher.ai engagement notes

How the index is computed

The index is a by-product of the OSES™ loop, not a separate research exercise. Every engagement emits the same outcome fields, which is what makes one organization's result comparable to another's instead of comparing one vendor's report format to another's.

01
Simulate
An orchestrated sequence runs against the real process, across the channels the organization actually uses
02
Observe Outcome
Action, escalation, detection, and the exact point at which the procedure gave way are recorded, not inferred
03
Normalize
Results are adjusted for scenario difficulty and channel mix so a hard sequence is not rewarded as a weak result
04
Position
The organization is placed against the median for its own vertical, not against a cross-industry average
05
Re-test Delta
Training runs against what the test found, then the same sequence runs again and the change is the evidence

Every engagement runs under signed authorization, within an agreed scope and window. Client names never appear in our public material, and index reporting is organizational and sector level only. The full computation is documented in the assessment methodology, and the scoring scale is published in the Social Engineering Risk Index.

Three findings that hold across sectors

Aggregate figures are context, not a benchmark. The number that matters to you is where you land against your own vertical.

Training moves it a third, then stops

13%
No training
8%
With training

Reported honestly: awareness training cuts susceptibility by roughly a third and then plateaus. The remaining exposure is not a training problem. It is a process problem, and no amount of additional content closes it.

Detection by eye is below a coin flip

38%
Average accuracy
63%
Could not distinguish

Average accuracy against current-generation synthetic media sits below chance, and it degrades as generation quality improves. Any control that depends on a person telling real from synthetic is a control with a declining success rate.

The variance is inside organizations

0%
Strongest observed
35%
Weakest observed

Functions that authorize payments or reset access run materially above their own organizational median, because the request that reaches them is the one worth making. The index reports this as a process finding at organizational level, never as a leaderboard.

A chain that walks around the controls

Adversaries are not defeating link protection. They are arriving somewhere it does not apply. This one exploits voicemail transcription on mobile to manufacture prior contact.

1
Establish

Ringless voicemail drop

Synthetic audio is deposited directly into the mailbox. No ring, no missed call notification, nothing for the target to decline.

2
Wait

The handset transcribes it

Transcription runs automatically within a few minutes and renders the message as readable text alongside the caller identity, which does most of the credibility work before any request is made.

3
Convert

Send the SMS

The link lands in a thread the device already treats as established contact. The target is not evaluating a stranger. They are continuing a conversation they believe they started.

The lesson is not that people should scrutinize voicemail transcripts more carefully. It is that a request to move money, credentials, or access should require verification regardless of how established the thread feels.

Procedural verification

Unglamorous, cheap, and the single thing separating the organizations that hold from the ones that do not. Four steps carry most of the value.

1

Out-of-band callback

To a number already on file, never a number supplied inside the request itself.

2

A second approver

For anything that moves money, credentials, or access. One person should not be able to complete the chain alone.

3

A named escalation path

Faster to use than to ignore, reachable while the call is still live, with no penalty for using it on a request that turns out to be real.

4

Urgency as a trigger, not an override

An explicit rule that pressure from an executive is a reason to verify, not a reason to skip verification.

None of these require a single person to correctly identify a deepfake. That is the point: the organizations sitting at the strong end of the index are not the ones with better eyes.

Where your vertical currently lands

Medians are computed per sector, because channel mix, verification procedure, and regulatory pressure differ enough that a cross-industry average is close to meaningless.

Banking
Wire authorization, service desk resets
Payments & Fintech
High-volume identity decisions
Insurance
Claims and policyholder verification
Digital Identity
Independent control validation
Healthcare
Records access, credential reset
Technology & SaaS
Privileged access, supply chain
Manufacturing
Supplier payment redirection
Energy & Utilities
Operational technology access
Retail & eCommerce
Vendor and refund process
Legal & Professional
Client fund transfers
Government & Public
Citizen identity and benefits
Education
Payroll and account recovery
Positioning against your own vertical is part of every engagement. See sector detail or read what an OSES™ engagement involves.

Common questions

What is the OSES Risk Index?

A benchmark of organizational exposure to orchestrated social engineering, computed from the outcomes of real simulation engagements rather than survey responses. It scores what happened when a credible multi-channel request arrived: whether the action completed, whether the process held, whether anything in the stack observed it, and how fast it was escalated.

How is the index computed?

Every engagement produces the same outcome fields across people, process, and technology. Those fields are normalized for scenario difficulty and channel mix, scored against the same scale, and rolled into a sector median. Because the inputs come from one platform running one methodology, the comparison holds across organizations instead of comparing one vendor's report to another's.

Why does the index measure action instead of clicks?

A click records that a person touched something once, on one channel, in a scenario that happened to contain a link. It does not record whether money moved, whether a credential was reset, or whether anyone raised it. Across our voice campaigns there was no link anywhere and the weighted mean action rate was 14.5 percent, so a phishing platform would have reported nothing at all.

What does a sector median tell me that my own numbers do not?

An isolated rate has no scale attached to it. Knowing that a result sits above or below the median for organizations with comparable process, channel mix, and regulatory exposure is what turns a number into a decision about where to spend. Aggregate cross-industry figures are context, not a benchmark.

Does the index identify individual employees or departments?

No. Reporting is organizational and sector level. There are no named individuals and no department leaderboards, because the failures that matter are procedural and naming people converts a process finding into a personnel one.

Does training move the index?

Partly, and then it stops. Measured against untrained baselines, training reduces the action rate by roughly a third, from about 13 percent to about 8 percent. The remainder is not an awareness gap. It is a process gap, and additional content does not close it.

How often is the index updated?

Continuously, as engagements complete. Sector medians shift as generation quality improves and as organizations re-test, which is the point of computing the index from live outcomes rather than publishing a fixed annual figure.

Can we see where we land without a full engagement?

A scoping call will place you approximately using channel mix, verification procedure, and industry. A position on the index requires a baseline engagement, because the index is computed from observed outcomes and not from a questionnaire.

Find out where your process breaks

Thirty minutes. We walk through a real OSES™ engagement, scenario design through findings, and you decide whether your process would have held.

No IT integration required Runs fully external Peer vertical benchmark
Book Your Demo

Or see a real simulation before you talk to anyone.