The OSES Risk Index
What we find when the simulation runs, published.
A benchmark computed from real engagement outcomes rather than survey answers: which orchestrated sequences work, where the process gives way first, and why the click was never the number that mattered. Every figure below came out of an OSES™ engagement against a live organization, scored the same way, so the comparison across sectors actually holds.
Organizational and sector reporting only. No named individuals, no department leaderboards.
Three things decide the outcome
Every consequential failure we have observed involved all three at once. The index scores each separately, because they fail for different reasons and get fixed by different budgets.
Did they act, push back, or raise it. Depending on the path, a person is the first line of defense, the last, or the only one. We score what they did with the request and how fast they escalated, not whether they could spot a rendering artifact.
Did the procedure hold under pressure. Callback rules, approval thresholds, identity checks at the service desk. This is where consequential failures actually occur, and it is different in every organization, which is why generic testing and generic training both miss it.
Did anything in the stack see it. These sequences land inside trusted communication platforms and travel through the seams between security tools. Each tool works as designed and nobody owns the gap. That gap is where the attack surface has moved.
The sequences that work best
Ranked by what people did, not by what they opened. Every one of these is a handoff between channels rather than a single message.
Click rate is the wrong measure
The most useful thing three years of engagements can tell you is that the industry's default metric does not describe the risk it claims to.
The strongest organizations we test are not the ones with the sharpest people. They are the ones where a request outside normal process hits a step the requester cannot talk past.
How the index is computed
The index is a by-product of the OSES™ loop, not a separate research exercise. Every engagement emits the same outcome fields, which is what makes one organization's result comparable to another's instead of comparing one vendor's report format to another's.
Every engagement runs under signed authorization, within an agreed scope and window. Client names never appear in our public material, and index reporting is organizational and sector level only. The full computation is documented in the assessment methodology, and the scoring scale is published in the Social Engineering Risk Index.
Three findings that hold across sectors
Aggregate figures are context, not a benchmark. The number that matters to you is where you land against your own vertical.
Training moves it a third, then stops
Reported honestly: awareness training cuts susceptibility by roughly a third and then plateaus. The remaining exposure is not a training problem. It is a process problem, and no amount of additional content closes it.
Detection by eye is below a coin flip
Average accuracy against current-generation synthetic media sits below chance, and it degrades as generation quality improves. Any control that depends on a person telling real from synthetic is a control with a declining success rate.
The variance is inside organizations
Functions that authorize payments or reset access run materially above their own organizational median, because the request that reaches them is the one worth making. The index reports this as a process finding at organizational level, never as a leaderboard.
A chain that walks around the controls
Adversaries are not defeating link protection. They are arriving somewhere it does not apply. This one exploits voicemail transcription on mobile to manufacture prior contact.
Ringless voicemail drop
Synthetic audio is deposited directly into the mailbox. No ring, no missed call notification, nothing for the target to decline.
The handset transcribes it
Transcription runs automatically within a few minutes and renders the message as readable text alongside the caller identity, which does most of the credibility work before any request is made.
Send the SMS
The link lands in a thread the device already treats as established contact. The target is not evaluating a stranger. They are continuing a conversation they believe they started.
Procedural verification
Unglamorous, cheap, and the single thing separating the organizations that hold from the ones that do not. Four steps carry most of the value.
Out-of-band callback
To a number already on file, never a number supplied inside the request itself.
A second approver
For anything that moves money, credentials, or access. One person should not be able to complete the chain alone.
A named escalation path
Faster to use than to ignore, reachable while the call is still live, with no penalty for using it on a request that turns out to be real.
Urgency as a trigger, not an override
An explicit rule that pressure from an executive is a reason to verify, not a reason to skip verification.
Where your vertical currently lands
Medians are computed per sector, because channel mix, verification procedure, and regulatory pressure differ enough that a cross-industry average is close to meaningless.
What the findings look like up close
A cloned voiceprint run against verbal verification controls, and what it revealed about the step-up path behind them.
Read the case study Agentic AIAutonomous agents driving generation and handoff end to end, which is what produced the top-ranked sequence above.
Read the case studyCommon questions
What is the OSES Risk Index?
A benchmark of organizational exposure to orchestrated social engineering, computed from the outcomes of real simulation engagements rather than survey responses. It scores what happened when a credible multi-channel request arrived: whether the action completed, whether the process held, whether anything in the stack observed it, and how fast it was escalated.
How is the index computed?
Every engagement produces the same outcome fields across people, process, and technology. Those fields are normalized for scenario difficulty and channel mix, scored against the same scale, and rolled into a sector median. Because the inputs come from one platform running one methodology, the comparison holds across organizations instead of comparing one vendor's report to another's.
Why does the index measure action instead of clicks?
A click records that a person touched something once, on one channel, in a scenario that happened to contain a link. It does not record whether money moved, whether a credential was reset, or whether anyone raised it. Across our voice campaigns there was no link anywhere and the weighted mean action rate was 14.5 percent, so a phishing platform would have reported nothing at all.
What does a sector median tell me that my own numbers do not?
An isolated rate has no scale attached to it. Knowing that a result sits above or below the median for organizations with comparable process, channel mix, and regulatory exposure is what turns a number into a decision about where to spend. Aggregate cross-industry figures are context, not a benchmark.
Does the index identify individual employees or departments?
No. Reporting is organizational and sector level. There are no named individuals and no department leaderboards, because the failures that matter are procedural and naming people converts a process finding into a personnel one.
Does training move the index?
Partly, and then it stops. Measured against untrained baselines, training reduces the action rate by roughly a third, from about 13 percent to about 8 percent. The remainder is not an awareness gap. It is a process gap, and additional content does not close it.
How often is the index updated?
Continuously, as engagements complete. Sector medians shift as generation quality improves and as organizations re-test, which is the point of computing the index from live outcomes rather than publishing a fixed annual figure.
Can we see where we land without a full engagement?
A scoping call will place you approximately using channel mix, verification procedure, and industry. A position on the index requires a baseline engagement, because the index is computed from observed outcomes and not from a questionnaire.
Find out where your process breaks
Thirty minutes. We walk through a real OSES™ engagement, scenario design through findings, and you decide whether your process would have held.
Or see a real simulation before you talk to anyone.
