The OSES™ Score: A Social Engineering Risk Score Explained

Measuring risk beyond the click.

Modern social engineering cannot be measured in clicks alone. OSES™ measures engagement, risk and consequence.

Depth Spread Escalation Ladder Four-Band Scale Vertical Rollups Department Rollups
Measure.
OSES™ Simulate
Train.
OSES™ Behave
Prove it changed.
OSES™ Index

What the OSES™ score is

Two independent terms, always reported separately, landing on a four-band scale.

A quick disambiguation first, because we use the mark for two related things. OSES™, Orchestrated Social Engineering Simulation, is the methodology: multi-stage, multi-channel sequences coordinated so each stage is informed by how the target behaved in the last one. The OSES™ score is what an engagement produces. This page is about the score.

It answers two questions that a single number cannot answer at the same time. How far did it get, and how many people went with it. One term for each, reported side by side, summed onto a band only after both are visible.

Depth Spread Score
Depth
How far did it get?

The deepest point any single person reached. Not divided by headcount. Floors at 70 the moment an outcome produces a consequence.

Spread
How many went with it?

The share of the population that moved. Per capita, and capped, so volume can never mask a contained breach.

Low0 to 39 Medium40 to 59 High60 to 79 Critical80 to 100

The band is what goes in front of a board. The two terms behind it are what tell you which problem you have, and the band is never reported without them.

The failure a single number cannot survive

Click rate is not a prediction of outcome. It is a data point.

The two-term design is not a stylistic preference. It came out of a specific failure that every single-number scoring model runs into. Two organizations, both from our own engagement book, are enough to show it.

Organization A

Barely anyone engaged at all. One person reached a synthetic login page and submitted credentials, and those credentials reached an external endpoint with nothing in the path stopping them.

Conventional report: an excellent result
Organization B

A large share of the workforce took a consequential action. Nothing downstream is known to have followed. Broad and shallow: many people moved, no consequence recorded.

Conventional report: a crisis

On a conventional report A is the good news and B is the emergency. In terms of what actually left the building, that ranking is backwards.

A score that ranks A above B is measuring the wrong thing. A score that ignores what happened in B is measuring the wrong thing in the other direction. Weighting cannot fix it either: for one narrow escalation to register as serious in a per capita formula, that single person would have to count as roughly five thousand people, and at that weight any result containing one escalation pins at the maximum regardless of size.

What Depth measures

Every person is placed at exactly one rung of the escalation ladder, the deepest one they touched. The highest rung anyone reached sets the term.

RungWhat it meansDepth
L0 No response to the lure at all.
L1 Opened the email, answered the call, listened to the voicemail.
L2 Clicked. Reached the page or payload and went no further.
L3 Engaged. Replied, conversed, called back, without acting.
L4 Took a consequential action. Credentials, payment, access, data.
L5 Escalation. The action produced a consequence.

Rungs are exclusive. Someone who clicked and then submitted credentials is L4 only, never both. L0 and L1 score zero, because opening an email is reach rather than failure, and if reach counted a score could rise simply by sending more messages. Not every rung exists on every vector: a voice campaign has no L2 because there is no click, and that is expected rather than an error.

What Spread measures

Per capita, and capped, so widespread failure can push a score toward Critical without ever being able to mask a contained breach.

Spread counts how much of the population moved, weighted by how far each person went. The weighting inside it is the model's position on click rate, stated numerically instead of rhetorically.

Weight inside Spread Relative · one consequential action = full weight
Action / escalation
Engagement
A click

A click sits at a fiftieth of a consequential action, which is why a program cannot look better or worse than it is by changing how many messages it sends.

A click is an indicator, and a data point. It doesn't necessarily predict an outcome.

Find out what your organization scores, and which of the two terms it comes from.

Book Your Demo

Where each vertical lands

Drawn from the engagement book and published to show how the two terms behave, not to state where a sector sits. Read the terms before the total.

Depth Spread Bars scaled 0 to 100
VerticalDepth + SpreadOSES™Band
Technology 72.7 High
Manufacturing 69.8 High
Financial services 56.8 Medium
Legal 46.7 Medium

Where each department lands

The same book cut by function instead of by industry. These are not additional engagements, so nothing here sums with the table above.

DepartmentDepth + SpreadOSES™Band
HR 88.4 Critical
All-staff 83.0 Critical
Finance 80.4 Critical
IT 66.5 High

What we would do first

Three things, in this order, and the first two cost nothing.

1

Split your number in two.Split whatever figure you report today into how far and how many, and put both on the same page. If you report a single number, you are almost certainly measuring one of the two and presenting it as both.

2

Audit your consequential actions.List every consequential action in the organization and mark which ones have a defined verification requirement that is reachable by the person who will be asked to perform it under time pressure. Outbound payment, vendor banking detail change, credential reset, MFA enrollment or reset, privileged access grant, software installation on request, data export, physical access grant. The coverage gap is usually visible in an afternoon and it does not require a simulation to find.

3

Run it against the real process.Then run the sequence with downstream observation authorized in the scope, so what comes back is an observation rather than a floor.

Jason ThatcherFounder and CEO of Breacher.ai and the creator of OSES™. Fifteen years in security operations and offensive testing, previously at ZeroFox, Deepwatch, and GuidePoint Security. He builds and runs orchestrated social engineering simulations against enterprise organizations.

Find out what your score is made of

We run the sequence against your real process and report both terms, so the number you take to the board says what actually happened.

Fully managed No integration required Depth and Spread reported separately
Book Your Demo

Or download the findings deck, every score in this piece set out by function and by sector.