The OSES™ Score: A Social Engineering Risk Score Explained
Measuring risk beyond the click.
Modern social engineering cannot be measured in clicks alone. OSES™ measures engagement, risk and consequence.
What the OSES™ score is
Two independent terms, always reported separately, landing on a four-band scale.
A quick disambiguation first, because we use the mark for two related things. OSES™, Orchestrated Social Engineering Simulation, is the methodology: multi-stage, multi-channel sequences coordinated so each stage is informed by how the target behaved in the last one. The OSES™ score is what an engagement produces. This page is about the score.
It answers two questions that a single number cannot answer at the same time. How far did it get, and how many people went with it. One term for each, reported side by side, summed onto a band only after both are visible.
The deepest point any single person reached. Not divided by headcount. Floors at 70 the moment an outcome produces a consequence.
The share of the population that moved. Per capita, and capped, so volume can never mask a contained breach.
The band is what goes in front of a board. The two terms behind it are what tell you which problem you have, and the band is never reported without them.
The failure a single number cannot survive
Click rate is not a prediction of outcome. It is a data point.
The two-term design is not a stylistic preference. It came out of a specific failure that every single-number scoring model runs into. Two organizations, both from our own engagement book, are enough to show it.
Barely anyone engaged at all. One person reached a synthetic login page and submitted credentials, and those credentials reached an external endpoint with nothing in the path stopping them.
A large share of the workforce took a consequential action. Nothing downstream is known to have followed. Broad and shallow: many people moved, no consequence recorded.
On a conventional report A is the good news and B is the emergency. In terms of what actually left the building, that ranking is backwards.
A score that ranks A above B is measuring the wrong thing. A score that ignores what happened in B is measuring the wrong thing in the other direction. Weighting cannot fix it either: for one narrow escalation to register as serious in a per capita formula, that single person would have to count as roughly five thousand people, and at that weight any result containing one escalation pins at the maximum regardless of size.
What Depth measures
Every person is placed at exactly one rung of the escalation ladder, the deepest one they touched. The highest rung anyone reached sets the term.
Rungs are exclusive. Someone who clicked and then submitted credentials is L4 only, never both. L0 and L1 score zero, because opening an email is reach rather than failure, and if reach counted a score could rise simply by sending more messages. Not every rung exists on every vector: a voice campaign has no L2 because there is no click, and that is expected rather than an error.
What Spread measures
Per capita, and capped, so widespread failure can push a score toward Critical without ever being able to mask a contained breach.
Spread counts how much of the population moved, weighted by how far each person went. The weighting inside it is the model's position on click rate, stated numerically instead of rhetorically.
A click sits at a fiftieth of a consequential action, which is why a program cannot look better or worse than it is by changing how many messages it sends.
A click is an indicator, and a data point. It doesn't necessarily predict an outcome.
Find out what your organization scores, and which of the two terms it comes from.
Book Your DemoWhere each vertical lands
Drawn from the engagement book and published to show how the two terms behave, not to state where a sector sits. Read the terms before the total.
Where each department lands
The same book cut by function instead of by industry. These are not additional engagements, so nothing here sums with the table above.
What we would do first
Three things, in this order, and the first two cost nothing.
Split your number in two.Split whatever figure you report today into how far and how many, and put both on the same page. If you report a single number, you are almost certainly measuring one of the two and presenting it as both.
Audit your consequential actions.List every consequential action in the organization and mark which ones have a defined verification requirement that is reachable by the person who will be asked to perform it under time pressure. Outbound payment, vendor banking detail change, credential reset, MFA enrollment or reset, privileged access grant, software installation on request, data export, physical access grant. The coverage gap is usually visible in an afternoon and it does not require a simulation to find.
Run it against the real process.Then run the sequence with downstream observation authorized in the scope, so what comes back is an observation rather than a floor.
Find out what your score is made of
We run the sequence against your real process and report both terms, so the number you take to the board says what actually happened.
Or download the findings deck, every score in this piece set out by function and by sector.
