Assessment Methodology
You cannot measure what you do not know. Step one is understanding your risk.
This is how the measurement works. What the assessment submits, what it records, how those outcomes become a single organizational score, and how the re-test establishes that the score moved. The same computation runs on every engagement, which is what makes the number comparable to your sector and defensible to an auditor.
Authorized, fully external, and reported at the organization level with no named individuals.
Why the baseline comes first
Most programs start by buying training and then look for a number to justify it. That order makes the number unfalsifiable.
A plan without a baseline is a guess
Choosing modules before measuring means choosing them from a catalogue rather than from evidence. The training may well be good. There is just no way to know whether it addresses the process that would actually fail.
Completion is not a control
A completion rate records that a person finished a module. It does not record whether the callback procedure held when someone with authority asked for an exception. Those are different measurements and only one of them is a control.
Improvement needs a starting point
A delta requires two measurements taken the same way. Without a baseline computed identically to the re-test, there is nothing to compare against and no defensible claim that anything improved.
How the assessment runs
OSES™, our Orchestrated Social Engineering Simulation framework. Five stages, run in order, producing one dataset that the training and the re-test both read from rather than three disconnected exercises that happen to share a vendor.
Every engagement runs under signed authorization within an agreed scope, environment, and window, with named approvers, agreed impersonated identities, and documented abort conditions. Nothing is installed and nothing touches your security stack. For the definition of the framework itself, see what OSES™ is.
Six dimensions, recorded on every engagement
These are the inputs to the index. Each one is a process property rather than a person property, which is what makes the score reproducible across engagements and across sectors.
How much of the organization is reachable without insider access: which roles are discoverable, which channels accept contact from outside, and how much public material exists to build a credible impersonation from.
Which delivery channel actually carried the request through. Organizations that hold firm on email frequently do not hold on a voice call, and a single-channel test would report that as a pass.
Whether the written verification step was performed as written when a request arrived from someone with the authority to make it. This is the dimension that carries the most weight, because it is the one that is actually a control.
Whether anyone raised it, who they raised it to, and whether that path led anywhere. An organization where the request was refused but never reported has a different exposure than one where it was refused and escalated within minutes.
How long the scenario ran before it met a step that slowed it down, and how long before it met one that stopped it. Time matters because the loss in a real intrusion is usually completed inside the first hour.
The change between the baseline and the re-test on the same paths, computed the same way. This is the dimension that turns a finding into evidence that something was fixed, rather than evidence that something was purchased.
Engagement metrics and risk are not the same measurement
The distinction is not stylistic. One set of numbers describes what people did with a message. The other describes whether a control functioned.
How the score is computed, defensibly
A number is only useful if someone outside your team can interrogate how it was produced. This is that explanation.
Outcomes are normalized for scenario difficulty
A bespoke scenario aimed at a named approver is harder to withstand than a generic pretext, so raw pass rates are not comparable across engagements. Normalization is what allows one organization's result to sit next to another's without flattering whoever received the easier test.
Dimensions are weighted by consequence
A bypassed payment authorization does not carry the same weight as a slow report of a suspicious message. Weighting is set by the severity of the process that failed, not by how visible the failure was.
The result is expressed as one organizational score
A single figure a board can hold, with the six contributing dimensions shown underneath it so the score can be taken apart rather than taken on trust. Nothing in the breakdown identifies an individual.
Position is set against your own sector
The comparison is drawn against the median for your industry rather than a general population, because the processes under test and the adversaries targeting them differ by sector. No client is identifiable in any published figure.
The re-test uses the identical computation
The same paths, normalized and weighted the same way, so the difference between the two runs is a property of your organization rather than a property of the measurement. That difference is the only thing that establishes the training worked.
The finding is what the training is built from
Measure risk, train for what you find, prove it changed. Three stages, one dataset, which is the part that makes the last one possible.
Baseline exposure
The assessment establishes where the process breaks and how far a synthetic identity gets before something stops it.
- Organization-level findings
- Six recorded dimensions
- Sector benchmark position
Built from your own data
Modules are assembled against the specific procedure that failed, which is what Gartner now describes as secure behavior management rather than awareness delivery.
- Aimed at the process that broke
- Role-relevant, not catalogue-wide
- Short enough to actually be completed
Re-test, not report
The same paths run again on the same computation, so the improvement is a measured delta on a control rather than a completion statistic.
- Identical normalization and weighting
- Delta on the failed procedure
- Evidence auditors accept as written
I think the entire company is already talking about voice cloning and the risks. It's been a huge win for us already.
Questions about the measurement
Why start with an assessment rather than training?
Because a training plan built without a baseline is a guess about which process is weak. The assessment establishes where the exposure actually is, which makes the training specific and makes the improvement measurable afterwards. You cannot measure what you do not know.
What does the assessment actually record?
Process outcomes rather than engagement metrics. Whether verification was performed, whether the request was escalated, how long it took, at which step the chain broke, and which channel carried it. A click is one input among several, not the finding.
How is the Social Engineering Risk Index computed?
Recorded outcomes are normalized so that scenarios of different difficulty are comparable, weighted by the severity of the process that failed, and expressed as a single organizational score with the contributing dimensions shown underneath it. The same computation runs on every engagement, which is what makes the sector median meaningful.
Do you report on individual employees?
No. Reporting is organizational. There are no named individuals and no department leaderboards. The objective is to test whether the procedure holds, not to identify a person who did not catch a simulation.
How is the benchmark comparison established?
Your score is positioned against the median for your own sector rather than a general population, because the processes being tested and the adversaries targeting them differ by industry. Comparisons are drawn from normalized outcomes, and no client is identifiable in any published figure.
How does the re-test prove anything changed?
The same paths are run again after remediation using the same computation, so the result is a measured delta on the process that previously failed. A course completion rate says a person finished a module. A re-test says the procedure held the second time.
Is the assessment authorized, and how is scope controlled?
Every engagement runs under signed authorization within an agreed scope, environment, and window, with named approvers, agreed impersonated identities, and documented abort conditions. Nothing is installed and nothing touches your security stack.
How long does a baseline take?
Two to three weeks from scoping call to readout, with the execution window itself usually a few days inside that. Re-tests are shorter because the scenario design already exists.
Step one is understanding your risk
Thirty minutes to scope a baseline: which process would hurt most if it failed, which channels reach the people who own it, and what the readout needs to prove.
You cannot measure what you do not know. Start with Breacher.ai.
