Assessment Methodology - Breacher.ai

Assessment Methodology

You cannot measure what you do not know. Step one is understanding your risk.

This is how the measurement works. What the assessment submits, what it records, how those outcomes become a single organizational score, and how the re-test establishes that the score moved. The same computation runs on every engagement, which is what makes the number comparable to your sector and defensible to an auditor.

Exposure Mapping Orchestrated Execution Process Outcomes Risk Index Sector Benchmark Re-test Delta

Authorized, fully external, and reported at the organization level with no named individuals.

One dataset, carried from the first measurement to the last

3
Stages run on a single dataset: measure risk, train for what you find, prove it changed
OSES™ program spine
6
Channels that can be orchestrated inside one scenario under a single objective
Voice, video, conference call, email, chat, and inbound callback
12
Sectors carried in the Social Engineering Risk Index for benchmark positioning
Current index scope

Why the baseline comes first

Most programs start by buying training and then look for a number to justify it. That order makes the number unfalsifiable.

A plan without a baseline is a guess

Choosing modules before measuring means choosing them from a catalogue rather than from evidence. The training may well be good. There is just no way to know whether it addresses the process that would actually fail.

Completion is not a control

A completion rate records that a person finished a module. It does not record whether the callback procedure held when someone with authority asked for an exception. Those are different measurements and only one of them is a control.

Improvement needs a starting point

A delta requires two measurements taken the same way. Without a baseline computed identically to the re-test, there is nothing to compare against and no defensible claim that anything improved.

How the assessment runs

OSES™, our Orchestrated Social Engineering Simulation framework. Five stages, run in order, producing one dataset that the training and the re-test both read from rather than three disconnected exercises that happen to share a vendor.

01
Reconnaissance
Open-source intelligence establishes who is reachable, through which channels, and which identities an adversary would impersonate
02
Scenario Design
Pretext, sequence, and synthetic voice, video, and identity assets built against the processes you are accountable for
03
Orchestrated Execution
The scenario runs across whichever channels it calls for, coordinated under one objective rather than one channel at a time
04
Outcome Analysis
What was requested, what was verified, where escalation happened, and at which step the chain broke
05
Index & Re-test
Outcomes normalized into the risk index, positioned against your sector, and re-measured after remediation

Every engagement runs under signed authorization within an agreed scope, environment, and window, with named approvers, agreed impersonated identities, and documented abort conditions. Nothing is installed and nothing touches your security stack. For the definition of the framework itself, see what OSES™ is.

Six dimensions, recorded on every engagement

These are the inputs to the index. Each one is a process property rather than a person property, which is what makes the score reproducible across engagements and across sectors.

Exposure surface

How much of the organization is reachable without insider access: which roles are discoverable, which channels accept contact from outside, and how much public material exists to build a credible impersonation from.

OSINT FootprintReachable RolesClonable Media
Channel susceptibility

Which delivery channel actually carried the request through. Organizations that hold firm on email frequently do not hold on a voice call, and a single-channel test would report that as a pass.

VoiceVideoMessage
Process hold rate

Whether the written verification step was performed as written when a request arrived from someone with the authority to make it. This is the dimension that carries the most weight, because it is the one that is actually a control.

Callback VerificationApproval ChainCredential Reset
Escalation behaviour

Whether anyone raised it, who they raised it to, and whether that path led anywhere. An organization where the request was refused but never reported has a different exposure than one where it was refused and escalated within minutes.

Reporting PathEscalation ReachSuppression
Time to friction

How long the scenario ran before it met a step that slowed it down, and how long before it met one that stopped it. Time matters because the loss in a real intrusion is usually completed inside the first hour.

Time to DetectTime to ReportDwell
Remediation delta

The change between the baseline and the re-test on the same paths, computed the same way. This is the dimension that turns a finding into evidence that something was fixed, rather than evidence that something was purchased.

Re-testSame ComputationMeasured Change

Engagement metrics and risk are not the same measurement

The distinction is not stylistic. One set of numbers describes what people did with a message. The other describes whether a control functioned.

Commonly reported
Engagement metrics
Click rate on a simulated message
Open and reply rates by department
Training completion percentage
Repeat clicker lists naming individuals
A single channel, tested in isolation
These describe behaviour toward one message. They do not establish whether the procedure behind it held.
What we record instead
Process outcomes
Whether verification was performed as written
Which step in the chain the request got through
Whether it was escalated, and how far it reached
How long the scenario ran before it met friction
The same paths re-measured after remediation
Reported at the organization level. No named individuals and no department leaderboards, on any engagement.

How the score is computed, defensibly

A number is only useful if someone outside your team can interrogate how it was produced. This is that explanation.

1

Outcomes are normalized for scenario difficulty

A bespoke scenario aimed at a named approver is harder to withstand than a generic pretext, so raw pass rates are not comparable across engagements. Normalization is what allows one organization's result to sit next to another's without flattering whoever received the easier test.

2

Dimensions are weighted by consequence

A bypassed payment authorization does not carry the same weight as a slow report of a suspicious message. Weighting is set by the severity of the process that failed, not by how visible the failure was.

3

The result is expressed as one organizational score

A single figure a board can hold, with the six contributing dimensions shown underneath it so the score can be taken apart rather than taken on trust. Nothing in the breakdown identifies an individual.

4

Position is set against your own sector

The comparison is drawn against the median for your industry rather than a general population, because the processes under test and the adversaries targeting them differ by sector. No client is identifiable in any published figure.

5

The re-test uses the identical computation

The same paths, normalized and weighted the same way, so the difference between the two runs is a property of your organization rather than a property of the measurement. That difference is the only thing that establishes the training worked.

The index is published as sector medians computed from real engagement outcomes. See the Social Engineering Risk Index for current positions, or the threat intelligence we publish from the field.

The finding is what the training is built from

Measure risk, train for what you find, prove it changed. Three stages, one dataset, which is the part that makes the last one possible.

Measure

Baseline exposure

The assessment establishes where the process breaks and how far a synthetic identity gets before something stops it.

  • Organization-level findings
  • Six recorded dimensions
  • Sector benchmark position
Train

Built from your own data

Modules are assembled against the specific procedure that failed, which is what Gartner now describes as secure behavior management rather than awareness delivery.

  • Aimed at the process that broke
  • Role-relevant, not catalogue-wide
  • Short enough to actually be completed
Prove

Re-test, not report

The same paths run again on the same computation, so the improvement is a measured delta on a control rather than a completion statistic.

  • Identical normalization and weighting
  • Delta on the failed procedure
  • Evidence auditors accept as written
Trusted by security leaders
G2
★★★★★
5.0 out of 5
Gartner Peer Insights
★★★★★
5.0 out of 5

I think the entire company is already talking about voice cloning and the risks. It's been a huge win for us already.

Questions about the measurement

Why start with an assessment rather than training?

Because a training plan built without a baseline is a guess about which process is weak. The assessment establishes where the exposure actually is, which makes the training specific and makes the improvement measurable afterwards. You cannot measure what you do not know.

What does the assessment actually record?

Process outcomes rather than engagement metrics. Whether verification was performed, whether the request was escalated, how long it took, at which step the chain broke, and which channel carried it. A click is one input among several, not the finding.

How is the Social Engineering Risk Index computed?

Recorded outcomes are normalized so that scenarios of different difficulty are comparable, weighted by the severity of the process that failed, and expressed as a single organizational score with the contributing dimensions shown underneath it. The same computation runs on every engagement, which is what makes the sector median meaningful.

Do you report on individual employees?

No. Reporting is organizational. There are no named individuals and no department leaderboards. The objective is to test whether the procedure holds, not to identify a person who did not catch a simulation.

How is the benchmark comparison established?

Your score is positioned against the median for your own sector rather than a general population, because the processes being tested and the adversaries targeting them differ by industry. Comparisons are drawn from normalized outcomes, and no client is identifiable in any published figure.

How does the re-test prove anything changed?

The same paths are run again after remediation using the same computation, so the result is a measured delta on the process that previously failed. A course completion rate says a person finished a module. A re-test says the procedure held the second time.

Is the assessment authorized, and how is scope controlled?

Every engagement runs under signed authorization within an agreed scope, environment, and window, with named approvers, agreed impersonated identities, and documented abort conditions. Nothing is installed and nothing touches your security stack.

How long does a baseline take?

Two to three weeks from scoping call to readout, with the execution window itself usually a few days inside that. Re-tests are shorter because the scenario design already exists.

Step one is understanding your risk

Thirty minutes to scope a baseline: which process would hurt most if it failed, which channels reach the people who own it, and what the readout needs to prove.

Fully managed Nothing installed Findings in 2 to 3 weeks

You cannot measure what you do not know. Start with Breacher.ai.