Social Engineering Simulation: The Complete Guide
Social Engineering Simulation: The Complete Guide
A simulation becomes a measurement at the point where it reaches something that matters. Everything before that point is engagement data. This is how the engagement is scoped, what the result has to contain, and the one structural number most programs have never calculated.
Most guides to this subject are written from the outside. This one is not. We design and run authorized social engineering simulations against enterprise organizations as paid engagements, and the material below is the working method rather than a description of the category.
The single decision that separates a useful engagement from an expensive one gets made before anything is built, and it is almost never the scenario. It is the answer to one question: what is the specific privileged act this simulation is trying to obtain? Programs that can answer that produce a finding. Programs that cannot produce a participation report.
What a Social Engineering Simulation Actually Is
A social engineering simulation is an authorized exercise in which an organization is contacted the way an adversary would contact it, in order to observe whether a defined privileged action can be obtained. It is not awareness training, which teaches. It is not a phishing test, which usually ends at a click. It continues past first contact into the procedure that was supposed to intervene.
The distinction matters because it changes the subject of the sentence. A phishing test produces a statement about a person: this user clicked. A simulation scoped to a consequential action produces a statement about an organization: the payment instruction changed, and the verification step that should have caught it either did not exist or was waived under pressure.
Those are different findings with different owners and different remediations. One belongs to a training coordinator. The other belongs to whoever owns the process.
Three things are formally in scope in a properly built engagement, and leaving any of them out is where most programs quietly narrow into a click test.
- People. Whether the request was complied with. This is the part every program already measures, and it is the smallest part.
- Process. Whether a documented verification step existed for the action being requested, and whether it ran. This is where the loss is usually permitted.
- Technology. Whether anything in the stack saw the sequence, alerted on it, or blocked the terminal action. Frequently nothing did, because the whole sequence arrived over channels the security stack does not inspect.
Scoping Starts With the Consequential Action
A consequential action is the terminal privileged outcome the engagement is scoped to reach, named in writing before anything is designed. Not "engaged with the caller." Not "responded to the message." A specific act with a real owner and a real system it lands in.
Eight cover most of the meaningful surface in an enterprise.
Naming the action first is not administrative tidiness. It determines the population, the pretext, the channels, the abort criteria, and every number in the report. Run the exercise the other way around, starting from a scenario somebody found interesting, and the result has nothing to anchor to. You end up reporting how many people found the caller convincing, which is a fact about the caller rather than about your controls.
It also sets the abort line. An authorized engagement should reach the point where the action would occur and stop there, with the evidence captured and the transaction never completed. That boundary is agreed in the authorization document, not improvised on the day.
The Six Stages of an Engagement
This is the sequence as we build and run it. None of it requires exotic capability. It requires the discipline to keep the terminal action fixed while everything upstream adapts.
Exposure mapping
What an adversary would want, who holds the authority to grant it, and how much public source material already exists on those people. Organizational charts get rebuilt from public professional profiles. Tooling, help desk hours, and approval structures get inferred from job postings and support documentation. None of this touches the perimeter, so none of it generates a detection.
Consequential action definition
The success condition, written down and agreed. Everything else is built backward from it. This stage also produces the abort criteria and the evidence standard, which is what makes the engagement safe to run against a live population.
Control layer map
Which control should see this, which documented procedure should catch it, and which role has standing to challenge it. If nobody in the organization can answer those three questions for the chosen action, that gap is already a finding, recorded before the simulation runs.
Orchestration
Two channels minimum, three is better, with each artifact referencing the last and crossing a different control boundary. A voicemail lands, a message references it, a live call confirms it. The distinction that matters is conditionality: each stage reads what the target did in the previous stage and decides what happens next. A target who verified independently at stage two never sees stage three, and that fact is itself a recorded result rather than a gap in the data.
Execution
Unannounced, in a real working window, originated entirely externally with nothing installed inside the environment. Announced drills have their place for practice, but an announced exercise measures a rehearsed process rather than the one that runs on an ordinary Tuesday.
Reporting and re-test
A stage-by-stage map of which controls held and which were bypassed, findings named against procedures rather than individuals, and remediation sequenced by layer. Then the same paths run again after the fixes land. A simulation with no re-test is an audit finding. The delta is the only thing that demonstrates a control changed.
We published a full worked example of stages four and five against a live population in the agentic voice engagement write-up, and the help desk variant in the remote support simulation breakdown.
Why Detection Is the Wrong Control to Build On
The instinct in most programs is to train people to recognize the fake. The reason this fails is not that people are careless. It is that the control has the wrong economics.
Detection is a decaying control. Its effectiveness is an inverse function of the adversary's generation quality. Generation quality improves continuously and cheaply. Human perceptual capability does not improve at all. Every artifact you teach people to listen for, the flat prosody, the odd pause, the breath that is not there, is a defect in the current generation of synthesis, and those defects are being engineered out. A control whose value declines from the day the training is delivered, at a rate set by the adversary rather than by the buyer, is a depreciating asset.
Procedural verification is an invariant control. A callback to a known-good number, a second-channel confirmation, or a dual-approval gate returns the same result against a crude imitation and against a convincing one. The control does not read the artifact. It does not care how good the artifact is.
The strategic point is not that process is better than detection in the abstract. It is that the gap between the two widens every quarter without anyone doing anything about it.
Two bounds on that argument are worth stating plainly, because overstating it is how the position gets dismantled by an informed reader.
First, this does not remove the human from the measurement. Procedures are executed by people. What changes is which human behavior gets measured, moving from perception to compliance under pressure, which is a better variable because it is trainable and observable. It is still a human variable, and anyone claiming users have stopped mattering is selling something.
Second, invariance is not total. Better synthetic media does not defeat a procedural control directly. It raises the pressure applied against it. A more convincing executive makes granting the exception feel more reasonable. What degrades as generation quality improves is not the control itself, it is the exception rate. That is measurable, and measuring it is considerably more defensible than claiming perfect invariance, which someone will eventually falsify in public.
First figure measured across more than 1,057 tested participants in Breacher.ai engagements. Second and third figures are shares of a single population: a multinational financial services organization of approximately 500 employees, cloned executive voice, autonomous outbound agent. Full write-up in the agentic voice case study.
Coverage Is the Limiting Factor, Not Adherence
Here is the part that reorders most remediation plans, and it is the reason the eight actions above are listed as a set rather than an example.
A procedural control only holds where a process exists. Most organizations have a documented verification requirement for outbound wire transfers, because that requirement was written after somebody lost money or an auditor insisted. Very few have anything comparable for a credential reset, a vendor banking detail change, an MFA enrollment, or a support request arriving over Teams from an internal-looking account.
Procedural invariance protects the covered paths. It says nothing at all about the uncovered ones, and the uncovered ones are where our engagements find the failures. An organization can have excellent adherence on the one path it covers and still lose, because the adversary was never going to use that path.
Verification coverage is the share of consequential actions that carry a defined verification requirement which is channel-independent and reachable by the person who will be asked to act under time pressure. All three conditions are load-bearing. A requirement that exists in a policy document nobody can find during a live call is declared coverage rather than effective coverage.
Coverage is structural, which means it can be assessed without running a simulation at all. That makes it the natural first step: it produces a finding in days rather than weeks, and it usually reveals that the remediation program is a documentation exercise before it is a training exercise.
What to Measure
Click rate is a delivery diagnostic for the email channel and little else. A phone call contains nothing to click, and in a multi-stage sequence the click is rarely where the loss occurs. Six measures replace it, and every one of them is anchored to the named consequential action.
| Measure | Definition | Why it matters |
|---|---|---|
| Coverage rate | Consequential actions with a defined verification requirement, over total consequential actions identified | The structural number most organizations have never calculated |
| Action rate | Targets who performed the consequential action, over the defined population | The exposure figure. Distinct from engagement, which is not a loss |
| Process hold rate | Verification executed, over verification required | Replaces click rate as the headline measure |
| Exception rate | Verification consciously waived or bypassed, over verification required | The variable that moves as generation quality improves |
| Time to verify | Elapsed time from request to verification action | Separates a working process from a slow one that fails under urgency |
| Uncovered path exposure | Consequential actions with no verification requirement, weighted by impact | Sizes the remediation program directly |
Process hold rate is the live-site term for the measure recorded as adherence rate in our internal methodology paper. The definition is identical.
Note what is absent from that set: there is no per-employee risk score in the headline. Per-user data has a real job, which is routing. We use individual results to aim remediation at the people and the process paths that actually failed. We do not put a score on a person in front of a board, because a score on a person does not tell you whether the payment went out.
If you already run a human risk management tool, this sits above it rather than against it. That tool measures who is likely to fail. This measures whether the organization survives when they do.
These are the questions a report should be able to answer directly.
The fifth one is worth dwelling on. A person who felt something was wrong and disengaged without reporting it is not a success. It is an adversary who gets to call the next name on the list with nobody watching. That is a process finding and it should be written as one.
What We Would Do First
Three things, all free, all executable this week, all of which will tell you more than a vendor demonstration will.
List your consequential actions and mark the covered ones. Take the eight above, add anything specific to your business, and for each one answer whether a documented verification requirement exists, whether it works regardless of which channel the request arrived on, and whether the person who would be asked to perform it could find that requirement inside two minutes under pressure. The ratio you get is your coverage rate. Most organizations who do this honestly are surprised by how low it is.
Audit one identity proofing script against a researched caller. Sit with a help desk analyst, take the script they follow for a password reset, and ask of every verification question whether the answer could be found on a professional network, in a data broker record, or in a public filing. Whatever survives that test is your actual control. It is usually shorter than the script.
Write the sequence, not the channel list. One paragraph describing how somebody would realistically reach something that matters in your organization: who they would call, what they would ask for, and which system the loss lands in. That paragraph is your evaluation script. Hand it to every vendor you talk to, including us, and ask them to reproduce it end to end.
How We Would Know We Are Wrong
We publish findings that cut against us, so the falsification conditions belong in the guide rather than in a footnote. This argument is weakened or broken if any of the following show up in our own data.
- High coverage and high process hold rates still produce comparable failure. If organizations with strong verification coverage lose the consequential action at rates similar to organizations without it, the control is not doing the work we claim for it.
- Exception rate does not rise with fidelity. If a live avatar on a video call produces the same exception rate as a scripted voice call, then the pressure mechanism described above is wrong, and invariance is stronger than we have claimed rather than weaker.
- Detection-trained populations outperform procedurally-trained populations on process hold rate. We consider this unlikely, and we have not tested it directly. It is the cleanest available test of the whole position, and we intend to run it as a pre-registered study and publish it either way.
One more thing that undercuts a good deal of vendor marketing, ours included. Across our engagements, the platform an organization runs has not reliably predicted how it performs. Run the same voicemail and callback technique across different organizations and the action rate ranges from zero to 34.5%. What correlates is whether a verification procedure exists and holds under pressure, and whether people were told what to do rather than what to spot. You run the simulation because you cannot remediate a failure you have never observed, not because the tool is the fix.
If you want the vendor landscape organized by what each architecture can actually measure, that is in our simulation platform generations guide. The scoring behind our benchmark figures is documented in the assessment methodology, and the engine that runs these engagements is described on the OSES™ platform page.
Frequently Asked Questions
What is a social engineering simulation?
A social engineering simulation is an authorized exercise in which an organization is contacted the way an adversary would contact it, in order to observe whether a defined privileged action can be obtained. It is distinct from awareness training, which teaches, and from a phishing test, which usually ends at a click. The simulation continues past first contact into the procedure that was supposed to intervene: the callback, the identity check, the second approver. The result is a statement about the organization rather than about the person who answered the phone.
What is a consequential action, and why does scoping start there?
A consequential action is the specific privileged outcome the simulation is trying to reach, named before anything runs. Common examples are an outbound payment, a vendor banking detail change, a credential reset, an MFA enrollment or reset, a privileged access grant, a software installation performed on request, a data export or disclosure, and a physical access grant. Scoping starts there because every other decision follows from it. If nobody can name the action in advance, the report has nothing to anchor to afterward and the finding collapses into engagement statistics.
Why is click rate the wrong measure for a social engineering simulation?
A click is a delivery diagnostic for one channel. A phone call contains nothing to click, and in a multi-stage sequence the click is rarely the point of loss. In one engagement against a multinational financial services organization of approximately 500 employees, 24.3 percent clicked, 21.6 percent called a spoofed number back and spoke to an autonomous voice agent, and 16.2 percent divulged credentials. A click-rate program reports the first figure and never records the other two, which is where the consequence actually sat.
Can employees be trained to detect synthetic voice and video?
Not durably, and it is the wrong control to build a program on. Across more than 1,057 tested targets, 63 percent could not distinguish synthetic voice or video from a real person while it was happening to them. The deeper problem is economic rather than statistical: detection effectiveness is an inverse function of the adversary's generation quality, and generation quality improves continuously. Procedural verification does not read the artifact, so it returns the same result against a crude imitation and a convincing one.
What does verification coverage mean?
Verification coverage is the share of consequential actions in an organization that carry a defined verification requirement which is channel-independent and reachable by the person who will be asked to act under time pressure. It is a structural property, so it can be assessed without running a simulation. Most organizations have a verification procedure for outbound wire transfers and nothing comparable for credential resets, vendor banking detail changes, MFA enrollment, or a support request arriving over Teams. Coverage, rather than adherence, is usually the limiting factor.
How often should a social engineering simulation be run?
Annually at minimum for a baseline, and quarterly where the population is large or turnover is high. The cadence matters less than the pairing: a simulation with no re-test produces an audit finding rather than a control. The re-test should run the same consequential action paths against the same population so the movement is attributable, rather than running a new scenario that produces a number nobody can compare to the last one.
Find Out Where Your Sequence Terminates
Thirty minutes. We walk a real engagement from scenario design through findings, and you decide whether your process would have held.
Book Your Demo
