Social Engineering Simulation: The Complete Guide

Categories: Deepfake,Published On: September 4th, 2026,
Social Engineering Simulation: The Complete Guide
Methodology

Social Engineering Simulation: The Complete Guide

A simulation becomes a measurement at the point where it reaches something that matters. Everything before that point is engagement data. This is how the engagement is scoped, what the result has to contain, and the one structural number most programs have never calculated.

By Jason Thatcher, Founder and CEO, Breacher.ai

Bring your own sequence. We will run it end to end and show you where it terminates.

Book Your Demo

Most guides to this subject are written from the outside. This one is not. We design and run authorized social engineering simulations against enterprise organizations as paid engagements, and the material below is the working method rather than a description of the category.

The single decision that separates a useful engagement from an expensive one gets made before anything is built, and it is almost never the scenario. It is the answer to one question: what is the specific privileged act this simulation is trying to obtain? Programs that can answer that produce a finding. Programs that cannot produce a participation report.

What a Social Engineering Simulation Actually Is

Definition

A social engineering simulation is an authorized exercise in which an organization is contacted the way an adversary would contact it, in order to observe whether a defined privileged action can be obtained. It is not awareness training, which teaches. It is not a phishing test, which usually ends at a click. It continues past first contact into the procedure that was supposed to intervene.

The distinction matters because it changes the subject of the sentence. A phishing test produces a statement about a person: this user clicked. A simulation scoped to a consequential action produces a statement about an organization: the payment instruction changed, and the verification step that should have caught it either did not exist or was waived under pressure.

Those are different findings with different owners and different remediations. One belongs to a training coordinator. The other belongs to whoever owns the process.

Three things are formally in scope in a properly built engagement, and leaving any of them out is where most programs quietly narrow into a click test.

  • People. Whether the request was complied with. This is the part every program already measures, and it is the smallest part.
  • Process. Whether a documented verification step existed for the action being requested, and whether it ran. This is where the loss is usually permitted.
  • Technology. Whether anything in the stack saw the sequence, alerted on it, or blocked the terminal action. Frequently nothing did, because the whole sequence arrived over channels the security stack does not inspect.
A phishing test tells you someone clicked. A simulation tells you whether the organization would have survived the thing that follows the click.

Scoping Starts With the Consequential Action

A consequential action is the terminal privileged outcome the engagement is scoped to reach, named in writing before anything is designed. Not "engaged with the caller." Not "responded to the message." A specific act with a real owner and a real system it lands in.

Eight cover most of the meaningful surface in an enterprise.

ACTION 01Outbound payment released
ACTION 02Vendor banking detail changed
ACTION 03Credential reset performed
ACTION 04MFA enrollment or reset completed
ACTION 05Privileged access granted
ACTION 06Software installed on request
ACTION 07Data exported or disclosed
ACTION 08Physical access granted

Naming the action first is not administrative tidiness. It determines the population, the pretext, the channels, the abort criteria, and every number in the report. Run the exercise the other way around, starting from a scenario somebody found interesting, and the result has nothing to anchor to. You end up reporting how many people found the caller convincing, which is a fact about the caller rather than about your controls.

It also sets the abort line. An authorized engagement should reach the point where the action would occur and stop there, with the evidence captured and the transaction never completed. That boundary is agreed in the authorization document, not improvised on the day.

The Six Stages of an Engagement

This is the sequence as we build and run it. None of it requires exotic capability. It requires the discipline to keep the terminal action fixed while everything upstream adapts.

STAGE 01

Exposure mapping

What an adversary would want, who holds the authority to grant it, and how much public source material already exists on those people. Organizational charts get rebuilt from public professional profiles. Tooling, help desk hours, and approval structures get inferred from job postings and support documentation. None of this touches the perimeter, so none of it generates a detection.

STAGE 02

Consequential action definition

The success condition, written down and agreed. Everything else is built backward from it. This stage also produces the abort criteria and the evidence standard, which is what makes the engagement safe to run against a live population.

STAGE 03

Control layer map

Which control should see this, which documented procedure should catch it, and which role has standing to challenge it. If nobody in the organization can answer those three questions for the chosen action, that gap is already a finding, recorded before the simulation runs.

STAGE 05

Execution

Unannounced, in a real working window, originated entirely externally with nothing installed inside the environment. Announced drills have their place for practice, but an announced exercise measures a rehearsed process rather than the one that runs on an ordinary Tuesday.

STAGE 06

Reporting and re-test

A stage-by-stage map of which controls held and which were bypassed, findings named against procedures rather than individuals, and remediation sequenced by layer. Then the same paths run again after the fixes land. A simulation with no re-test is an audit finding. The delta is the only thing that demonstrates a control changed.

We published a full worked example of stages four and five against a live population in the agentic voice engagement write-up, and the help desk variant in the remote support simulation breakdown.

Why Detection Is the Wrong Control to Build On

The instinct in most programs is to train people to recognize the fake. The reason this fails is not that people are careless. It is that the control has the wrong economics.

Detection is a decaying control. Its effectiveness is an inverse function of the adversary's generation quality. Generation quality improves continuously and cheaply. Human perceptual capability does not improve at all. Every artifact you teach people to listen for, the flat prosody, the odd pause, the breath that is not there, is a defect in the current generation of synthesis, and those defects are being engineered out. A control whose value declines from the day the training is delivered, at a rate set by the adversary rather than by the buyer, is a depreciating asset.

Procedural verification is an invariant control. A callback to a known-good number, a second-channel confirmation, or a dual-approval gate returns the same result against a crude imitation and against a convincing one. The control does not read the artifact. It does not care how good the artifact is.

The strategic point is not that process is better than detection in the abstract. It is that the gap between the two widens every quarter without anyone doing anything about it.

No bank measures whether a teller can eyeball a counterfeit note. They measure whether the teller ran the check. The bill got better. The check did not have to.

Two bounds on that argument are worth stating plainly, because overstating it is how the position gets dismantled by an informed reader.

First, this does not remove the human from the measurement. Procedures are executed by people. What changes is which human behavior gets measured, moving from perception to compliance under pressure, which is a better variable because it is trainable and observable. It is still a human variable, and anyone claiming users have stopped mattering is selling something.

Second, invariance is not total. Better synthetic media does not defeat a procedural control directly. It raises the pressure applied against it. A more convincing executive makes granting the exception feel more reasonable. What degrades as generation quality improves is not the control itself, it is the exception rate. That is measurable, and measuring it is considerably more defensible than claiming perfect invariance, which someone will eventually falsify in public.

63%
Of more than 1,057 tested participants could not distinguish synthetic voice or video from a real person while it was happening to them
21.6%
Of one 500-employee population called a spoofed number back and spoke to an autonomous voice agent, unprompted
16.2%
Of that same population divulged credentials, after a 24.3% click rate on the opening message

First figure measured across more than 1,057 tested participants in Breacher.ai engagements. Second and third figures are shares of a single population: a multinational financial services organization of approximately 500 employees, cloned executive voice, autonomous outbound agent. Full write-up in the agentic voice case study.

Coverage Is the Limiting Factor, Not Adherence

Here is the part that reorders most remediation plans, and it is the reason the eight actions above are listed as a set rather than an example.

A procedural control only holds where a process exists. Most organizations have a documented verification requirement for outbound wire transfers, because that requirement was written after somebody lost money or an auditor insisted. Very few have anything comparable for a credential reset, a vendor banking detail change, an MFA enrollment, or a support request arriving over Teams from an internal-looking account.

Procedural invariance protects the covered paths. It says nothing at all about the uncovered ones, and the uncovered ones are where our engagements find the failures. An organization can have excellent adherence on the one path it covers and still lose, because the adversary was never going to use that path.

Adherence measured on one covered path is a number about the path, not about the organization.

Verification coverage is the share of consequential actions that carry a defined verification requirement which is channel-independent and reachable by the person who will be asked to act under time pressure. All three conditions are load-bearing. A requirement that exists in a policy document nobody can find during a live call is declared coverage rather than effective coverage.

Coverage is structural, which means it can be assessed without running a simulation at all. That makes it the natural first step: it produces a finding in days rather than weeks, and it usually reveals that the remediation program is a documentation exercise before it is a training exercise.

What to Measure

Click rate is a delivery diagnostic for the email channel and little else. A phone call contains nothing to click, and in a multi-stage sequence the click is rarely where the loss occurs. Six measures replace it, and every one of them is anchored to the named consequential action.

MeasureDefinitionWhy it matters
Coverage rateConsequential actions with a defined verification requirement, over total consequential actions identifiedThe structural number most organizations have never calculated
Action rateTargets who performed the consequential action, over the defined populationThe exposure figure. Distinct from engagement, which is not a loss
Process hold rateVerification executed, over verification requiredReplaces click rate as the headline measure
Exception rateVerification consciously waived or bypassed, over verification requiredThe variable that moves as generation quality improves
Time to verifyElapsed time from request to verification actionSeparates a working process from a slow one that fails under urgency
Uncovered path exposureConsequential actions with no verification requirement, weighted by impactSizes the remediation program directly

Process hold rate is the live-site term for the measure recorded as adherence rate in our internal methodology paper. The definition is identical.

Note what is absent from that set: there is no per-employee risk score in the headline. Per-user data has a real job, which is routing. We use individual results to aim remediation at the people and the process paths that actually failed. We do not put a score on a person in front of a board, because a score on a person does not tell you whether the payment went out.

If you already run a human risk management tool, this sits above it rather than against it. That tool measures who is likely to fail. This measures whether the organization survives when they do.

These are the questions a report should be able to answer directly.

What share of the defined population performed the consequential action, as distinct from merely engaging with the caller?
Where a verification requirement existed, did it run, or was it waived under pressure?
Which consequential actions in scope had no verification requirement at all?
How long from first contact to the first report reaching security, and by which route?
Where a person resisted, did the process give them a defined next step, or did they simply hang up?
How does the action rate compare against a peer benchmark with a stated number of organizations behind it?

The fifth one is worth dwelling on. A person who felt something was wrong and disengaged without reporting it is not a success. It is an adversary who gets to call the next name on the list with nobody watching. That is a process finding and it should be written as one.

What We Would Do First

Three things, all free, all executable this week, all of which will tell you more than a vendor demonstration will.

List your consequential actions and mark the covered ones. Take the eight above, add anything specific to your business, and for each one answer whether a documented verification requirement exists, whether it works regardless of which channel the request arrived on, and whether the person who would be asked to perform it could find that requirement inside two minutes under pressure. The ratio you get is your coverage rate. Most organizations who do this honestly are surprised by how low it is.

Audit one identity proofing script against a researched caller. Sit with a help desk analyst, take the script they follow for a password reset, and ask of every verification question whether the answer could be found on a professional network, in a data broker record, or in a public filing. Whatever survives that test is your actual control. It is usually shorter than the script.

Write the sequence, not the channel list. One paragraph describing how somebody would realistically reach something that matters in your organization: who they would call, what they would ask for, and which system the loss lands in. That paragraph is your evaluation script. Hand it to every vendor you talk to, including us, and ask them to reproduce it end to end.

How We Would Know We Are Wrong

We publish findings that cut against us, so the falsification conditions belong in the guide rather than in a footnote. This argument is weakened or broken if any of the following show up in our own data.

  • High coverage and high process hold rates still produce comparable failure. If organizations with strong verification coverage lose the consequential action at rates similar to organizations without it, the control is not doing the work we claim for it.
  • Exception rate does not rise with fidelity. If a live avatar on a video call produces the same exception rate as a scripted voice call, then the pressure mechanism described above is wrong, and invariance is stronger than we have claimed rather than weaker.
  • Detection-trained populations outperform procedurally-trained populations on process hold rate. We consider this unlikely, and we have not tested it directly. It is the cleanest available test of the whole position, and we intend to run it as a pre-registered study and publish it either way.

One more thing that undercuts a good deal of vendor marketing, ours included. Across our engagements, the platform an organization runs has not reliably predicted how it performs. Run the same voicemail and callback technique across different organizations and the action rate ranges from zero to 34.5%. What correlates is whether a verification procedure exists and holds under pressure, and whether people were told what to do rather than what to spot. You run the simulation because you cannot remediate a failure you have never observed, not because the tool is the fix.

Measure your risk. Train for what you find. Prove it changed.

If you want the vendor landscape organized by what each architecture can actually measure, that is in our simulation platform generations guide. The scoring behind our benchmark figures is documented in the assessment methodology, and the engine that runs these engagements is described on the OSES™ platform page.

Frequently Asked Questions

What is a social engineering simulation?

A social engineering simulation is an authorized exercise in which an organization is contacted the way an adversary would contact it, in order to observe whether a defined privileged action can be obtained. It is distinct from awareness training, which teaches, and from a phishing test, which usually ends at a click. The simulation continues past first contact into the procedure that was supposed to intervene: the callback, the identity check, the second approver. The result is a statement about the organization rather than about the person who answered the phone.

What is a consequential action, and why does scoping start there?

A consequential action is the specific privileged outcome the simulation is trying to reach, named before anything runs. Common examples are an outbound payment, a vendor banking detail change, a credential reset, an MFA enrollment or reset, a privileged access grant, a software installation performed on request, a data export or disclosure, and a physical access grant. Scoping starts there because every other decision follows from it. If nobody can name the action in advance, the report has nothing to anchor to afterward and the finding collapses into engagement statistics.

Why is click rate the wrong measure for a social engineering simulation?

A click is a delivery diagnostic for one channel. A phone call contains nothing to click, and in a multi-stage sequence the click is rarely the point of loss. In one engagement against a multinational financial services organization of approximately 500 employees, 24.3 percent clicked, 21.6 percent called a spoofed number back and spoke to an autonomous voice agent, and 16.2 percent divulged credentials. A click-rate program reports the first figure and never records the other two, which is where the consequence actually sat.

Can employees be trained to detect synthetic voice and video?

Not durably, and it is the wrong control to build a program on. Across more than 1,057 tested targets, 63 percent could not distinguish synthetic voice or video from a real person while it was happening to them. The deeper problem is economic rather than statistical: detection effectiveness is an inverse function of the adversary's generation quality, and generation quality improves continuously. Procedural verification does not read the artifact, so it returns the same result against a crude imitation and a convincing one.

What does verification coverage mean?

Verification coverage is the share of consequential actions in an organization that carry a defined verification requirement which is channel-independent and reachable by the person who will be asked to act under time pressure. It is a structural property, so it can be assessed without running a simulation. Most organizations have a verification procedure for outbound wire transfers and nothing comparable for credential resets, vendor banking detail changes, MFA enrollment, or a support request arriving over Teams. Coverage, rather than adherence, is usually the limiting factor.

How often should a social engineering simulation be run?

Annually at minimum for a baseline, and quarterly where the population is large or turnover is high. The cadence matters less than the pairing: a simulation with no re-test produces an audit finding rather than a control. The re-test should run the same consequential action paths against the same population so the movement is attributable, rather than running a new scenario that produces a number nobody can compare to the last one.

JT

Jason ThatcherFounder and CEO of Breacher.ai and creator of OSES™. Fifteen years in security operations and offensive testing, previously at ZeroFox, Deepwatch, and GuidePoint Security. He builds and runs orchestrated social engineering simulations against enterprise organizations.

Find Out Where Your Sequence Terminates

Thirty minutes. We walk a real engagement from scenario design through findings, and you decide whether your process would have held.

Book Your Demo

Latest Posts

  • Social Engineering Simulation: The Complete Guide

  • Voice Phishing Assessment: What to Scope, and What the Report Has to Prove

  • Protecting Against Deepfake Scams: Deepfake Phishing Is a Campaign, Not just a Call

Table Of Contents

About the Author: Jason Thatcher

Jason Thatcher is the Founder of Breacher.ai and comes from a long career of working in the Cybersecurity Industry. His past accomplishments include winning Splunk Solution of the Year in 2022 for Security Operations.

Share this post