The Callback Is the Simulation | Breacher.ai

Categories: Deepfake,Published On: August 18th, 2026,
  • Dark Breacher.ai banner reading THE CALLBACK IS THE SIMULATION over a loop diagram in which the outbound call and voicemail arc is dashed and dimmed while the return callback arc is drawn solid bright green with an arrowhead.
Orchestrated Deepfake Phishing Simulation: The Callback
Methodology

The Callback Is the Simulation

A deepfake phone call on its own is a weak test. The voicemail, and the number the target dials back, is where the credentials actually move. If your deepfake phishing simulation does not orchestrate the return leg, you have measured the easy half.

By Jason Thatcher, Founder and CEO, Breacher.ai

See an orchestrated sequence run end to end against your own process, callback included.

Book Your Demo

There is a version of deepfake testing that has become the industry default, and it does not work very well. A cloned voice places a call. Someone answers or does not. A number gets written down. The report says a percentage of the population engaged with a synthetic caller, and everyone agrees the technology is impressive.

We run these engagements for a living, and the single most useful thing three years of them has taught us is that the outbound call is the weakest leg of the sequence. It is the leg an adversary has the least confidence in, the leg most likely to be refused, and the leg most likely to be reported. It is also, almost universally, the only leg anybody simulates.

The consequences we see land on the inbound. A voicemail goes out. Some hours later the target picks up their own phone and dials the number in it. That return call is where the credential goes across. In our red team work, the return path is where the majority of credential handovers happen, and it is the path with the least procedural protection standing in front of it.

This post is the argument for why orchestration is not a feature of a deepfake phishing simulation. It is the thing that determines what the simulation is capable of observing at all.

The Voice Channel Produces Action That Nothing Records

Before the callback argument, the baseline. Voice campaigns produce consequential outcomes in populations where an email-first tool would have reported an empty result, because there was no link anywhere in the sequence for it to count.

14.5%
Weighted mean action rate in voice campaigns that contained no link at all. A phishing platform would have reported nothing.
63%
Of people tested could not distinguish synthetic voice from a real person.
78%
Of assessed organizations rated highly vulnerable, scored on the verification process failing rather than on one person slipping.

Source: Breacher.ai OSES™ engagement outcomes, published in the OSES Risk Index. Detection figures are measured across engagements covering more than one thousand individual targets. Reporting is organizational and sector level only.

Those figures describe the channel. They do not tell you where inside the channel the failure sits, and that is the question orchestration exists to answer.

The Sequence We Actually Run

OSES™, Orchestrated Social Engineering Simulations™, sequences each contact against the state left by the one before it. Below is the IT support variant, which is the one that produces our most serious findings. Note where the featured card sits.

PHASE 01

Reconnaissance and pretext construction

Reporting lines, help desk hours, ticketing system, device fleet, and VPN vendor are assembled from public sources. Voice samples come from webinars, recorded panels, and voicemail greetings. Nothing in this phase touches the perimeter, so nothing in this phase generates a detection.

PHASE 02

The outbound call, which is expected to miss

The call is placed in a cloned voice. If it is answered, the agent holds a live conversation. Most of the time it is not answered, and that is not a failure of the sequence. In practice the majority of outbound calls in a campaign go to voicemail, which makes the next phase the dominant route through the population rather than the fallback.

PHASE 03

The voicemail that manufactures a prior relationship

A short message references a ticket the target may or may not remember opening, and leaves a number. On a mobile handset the message is transcribed automatically within minutes and rendered as readable text beside the caller identity. The target now has a written record of a support contact. No request has been made yet, and the credibility work is already done.

PHASE 05

Continuation into the consequential action

The sequence does not end with the conversation. It follows through into the thing that matters: the reset ticket raised, the approval requested, the remote session joined, the confirmation email that lands carrying the credibility of everything before it. A call that ends in a promise is not a finding. The finding is whether the downstream process caught what the call did not.

Why the Return Leg Is Where It Breaks

Three mechanisms, and they compound. None of them is about how good the audio was.

Self-initiation is mistaken for authentication. Every instinct a user has about verifying a caller is built around receiving a call. Dialling a number yourself feels like the verification step, because in the user's mental model it is the same motion as looking up the help desk and calling them. The distinction between a number you looked up and a number you were given collapses the moment several hours have passed and the transcript is sitting in the phone.

Absence of pressure suppresses the report, not the risk. This is counterintuitive and it is the finding people push back on. A high-pressure call leaves a mark. The target remembers being rushed, mentions it to a colleague, and sometimes reports it. A calm return call on the target's own schedule leaves no mark at all, because nothing about it registered as an event. The handover happens and then nothing happens, which is the worst possible outcome for the organization and the reason we treat report rate on the inbound leg as a separate measure.

The path is uncovered, not merely unheld. Read your own verification procedures and count how many are written for an inbound contact. Nearly all of them. A self-initiated call is treated as self-authenticating by default, so there is no requirement to skip, no exception to grant, and no step for anyone to fail. This is not a discipline problem in your people. It is a hole in the control surface, and it is invisible to any test that never traverses it.

The callback is not the leg where users are weak. It is the leg where no procedure was ever written, so there was nothing standing there to be strong.

This Is a Coverage Problem, Not an Adherence Problem

The distinction is worth being precise about, because the remediation is completely different depending on which one you have.

Adherence is whether a documented verification requirement was executed when it applied. You improve adherence with training, with a lower-friction procedure, and with removing the exception path that lets a helpful person skip it. Coverage is whether a verification requirement exists for a given consequential action in the first place. You cannot train your way to coverage. Someone has to write the procedure.

Most organizations have a verification requirement for outbound payments and nothing whatsoever for credential resets, vendor banking detail changes, multi-factor enrollment, privileged access grants, software installation on request, or a support conversation that arrives over Teams. The consequential actions worth listing are short and every one of them should have a named requirement attached to it. On the return leg, the coverage gap is close to total.

It is also worth stating the bound honestly. Better synthetic media does not defeat a callback procedure directly. A callback to a number the organization already holds returns the same result against a crude clone and against a perfect one, because the procedure never listens to the audio. What improves with generation quality is the pressure applied against the procedure, which raises the rate at which someone waives it. We track that as exception rate rather than claiming the control is untouchable, because the stronger claim is the one somebody will eventually falsify in front of us.

What Each Architecture Can Observe

Simulation platforms fall into three bands. The useful question is not which channels a platform sends on. It is what the architecture is able to see.

Dimension Gen 1: email-first Gen 2: multi-channel Gen 3: orchestrated
What it sends A templated email Email plus voice or SMS, each sent independently A sequence in which each contact references the last
Voicemail Not applicable A recording, with no return path behind it A pretext and a number that is answered when dialled
Inbound callback Not possible Not handled. The number rings out Answered autonomously and run from the receiving side
Headline measure Click rate Action rate, per channel Coverage and process hold rate across the whole path
What it can observe Whether a link was touched, once Whether a person acted on one channel The point in the sequence at which the procedure gave way
Blind spot Every channel with no link in it The handoff between channels, and the return leg Consequential actions outside the agreed scope, which is why coverage is assessed separately

The uncomfortable middle band is the one to watch. Adding a voice call to an email tool is real progress and it is better than not testing voice. But partial simulation produces partial confidence, and a workforce that passed a standalone vishing test can leave a security leader believing the organization is covered on a sequence it has never been tested against.

What an Orchestrated Simulation Lets You Measure

Click rate does not survive contact with this channel, and neither does any single headline number. These are the measures we report, and the two at the top are the ones most organizations have never calculated.

  • Coverage rate. Consequential actions that carry a defined verification requirement, over consequential actions identified. This can be assessed before a single call is placed, which makes it the cheapest place to start.
  • Process hold rate. Verification executed, over verification required. The headline figure, and the one that replaces click rate.
  • Exception rate. Verification consciously waived or talked past, over verification required. This is the number that moves as generation quality improves, so it is the one to trend over time.
  • Callback verification rate. On the return leg specifically, how often the target confirmed the number against the directory before acting. Reported separately, because a blended figure hides the gap.
  • Report rate and time to first report. Measured per leg. A refusal that nobody escalated leaves the next target facing the same caller with no warning in front of them.
  • Action rate, kept as a secondary figure. It moves with scenario difficulty and it does not name the control that failed, so it belongs in the report rather than at the top of it.

Findings are named against procedures rather than individuals, and results roll up to the organization. A per-user result is useful for deciding who receives which training module. It is not an outcome metric, because a score on a person does not tell you whether the wire went out.

Which of your consequential actions have a written verification requirement, and which have none at all?
Does that requirement apply to a call the employee placed themselves, or only to one they received?
If a voicemail leaves a number, is there any step that compares it against the directory?
When verification was skipped, was it never required or was it consciously waived?
How long from the first contact to the first report reaching security, measured per leg?
On re-test, did the same control hold better, or did you simply run an easier scenario?

What We Would Do First

Three things, none of which requires a purchase order, in this order.

List your consequential actions and mark the uncovered ones. Outbound payment, vendor banking detail change, credential reset, multi-factor enrollment or reset, privileged access grant, software installation on request, data export, physical access grant. For each one, write down whether a verification requirement exists, whether it is independent of the channel the request arrived on, and whether the person who will be asked at four on a Friday can actually reach it. Most teams find this exercise takes an afternoon and produces the most useful page in their programme.

Extend one rule to cover self-initiated contact. The standing rule is usually written as a rule about incoming calls. Rewrite it so it covers the direction that is currently exempt: a number obtained from a voicemail, a text, or a caller is not a verified number, and the only verified number is one looked up independently. That single amendment closes the largest uncovered path we find.

Then run the return leg, not just the call. Place the calls, leave the voicemails, answer the callbacks, and follow through into the reset or the approval. Report the coverage rate, the process hold rate, and the callback verification rate separately. The gap between what people say they would do and what they do on a call they placed themselves is the entire finding.

If you want the channel detail behind this, the voice mechanics are covered under AI vishing simulation, and the full multi-channel engagement under deepfake phishing simulation.

Measure your risk. Train for what you find. Prove it changed.

Frequently Asked Questions

What is an orchestrated deepfake phishing simulation?

A simulation in which every contact is sequenced against the ones before it rather than sent independently. A voice call places a pretext, a voicemail leaves a callback number, the number is answered when the target dials it, and email or SMS follows carrying the credibility of everything that came earlier. The measurement is not the response to any single message. It is the point in the sequence at which the verification procedure gave way.

Why is the callback the most important part of a deepfake phishing simulation?

Because most outbound calls go to voicemail, which makes the return call the dominant route through any population. It is also the leg where verification procedures do not apply. Almost every documented verification requirement is written for an inbound contact, so a call the target placed themselves is treated as self-authenticating and no procedure is triggered. The return leg carries the highest yield and the lowest control coverage at the same time.

Is a single deepfake voice call enough to test our organization?

It tests one leg of a sequence that adversaries run in three or more. A single unannounced call arrives with no prior context, so it has to carry all of its own credibility, which makes it easier to refuse and easier to report. It also cannot observe what happens on the return path. A test that ends when the outbound call ends measures the half of the sequence that performs worst for the adversary.

What should an orchestrated simulation measure instead of click rate?

Coverage rate, meaning the share of consequential actions that have a defined verification requirement at all. Process hold rate, meaning how often that requirement was executed under pressure. Exception rate, meaning how often it was consciously waived. Report rate and time to first report. Action rate belongs in the report as a secondary figure, because it moves with scenario difficulty and does not tell you which control failed.

Does better synthetic media defeat a callback verification procedure?

Not directly. A callback to a number the organization already holds returns the same result against a crude clone and a perfect one, because the procedure never evaluates the audio. What improves with generation quality is the pressure applied against the procedure, which raises the rate at which someone waives it. That is why exception rate is the figure worth tracking over time, and why claiming total invariance would be overstating the case.

How is orchestration different from running several simulations at once?

Running a phishing test, a voice test, and an SMS test in the same month produces three independent results. Orchestration produces one result about a sequence, because each contact is built on the state left by the previous one. The difference shows up in the finding: parallel tests tell you which channel people responded to, and an orchestrated engagement tells you which step in a chain your process failed to interrupt.

JT

Jason ThatcherFounder and CEO of Breacher.ai. Fifteen years in security operations and offensive testing, previously at ZeroFox, Deepwatch, and GuidePoint Security. He builds and runs orchestrated social engineering simulations against enterprise organizations.

Find Out What Happens on the Return Call

We run the whole sequence, answer the callback, follow it into the reset or the approval, and report coverage and process hold rate against your vertical.

Book Your Demo

Latest Posts

  • The Callback Is the Simulation | Breacher.ai

  • IT Support Impersonation: The Help Desk Is the Front Door

  • AI Spear Vishing: The IT Support Impersonation Threat

Table Of Contents

About the Author: Jason Thatcher

Jason Thatcher is the Founder of Breacher.ai and comes from a long career of working in the Cybersecurity Industry. His past accomplishments include winning Splunk Solution of the Year in 2022 for Security Operations.

Share this post