The Callback Is the Simulation | Breacher.ai
The Callback Is the Simulation
A deepfake phone call on its own is a weak test. The voicemail, and the number the target dials back, is where the credentials actually move. If your deepfake phishing simulation does not orchestrate the return leg, you have measured the easy half.
There is a version of deepfake testing that has become the industry default, and it does not work very well. A cloned voice places a call. Someone answers or does not. A number gets written down. The report says a percentage of the population engaged with a synthetic caller, and everyone agrees the technology is impressive.
We run these engagements for a living, and the single most useful thing three years of them has taught us is that the outbound call is the weakest leg of the sequence. It is the leg an adversary has the least confidence in, the leg most likely to be refused, and the leg most likely to be reported. It is also, almost universally, the only leg anybody simulates.
The consequences we see land on the inbound. A voicemail goes out. Some hours later the target picks up their own phone and dials the number in it. That return call is where the credential goes across. In our red team work, the return path is where the majority of credential handovers happen, and it is the path with the least procedural protection standing in front of it.
This post is the argument for why orchestration is not a feature of a deepfake phishing simulation. It is the thing that determines what the simulation is capable of observing at all.
The Voice Channel Produces Action That Nothing Records
Before the callback argument, the baseline. Voice campaigns produce consequential outcomes in populations where an email-first tool would have reported an empty result, because there was no link anywhere in the sequence for it to count.
Source: Breacher.ai OSES™ engagement outcomes, published in the OSES Risk Index. Detection figures are measured across engagements covering more than one thousand individual targets. Reporting is organizational and sector level only.
Those figures describe the channel. They do not tell you where inside the channel the failure sits, and that is the question orchestration exists to answer.
The Sequence We Actually Run
OSES™, Orchestrated Social Engineering Simulations™, sequences each contact against the state left by the one before it. Below is the IT support variant, which is the one that produces our most serious findings. Note where the featured card sits.
Reconnaissance and pretext construction
Reporting lines, help desk hours, ticketing system, device fleet, and VPN vendor are assembled from public sources. Voice samples come from webinars, recorded panels, and voicemail greetings. Nothing in this phase touches the perimeter, so nothing in this phase generates a detection.
The outbound call, which is expected to miss
The call is placed in a cloned voice. If it is answered, the agent holds a live conversation. Most of the time it is not answered, and that is not a failure of the sequence. In practice the majority of outbound calls in a campaign go to voicemail, which makes the next phase the dominant route through the population rather than the fallback.
The voicemail that manufactures a prior relationship
A short message references a ticket the target may or may not remember opening, and leaves a number. On a mobile handset the message is transcribed automatically within minutes and rendered as readable text beside the caller identity. The target now has a written record of a support contact. No request has been made yet, and the credibility work is already done.
The callback, answered from the other side
The target dials the number. The agent answers and runs the scenario from the receiving side. This is the leg almost no simulation platform can execute, and it is the leg that converts. The target is calm, unhurried, and operating on the belief that a call they placed themselves cannot be an impersonation. There is no urgency to resist and no pressure to notice, which is precisely the problem.
Continuation into the consequential action
The sequence does not end with the conversation. It follows through into the thing that matters: the reset ticket raised, the approval requested, the remote session joined, the confirmation email that lands carrying the credibility of everything before it. A call that ends in a promise is not a finding. The finding is whether the downstream process caught what the call did not.
Why the Return Leg Is Where It Breaks
Three mechanisms, and they compound. None of them is about how good the audio was.
Self-initiation is mistaken for authentication. Every instinct a user has about verifying a caller is built around receiving a call. Dialling a number yourself feels like the verification step, because in the user's mental model it is the same motion as looking up the help desk and calling them. The distinction between a number you looked up and a number you were given collapses the moment several hours have passed and the transcript is sitting in the phone.
Absence of pressure suppresses the report, not the risk. This is counterintuitive and it is the finding people push back on. A high-pressure call leaves a mark. The target remembers being rushed, mentions it to a colleague, and sometimes reports it. A calm return call on the target's own schedule leaves no mark at all, because nothing about it registered as an event. The handover happens and then nothing happens, which is the worst possible outcome for the organization and the reason we treat report rate on the inbound leg as a separate measure.
The path is uncovered, not merely unheld. Read your own verification procedures and count how many are written for an inbound contact. Nearly all of them. A self-initiated call is treated as self-authenticating by default, so there is no requirement to skip, no exception to grant, and no step for anyone to fail. This is not a discipline problem in your people. It is a hole in the control surface, and it is invisible to any test that never traverses it.
This Is a Coverage Problem, Not an Adherence Problem
The distinction is worth being precise about, because the remediation is completely different depending on which one you have.
Adherence is whether a documented verification requirement was executed when it applied. You improve adherence with training, with a lower-friction procedure, and with removing the exception path that lets a helpful person skip it. Coverage is whether a verification requirement exists for a given consequential action in the first place. You cannot train your way to coverage. Someone has to write the procedure.
Most organizations have a verification requirement for outbound payments and nothing whatsoever for credential resets, vendor banking detail changes, multi-factor enrollment, privileged access grants, software installation on request, or a support conversation that arrives over Teams. The consequential actions worth listing are short and every one of them should have a named requirement attached to it. On the return leg, the coverage gap is close to total.
It is also worth stating the bound honestly. Better synthetic media does not defeat a callback procedure directly. A callback to a number the organization already holds returns the same result against a crude clone and against a perfect one, because the procedure never listens to the audio. What improves with generation quality is the pressure applied against the procedure, which raises the rate at which someone waives it. We track that as exception rate rather than claiming the control is untouchable, because the stronger claim is the one somebody will eventually falsify in front of us.
What Each Architecture Can Observe
Simulation platforms fall into three bands. The useful question is not which channels a platform sends on. It is what the architecture is able to see.
| Dimension | Gen 1: email-first | Gen 2: multi-channel | Gen 3: orchestrated |
|---|---|---|---|
| What it sends | A templated email | Email plus voice or SMS, each sent independently | A sequence in which each contact references the last |
| Voicemail | Not applicable | A recording, with no return path behind it | A pretext and a number that is answered when dialled |
| Inbound callback | Not possible | Not handled. The number rings out | Answered autonomously and run from the receiving side |
| Headline measure | Click rate | Action rate, per channel | Coverage and process hold rate across the whole path |
| What it can observe | Whether a link was touched, once | Whether a person acted on one channel | The point in the sequence at which the procedure gave way |
| Blind spot | Every channel with no link in it | The handoff between channels, and the return leg | Consequential actions outside the agreed scope, which is why coverage is assessed separately |
The uncomfortable middle band is the one to watch. Adding a voice call to an email tool is real progress and it is better than not testing voice. But partial simulation produces partial confidence, and a workforce that passed a standalone vishing test can leave a security leader believing the organization is covered on a sequence it has never been tested against.
What an Orchestrated Simulation Lets You Measure
Click rate does not survive contact with this channel, and neither does any single headline number. These are the measures we report, and the two at the top are the ones most organizations have never calculated.
- Coverage rate. Consequential actions that carry a defined verification requirement, over consequential actions identified. This can be assessed before a single call is placed, which makes it the cheapest place to start.
- Process hold rate. Verification executed, over verification required. The headline figure, and the one that replaces click rate.
- Exception rate. Verification consciously waived or talked past, over verification required. This is the number that moves as generation quality improves, so it is the one to trend over time.
- Callback verification rate. On the return leg specifically, how often the target confirmed the number against the directory before acting. Reported separately, because a blended figure hides the gap.
- Report rate and time to first report. Measured per leg. A refusal that nobody escalated leaves the next target facing the same caller with no warning in front of them.
- Action rate, kept as a secondary figure. It moves with scenario difficulty and it does not name the control that failed, so it belongs in the report rather than at the top of it.
Findings are named against procedures rather than individuals, and results roll up to the organization. A per-user result is useful for deciding who receives which training module. It is not an outcome metric, because a score on a person does not tell you whether the wire went out.
What We Would Do First
Three things, none of which requires a purchase order, in this order.
List your consequential actions and mark the uncovered ones. Outbound payment, vendor banking detail change, credential reset, multi-factor enrollment or reset, privileged access grant, software installation on request, data export, physical access grant. For each one, write down whether a verification requirement exists, whether it is independent of the channel the request arrived on, and whether the person who will be asked at four on a Friday can actually reach it. Most teams find this exercise takes an afternoon and produces the most useful page in their programme.
Extend one rule to cover self-initiated contact. The standing rule is usually written as a rule about incoming calls. Rewrite it so it covers the direction that is currently exempt: a number obtained from a voicemail, a text, or a caller is not a verified number, and the only verified number is one looked up independently. That single amendment closes the largest uncovered path we find.
Then run the return leg, not just the call. Place the calls, leave the voicemails, answer the callbacks, and follow through into the reset or the approval. Report the coverage rate, the process hold rate, and the callback verification rate separately. The gap between what people say they would do and what they do on a call they placed themselves is the entire finding.
If you want the channel detail behind this, the voice mechanics are covered under AI vishing simulation, and the full multi-channel engagement under deepfake phishing simulation.
Frequently Asked Questions
What is an orchestrated deepfake phishing simulation?
A simulation in which every contact is sequenced against the ones before it rather than sent independently. A voice call places a pretext, a voicemail leaves a callback number, the number is answered when the target dials it, and email or SMS follows carrying the credibility of everything that came earlier. The measurement is not the response to any single message. It is the point in the sequence at which the verification procedure gave way.
Why is the callback the most important part of a deepfake phishing simulation?
Because most outbound calls go to voicemail, which makes the return call the dominant route through any population. It is also the leg where verification procedures do not apply. Almost every documented verification requirement is written for an inbound contact, so a call the target placed themselves is treated as self-authenticating and no procedure is triggered. The return leg carries the highest yield and the lowest control coverage at the same time.
Is a single deepfake voice call enough to test our organization?
It tests one leg of a sequence that adversaries run in three or more. A single unannounced call arrives with no prior context, so it has to carry all of its own credibility, which makes it easier to refuse and easier to report. It also cannot observe what happens on the return path. A test that ends when the outbound call ends measures the half of the sequence that performs worst for the adversary.
What should an orchestrated simulation measure instead of click rate?
Coverage rate, meaning the share of consequential actions that have a defined verification requirement at all. Process hold rate, meaning how often that requirement was executed under pressure. Exception rate, meaning how often it was consciously waived. Report rate and time to first report. Action rate belongs in the report as a secondary figure, because it moves with scenario difficulty and does not tell you which control failed.
Does better synthetic media defeat a callback verification procedure?
Not directly. A callback to a number the organization already holds returns the same result against a crude clone and a perfect one, because the procedure never evaluates the audio. What improves with generation quality is the pressure applied against the procedure, which raises the rate at which someone waives it. That is why exception rate is the figure worth tracking over time, and why claiming total invariance would be overstating the case.
How is orchestration different from running several simulations at once?
Running a phishing test, a voice test, and an SMS test in the same month produces three independent results. Orchestration produces one result about a sequence, because each contact is built on the state left by the previous one. The difference shows up in the finding: parallel tests tell you which channel people responded to, and an orchestrated engagement tells you which step in a chain your process failed to interrupt.
Find Out What Happens on the Return Call
We run the whole sequence, answer the callback, follow it into the reset or the approval, and report coverage and process hold rate against your vertical.
Book Your Demo
