Orchestrated Social Engineering Simulations: Sequence, Not Send

Categories: Deepfake,Published On: September 12th, 2026,
  • Dark Breacher.ai banner reading ORCHESTRATED SOCIAL ENGINEERING SIMULATIONS above a green five-step chain, voicemail to SMS to Teams call to help desk to the action, with each link brighter than the last.
Orchestrated Social Engineering Simulations | OSES Platform

Orchestrated Social Engineering Simulations: Sequence, Not Send

A phishing test sends one message and counts who clicked it. An orchestrated simulation runs a sequence across channels, holds a live conversation, answers the callback, and does not stop until the target either performs the consequential action or verifies. That difference is not a feature. It is the entire measurement.

By Jason Thatcher, Founder and CEO, Breacher.ai

See an orchestrated sequence run against your own process, start to finish.

Book Your Demo

Most organizations have tested their people. Very few have tested their process. The distinction sounds like semantics until you watch both happen in the same week, against the same population, and get two completely different numbers.

We build and run orchestrated social engineering simulations as paid engagements against enterprise organizations. That is the whole business, and it is the reason this piece is written in the first person rather than assembled from threat reporting. What follows is what orchestration means operationally, what an orchestration platform has to be able to do before the word is honest, and what the output looks like when it is done properly.

What an orchestrated social engineering simulation actually is

An orchestrated social engineering simulation is a controlled, authorized exercise in which synthetic voice, video and messaging are coordinated into a single campaign, and the response is observed from first contact through to a consequential action. A consequential action is the thing that would have cost you something: an outbound payment, a vendor banking-detail change, a credential reset, an MFA enrollment, a privileged access grant, a data export.

The defining property is the sequence. Each artifact references the one before it, and each one crosses a different control boundary. A target who receives a single unexpected message evaluates it cold. A target who receives a voicemail, then a text confirming the voicemail, then a video call from the person in the voicemail is not evaluating anything by the third step. They are continuing something already in progress.

Each step borrows credibility from the one before it. The sequence, not the artifact, is what a single-channel test cannot reproduce.

A phishing test asks whether someone was suspicious of a stranger. An orchestrated simulation asks whether your process survives someone who is no longer a stranger by the time they ask.
The Engine

Five capabilities an orchestration platform has to have

These five are the capabilities that decide whether a product can actually run the sequence above, and every one of them is a question you can ask in a demo and watch someone either answer or deflect.

CAPABILITY 01

A contextual layer built from your organization

Who your executives are, which approval chains carry real money or access, which vendors your finance team already expects to hear from, and how much public source material exists on each of them. A generic lure measures whether your people are suspicious of strangers. Nobody needs that number. Scenario design has to start from research, and every pretext should trace back to the public material it was built from.

CAPABILITY 02

Channels that hand off to each other

Two channels minimum, three is better, with each artifact referencing the last. Sending an email campaign and separately running a phone campaign is not orchestration, it is two tests with a shared invoice. The platform has to hold the campaign as one object, so the voicemail knows what the text said and the caller knows which ticket the target thinks exists.

CAPABILITY 04

A live conversation that survives being questioned

A recording fails the first time someone asks an unexpected question. A conversational voice agent, or a live avatar joining a Teams, Zoom or Meet session with the camera on, answers, tolerates being put on hold, adjusts when challenged, and stays in character. Without this, the simulation stops exactly where a real operator would push, which means it under-reports every time.

CAPABILITY 05

A run through to the consequential action

The scenario terminates at the payment, the reset, the enrollment or the access grant. Not at the click. If the product cannot express a success condition other than an interaction, then everything downstream of it, including the entire report, is measuring the wrong event.

The Gap

Why a single-channel send under-reports your risk

It tests the one channel your controls already cover. The email gateway is the most mature filter in the enterprise. A program that only sends mail is sampling the best-defended path and reporting the result as an organizational number. Meanwhile the phone, the meeting invite, the chat message and the help desk queue have nothing in front of them.

It terminates before the expensive part. A click is where a phishing test ends. It is where an incident starts. Everything that determines whether the click costs anything, the callback, the approval, the reset, the second signature, happens after the point at which most programs stop recording.

It produces a metric that cannot be acted on. A click rate tells you a number went up or down against yourself. It does not tell you which verification requirement was missing, which process had no defined step, or which of those two failures you are actually looking at. Those are different remediations, and a send cannot distinguish them.

Evidence

What the sequence produces that a send does not

21.6%
Called the spoofed number back and spoke to the voice agent unprompted. Multinational financial services, approximately 500 employees.
63%
Could not distinguish synthetic voice or video from a real person while it was happening, across engagements covering more than 1,057 individual targets.
16%
Entered credentials on the landing page, in both published engagements. Passwords are never captured.

Figures are from published Breacher.ai engagement reports. Each carries the population it was measured against, and none of them should be read as an industry rate.

The middle figure is the one people quote back to us, and it is also the one we are careful about. It says detection did not work. It does not say training is worthless, and we do not make that claim. It says that a control whose effectiveness is set by the quality of the adversary's generation model is a control with a declining value, because generation quality improves continuously and human perception does not improve at all.

The first figure is the one that should change what you buy. More than one in five people in that population rang a spoofed number back and talked to a machine, having already decided the counterparty was real. That path does not exist in an email test. There is no version of a send that finds it.

Measurement

What we report, and what we deliberately do not

Verification is the control that holds, because it does not read the artifact. A callback placed to a number from your own directory returns the same result against a crude fake and a perfect one. That property is why the reporting is built around the process rather than the person.

  • Action rate. The share of targets who performed the consequential action. This is the headline, and it replaces click rate.
  • Verification rate. The share who checked through an independent channel before acting. This is the control actually working, and it is the number that should move after remediation.
  • Report rate and time to first report. A person who felt something was wrong and hung up without reporting is not a success. It is an adversary calling the next name on the list with nobody watching.
  • Where the sequence survived. Which step crossed which control boundary untouched, so remediation is sequenced by layer rather than by scenario.
  • A peer vertical benchmark. Your own last quarter is a low bar when you do not know where you started relative to comparable organizations. The Social Engineering Risk Index is where the sector medians are published.

This is the practical case for orchestration, and it is a measurement case rather than a technology one. An orchestrated simulation measures risk far better than a send can, because it puts the request in front of the process that would actually have to stop it. A single-channel test can only report that somebody interacted with something. A sequence carries the request down the path a real one would take, through the channel it would arrive on, to the person who holds the authority, and then records what the process did with it. The output is not a bigger number or a more alarming one. It is a number that names a control.

That changes what the report is for. A click rate gives you a trend line against yourself. An action rate paired with a verification rate gives you a specific failed step, on a specific path, with a defined owner, which is the difference between a finding somebody can act on this quarter and a metric that gets presented and then filed.

Evaluation

Six questions for any orchestration vendor

Including us. These are the questions that separate a product that runs sequences from a product that sends mail and has added the word orchestrated to its homepage.

Can one campaign run a voicemail, an SMS and a live video call against the same target, as one object, with each step referencing the last?
What happens when the target calls the number back? Who answers, and is a human operator required?
Can a scenario express a success condition that is not a click, such as an MFA reset completed at the help desk?
Is the persona a reusable object with a voice, a likeness and an authority level, or is every campaign rebuilt from scratch?
Does the report distinguish verification attempted and waived from verification never attempted?
Is the benchmark your own previous quarter, or comparable organizations in your sector?

We wrote up the single most reliable sequence we run, help desk impersonation, in detail in the IT support impersonation breakdown. The engine and the capability list behind all of this live on the OSES™ platform page, and the synthetic media layer is detailed under deepfake simulation.

Start Here

What we would do first, at no cost

Before running anything, list the consequential actions in your organization. Outbound payment, vendor banking-detail change, credential reset, MFA enrollment or reset, privileged access grant, software installation on request, data export, physical access grant. That list is short, it is specific to you, and almost nobody has written it down.

Then build a risk register from it. One row per consequential action, with these columns filled in:

  • Who can perform it. The roles that actually hold the authority, not the policy owner. It is usually more people than the policy suggests.
  • Which channels a request can arrive on. Email, phone, Teams or Slack, a ticket, a meeting, someone at a desk. Every channel on that list is a path an adversary can choose.
  • Whether a defined verification requirement exists. Written down, not assumed. If two people describe it differently, it does not exist.
  • Whether that requirement is independent of the channel. A callback to a number supplied during the request is not verification, it is the same channel wearing a different hat.
  • Whether the person under pressure can reach it. If the directory number is three clicks into an intranet the caller is telling them not to open, the control is theoretical.
  • What it costs when it goes wrong. Money moved, access granted, records disclosed. This is what sequences the remediation, not the likelihood.

Then map each row to the layers that should have caught it: which technical control should see the request, which process should stop it, and which role is expected to challenge it. Where nobody can name all three, that is the finding, and you found it without running anything.

That is a structural assessment. It requires no simulation, no vendor and no budget, and in our experience most organizations have never done it. The answer is usually that a verification requirement exists for wire transfers and for almost nothing else. If that is what your register shows, you have the remediation program before a single test has run, and you know exactly which rows a simulation should be pointed at first.

Measure your risk. Train for what you find. Prove it changed.

Frequently Asked Questions

What is an orchestrated social engineering simulation?

An orchestrated social engineering simulation is a controlled, authorized exercise in which synthetic voice, video and messaging are coordinated into a single campaign, and the response is observed from first contact through to a consequential action such as a payment, a credential reset, an MFA enrollment or an access grant. The defining property is the sequence. Each artifact references the one before it and crosses a different control boundary, which is what a single-channel test cannot reproduce.

How is a phishing orchestration platform different from a phishing simulation tool?

A phishing simulation tool sends messages on one channel and records who interacted with them. A phishing orchestration platform holds the campaign as a single object across channels, carries a persona with a voice and a likeness through every step, runs live conversations, answers inbound callbacks, and terminates at a privileged action rather than a click. The practical test is whether the product can run a voicemail, a text and a video call as one sequence against the same target, and report what the process did at the end of it.

Why are orchestrated phishing simulations better than an email-only test?

Because the email gateway is the one channel with mature defenses, and a test that only sends mail measures the channel your controls already cover. An orchestrated phishing simulation crosses the channels where nothing is filtering, and it reaches the places where real losses happen: the help desk, the approval step, the callback. It also produces a different class of finding. An email test tells you who clicked. An orchestrated simulation tells you whether a verification requirement existed and whether it held under pressure.

Can our own team run orchestrated social engineering simulations?

Yes. The first cycle typically runs fully managed so your team sees how scenarios are designed and how findings are written, and after that the platform is yours to operate through the console, the API or the CLI. Enterprises with a staffed program tend to move in house quickly. Organizations below that scale generally get more from managed delivery, because orchestration is an operational skill and running it badly produces findings nobody trusts.

How should orchestrated social engineering simulations be measured?

On the action rate, meaning the share of targets who performed the consequential action, the verification rate, meaning the share who checked through an independent channel, the report rate and time to first report, and a peer vertical benchmark rather than your own previous quarter. Click rate belongs nowhere near the headline, and neither does a per-employee risk score. Per-user results are useful for routing training to the people and processes that failed, but a score on a person does not tell you whether the payment went out.

JT

Jason ThatcherFounder and CEO of Breacher.ai. Fifteen years in security operations and offensive testing, previously at ZeroFox, Deepwatch, and GuidePoint Security. He builds and runs orchestrated social engineering simulations against enterprise organizations.

Find Out Whether Your Process Holds

We design the sequence, run it end to end, and report the action rate, the verification rate and where you sit against your sector.

Book Your Demo

Latest Posts

  • Orchestrated Social Engineering Simulations: Sequence, Not Send

  • Stop Training People to Spot Deepfakes. Train the Procedure, Then Test It.

  • AI Vishing Red Team: What We Are Seeing in the Field

Table Of Contents

About the Author: Jason Thatcher

Jason Thatcher is the Founder of Breacher.ai and comes from a long career of working in the Cybersecurity Industry. His past accomplishments include winning Splunk Solution of the Year in 2022 for Security Operations.

Share this post