AI Vishing: Spear Phishing Didn’t Scale. Now It Does.

Categories: Deepfake,Published On: August 2nd, 2026,
Spear Phishing Didn't Scale. AI Vishing Made It Cheap and Easy.

Spear Phishing Didn't Scale.
AI Vishing Made It Cheap and Easy.

Real voice · 00:30 sampleAI Spear Phishing

Cloned in 30 seconds. Trusted in 90.

For twenty years, two constraints protected your organization better than any control you deployed. Neither was technical. Both were economic, and both are gone. What replaced them is the channel most organizations still are not testing.

Voice-based deepfake campaigns
Action Rate by Engagement
BREACHER.AI
Voicemail / callback / SMSn ~40
0.0%
Voicemail / agentic call / emailn ~250
4.0%
Phone call / agentic AIn ~25
8.9%
Phone call / SMSn ~80
13.3%
Phone call / SMSn ~35
16.2%
Phone call / SMSn ~50
16.7%
Voicemail / callbackn ~350
19.6%
Voicemail / callbackn ~70
34.5%
Sample · 8 engagements · 900 personnel
Action rate as share of personnel targeted

A sample of voice-involving engagements from the Breacher.ai benchmark log with directly comparable action-rate data. Action rate is the share of targeted personnel who completed the requested action. Population counts are rounded.

Scope an engagement

Find out where your organization would rank against peer benchmarks.

Book a 30-minute walkthrough. We will go through how an orchestrated voice engagement is scoped and executed, including the voicemail and callback path, and the reporting your team would receive. No marketing slides, no IT integration required.

Book a Demo Free 30-min consultation
Part One · The Threat: AI Spear Vishing

The Constraint Was Never Technical

Spear phishing was accurate but slow. Someone had to research the target, learn the org chart, understand the vendor relationships, and write something that landed. That work does not parallelize. An attacker who wanted precision paid for it in hours.

Voice phishing was persuasive but serialized. One caller, one call, one target. A skilled operator on the phone beats any email ever written, because voice carries urgency, authority, and improvisation that text cannot. But that operator could only be on one call at a time.

So the market split cleanly. Volume attacks were generic and your filters caught most of them. Targeted attacks were devastating and rare, because they were expensive. That tradeoff between precision and volume was never a security control you built. It was an economic accident you benefited from.

Nothing in your stack was holding the line. Arithmetic was. And the arithmetic changed.

Three Things Collapsed at Once

None of these is speculative, and none of them required capability beyond what is commercially available today.

01
Voice synthesis became real-time and cheap

Not "convincing in a demo" cheap. Convincing enough to hold a live phone conversation with someone who has no reason to suspect anything, at a latency low enough that the natural rhythm of a call never breaks. The moment a caller pauses a beat too long before answering, a human on the other end starts wondering why. That gap closed, and it closed quietly.

Why it matters: synthetic voice stopped being the hard part of the attack roughly two years ago. Everything since has been orchestration.
02
Conversational agents became autonomous

The meaningful shift is not that a machine can read a script in a believable voice. It is that a machine can place the outbound call, recognize that it reached voicemail, leave a message with a callback number, then answer that callback and improvise through objections it was never scripted for. Voice social engineering has always been rate-limited by human labor. That ceiling is gone, and what used to require a team of skilled callers now requires compute.

Why it matters: a recorded clip tests whether someone notices a fake. A live agent tests whether they hold the line when the fake answers back.
03
Personalization became free

Public profiles, org charts, press releases, job postings, and credential dumps supply enough context to make every call specific to the person answering it. The research step that used to gate spear phishing now costs effectively nothing per target. The attacker no longer chooses between reach and relevance, which is why autonomous social engineering testing moved from novelty to baseline assurance requirement in under two years.

Why it matters: targeting an entire workforce by phone is now cheaper than targeting one executive was three years ago.

What It Looks Like in the Field

The chart above is a sample of voice-involving engagements from our benchmark log: phone call, voicemail, and callback campaigns run against enterprise populations. It is a sample, not the full log, and not a random one, so read it as illustration of range rather than as a population estimate. Eight engagements, 900 people in total, delivered through pretexts ranging from a straightforward impersonated call to fully autonomous agents handling outbound dialing, voicemail, and inbound callbacks.

14.5%weighted mean action rate across the sample
34.5%worst-performing organization, more than one in three complied
0%best-performing organization, same threat, nobody complied

The mean is the least useful number on that chart. Two organizations facing the same category of attack landed at 0% and 34.5%. A fourteen-point average describes neither of them, and it would not have told either one anything actionable about its own exposure. This is the entire case for peer vertical benchmarking over industry aggregates: the question worth answering is not what the market average is, it is whether you are the 0% or the 34.5%, and nothing published in a vendor report can tell you that.

The spread is the finding. Same threat, same category of pretext, and outcomes that run from nobody complying to more than one in three.

The callback inversion

Look at what sits at the top of that chart. The two highest results were both voicemail-and-callback campaigns, and in the largest autonomous engagements the overall majority of outbound calls went unanswered and landed in voicemail. That looks like a miss until you follow what happened next. The voicemail with a callback number became the primary route the population used to reach the agent. Those campaigns did not run on people picking up the phone. They ran on people calling back.

When a target calls you back, they have already authenticated you in their own mind. Every instinct that says "be suspicious of unsolicited callers" is disarmed, because from their perspective the call was not unsolicited. It was theirs.

Detection is not what separated them

Some targets do recognize that a voice is synthetic and refuse. That is worth crediting, and it is also precisely why detection cannot be the control: recognition is distributed unevenly inside any workforce, varying by person, by role, and by how busy someone happened to be when the phone rang. What separated the top of that chart from the bottom was process, not perception. The organizations that held had a verification step the caller could not satisfy, and it applied no matter who answered the phone.

The wider benchmark

Widening out from voice to every simulation we have run: across 1,000+ simulations, 63% of users have been unable to distinguish synthetic media from real. The two figures beside it are our own classification of engagement outcomes rather than an industry-standard scale, so read them as our scoring rather than a measured constant.

92%of organizations tested have shown vulnerability to deepfake social engineering
78%have been highly vulnerable
63%of users could not distinguish synthetic media from real

Treat AI Spear Vishing as an Active Threat

AI spear vishing is not phishing with better grammar. It is a live conversation, personalized to whoever answers, run by something that never tires, never breaks character, and costs close to nothing to aim at your entire staff at once.

Four properties make it more dangerous than the threat your awareness program was designed around:

PropertyWhat it means for defenders
There is no artifact to inspect No headers, no sender domain, no link to hover over. The evidence a person would normally judge on does not exist, and the decision is made out loud, in seconds, while someone waits.
The call itself does not pass through your email security stack Secure email gateways, link rewriting, attachment detonation, impersonation filters. None of it sits in the path of a phone call, even when an earlier stage arrived by email.
There is no click, so most programs record nothing A campaign can run against your whole workforce and leave no trace in the metrics your reporting is built on.
Voice carries authority that text never did Urgency, a familiar-sounding person, and a plausible reason, on a channel staff have been trained for years to trust more than email.

In the sample above, the worst-performing organization had more than one in three targeted people complete the requested action. Not a projection, not a lab result. The planning assumption is not whether this gets attempted against you, but what your number will be when it does.

The economics that kept targeted voice attacks rare are gone, and nothing replaced them. What used to be a boutique attack against one executive is now a commodity attack against everyone who answers a phone.

Will attested calling solve this?

Email spoofing looked unsolvable until SPF, DKIM, and DMARC largely settled it, so the fair question is whether caller authentication does the same for voice and makes this a five-year problem rather than a permanent one. It will help, and it is worth adopting. It does not close this attack.

Attestation authenticates the line, not the person speaking on it. A verified number tells you the call came from where it claims, not that the human voice on it is human, or that it is who it says. And it does nothing at all for the dominant path in our data, where the target dials a number they were given in a voicemail. That call is legitimately placed, correctly attested, and completely attacker-controlled. Plan for controls that survive a caller who passes every check.

The encouraging part is in the same chart. One organization in that sample recorded no successful compromises at all, against the same class of pretext that beat everyone else. Whatever produced that result, it is reproducible in a way that individual detection skill is not: a rule that says consequential actions cannot complete on voice authority alone is a control you can build this quarter, and it does not depend on anyone recognizing anything.

Part Two · How to Counter It

Teaching People to Spot Deepfakes Is the Wrong Approach

We hold this position strongly, and it is the one thing we would change about how most programs are built: detection of synthetic media is the wrong thing to build a program on. Some people will always notice, and that is worth having. It is not worth relying on. The obvious counter is to attack that 63% directly, teaching people to listen for flat affect, odd cadence, unnatural pauses, the tells. Do not do it. Detection training is a losing position by construction, for one reason: generation quality improves every quarter and human perception does not. You are training a fixed capability against an improving one, and every artifact you teach people to listen for is a bug someone is already fixing.

It is worse than neutral, because it manufactures confidence: a workforce that believes it can hear a fake stops verifying when it does not hear one.

What a click-rate test measures

Whether a person recognized a suspicious message and declined to interact with it. One decision, made alone, in front of a screen, with unlimited time to think and no social pressure applied. A clean click rate would have told this organization it was in good shape.

What the voice engagement measured

Whether the process around credential handling holds when a plausible authority is on the phone, in real time, asking a reasonable-sounding question and waiting. Half of the people who reached that moment had no process to fall back on. So they answered.

The organization did not fail because its people could not spot a deepfake. It failed because nothing in the process required them to.

The better approach is not a sharper detection module. It is teaching process, policy, and procedure, and making them rigorous enough to thwart the attack on their own. Secure behavior is not whether someone recognized a fake. It is whether they followed the procedure. Which carries a consequence most vendors will not say out loud: process and policy are specific to your organization, so generic or canned content cannot carry this load. No module library teaches your wire approval path, your helpdesk rules, or your escalation route.

It also changes what your people are for. Depending on the attack path, a person is the first line of defense, the last line of defense, or the only one. A procedure they can actually execute under pressure is what makes that position survivable. Asking them to work as a forensic audio analyst is not.

What Holds Instead: Process, Not Perception

Train process, policy, and procedure. That content is specific to your organization, which is exactly why generic, canned deepfake awareness training does not move outcomes against this threat. None of the following requires buying anything. All of it requires deciding something.

ControlWhat it means in practice
Break the callback assumption Verification means a number your people look up in a directory you control, never a number supplied by the caller, a voicemail, an email signature, or a chat message.
Make the helpdesk incapable of asking If support will never request credentials, MFA codes, or push approvals by phone, publish that to the entire workforce, not just to the helpdesk, so every employee becomes an enforcement point instead of a target.
Put a non-voice gate on consequential actions Wire transfers, payment detail changes, credential resets, privilege grants, vendor bank changes. This is the practical core of CEO fraud prevention: if a sufficiently convincing call can complete the action, the call is your control.
Make it culturally safe to hang up on an executive If an analyst who refused a CFO's urgent request gets a hard conversation afterward, your verification policy is decorative and the pretext is exploiting exactly that.
Own the seam These campaigns cross email, phone, and collaboration platforms in one sequence. Each tool works as designed inside its own lane and nobody owns the gaps. That is the real argument for multi-channel deepfake protection platforms: not more coverage, but assigned ownership of the handoffs.
Retest A one-time result is a snapshot, and this threat is moving considerably faster than annual training cycles.

To prevent CEO fraud reliably, the approval has to require something a voice or video call cannot satisfy on its own. Everything else is a preference.

Where Deepfake Phishing Simulation Fits

You do not validate those controls with a questionnaire. You validate them by running the attack. That is what deepfake simulation is for: not to prove synthetic voice is convincing, but to find out whether your procedures survive contact with a convincing request. What are the odds this action completes here when someone asks for it persuasively.

What to require of deepfake simulation platforms

Whether you are evaluating deepfake simulation platforms for enterprises or deepfake simulation platforms for IT teams running the program in-house, the differentiators are narrower than vendor pages suggest. Most deepfake attack simulation platforms clear the first requirement below and fail the second.

RequirementWhy it matters
Multi-channel by default Real campaigns are orchestrated: email into a collaboration platform into a live call. Deepfake phishing simulation software that produces a single-channel artifact tests one link in a chain that has four.
Autonomous voice, including inbound Most deepfake content simulation platforms can generate a convincing outbound message. Far fewer can handle the callback, which the data above shows is the dominant path. That is the practical bar for any tools for simulating deepfake voice phishing.
Customer-controlled scenario authoring Your process is yours. AI-based deepfake simulation software offering only a fixed library tests a generic organization, not yours. You want deepfake scenario-based training tools where your own security team defines the pretext, the pressure, and the ask.
Peer vertical benchmarking This matters more than everything else on the list and almost nobody does it. Aggregate industry figures are context, not a benchmark. The report that changes decisions is "you scored X and organizations like you score Y," because that is what separates a genuine outlier from a sector median that is itself unacceptable. Ask any vendor how many engagements sit behind the comparison in your vertical: a benchmark built on one or two is a sample of one client's culture, not a sector.
On-demand execution On-demand deepfake simulation platforms let you re-test after a process change instead of waiting for an annual window. If your remediation cycle is faster than your testing cycle, you are flying blind between tests.

Measuring when there is no click

Click rate does not apply. There is no link and no click, so a program built on click and report rates records a voice campaign as nothing at all. That is a fair summary of why this category blindsides people.

Score at the organizational level, not per user. The useful question is not who failed. It is whether the process holds. Knowing which roles sit on impersonation paths is worth having, because it tells you where to harden first. Scoring individuals as the outcome is not, because what failed was a procedure.

Score all three layers of the chain, not just the person. When a consequential incident happens, three things went wrong: the technology failed to prevent it, the process failed to hold, and a person took the wrong action. Measurement has to span all three, because a single behavioral metric captures one of them and hides the two you can actually fix.

Practice announced, assess unannounced

People want drills, not gotchas; survey your own workforce and you will find this within a day. Run practice as an announced exercise with a stated window and a stated procedure. Interactive deepfake training tools work well here, because the point is rehearsal and knowing it is a rehearsal does not reduce the value. Then assess unannounced, separately, to see whether the process holds under real conditions.

Running unannounced tests repeatedly and calling it training is the common mistake. It produces a failure rate and a resentful workforce, and it does not teach the procedure someone needs at 4:45 on a Friday when a voice that sounds exactly like the CFO asks for an exception. That split is how to run deepfake simulation exercises without burning goodwill on the first cycle.

The Category Has Three Shapes

Tooling here falls into three architectures that get compared as though they were interchangeable. They are not. If you are shortlisting platforms for deepfake awareness training alongside adversarial testing, shape matters more than any feature list.

ArchitectureWhat it does, and where it stops
Human risk management Correlates stack signals into a per-user risk index; simulation is one input. Coherent if you want continuous per-user measurement, and the stronger products here are among the more capable AI-powered social engineering training platforms available. Our disagreement: when a synthetic CFO calls the AP clerk, the question is not whether that clerk is high-risk. It is whether a wire can leave on voice authority alone. That is a process property, not a person property.
Automated awareness Progressive difficulty, behavioral scoring, an academy curriculum. Good and cheap in staff time for reducing email click rates. The gap is channel: an email-centered program does not test a live inbound call. Ask any vendor here whether the platform can answer a callback and hold a conversation. Many cannot.
Adversarial simulation Where Breacher.ai sits. Unannounced, orchestrated, multi-channel social engineering simulation, autonomous voice across outbound, voicemail, and inbound callbacks, with peer vertical benchmarking in the deliverable. Not the fit if you need SCIM, a large module library, and completion reporting for an audit; plenty of deepfake awareness training platforms do that well, and deepfake security awareness training of that kind belongs in a program.

One caution applies to every vendor here, us included. Be skeptical of anyone selling deepfake analysis security training solutions premised on teaching staff to spot synthetic media. That capability degrades every quarter by design.

AI Vishing Spear Phishing Deepfake Phishing Simulation Callback Phishing CEO Fraud Prevention OSES™ Peer Vertical Benchmarking

Frequently Asked Questions

Direct answers to what security teams ask when they start scoping this.

Q
What is deepfake phishing simulation?

A controlled exercise that reproduces an AI-generated voice or video social engineering attack against your own workforce, under written authorization, to determine whether your verification procedures hold. Unlike email phishing tests, deepfake phishing simulations usually span multiple channels and often have no click to measure, so they are scored on whether the requested action completed.

Q
How do I run deepfake simulation exercises without damaging trust?

Split them. Run announced drills for practice with a stated window and a stated procedure, and run unannounced assessments separately to measure whether the process holds under real conditions. Publish results at the organizational level, never as a list of who failed. Using deepfake simulation for security awareness works when people know rehearsal is rehearsal.

Q
What are the best tools for deepfake simulation training?

The differentiators worth paying for are multi-channel orchestration, autonomous inbound voice handling, customer-controlled scenario authoring, and peer vertical benchmarking. Module count and template library size have not tracked with outcomes in our engagement data.

Q
Which top AI social engineering training platforms should we shortlist?

Shortlist by shape rather than brand. Pick one human risk management platform, one automated awareness platform, and one adversarial simulation provider, then require each to demonstrate a live inbound callback scenario. Most shortlists collapse quickly at that step.

Q
How do we evaluate a deepfake simulation vendor's actual capability?

Put one scenario in front of them and watch. Ask the vendor to demonstrate an inbound callback: the agent leaves a voicemail, you call the number back, and it answers and holds a conversation it did not initiate. Capability claims collapse quickly at that step, and it separates orchestration from a library of pre-rendered clips faster than any feature matrix.

Q
Is deepfake phishing simulation for regulated industries viable?

Yes, and it is usually easier to authorize than teams expect, because no real credential material needs to be captured to produce the finding. Deepfake simulation training for compliance purposes should be scoped with legal and HR up front, with data handling, retention, and the non-capture of credential material stated in writing before execution.

Q
How is deepfake risk simulation for executives different?

Executives are both the most impersonated population and the least tested one. Executive scenarios test two directions at once: whether the executive can be compromised, and whether the organization will comply with a request that appears to come from them. The second is almost always the larger exposure.

Q
Does vishing simulation require a human caller?

No longer. Autonomous voice agents now handle outbound calls, voicemail, and live inbound callbacks without an operator on the line, which is precisely why this threat scales in a way traditional vishing never did.

Distribution figures represent a sample of eight voice-involving engagements drawn from the Breacher.ai benchmark log, covering 900 targeted personnel in total, selected because they carry directly comparable action-rate data. The sample is deliberately not exhaustive. Further voice engagements are withheld either because the available measure was not directly comparable or because the client has not consented to publication; the withheld results fall within the range shown rather than outside it. Action rate is defined as the share of targeted personnel who completed the requested action. Further voice engagements were excluded where the available measure was not directly comparable. Population counts are rounded. Benchmark percentages elsewhere on this page reflect aggregate results across 1,000+ simulations. Client identity, sector, and scope details are withheld under engagement confidentiality.

Author
JT

Jason Thatcher

Founder & CEO, Breacher.ai

Jason Thatcher is the Founder and CEO of Breacher.ai and creator of OSES™ (Orchestrated Social Engineering Simulations). He has 15+ years in cybersecurity spanning security operations, threat intelligence, and executive leadership, with prior roles at ZeroFox, Deepwatch, and GuidePoint Security. He built Breacher.ai from a practitioner's view of defender blind spots and writes about the gap between the deepfake threat that incident reports document and the simulation tooling most enterprises actually buy. Connect on LinkedIn.

Find Out What Your Number Is

Book a 30-minute walkthrough. We will demonstrate the live AI voice agent in conversation, show the voicemail and callback path end to end, and scope what an orchestrated engagement against your population would look like.

Live agent demonstration
No IT integration required
Free 30-minute consultation
Peer benchmark preview
Book Your Engagement Walkthrough

Latest Posts

  • AI Vishing: Spear Phishing Didn’t Scale. Now It Does.

  • Best Deepfake Simulation Platforms | Breacher.ai 2026

  • Deepfake Benchmark 2026: Click Rate Is the Wrong Metric | Breacher.ai

Table Of Contents

About the Author: Jason Thatcher

Jason Thatcher is the Founder of Breacher.ai and comes from a long career of working in the Cybersecurity Industry. His past accomplishments include winning Splunk Solution of the Year in 2022 for Security Operations.

Share this post