AI Vishing Red Team: What We Are Seeing in the Field

Categories: Deepfake,Published On: September 7th, 2026,
AI Vishing Red Team: Field Findings From the Callback Leg
Threat Intelligence

AI Vishing Red Team: What We Are Seeing in the Field

What follows is some of what we are seeing. There is practical guidance at the bottom.

By Jason Thatcher, Founder and CEO, Breacher.ai

See what an orchestrated voice sequence produces against your own population.

Book Your Demo

As a preface, we founded Breacher.ai Corp. in late 2023 as a red team service to test organizations against deepfake based social engineering. We have long believed that voice is the most dangerous vector, not video.

Secondly, we are security practitioners first and foremost, vendors second. We execute red team engagements in the field on a weekly basis. It keeps us sharp, it helps us learn, and it ensures we stay current with the threat landscape. The latest model drops, and we are likely using it in the field in short order.

We translate our findings into research, threat intelligence, simulations, and security awareness training. One of our strongest motivations for doing this was how painfully behind the curve awareness training was for deepfakes. We saw a gap, and we acted.

The Threat That Keeps Us Up at Night

For the past few years we have been testing voice agents in the field to see whether AI can social engineer its way into an organization. A year ago the results were mixed. Roughly 1 to 3 percent of users contacted would fall for it. Most people could still tell they were talking to a machine.

That has changed. The models have matured at a remarkable pace, and voice quality is no longer the limiting factor.

Across our most recent agentic voice simulations, the weighted mean action rate is 11.5 percent of users contacted, with a range from 0 percent to 34.5 percent of users contacted. Action means the user did something that could put the organization at risk, most often divulging information. In one scenario we had an AI agent ask employees to read back their badge ID. About half of the users who engaged complied.

Action rate, then and now
0% 10% 20% 30% 35% A year ago 1 to 3% of users contacted Most recent simulations 11.5% weighted mean Range 0% to 34.5% of users contacted Scale: action rate as a percent of users contacted, 0 to 35.

IT Helpdesk Impersonation: The Scenario to Watch

We have run these scenarios against enterprises, and statistically they are higher in our engagement data. Significantly.

Our working theory is the red flag indicators of what awareness training has taught users. Pressure, urgency, and suspicion do not exist with these campaigns. It is a routine task being asked for, one that may not trip suspicion. This is why we believe threat actors like Scattered Spider are successful.

The pattern we keep seeing is that trusted workflows are being exploited to compromise organizations. Teams, voice, IT impersonation, HR hiring, and the rest. It is easy to do and it is effective. This was the progression we predicted. As defensive tools improve, threat actors are exploiting the trusted seams in between workflows, because it is easy to do and effective.

We wrote up the pretext itself, in both directions, in our breakdown of IT support impersonation.

The Scale Problem: No Human Operator

The silent defender: scale

Human operators have been the limiting factor. A human operator was required to run these campaigns and dial users, converse, and operate the campaign.

Today, AI can be used for this, at scale.

We built an agent based simulation platform for this reason, to help us scale and remove the human operator. The system calls a target population. If someone answers, it converses live and adapts. If nobody answers, it leaves a voicemail with a callback number. When someone dials back, it answers and converses again.

Nobody is reading a script, and the whole thing executes with the click of a button, at scale.

An agent has no such ceiling. An entire population of an organization is reachable, in a day.

Why This Is Different: Orchestration

What often gets overlooked is the orchestration component of this whole scenario and how important it is. Modern social engineering is not simply sending a phone call. It is conversational, it is informed by open-source intelligence, and it combines different media. Email, phone call, and the rest. The key is that there is a contextual layer that holds each of these steps together to hit an objective. Without that contextual layer and orchestration, the scenario will fall apart. It is why we built an orchestration engine rather than another simulation platform.

Our orchestration for a typical engagement
STEP 01 Outbound call STEP 02 Voicemail drop STEP 03 Callback request STEP 04 Agent inbound Where the damage occurs and where most simulations stop STEP 05 Converses STEP 06 Objective hit
Outbound phone call, then voicemail drop, then callback request, then agent inbound, then converses, then objective hit. The contextual layer runs across all six steps and is what holds them together. The fourth step is highlighted because it is where the worst outcomes happen and where most simulations stop.

The Callback Is Where the Danger Lies

Most outbound calls go unanswered. And the callback is where the worst outcomes happen.

We orchestrate some very sophisticated social engineering campaigns. The only complaint we have received in three years of operations is that it was too hard and advanced. We are flattered. Adversaries do not have an easy mode, and we fundamentally believe that simulations should match what adversaries are doing. If they do not, then what you are measuring is wrong.

Most simulations only measure the outbound leg. They are missing the interaction where the damage occurs.

About 11 percent of a user base will call back. Of that population, about 34 percent will take an action.

The callback leg, in two denominators
Of the user base 100 marks = 100 users contacted About 11% of a user base calls back Of the callers only New denominator: the people who called back About 34% of those who call back take an action Scale: 0 to 100 percent of callers. Full bar = every caller. Two different denominators. One is a share of the user base. The other is a share of callers. Do not multiply them together.
Breacher.ai engagement data. On the left, each mark is one user in a contacted user base of 100, and about 11 of them call back. On the right, the bar is scaled against a different denominator, the people who called back, of whom about 34 percent take an action.

What Actually Holds

In a few of our engagements we have stumbled, and we are sharing what works and what does not.

Adding friction between the phone number and the human

Leaving a phone number wide open that can be answered is inviting trouble. Simple steps like press four to connect, or leave a message to connect to the caller, add minor friction. But in some cases these have tripped up our agents. Having a complex extension tree can add friction as well. While not a permanent solution, it is a decent speed bump for now.

Follow the process

Support requests, remote access, and information are only done through a support initiated ticket or a rigid process that is always followed. The callback number is static, it is documented, and users follow the process. The process for IT support is well documented and followed. No exceptions.

Security is not an afterthought

In our engagement data, the platform an organization uses matters less than how seriously the organization takes security measures. It is somewhat an obvious statement. But take security seriously and instill it in your organization.

  • Friction on the number. Press four to connect, a message to connect step, or a complex extension tree. A decent speed bump for now, not a permanent solution.
  • A support initiated ticket, every time. Support requests, remote access, and information handoffs happen only through the ticket or the rigid process.
  • A static, documented callback number. One number, written down, and users follow the process to reach it.
  • No exceptions. The process for IT support is well documented, and it is followed.

Enterprise Risk

There is a finding in our data between the organization size and how effective IT impersonation is. Smaller organizations are less prone. It is likely because of an in-house IT support where they have engaged with the person, or are more intimately involved in day to day support. Enterprise is where the real risk is. Outsourced IT support as well. The technician may not be known. Enterprise fares the worst in our engagements.

Why size changes the outcome
Smaller organizations IT support is in-house Users have engaged with the person Involved in day to day support Less prone in our engagements Enterprise IT support is often outsourced The technician may not be known No prior contact with the person calling Fares the worst in our engagements
Smaller organizations are less prone, likely because IT support is in-house and users have engaged with the person. Enterprise is where the real risk is, and outsourced IT support compounds it because the technician may not be known.
Measure your risk. Train for what you find. Prove it changed.

The voice channel is covered end to end, including the inbound leg, in AI vishing simulation. A full engagement write-up, where autonomous agents ran the research, the voice, and the conversation, is here: agentic AI and a cloned executive. Where an organization sits relative to comparable ones is covered in the Social Engineering Risk Index.

Frequently Asked Questions

What action rate do AI vishing engagements produce?

Across our most recent agentic voice simulations, the weighted mean action rate is 11.5 percent of users contacted, with a range from 0 percent to 34.5 percent of users contacted. Action means the user did something that could put the organization at risk, most often divulging information. A year ago the same work produced roughly 1 to 3 percent of users contacted, and most people could still tell they were talking to a machine.

Why is IT helpdesk impersonation the scenario to watch?

Statistically it is higher in our engagement data, significantly so. Our working theory is that the red flag indicators awareness training has taught users, meaning pressure, urgency, and suspicion, do not exist in these campaigns. It is a routine task being asked for, and a routine task may not trip suspicion. This is why we believe threat actors like Scattered Spider are successful.

Why does the callback leg matter more than the outbound call?

Most outbound calls go unanswered, and the callback is where the worst outcomes happen. About 11 percent of a user base will call back, and of that population about 34 percent will take an action. Most simulations only measure the outbound leg, so they are missing the interaction where the damage occurs.

What is orchestration in a social engineering simulation?

Modern social engineering is not simply placing a phone call. It is conversational, it is informed by open-source intelligence, and it combines different media. The key is a contextual layer that holds each of those steps together to hit an objective. Without that contextual layer the scenario falls apart, which is why we built an orchestration engine rather than another simulation platform.

Does organization size change how well IT impersonation works?

Yes. There is a finding in our data between organization size and how effective IT impersonation is. Smaller organizations are less prone, likely because support is in-house and users have engaged with the person or are more intimately involved in day to day support. Enterprise is where the real risk is, and outsourced IT support compounds it because the technician may not be known. Enterprise fares the worst in our engagements.

What actually holds against an AI vishing campaign?

Friction between the phone number and the human, such as press four to connect or a complex extension tree, which is a decent speed bump rather than a permanent solution. A rigid process in which support requests, remote access, and information handoffs happen only through a support initiated ticket, with a static documented callback number and no exceptions. And an organization that treats security seriously rather than as an afterthought, which in our engagement data matters more than the platform it runs.

JT

Jason ThatcherFounder and CEO of Breacher.ai Corp. Fifteen years in security operations and offensive testing, previously at ZeroFox, Deepwatch, and GuidePoint Security. He builds and runs orchestrated social engineering simulations against enterprise organizations.

Find Out What Your Callback Leg Produces

We run the full sequence, including the inbound leg, and report the action rate against your real process.

Book Your Demo

Latest Posts

  • AI Vishing Red Team: What We Are Seeing in the Field

  • Social Engineering Simulation: The Complete Guide

  • Voice Phishing Assessment: What to Scope, and What the Report Has to Prove

Table Of Contents

About the Author: Jason Thatcher

Jason Thatcher is the Founder of Breacher.ai and comes from a long career of working in the Cybersecurity Industry. His past accomplishments include winning Splunk Solution of the Year in 2022 for Security Operations.

Share this post