AI Voice Agent Social Engineering: Video Is the Wrong Vector to Watch. It Is Audio.
AI Voice Agent Social Engineering: Video Is the Wrong Vector to Watch. It Is Audio.
We build these engagements for a living, so we tested the market’s picture of the deepfake threat against what we see in the field. The pre-recorded deepfake video of your CEO is the weakest thing we run. The credentials move over audio: an interactive synthetic voice that calls back, holds the conversation, and now runs that loop against hundreds of people at once.
Everyone is watching the video. The credentials are moving over the phone.
We build and run authorized social engineering simulations against enterprise organizations as paid engagements. So when the market settled on a particular image of the deepfake threat, we tested it against what we actually see in the field. The picture that emerged is sharper, less cinematic, and far more useful than the pitch most of the category is running.
This post is about the finding that reorganized how we scope every engagement. The dangerous deepfake is not a video. It is a voice, in a loop, at scale.
Each grid is 100 people hit with the same credential-expiry pretext on a different channel. The one-way clip pulls plenty of views but stalls. The interactive channels are where people actually reach the ask, and the voice loop is the one an agent runs against everyone at once.
Pre-recorded video
Plenty of views. A one-way clip gives nobody anything to answer back, so the trail goes cold long before the ask.
Live video call
Converts well once someone shows up, because it is a conversation. But every target still needs an invite, a slot and a reason to join.
AI voice loop
The full loop on a channel with no gateway in front of it, run by one autonomous agent against everyone at once.
Representative 100-person sample, illustrative, drawn to the shape we see in the field. The voice loop is scaled from a single autonomous-agent engagement. Not a controlled head-to-head: each vector ran against different populations, so read the pattern, not the decimal.
How we measure
We score every simulation on a depth ladder: did the target engage, did they click, did they take an action such as entering a credential, and did that action reach a real consequence. A click alone never counts as a finding. We only credit a result when the action reaches something that matters, and we state the definition in play every time. That discipline is what separates an assessment from a vanity metric, and it is what lets us say with confidence where the risk actually sits.
Video is the wrong vector to watch
The image the market sells is a flawless pre-recorded video of your CEO that instantly moves money. It is the clip that photographs well for a conference keynote, and it is where most of the attention goes: video detection, spot-the-glitch training, the cinematic demo. In our engagements that one-way clip is consistently the weakest vector we run. Pre-recorded broadcasts produce clicks and meeting-joins, and very little that reaches a real consequence.
Read that carefully, because it is easy to misread. The point is not that synthetic video is harmless, and it is not that a live video call is harmless. A live, interactive video call is a conversation, and it converts for the same reason the voice loop does. The vector that fails is the pre-recorded, one-way clip. A convincing face delivered as a one-way blast gives the target nothing to respond to, so the sequence stalls before it reaches anything that matters. If your threat model is built around a pre-recorded video of an executive, you are watching the weakest vector while the one that works runs over audio.
Audio is the vector: a voice agent that calls back
The simulations that convert share a different shape: an interactive voice loop. The target dials a number back or holds a conversation with a live agent on the phone. That loop is where the consequences land in our engagements, by a wide margin. Our single largest result came from an autonomous AI voice agent running a credential-expiry pretext against hundreds of people at once, with no human operator on any call.
That result breaks a comfortable assumption: that deepfakes are surgical, deep but not wide. A human impersonating an executive can go deep or wide, not both, because nobody can staff hundreds of personalized calls. The AI agent does both at once. That is the categorically new thing. Not that a fake can be convincing, people have always been convincing, but that a convincing, interactive, authoritative voice now scales to hundreds of targets at the same time. The synthetic voice is doing real work here. What agentic autonomy adds is reach at a depth a human could only reach one call at a time.
Why audio wins
The field results say where the consequences land. The mechanics explain why they land on the phone and not on the screen.
The loop needs a two-way channel, and the phone is the native one
A voicemail invites a callback. The callback opens a conversation. The conversation carries the ask. Every step is ordinary phone behavior, so nothing about the sequence looks unusual to the person inside it. A one-way clip cannot do any of that.
The phone has no gateway
Email passes through filters, sandboxes and warning banners before a person ever sees it. A call rings straight through to a human, with nothing in between to score it, quarantine it or flag it. That unguarded channel is exactly where the loop does its work.
Audio is what an autonomous agent runs live, at scale
A single agent can hold real-time voice conversations with hundreds of people at once. The synthetic voice answers, adapts and presses toward the request in every one of those calls simultaneously. That is reach and depth together, delivered over audio.
The help desk beats the CEO
Conventional wisdom says authority bias tracks seniority: the more senior the impersonated title, the better the result. What we see in the field is close to the opposite. Ranked by consequences, the help desk persona outperforms every executive we impersonate, and the most senior titles, the CISO above all, are the hardest sell.
The driver is not altitude, it is routine plausibility. A help desk call about an expiring credential is something staff are trained to comply with. A personal call from the CISO is rare enough to trigger suspicion. The most dangerous impersonation is not the most powerful person in the building. It is the most ordinary request, the one that sounds like a Tuesday.
Authority climbs as you read down. Success falls the whole way. The impersonation that lands is the ordinary one, not the powerful one.
Relative to each other, illustrative, ranked by the pattern we see in the field. The scale is the story: lower authority, higher success.
Synthetic voice wins the second move
When we run synthetic voice alongside conventional lures, email still reaches more people at first contact. That is expected. Where synthetic voice separates itself is what happens next: it is the vector that converts engaged people to credential disclosure.
The lesson is precise, and it credits the synthesis rather than dismissing it. Synthetic voice does not win by making the first touch more tempting. It wins by changing what happens after someone is already on the hook, inside a live conversation where a human-sounding, authoritative agent can press, reassure and redirect in real time. Reach is an email problem. Conversion is where the synthetic voice earns its reputation, and agentic autonomy is what lets that conversion run at scale.
- Hundreds of people dialed by a single agent. Reach no human caller could staff one line at a time.
- No human operator on any call. The agent left the voicemail, took the callback and held the conversation itself.
- Valid credential disclosures captured end to end. Each one reached only because the loop ran from the first call to the ask.
Credentials are captured for measurement only and never retained. For the architecture of the return leg, see our callback simulation breakdown and the wider field findings from the callback log.
Deepfake social engineering is not overhyped. The industry is just watching the wrong channel.
The fix follows the finding
If the threat is an interactive voice loop that converts a routine request into a disclosed credential, then the defense is not teaching people to hear the fake. That control decays with every model release, and we wrote out why in why detection training is a decaying control. The voice will keep getting better. Human perception will not.
What holds is a procedure that does not depend on anyone noticing anything. Three moves, in order:
- Add friction to the voice channel. Any consequential request that arrives by phone, a credential reset, an MFA change, a payment instruction, a data export, gets a verification step before anyone acts on it. The channel most organizations have left wide open is the one the loop is built to exploit.
- Require an out-of-band callback to a known-good number. Not the number that called you, and not the one in the voicemail. The directory number. A callback returns the same answer against a crude fake and a perfect one, because it does not read the audio. We walk through how to scope that in our voice phishing assessment guide, and the training logic in process over detection.
- Rehearse the ordinary call, not just the cinematic clip. Defend the help desk request about an expiring credential, because that is the impersonation that converts. Then test whether the procedure actually ran under pressure, rather than assuming it did.
Stop watching the video. Listen to the phone. The boring help desk call is the attack that works.
This is what the real-time voice agent changed, and it is exactly what Breacher.ai is built to reproduce. We run the whole loop as one orchestrated, fully automated engagement, the outbound call, the voicemail, the inbound callback and the live two-way conversation, driven by an autonomous agent rather than a person on the line. Every scenario terminates at a named consequential action, and the report tells you whether your verification procedure executed, separating a check that was skipped from one that was consciously waived. You cannot fix a failure you have never observed.
“Kudos to your entire team. We haven't even seen the report and the whole company is talking about the risks of voice cloning. It's been a huge win for us already.”
Measure your risk.
Train for what you find.
Prove it changed.
Bring the voice scenario you actually fear. We will run the whole loop end to end and show you where it terminates.
Book Your DemoFrequently asked questions
It is an impersonation technique in which an autonomous AI voice agent, not a human operator, places and answers calls using synthetic speech, holds a live two-way conversation, and drives the person toward a consequential action such as disclosing a credential or approving a payment. The agent does not read from a script and hang up. It conducts an interactive loop: it can leave a voicemail, take an inbound callback to a spoofed number, answer off-script questions, and keep the conversation moving toward the request. The autonomy is what lets a single agent run that loop against hundreds of people at once.
Because that is where the consequences land. In our engagements, one-way broadcasts such as the pre-recorded synthetic video clip are consistently the weakest vector we run. The interactive voice loop, where the target calls a number back and talks to a live agent, is where credentials actually move. The phone is a two-way channel with no gateway between the caller and the person, which is exactly what the loop needs. Deepfakes are dangerous, but the danger is not the cinematic clip. It is an interactive synthetic voice with a credential request waiting at the end, now driven by an agent that can hold hundreds of those conversations at once.
Ranked by realized consequences, the help desk persona outperforms every executive persona we run, and the most senior titles perform worst of all. The driver is not authority, it is routine plausibility. A help desk call about an expiring credential is something staff are trained to comply with, so it does not trigger the suspicion a rare personal call from a senior executive does. The most dangerous impersonation is the most ordinary request, not the most powerful person.
Both matter, at different moments. Conventional email still reaches more people at first contact. Synthetic voice is the vector that converts engaged people to credential disclosure, inside a live conversation where a human-sounding agent can press, reassure and redirect in real time. Reach is an email problem. Conversion is where the synthetic voice earns its reputation, and agentic autonomy is what lets that conversion scale.
Not training people to hear the fake, because that control decays as synthesis improves. What holds is a procedure that does not depend on anyone noticing anything: add friction to the voice channel, require an out-of-band callback to a known-good number before any credential reset, MFA change, payment instruction or data export requested by phone, and rehearse that step until it runs under pressure. The callback does not read the audio. It does not care how good the voice is.
Breacher.ai runs the full interactive loop as an authorized, fully automated engagement: the outbound call, the voicemail, the inbound callback, and the live two-way conversation, driven by an autonomous agent rather than a person on the line. Every scenario terminates at a named consequential action, and the report measures whether the verification procedure that was supposed to stop it executed, separating a check that was skipped from one that was consciously waived. Credentials are never retained.
Run the Loop Against Your Own People
Thirty minutes. Bring the voice scenario you actually fear, and we will walk it from the first call through the callback to the findings, so you can decide whether your process would have held.
Book Your Demo
