Deepfake and AI voice phishing assessmentRun the way organizations are actually being breached
A fully managed assessment that replicates how organizations are being breached right now. An autonomous AI voice agent calls your people as their IT support desk. If they don't answer, it leaves a voicemail with a callback number. When they call back, it answers and holds the conversation. We measure every leg of that sequence and hand you an OSES™ risk score with your position against your peer vertical.
$4,800
Flat fee, up to 250 users. One assessment, no commitment.
- Outbound calls, voicemail drops and inbound callbacks, all measured
- Fully managed by Breacher.ai from scoping to readout
- Externally originated. No agents, no integration, no access to your environment
- Voice cloning of a named individual available on request
- Your OSES risk score, banded Low to Critical
- Your position against your peer vertical
No subscription, no platform to administer, no renewal to cancel. Populations above 250 users are scoped separately.
Why this vector
Voice has overtaken email as the way people get in
Mandiant's M-Trends 2026, built on more than 500,000 hours of incident response, put highly interactive voice phishing at 11% of initial infection vectors, second only to exploitation. Email phishing fell to 6%. In cloud-related compromises voice was the number one vector outright, at 23%, ahead of third-party compromise, stolen credentials and email.
The reason is simple. Email security got better. Human voices did not get easier to distrust. Attackers moved to the channel where the automated controls you bought do not operate, and where the thing being exploited is a routine workflow rather than a suspicious message.
The 2026 breach record reads accordingly: a telecoms operator through a vished employee's single sign-on account, an identity protection company through one employee, a networking vendor through one employee, each of them ending in mass customer data exposure. And the tooling has industrialized. High-volume automated voice campaigns now run this at a scale that used to require a room full of native English speakers.
Detection is a decaying control. Generation quality improves every quarter and human perception does not. Procedural verification is an invariant one. This assessment does not ask whether your people can hear that a voice is synthetic. It asks whether your process holds when they can't.
It does not need a deepfake to work
Some of the most effective campaigns in our book used a conversational AI voice agent with no cloned voice at all. The pretext carries it. Cloning raises the ceiling, and we will run it where you want it tested, but the absence of a cloned voice is not a defense and should not be treated as one.
Nor does it need a click. The call is the payload.
What runs
One agent, your help desk's identity, your whole population
The pretext is internal IT support. It is deliberately mundane, because that is what makes it work, and it maps directly onto what is landing in the field right now: a call, a compliant employee, a single sign-on account, and from there whatever that account can reach.
Why nothing about it feels wrong
A call from support is a routine workflow. Remote access is a routine workflow. Support is frequently outsourced, so an unfamiliar voice is expected rather than anomalous, and there is no deadline attached to the request.
Every signal your awareness training taught people to read comes back clean. Instead of a red flag, they get a false green flag. That is why urgency-based training does not fire on this, and why measuring whether people noticed is the wrong question.
- Credential and identity verification
- Access or account recovery assistance
- Screen share and remote support session
- Optional: cloned voice of a named individual, with documented consent
The call is the beginning, not the incident
One compliant employee is now one working identity. That identity reaches single sign-on, and single sign-on reaches the SaaS estate where the customer data actually lives. The recent breaches did not need lateral movement or malware to become nine-figure disclosure events.
Mandiant measured the hand-off from initial access to a second actor collapsing from roughly eight hours in 2022 to 22 seconds. There is no realistic window in which a person reports the call in time.
- Identity compromise, not endpoint compromise
- SaaS and cloud data exposure downstream
- No malware, no exploit, nothing for the endpoint stack to catch
- Detection sits with your people and your process, or nowhere
Available as an add-on: we call your service desk posing as one of your employees, locked out, and request a reset. That direction tests your account recovery procedure rather than your population, and needs its own authorization and a direct-dial number. Ask for it on the scoping call if your recovery workflow is what you want measured.
The sequence
Three legs, measured separately, each with its own denominator
Most voice testing measures the outbound call and stops there. That discards the majority of the sample and misses the leg where, in our data, the worst outcomes happen.
The agent places the call
It opens with a routine pretext and holds a live, unscripted conversation if the call connects, adapting to what the person says rather than reading a decision tree.
Most people don't pick up
The agent leaves a voicemail with a callback number. Across our voice work this is the dominant route through a population, not the exception, so a test that treats voicemail as a miss is throwing away most of its own sample.
They call the number back
The agent answers live and continues the conversation. This is where the most consequential outcomes occur, for a structural reason: the person initiated the contact, so there is no cold open, no urgency and no reason for suspicion.
Orchestrating those legs so a single agent carries one identity across all of them is the technically hard part, and it is what determines whether the number you get back means anything. A voice simulation that isn't built to capture the callback is missing its own most important finding.
How OSES scores it
One person, counted once, at the deepest point they reached
OSES™ places every target on exactly one rung of an escalation ladder, the deepest one they got to across all three legs. The rungs are mutually exclusive, so nobody is double counted and the score cannot be inflated by activity.
| Rung | Definition | What it looks like in this assessment |
|---|---|---|
| L0 | No action | Never answered, never returned the voicemail, no engagement of any kind. |
| L1 | Small engagement | Answered the call or listened to the voicemail, then disengaged without conversing. |
| L2 | Clicked | Followed a link where the sequence includes one. A voice-led run shows few or none, and that is the point: a click metric cannot see this tradecraft at all. |
| L3 | Engaged | Held a conversation with the agent, or rang the voicemail number back, without taking the requested action. |
| L4 | Took an action | Complied. Disclosed credential or identity information, approved access, or agreed to a remote support session. Also covers a process failure: a verification step that existed on paper and was not performed. |
| L5 | Escalation | A consequential outcome landed downstream of the action. Recorded separately from L4 so no target is counted twice. |
The score has two terms
DEPTH
How far did the agent get?
Set by the deepest rung any single person reached. Not averaged, not divided by headcount. One reset performed for a caller who could not prove who they were is an organizational failure regardless of how many colleagues held the line, and DEPTH is a floor that keeps it visible in the score.
SPREAD
How much of the workforce went with it?
Per-capita and weighted by rung, so engagement counts for a little and action counts for a lot. Capped, so it adjusts the reading rather than dominating it. This is the term that tells you whether you saw one outlier or a population-wide pattern.
Low
0 to 39
Medium
40 to 59
High
60 to 79
Critical
80 to 100
The deliverable that matters
Your score is context. Your peer vertical is the finding.
A number on its own tells you almost nothing. Knowing 14% of your people complied does not tell you whether you are an outlier or whether your entire sector sits at the same figure, and those two situations call for completely different responses.
This matters more here than anywhere, because these crews work a sector at a time. Retail, then insurance, then aviation. Knowing where you sit inside your own vertical is knowing whether you are the soft target in the set they are about to work through.
Peer medians are computed on SPREAD only. One organization's incident cannot set a floor for its whole sector, and your worst single outcome cannot be diluted by someone else's good quarter.
11%
Of initial infection vectors were interactive voice phishing, the second most common overall
6%
Email phishing, down sharply as automated controls improved
23%
Voice phishing in cloud-related compromises, the single most common vector
22 sec
Median hand-off from initial access to a second actor, down from about eight hours in 2022
Source: Mandiant M-Trends 2026, based on more than 500,000 hours of incident response investigations conducted in 2025.
1,000+
Simulated targets in the benchmark
92%
Of organizations tested were vulnerable to deepfake social engineering
78%
Were highly vulnerable
63%
Of users could not distinguish synthetic from real
Drawn from Breacher.ai's own engagement book.
What you get back
A readout you can take to a board, not a CSV of click events
Your OSES risk score and band
The composite 0 to 100 figure with DEPTH and SPREAD shown separately, so you can see whether you are carrying a depth problem, a spread problem, or both.
Peer vertical position
Where you sit against organizations like yours, with the suppression rules stated on the page so you know exactly what the comparison rests on.
Named process failures
The specific verification step that should have interrupted the call and didn't, and what the compliant identity would have reached. Written against your own policy where you supply one.
Per-leg breakdown
Outbound, voicemail and inbound reported separately with their own denominators. Callback rate against the unanswered pool, never against total headcount.
Procedural recommendations
What to change in your verification and support policy, written against what actually failed in your run rather than a generic checklist.
How it runs
Roughly three weeks, and almost none of it is yours
STEP 1
Authorization and scope
Signed authorization, named contact, escalation path, and the roster of up to 250 users with phone numbers. We agree the pretext, the calling window and who gets told in advance. One call and a few days.
STEP 2
Execution
We run the sequence across your agreed window, business hours only unless you ask otherwise. Outbound, voicemail and inbound all run autonomously. Your team does nothing except stay reachable in case something needs stopping.
STEP 3
Scoring and readout
We score the run, place you against your peer vertical and walk your team through it live. You keep the report.
What we need from you
- Written authorization from someone who can give it
- A roster of up to 250 users with phone numbers
- A named contact who can halt the run
- Your support and verification policy, if one exists, so we can test against it rather than against a generic standard
What we don't need
- Any access to your environment. The assessment originates entirely externally
- Any software, agent or integration installed anywhere
- Allowlisting, mail flow changes or directory sync
- Time from your security team beyond the scoping call and the readout
No credential material is captured
We stop at the point of demonstrated compliance. We record that the person disclosed, not what they disclosed. No account is accessed, nothing usable is stored, and nothing is passed back to you.
Nobody is named to their manager
The readout is organizational. Individual identifiers go to your named contact only if you need them for follow-up, and are not part of the standard report. This measures the control, not the person.
Consent and likeness are handled
Call recording runs against the consent requirements of the jurisdictions your users sit in. Voice cloning of a named individual requires that person's documented consent, without exception.
Questions
Before you buy
What is included in the $4,800?
Scoping, authorization, pretext design, execution of all three legs against up to 250 users, OSES scoring, the peer vertical comparison, the written report and a live readout. Flat fee, no commitment. Above 250 users we scope separately.
Is this really the same tradecraft?
It is the initial access half, which is the half that decides whether the rest happens. We do not deploy remote access tooling, move laterally or touch your data. If you want the chain tested past the point of access, that is a bespoke red team engagement, not this.
Does it use a cloned voice?
Not by default. The standard run uses an AI voice agent, because the pretext is what carries this vector and we want the result to be clean rather than flattering to us. Cloning a specific person is available where you want it tested and consent for that likeness is in place.
Our IT support is outsourced. Does that change anything?
It makes the result worse, which is the point of measuring it. Your people do not know their own support staff, so an unfamiliar voice is not anomalous and the pretext gets more room. If you also want the provider's own desk tested as an add-on, that needs their authorization alongside yours.
How is this different from a phishing simulation?
A phishing simulation measures whether someone clicked. Here there is usually nothing to click; the call is the payload. What is being measured is whether your verification process held, which is a different question with a different answer.
What if we score badly?
Most organizations do on first run. The value is in knowing where you sit relative to your peers and which procedural step failed, not in passing. Nobody has ever been sold a control by being told everything was fine.
Can we re-run it to show improvement?
Yes, scored identically with no adjustment for having been tested before. The baseline-versus-retest delta is reported alongside the new score rather than folded into it, so the improvement is real and not manufactured by the model.
Do our users need to be warned?
Your call, and we support either. Announced runs with a stated window are a legitimate drill and produce useful process data. Unannounced runs produce a truer measurement. We recommend practicing announced and assessing unannounced.
Find out how far a phone call gets in your organization
Outbound, voicemail and callback, all measured. Up to 250 users. $4,800 flat, no commitment. You get your OSES risk score, your peer vertical position and the specific procedural step that let the call through.
