Voice Phishing Assessment: What to Scope, and What the Report Has to Prove

Categories: Deepfake,Published On: August 31st, 2026,
Voice Phishing Assessment: Scope It So It Proves Something
Assessment Guide

Voice Phishing Assessment: What to Scope, and What the Report Has to Prove

Most voice phishing assessments are scoped against a pretext and scored on whether the call sounded convincing. Both are the wrong end of the problem. Scope it against the actions a call could cause, and the report tells you something you can fix.

By Jason Thatcher, Founder and CEO, Breacher.ai

See how a scoped voice phishing assessment runs against your own verification procedures.

Book Your Demo

We run authorized voice engagements against enterprise organizations for a living, which means we also read a lot of other people's assessment reports. The common failure is not that the calls were unrealistic. It is that the assessment was scoped against a story rather than against a set of actions, so the report describes what happened on the phone and cannot tell the buyer which procedure to change on Monday.

A voice phishing assessment is worth its cost when it answers a narrow question: if a credible caller asks for one of the things that would actually hurt us, does the verification step happen. Everything else in the engagement is machinery in service of that question. This is the scoping and reading guide we give buyers, including the parts that make our own findings less flattering.

What a Voice Phishing Assessment Is, and What It Is Not

The term covers three different products in this market, and the differences matter more than the shared label.

It is not a questionnaire. A maturity survey records what a policy says and what people believe they would do. Every organization we have tested had a callback policy on paper. The gap between the policy and the phone call is the entire subject of the assessment, so a method that reads the policy cannot measure it.

It is not an awareness exercise with audio bolted on. Playing a recorded message at a population and counting who stayed on the line measures availability. A verification procedure is a conversation, and it can only be tested by something that answers a question, tolerates being challenged, and handles the callback when one is requested.

It is an authorized measurement of a control. Calls run under signed authorization against your own organization, targeting the procedures that stand between a request and a consequential action. The finding is which procedure held, which one was skipped, and which one never existed. The mechanics of how the calls are delivered are covered on our AI vishing simulation page. This piece is about scoping and reading the result.

A voice phishing assessment does not measure whether people can tell. It measures whether it mattered that they could not.

The Five Decisions Made Before Anyone Dials

These are decided in scoping, and they determine whether the report is actionable. Getting them wrong is not recoverable by running better calls.

DECISION 03

Choose the direction and the population

Outbound calls impersonate internal support or an executive to an employee. Inbound calls impersonate an employee to the service desk. They fail differently and they need separate scoping, because a successful outbound call typically yields one session while a successful inbound call yields a working identity. Size the population so the result is not a sample of whoever happened to answer, and include the shift patterns where the desk is thinnest.

DECISION 04

Fix authorization, pacing, and abort conditions

Signed authorization against your own organization, named approvers, agreed call windows and concurrency, written approval for any likeness used, retention and destruction terms for voice assets, and documented conditions under which the engagement stops. Pacing belongs in scope rather than in delivery, because a campaign that floods a queue changes the behavior it was meant to observe.

DECISION 05

Define continuation and the re-test

A call that ends in a promise is not a finding. Decide in advance how far the simulation follows the request into the process behind it: the reset ticket raised, the approval requested, the follow-up message sent. Then decide which paths get run again after remediation. A single measurement is a baseline. The delta on the same paths is the evidence that anything changed.

Why an Assessment Scored on Detection Tells You Nothing You Can Fix

A large share of voice assessments still report a version of how many people noticed. It reads well and it does not survive contact with the next model release.

Detection is a decaying control. Its effectiveness is a function of how good the synthetic audio is, and generation quality improves continuously and cheaply while human hearing does not improve at all. A finding expressed as a detection rate has a shelf life set by the adversary rather than by you.

Procedural verification is invariant to the same variable. A callback to a number taken from your own directory returns the same result against a crude clone and against a perfect one. The procedure never listens to the audio. That is the whole reason it holds.

The honest bound on that claim. Better synthetic media does not defeat a verification requirement directly. It raises the pressure applied against it, which shows up as a rise in the exception rate, meaning the number of times someone consciously waived a step that they knew existed. That is a measurable variable and we report it separately, because claiming perfect invariance is the kind of statement someone eventually falsifies in front of you.

None of this removes people from the measurement. Procedures are executed by people under pressure, so the assessment is still measuring human behavior. It has simply moved from a variable that cannot be trained to one that can. We covered how the same argument plays out at the service desk in our breakdown of IT support impersonation.

63%
Of the people tested across our engagements could not distinguish synthetic voice or video from a real person while it was happening to them
14.5%
Weighted mean action rate across voice campaigns that contained no link anywhere, which a phishing platform would have recorded as nothing at all
0
Tools in a normal enterprise stack that inspect, filter, or score an inbound phone call before it reaches a person

Figures from Breacher.ai OSES™ engagement data, published in the OSES Risk Index.

What the Report Has to Contain

Read a sample report before you sign anything. These are the elements that make findings usable, and their absence is the most common reason an assessment produces a slide instead of a remediation plan.

  • Coverage, stated as a figure. How many of your consequential actions carry a verification requirement, and which ones do not. The uncovered paths are the remediation program, and they can be sized before a single call is placed.
  • Action rate. How often the call produced the outcome it asked for: a reset completed, a code read back, an approval granted, a payment detail changed. This is the closest equivalent to a click and it is a far more serious event.
  • Process hold rate. How often the documented verification procedure was executed end to end under pressure. A low action rate with a low hold rate means you were lucky rather than protected, and the report should say so in those words.
  • Exception rate, separated from omission. Verification not attempted and verification attempted then waived are different failures with different fixes. A report that collapses them into one number has thrown away the distinction that decides your remediation.
  • Report rate and time to first report. Including the calls that were refused and never escalated. A person who resisted and said nothing leaves the next person facing the same caller with no warning.
  • Findings written against procedures, not people. Organizational and process level reporting, with no named individuals and no departmental leaderboards. Per-user results are useful for routing training to the path that failed. They are not the outcome metric, because a score on a person does not tell you whether the payment goes out.
  • A peer position, not just your own trend line. An internal comparison tells you whether you improved against yourself, which is a low bar when you do not know where you started relative to comparable organizations.

Questions Worth Asking Before You Sign

Ask these of any vendor, including us. The answers separate an assessment from a demonstration.

Which consequential actions will this engagement be scoped against, and who writes that list?
Does the call handle an inbound callback, or does the scenario end when the target says they will call back?
Will the simulation continue past the call into the reset, the approval, or the payment step?
Does the report distinguish verification skipped from verification waived as an exception?
Is the finding reported against a procedure, or against a list of employees who failed?
What exactly gets re-tested after remediation, and is the same path used so the delta means something?

The callback question is the one that most reliably separates vendors. Real voice fraud converts on the return call, because a number the target dialed themselves feels verified before anyone speaks. An assessment that cannot answer the phone has skipped the part where the loss usually happens.

What We Would Do First

If you have one week and no budget, do these three things in this order. None of them requires a vendor.

Write the consequential action list. One page. Every action in your organization that moves money, grants access, changes a payment destination, or discloses data on request. Do it with finance, the service desk, and identity in the room, because the list is always longer than any one of them expects.

Mark which ones have a verification requirement. Not a policy that mentions verification in general terms, a specific requirement attached to that specific action, reachable by the person who performs it. Count the marks. That fraction is your coverage, and for most organizations the first calculation of it is uncomfortable.

Then run the calls against the uncovered paths first. Testing a well-covered wire process proves a control you already have. The value is in the paths where nothing stands between a plausible request and a consequential action, because those are the ones an adversary reaches without a fight.

Measure your risk. Train for what you find. Prove it changed.

Frequently Asked Questions

What is a voice phishing assessment?

An authorized assessment that places voice calls against your own people and processes to establish whether the verification requirements protecting consequential actions hold under pressure. There is no link and no landing page, so the measurement is behavioral: whether the requested action completed, whether the documented verification step was executed, and whether anyone reported the call afterwards. A modern engagement uses cloned voice and a conversational agent because that is what the adversary uses, but the subject of the measurement is the procedure, not the audio.

How is a voice phishing assessment different from a phishing simulation?

A phishing simulation measures a click on a link. A voice call contains no link, so there is nothing to count and the metric has to be the action taken and the procedure that was supposed to prevent it. The second difference is interactivity. The call answers questions, absorbs hesitation, and adapts, which is the only way to test a verification step, because verification is itself a conversation. A recorded message played at a person cannot test a procedure that a person has to talk their way through.

What should a voice phishing assessment measure?

Three organizational measures. Action rate, meaning how often the call produced the outcome it asked for. Process hold rate, meaning how often the documented verification procedure was executed end to end under pressure. Report rate, including how long it took for the first report to reach security. Alongside those, the coverage figure computed during scoping: how many of your consequential actions have a verification requirement at all. Coverage is the number that sizes the remediation program, and most organizations have never calculated it.

How long does a voice phishing assessment take?

Scoping and authorization usually take longer than delivery. Building the consequential action list and confirming which of those actions carry a verification requirement is the work that determines whether the findings are useful, and it is desk work that can start immediately. Call delivery for a defined population typically runs inside an agreed window of days rather than weeks, with reporting to follow. A re-test on the same paths is scheduled after remediation, because a single measurement is a baseline and not evidence of change.

Is a voice phishing assessment legal, and how is authorization handled?

Every engagement runs under signed authorization against your own organization, inside an agreed scope, window, and call volume, with named approvers and documented abort conditions. Voice assets are generated only for individuals whose likeness use has been approved in writing, and they are destroyed at the close of the engagement. Recording is decided during scoping according to your jurisdiction and your own policy. Reporting is organizational, with no named individuals and no departmental leaderboards.

JT

Jason ThatcherFounder and CEO of Breacher.ai Corp. Fifteen years in security operations and offensive testing, previously at ZeroFox, Deepwatch, and GuidePoint Security. He builds and runs orchestrated social engineering simulations against enterprise organizations.

Scope One Against Your Own Actions

Thirty minutes. We walk your consequential action list, show where coverage is missing, and pick the path worth testing first.

Book Your Demo

Latest Posts

  • Voice Phishing Assessment: What to Scope, and What the Report Has to Prove

  • Protecting Against Deepfake Scams: Deepfake Phishing Is a Campaign, Not just a Call

  • Case Study: When the Hire Is the Intrusion | Breacher.ai

Table Of Contents

About the Author: Jason Thatcher

Jason Thatcher is the Founder of Breacher.ai and comes from a long career of working in the Cybersecurity Industry. His past accomplishments include winning Splunk Solution of the Year in 2022 for Security Operations.

Share this post