AI Social Engineering Red Team Engagements | Breacher.ai

AI Social Engineering Red Team Engagements

It takes one person. We find out who, how far a synthetic identity got, and which process let it through.

Agentic AI voice, live executive deepfakes and synthetic candidates, composed into a scenario built from your threat model and run unannounced from outside your environment. We follow it past the click into the callback, the reset, the hire and the payment, then score it on the OSES™ Score and put it next to your peers.

Agentic AI Vishing Live Executive Deepfakes Fake Candidates Help Desk Impersonation Orchestrated Kill Chains

Unannounced, fully managed and externally originated. Nothing installed.

92%
Of organizations we test are vulnerable to deepfake social engineering
Source: Breacher.ai engagement findings
78%
Were highly vulnerable, meaning the process itself failed, not one person
Source: Breacher.ai engagement findings
11.5%
Weighted mean action rate in recent agentic voice simulations, up from 1 to 3% a year earlier
Source: Breacher.ai agentic voice simulations

What we compose into your scenario

Not a menu. These are the ways a synthetic identity gets inside an organization, and your engagement uses whichever ones your threat model calls for. A real adversary moves between them, so most engagements use more than one.

Agentic AI Vishing

An autonomous voice agent places the call, converses live if answered, leaves a callback voicemail if not, and handles the inbound callback itself. No operator per call, up to 150 concurrent live sessions, so an entire population is reachable in a day.

OutboundVoicemail DropInbound Callback
Help Desk Impersonation

The pretext that never trips a red flag. A routine request from IT support, scripted against your actual reset and MFA re-enrollment procedure, landing where users are least likely to know the voice on the line.

Credential ResetMFA Re-enrollmentRemote Access
Live Executive Deepfakes

A consenting executive, cloned in face and voice, joining a real call on Teams, Zoom or Meet and holding the pretext under questioning. The test is not whether anyone notices. It is whether your approval procedure holds when authority asks.

Live AvatarCloned VoiceApproval Chains
Fake Candidates

A synthetic applicant built for a role you are actually hiring, with a history that survives a reference check and a live face on the interview. Reported as how far the candidate got, and which control stopped them, if any did.

Hiring PipelineLive InterviewIdentity Checks
Orchestrated Kill Chains

A voicemail lands, a text references it, a Teams call confirms it. Each step borrows credibility from the last and crosses a different control boundary. Coordinated under one objective by OSES™, not sent one channel at a time.

VoiceSMSTeams & Email
Contextual OSINT Layer

Before anything runs, we map who an adversary would impersonate, which approval chains carry real money or access, and which vendors your people already expect to hear from. Every source is recorded, so each pretext traces back to public material.

Reporting LinesVendor SignalsSource Material

We follow it past the click

A phishing test ends at L2. That is where a real incident starts. Every person in an engagement is placed on exactly one rung of the escalation ladder, the deepest one they reached, and the engagement keeps going until the scenario hits its objective or a control stops it.

Depth is set by the furthest any single person got. It is not divided by headcount, because one credential handed over in a population of thousands is not a small result. A breached organization cannot score low.

One person taking a consequential action is a failure for the organization. It only takes one.
L0
No response
Nothing happened at all
L1
Reached
Opened the email, answered the call, heard the voicemail
L2
Clicked Most tests stop here
Reached the page or payload and went no further
L3
Engaged
Replied, conversed or called back, without acting
L4
Acted
Credentials, payment, access granted or data disclosed
L5
Escalation
The action landed: funds moved, access used, candidate hired

Two questions a click rate cannot answer

How far did it get, and how many people went with it. One term for each, reported side by side, and summed onto a band only after both are visible. See how the OSES™ Score works

Depth
How far did it get?

The deepest rung any single person reached. Floors at 70 the moment an action produces a consequence.

Spread
How many went with it?

Weighted share of the population that moved. Per capita and capped, so volume can never mask a contained breach.

Score
Which band, and why?

A single figure for the board, never shown without the two terms that tell you which problem you have.

Low
0 to 39
Medium
40 to 59
High
60 to 79
Critical
80 to 100

A click carries a fiftieth of the weight of a consequential action, so a program cannot look better or worse simply by changing how many messages it sends. Reporting is organizational: no named individuals, no department leaderboards.

You scored X. Your peers scored Y.

A number with no reference point is an opinion. Your result arrives next to the median for your own vertical, so you can tell an outlier from a sector median that is itself unacceptable.

Benchmark readout Sample report. Illustrative values.
Your organizationOSES™ 68 · High
Vertical median (Spread)23 of 30
Your Spread17 of 30

Read it like this: fewer of your people moved than your peers, but one case reached a consequential action. The problem is a single process gap, not the population, and that changes what you fix first.

✓
Medians on Spread only

Depth is excluded from sector medians, so one incident at one company can never set the floor for everyone else in its vertical.

✓
Weighted events, never averaged rates

Rollups are total weighted events over total targets. Averaging per-engagement rates lets small samples distort the picture.

✓
Insufficient means insufficient

Where a vertical does not yet meet our minimum sample thresholds, we say so and show no number. One organization is not a sector.

✓
Re-test, then show the delta

After remediation the same paths run again. The improvement is a measured change in the score, reported alongside it rather than folded in.

From your threat model to a signed finding

The bespoke variant of OSES™. The first two stages are what make no two engagements alike. Your side of it is a scoping call, an authorization letter and a readout.

1Design

Threat model intake

Which process would hurt most if it failed, who has authority over it, and what leadership needs to see. That sets the objective.

2Design

Build the scenario

OSINT establishes who to impersonate and how to reach your people. Pretext, sequence and synthetic assets are built for this one engagement.

3Execution

Run it unannounced

Across whichever channels the scenario needs, in a real working window, coordinated under one objective and fully external.

4Execution

Follow it downstream

Into the help desk approval, the callback, the hiring decision or the payment. Downstream observation is authorized in scope, so Depth is observed, not assumed.

5Evidence

Score and benchmark

OSES™ Score with Depth and Spread, peer vertical position, and process failure findings written for a board and an auditor.

6Evidence

Train, then re-test

Training built from what failed, then the same paths run again. A course completion is not a control. A measured delta is.

Every engagement runs under signed authorization, with named approvers, agreed impersonated identities and documented abort conditions. Client names never appear in our public material. For the deepfake methodology in depth, see deepfake red teaming. To run simulations yourself, see the OSES™ Platform.

Real engagements, real outcomes

Published with numbers where we have permission, and as outcomes only where we do not. Client names never appear.

Kudos to your entire team. We haven't even seen the report and the whole company is talking about the risks of voice cloning. It's been a huge win for us already.

CISO, Large Financial Enterprise

Published, because we have nothing to hide

The price sets the scope. Your threat model sets the scenario. Every tier is unannounced, fully managed and externally originated. See all pricing

OSES™ Assessment

$4,800

One vector, such as IT support impersonation. Up to 250 users. OSES™ Score and peer vertical rank.

Multi-Vector Engagement

$12,500

Multiple vectors, orchestrated end to end. Up to 1,000 users.

Full Scope

from$25,000

Executive impersonation, hiring pipeline, fake candidates and custom kill chains, designed per engagement.

Start with a red team. Keep half of it.
50% of any red team fee is credited toward your first year of the OSES™ Platform, capped at half the first-year fee.
Scope an Engagement

Executive voice and likeness are cloned only with the consent of the person being cloned, confirmed by a biometric check and voice verification before anything is generated. An administrator cannot authorize cloning someone else.

What CISOs ask before the first engagement

What is an AI social engineering red team engagement?

An authorized, unannounced assessment in which AI-generated voice, video and identity are used to impersonate people your organization trusts, to find out whether your people and processes can be moved to a consequential action. Unlike a phishing test, which ends at the click, the engagement follows the scenario into the callback, the help desk reset, the hiring decision or the payment approval.

How is the OSES™ Score calculated?

The OSES™ Score is two terms added together. Depth is how far the deepest single case progressed, and it is not divided by headcount. Spread is how much of the population moved, weighted by how far each person went, and capped so volume cannot mask a contained breach. The total lands on one of four bands: Low, Medium, High or Critical. The band is never reported without the two terms behind it.

How does the peer benchmark work?

Your result is reported next to the median for your own vertical, so you can tell an outlier from a sector median that is itself unacceptable. Vertical medians are computed on Spread only, so one incident at one company cannot set a sector floor. Rollups are total weighted events over total targets, and where a group does not yet meet our minimum sample thresholds we report it as insufficient rather than publish a number.

Is this a fixed package?

No. The published price sets the scope: how many vectors and how many users. The scenario itself is designed from your threat model, so the adversary, the pretext, the channels, the identities impersonated and the objective are decided for your organization specifically. Two clients in the same sector get materially different engagements.

Is it safe to run against real employees?

Yes. Every engagement runs under signed authorization with named approvers and documented abort conditions. Executive voice and likeness are cloned only with the consent of the person being cloned, confirmed by a biometric check and voice verification before anything is generated. Findings are reported at the organization level, with no named individuals and no department leaderboards.

Do we need to install anything?

No. Engagements are fully managed and externally originated. Nothing is installed and nothing touches your security stack, which removes most of the change-control burden that blocks testing in regulated environments.

How does the red team credit work?

50% of any red team fee is credited toward your first year of the OSES™ Platform, capped at half the first-year fee. Run an engagement to establish where you stand, then apply the credit when your team takes over running simulations itself.

Find out how far a synthetic identity gets inside

Thirty minutes on your threat model: which process would hurt most if it failed, who has authority over it, and what the report needs to prove.

✓ Designed per engagement ✓ OSES™ Score and peer benchmark ✓ From $4,800
Book Demo

MSPs and MSSPs: see the partner program.