AI Social Engineering Red Team Engagements
It takes one person. We find out who, how far a synthetic identity got, and which process let it through.
Agentic AI voice, live executive deepfakes and synthetic candidates, composed into a scenario built from your threat model and run unannounced from outside your environment. We follow it past the click into the callback, the reset, the hire and the payment, then score it on the OSES™ Score and put it next to your peers.
Unannounced, fully managed and externally originated. Nothing installed.
What we compose into your scenario
Not a menu. These are the ways a synthetic identity gets inside an organization, and your engagement uses whichever ones your threat model calls for. A real adversary moves between them, so most engagements use more than one.
An autonomous voice agent places the call, converses live if answered, leaves a callback voicemail if not, and handles the inbound callback itself. No operator per call, up to 150 concurrent live sessions, so an entire population is reachable in a day.
The pretext that never trips a red flag. A routine request from IT support, scripted against your actual reset and MFA re-enrollment procedure, landing where users are least likely to know the voice on the line.
A consenting executive, cloned in face and voice, joining a real call on Teams, Zoom or Meet and holding the pretext under questioning. The test is not whether anyone notices. It is whether your approval procedure holds when authority asks.
A synthetic applicant built for a role you are actually hiring, with a history that survives a reference check and a live face on the interview. Reported as how far the candidate got, and which control stopped them, if any did.
A voicemail lands, a text references it, a Teams call confirms it. Each step borrows credibility from the last and crosses a different control boundary. Coordinated under one objective by OSES™, not sent one channel at a time.
Before anything runs, we map who an adversary would impersonate, which approval chains carry real money or access, and which vendors your people already expect to hear from. Every source is recorded, so each pretext traces back to public material.
We follow it past the click
A phishing test ends at L2. That is where a real incident starts. Every person in an engagement is placed on exactly one rung of the escalation ladder, the deepest one they reached, and the engagement keeps going until the scenario hits its objective or a control stops it.
Depth is set by the furthest any single person got. It is not divided by headcount, because one credential handed over in a population of thousands is not a small result. A breached organization cannot score low.
Two questions a click rate cannot answer
How far did it get, and how many people went with it. One term for each, reported side by side, and summed onto a band only after both are visible. See how the OSES™ Score works
The deepest rung any single person reached. Floors at 70 the moment an action produces a consequence.
Weighted share of the population that moved. Per capita and capped, so volume can never mask a contained breach.
A single figure for the board, never shown without the two terms that tell you which problem you have.
A click carries a fiftieth of the weight of a consequential action, so a program cannot look better or worse simply by changing how many messages it sends. Reporting is organizational: no named individuals, no department leaderboards.
You scored X. Your peers scored Y.
A number with no reference point is an opinion. Your result arrives next to the median for your own vertical, so you can tell an outlier from a sector median that is itself unacceptable.
Read it like this: fewer of your people moved than your peers, but one case reached a consequential action. The problem is a single process gap, not the population, and that changes what you fix first.
Depth is excluded from sector medians, so one incident at one company can never set the floor for everyone else in its vertical.
Rollups are total weighted events over total targets. Averaging per-engagement rates lets small samples distort the picture.
Where a vertical does not yet meet our minimum sample thresholds, we say so and show no number. One organization is not a sector.
After remediation the same paths run again. The improvement is a measured change in the score, reported alongside it rather than folded in.
From your threat model to a signed finding
The bespoke variant of OSES™. The first two stages are what make no two engagements alike. Your side of it is a scoping call, an authorization letter and a readout.
Threat model intake
Which process would hurt most if it failed, who has authority over it, and what leadership needs to see. That sets the objective.
Build the scenario
OSINT establishes who to impersonate and how to reach your people. Pretext, sequence and synthetic assets are built for this one engagement.
Run it unannounced
Across whichever channels the scenario needs, in a real working window, coordinated under one objective and fully external.
Follow it downstream
Into the help desk approval, the callback, the hiring decision or the payment. Downstream observation is authorized in scope, so Depth is observed, not assumed.
Score and benchmark
OSES™ Score with Depth and Spread, peer vertical position, and process failure findings written for a board and an auditor.
Train, then re-test
Training built from what failed, then the same paths run again. A course completion is not a control. A measured delta is.
Every engagement runs under signed authorization, with named approvers, agreed impersonated identities and documented abort conditions. Client names never appear in our public material. For the deepfake methodology in depth, see deepfake red teaming. To run simulations yourself, see the OSES™ Platform.
Real engagements, real outcomes
Published with numbers where we have permission, and as outcomes only where we do not. Client names never appear.
A cloned executive, driven by an agent
A cloned voice of a local executive, driven by an agent that called each user and held a live conversation, followed by an SMS. More than one in five rang the number back and talked to a machine unprompted.
Deepfake voicemail drop, then SMS
Public video supplied enough material to clone the CEO's voice. A personalized voicemail landed, an SMS followed, and the callback line was answered by the clone itself.
The candidate who did not exist
A synthetic applicant with a constructed history and a live face on the interview cleared hiring screening. HR and recruiting remain one of the easiest paths into an organization we test.
Kudos to your entire team. We haven't even seen the report and the whole company is talking about the risks of voice cloning. It's been a huge win for us already.
CISO, Large Financial Enterprise
Published, because we have nothing to hide
The price sets the scope. Your threat model sets the scenario. Every tier is unannounced, fully managed and externally originated. See all pricing
OSES™ Assessment
One vector, such as IT support impersonation. Up to 250 users. OSES™ Score and peer vertical rank.
Multi-Vector Engagement
Multiple vectors, orchestrated end to end. Up to 1,000 users.
Full Scope
Executive impersonation, hiring pipeline, fake candidates and custom kill chains, designed per engagement.
Executive voice and likeness are cloned only with the consent of the person being cloned, confirmed by a biometric check and voice verification before anything is generated. An administrator cannot authorize cloning someone else.
What CISOs ask before the first engagement
What is an AI social engineering red team engagement?
An authorized, unannounced assessment in which AI-generated voice, video and identity are used to impersonate people your organization trusts, to find out whether your people and processes can be moved to a consequential action. Unlike a phishing test, which ends at the click, the engagement follows the scenario into the callback, the help desk reset, the hiring decision or the payment approval.
How is the OSES™ Score calculated?
The OSES™ Score is two terms added together. Depth is how far the deepest single case progressed, and it is not divided by headcount. Spread is how much of the population moved, weighted by how far each person went, and capped so volume cannot mask a contained breach. The total lands on one of four bands: Low, Medium, High or Critical. The band is never reported without the two terms behind it.
How does the peer benchmark work?
Your result is reported next to the median for your own vertical, so you can tell an outlier from a sector median that is itself unacceptable. Vertical medians are computed on Spread only, so one incident at one company cannot set a sector floor. Rollups are total weighted events over total targets, and where a group does not yet meet our minimum sample thresholds we report it as insufficient rather than publish a number.
Is this a fixed package?
No. The published price sets the scope: how many vectors and how many users. The scenario itself is designed from your threat model, so the adversary, the pretext, the channels, the identities impersonated and the objective are decided for your organization specifically. Two clients in the same sector get materially different engagements.
Is it safe to run against real employees?
Yes. Every engagement runs under signed authorization with named approvers and documented abort conditions. Executive voice and likeness are cloned only with the consent of the person being cloned, confirmed by a biometric check and voice verification before anything is generated. Findings are reported at the organization level, with no named individuals and no department leaderboards.
Do we need to install anything?
No. Engagements are fully managed and externally originated. Nothing is installed and nothing touches your security stack, which removes most of the change-control burden that blocks testing in regulated environments.
How does the red team credit work?
50% of any red team fee is credited toward your first year of the OSES™ Platform, capped at half the first-year fee. Run an engagement to establish where you stand, then apply the credit when your team takes over running simulations itself.
Find out how far a synthetic identity gets inside
Thirty minutes on your threat model: which process would hurt most if it failed, who has authority over it, and what the report needs to prove.
MSPs and MSSPs: see the partner program.
