Deepfake Attacks in Teams and Google Meet
Neither platform checks that a camera feed came from a camera.
Your finance team verifies identity the same way everyone does: they look at the face on the call and they recognise the voice. Both are now trivially reproducible, and both arrive through the exact same pipe as the real thing. Below is how the attack runs inside Microsoft Teams and Google Meet, vector by vector. Then you can build one of yourself and watch it happen.
Record a short sample. We generate your synthetic double and join it to a sandboxed meeting you control, so you see exactly what your team would see. Your likeness only, deleted after the session.
Built for banks, payments, insurance, and enterprise finance teams.
How it runs, platform by platform
The two platforms differ in their default settings and their weak points. Attackers pick accordingly.
What the call actually looks like
The technology is the least interesting part. The sequence around it is what makes it work.
Collection
Conference talks, earnings calls, podcast appearances, and webinar recordings supply the face and the voice. An executive who has ever presented publicly has already published the training data. Org charts on LinkedIn supply the approval hierarchy.
The pretext lands in a trusted channel
A Teams message from a federated account, or a calendar invitation for a short confidential call. Urgency plus confidentiality is the standard pairing, because confidentiality is what suppresses the one behaviour that defeats the attack: asking someone else.
Familiar faces, plural
The strongest version of the attack does not put one synthetic executive on the call. It puts several, so the target is outnumbered by people who appear to agree with each other. Social proof does the work that the video quality does not have to.
The camera conveniently fails
If video is degrading, the synthetic participant apologises for a bad connection and switches to audio only. This is not a failure of the attack. It is a planned downgrade to the channel where the clone is strongest and the target has already accepted who they are talking to.
An instruction that fits the process
Rarely an implausible sum. Usually a payment that resembles a normal one, to an account that changed for a plausible reason, approved inside the call so the usual out-of-band step is skipped because a senior person is on the line asking for it.
Six vectors, run in your own tenant
Run the full set, or scope down to the one your last tabletop exercise could not answer.
A real-time rendered face delivered through a virtual camera device into Teams or Meet. It blinks, turns, and responds in conversation. Tests whether anyone on the call questions a feed that behaves correctly but is not a person.
A guest or federated participant presenting as a named executive, with matching avatar and a domain that survives a glance. Tests whether external participant warnings are read or dismissed.
The planned downgrade. Camera fails, the call continues on a cloned voiceprint alone, and the identity decision was already made when the video was still up. Also runs standalone against the call centre and the conference bridge.
Several synthetic attendees on one call, agreeing with each other. Isolates how much of the target's judgement was independent and how much was deference to an apparent consensus of seniors.
The instruction is carried through to the point of approval, then stopped. Establishes exactly how far a request travels before a second channel is required, and whether that requirement is enforced or waived under seniority.
What happened after. Who raised it, how quickly, through which channel, and whether the report reached anyone who could act. A caught attack that is never reported leaves the campaign running.
Build your own, in four steps
Free, and it uses your likeness only. There is no version of this that points at anyone else.
Pick a channel
Microsoft Teams, Google Meet, or voice only. The Teams and Meet options produce a synthetic video participant. The voice option produces a cloned voiceprint you can play back against your own ear.
Record a short sample of yourself
Around thirty seconds of your face and voice, captured in the browser. No upload of anyone else's media is accepted, and the sample is bound to the session that created it.
We generate and join the meeting
Your synthetic double is built and joined to a sandboxed meeting we host on the platform you chose. You join as yourself and sit across from it. It is not connected to your tenant and it cannot be pointed at one.
Watch yourself say something you never said
You type a line, your double delivers it. Most people stop arguing about whether this is a real problem somewhere in the first fifteen seconds. Your sample and the generated media are deleted when the session ends.
How a live simulation runs
The video conferencing variant of OSES™, our orchestrated social engineering simulation framework. Same discipline on scoping, authorization, and evidence as every other OSES™ engagement, run in the channel where your people actually make trust decisions.
Every engagement runs under signed authorization inside your own tenant, with an agreed participant list, named approvers, and documented abort conditions. No transaction is ever executed. Reporting is organizational, with no named individuals. For engagements targeting verification systems rather than people, see deepfake penetration testing.
Where a call moves money
Organizations where a decision made on a video call has a payment, a credential, or a client account behind it.
What makes this different
Live, not pre-recorded
The synthetic participant holds a conversation. It answers questions, reacts to interruption, and adapts. A pre-rendered clip tests nothing, because nobody in a real attack is watching a video that cannot respond.
Inside your real platform
The simulation runs in your own Teams or Meet environment, with your settings, your external participant labelling, and your approval workflow. Not a lab, and not a slide about what could happen.
Behaviour, then training
The recording of your own team on the call becomes the training material in the same platform. Measure, train on exactly what you found, then run it again and show the change.
Real threats, real simulations, real findings
A cloned voiceprint submitted against verbal verification controls, and what it revealed about the step-up path behind them.
Read the case study Agentic AIAutonomous agents driving synthetic media generation and delivery end to end, with no human operator in the loop.
Read the case studyCommon questions
Can a deepfake actually join a Microsoft Teams or Google Meet call?
Yes. Neither platform verifies that a camera feed originated from a physical camera. Both accept any device registered with the operating system as a webcam, which means a rendered face can be presented as a normal video feed. The same applies to the microphone input, so a cloned voice arrives through the same path as a real one.
Does Microsoft Teams or Google Meet detect deepfakes?
Neither platform ships a synthetic media detection control as standard. Their protections are about access and identity at the door: who is allowed to join, whether the participant is external, and whether the meeting is locked. Once a participant is in the meeting, the video and audio they send is treated as authentic by default.
What does the free demo actually do?
It builds a synthetic version of you, from a short sample you record of your own face or voice, and joins it to a sandboxed meeting we host. You watch your own likeness say something you never said. It only ever uses your own likeness, and it never joins a meeting in your organization or anyone else's.
Is this the same as deepfake penetration testing?
No. Deepfake penetration testing targets systems: liveness detection, document verification, and voice biometrics. This targets the human and process layer inside a meeting, where the control is a person deciding whether to trust a face. Many organizations run both, because a failure in either one produces the same outcome.
How is a live simulation authorized and scoped?
Every engagement runs under signed authorization inside your own tenant, with an agreed participant list, meeting windows, named approvers, and documented abort conditions. Nothing is generated before scope is signed, and no real payment or transaction is ever executed.
Do you report on individual employees?
No. Reporting is organizational. We report on verification behaviour, escalation paths, and where the approval process held or gave way. There are no named individuals and no department leaderboards.
What controls actually stop this?
Out-of-band verification on a known channel, a payment approval process that cannot be completed inside a single call, external participant labelling that is enforced rather than advisory, and a culture where asking a senior person to verify carries no cost. The report tells you which of these you already have and which failed under pressure.
How long does an engagement take?
Two to three weeks from scoping call to final report, with the live meeting window itself usually a single day inside that.
Find out what your team would do on that call
Thirty minutes. We will walk through your meeting settings and approval workflow and identify which scenario is worth running first.
Or read the full assessment methodology.
