The face and voice on the screen seem completely authentic.
A video call comes in, and the screen shows a familiar boss or family member, urgently stating that they are in an emergency and need you to transfer some money immediately. They can even interact with you normally and respond to a few questions you casually ask. This situation has become increasingly prevalent in various fraud cases in recent years, driven by real-time deepfake technology—combining real-time face swapping and voice modulation techniques to fabricate a seemingly genuine video call.
What real-time deepfake technology is actually capable of.
The core principle behind deepfake technology lies in collecting a substantial number of facial images and voice samples of the target individual, training a system that can instantly convert the attacker’s facial movements and voice into the appearance and voice of the target. Recent advancements have significantly enhanced the real-time processing capabilities of such technologies, enabling the generated visuals to keep pace with the interactive nature of video calls, unlike in the past when they could only be applied to pre-recorded footage. However, real-time deepfake still has technical limitations, especially under conditions of unstable internet signals, complex lighting, or when unexpected actions are required from the target, making flaws easier to spot. This indicates that although the technology does possess a degree of deception capability, it is not flawless. Understanding these limitations is a crucial entry point for discerning authenticity.
A few verification methods to try during the call.
1. Ask the other person to perform an unexpected action that hasn't been premeditated, such as suddenly turning to show a side profile or quickly waving their hand in front of their face. Deepfake technology struggles with real-time changes involving unanticipated movements, often resulting in distorted visuals, misaligned features, or synchronization delays. 2. Inquire about a specific detail known only to you, which doesn’t need to be confidential. It can be something trivial known only between the two of you, such as what you ate during your last meeting or the recent status of a mutual friend. Deepfake technology can imitate appearances and voices but cannot conjure knowledge of these types of details that only exist in actual memories. 3. Hang up the call and reconnect through a separate, independent channel, such as directly dialing a number you know the person well by, rather than returning the call to the number that just displayed. If the person on the other end picks up and is completely unaware of the previous video call, you can confirm that the latter was fabricated.
If you suspect that you or a family member has fallen victim to such deepfake video scams, or if a transfer has already been completed, quickly retain screenshots of the call and transfer records, and report to your bank and the police as soon as possible. VexelOps can also assist in clarifying the overall situation.
Frequently Asked Questions About Deepfake Video Scams
How can scammers acquire enough facial and voice data to train this system?
In most victim cases, the materials attackers use often come from the target's own publicly shared content on social media platforms, such as video posts, live recordings, public speeches, or interview clips. As long as there is sufficient duration and variety in angles in the publicly available material, it can potentially train a deepfake model. This is why there have been security advisories in recent years urging public figures or high-level executives to consider the risks of sharing extensive, clear, multi-angle video content on social media. This doesn’t discourage sharing entirely but reminds individuals to evaluate the scope and frequency of their sharing.
Is it possible for voice phone calls to also be deepfaked, not just video calls?
Yes, and voice deepfakes typically have a lower technical barrier than real-time video face-swapping since voice samples are relatively easier to obtain, and the computational resources required for training are also lower. Recently, many fraud cases have been conducted through pure voice calls, impersonating family members or bosses in urgent requests, with voices that sound almost identical to the real person. The same verification logic mentioned earlier applies to voice calls: hanging up and reconfirming through the originally known contact method is the most reliable approach, with no need to distinguish between video and pure voice calls as the same verification logic remains applicable.
What kind of response mechanisms should companies establish to reduce the risk of employees being scammed?
From a corporate perspective, establishing a clear verification process for fund transfer requests, particularly for urgent or large sums, is advisable. It should be mandated that regardless of the communication method used—including video, phone, or instant messaging—confirmations must be conducted through another independent and pre-agreed channel before executing the related operations. Additionally, regular training of employees on awareness regarding deepfake scams is essential, enabling the team to understand the existence and operation of such technologies, and thus effectively reducing overall risk more than simply relying on personal vigilance.
One Key Takeaway: Real-time deepfake technology can replicate appearance and voice, but it cannot cope with sudden unexpected actions, nor does it know details that exist only in real memories. When faced with emergency transfer requests, hanging up and re-confirming through familiar channels is the most reliable way to identify such fraud.