An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

Ganglai Wang,Peng Zhang,Lei Xie,Wei Huang,Yufei Zha,Yanning Zhang

DOI: https://doi.org/10.48550/arXiv.2203.05178

2022-03-10

Abstract:DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake video detection is further improved. By only changing lip shape to match the given speech, the facial features of identity is hard to be discriminated in such fake talking face videos. Together with the lack of attention on audio stream as the prior knowledge, the detection failure of fake talking face generation also becomes inevitable. Inspired by the decision-making mechanism of human multisensory perception system, which enables the auditory information to enhance post-sensory visual evidence for informed decisions output, in this study, a fake talking face detection framework FTFDNet is proposed by incorporating audio and visual representation to achieve more accurate fake talking face videos detection. Furthermore, an audio-visual attention mechanism (AVAM) is proposed to discover more informative features, which can be seamlessly integrated into any audio-visual CNN architectures by modularization. With the additional AVAM, the proposed FTFDNet is able to achieve a better detection performance on the established dataset (FTFDD). The evaluation of the proposed work has shown an excellent performance on the detection of fake talking face videos, which is able to arrive at a detection rate above 97%.

Computer Vision and Pattern Recognition,Sound,Audio and Speech Processing

What problem does this paper attempt to address?

The problem that this paper attempts to solve is the public media security threats brought by DeepFake technology when generating fake talking - face videos, especially when lip manipulation is used to generate talking faces, the detection difficulty of such fake videos is further increased. Since only the lip shape is changed to match the given voice without changing the identity features, the identity features of such fake talking - face videos are difficult to distinguish. In addition, the existing methods pay insufficient attention to the audio stream as prior knowledge, resulting in the inevitable problem of failure in detecting fake talking - face videos. For this reason, the paper proposes a fake - talking - face detection framework FTFDNet that combines audio and visual representations, as well as a novel audio - video attention mechanism (AVAM), aiming to improve the detection performance, especially achieving a detection rate of over 97% on the established dataset (FTFDD).

An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

Refining Localized Attention Features with Multi-Scale Relationships for Enhanced Deepfake Detection in Spatial-Frequency Domain

A defensive attention mechanism to detect deepfake content across multiple modalities

Multimodal Deepfake Detection for Short Videos

AVoiD-DF: Audio-Visual Joint Learning for Detecting Deepfake

FST-Net - Exploiting Frequency Spatial Temporal Information for Low-Quality Fake Video Detection.

A Unified Framework for Modality-Agnostic Deepfakes Detection

AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Video Deepfake Detection

Multi-attentional Deepfake Detection

AVForensics: Audio-driven Deepfake Video Detection with Masking Strategy in Self-supervision.

Video Detection Method Based on Temporal and Spatial Foundations for Accurate Verification of Authenticity

DeepFake detection method based on multi-scale interactive dual-stream network

Exploiting Complementary Dynamic Incoherence for DeepFake Video Detection

Audio-Visual Temporal Forgery Detection Using Embedding-Level Fusion and Multi-Dimensional Contrastive Loss

FakeTransformer: Exposing Face Forgery From Spatial-Temporal Representation Modeled By Facial Pixel Variations

A Robust Approach to Multimodal Deepfake Detection

Multi-feature fusion based face forgery detection with local and global characteristics

Audio-Visual Contrastive Pre-train for Face Forgery Detection

Interactive Two-Stream Network Across Modalities for Deepfake Detection

Lightweight detection method for deepfake face video