Guided and Fused: Efficient Frozen CLIP-ViT with Feature Guidance and Multi-Stage Feature Fusion for Generalizable Deepfake Detection

Yingjian Chen,Lei Zhang,Yakun Niu,Pei Chen,Lei Tan,Jing Zhou

2024-08-25

Abstract:The rise of generative models has sparked concerns about image authenticity online, highlighting the urgent need for an effective and general detector. Recent methods leveraging the frozen pre-trained CLIP-ViT model have made great progress in deepfake detection. However, these models often rely on visual-general features directly extracted by the frozen network, which contain excessive information irrelevant to the task, resulting in limited detection performance. To address this limitation, in this paper, we propose an efficient Guided and Fused Frozen CLIP-ViT (GFF), which integrates two simple yet effective modules. The Deepfake-Specific Feature Guidance Module (DFGM) guides the frozen pre-trained model in extracting features specifically for deepfake detection, reducing irrelevant information while preserving its generalization capabilities. The Multi-Stage Fusion Module (FuseFormer) captures low-level and high-level information by fusing features extracted from each stage of the ViT. This dual-module approach significantly improves deepfake detection by fully leveraging CLIP-ViT's inherent advantages. Extensive experiments demonstrate the effectiveness and generalization ability of GFF, which achieves state-of-the-art performance with optimal results in only 5 training epochs. Even when trained on only 4 classes of ProGAN, GFF achieves nearly 99% accuracy on unseen GANs and maintains an impressive 97% accuracy on unseen diffusion models.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The paper aims to address the generalization problem in deepfake image detection. Specifically, existing methods based on the frozen pre-trained CLIP-ViT model include too much task-irrelevant information when extracting visual general features, leading to limited detection performance. To solve this problem, the paper proposes an efficient method—Guided Fusion of Frozen CLIP-ViT (GFF), which includes two simple but effective modules: 1. **Deepfake-Specific Feature Guidance Module (DFGM)**: This module guides the frozen pre-trained model to extract features specifically for deepfake detection, reducing irrelevant information while retaining its generalization ability. 2. **Multi-Stage Fusion Module (FuseFormer)**: This module captures both low-level and high-level information by fusing features extracted at various stages of ViT, fully leveraging the advantages of CLIP-ViT. Through these two modules, GFF significantly improves the effectiveness of deepfake image detection and demonstrates excellent performance in experiments, achieving state-of-the-art performance with only 5 training epochs. Even when trained with only 4 types of data from ProGAN, GFF can achieve nearly 99% accuracy on unseen GAN and diffusion models.

Guided and Fused: Efficient Frozen CLIP-ViT with Feature Guidance and Multi-Stage Feature Fusion for Generalizable Deepfake Detection

Towards More General Video-based Deepfake Detection through Facial Feature Guided Adaptation for Foundation Model

DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion

CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection

C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake Detection

MFCLIP: Multi-modal Fine-grained CLIP for Generalizable Diffusion Face Forgery Detection

UCF: Uncovering Common Features for Generalizable Deepfake Detection

FFR_FD: Effective and fast detection of DeepFakes via feature point defects

Multiclass AI-Generated Deepfake Face Detection Using Patch-Wise Deep Learning Model

Multi-Modal Generative DeepFake Detection via Visual-Language Pretraining with Gate Fusion for Cognitive Computation

GM-DF: Generalized Multi-Scenario Deepfake Detection

FFR_FD: Effective and Fast Detection of DeepFakes Based on Feature Point Defects

Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning

Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer

Learning spatial‐frequency interaction for generalizable deepfake detection

DeepFake detection algorithm based on improved vision transformer

Fusing Global and Local Features for Generalized AI-Synthesized Image Detection

Combining EfficientNet and Vision Transformers for Video Deepfake Detection

Deepfake Detection Scheme Based on Vision Transformer and Distillation

Fake It till You Make It: Curricular Dynamic Forgery Augmentations towards General Deepfake Detection

Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection