Abstract:With the rapid progress of generative models, the current challenge in face forgery detection is how to effectively detect realistic manipulated faces from different unseen domains. Though previous studies show that pre-trained Vision Transformer (ViT) based models can achieve some promising results after fully fine-tuning on the Deepfake dataset, their generalization performances are still unsatisfactory. One possible reason is that fully fine-tuned ViT-based models may disrupt the pre-trained features [1, 2] and overfit to some data-specific patterns [3]. To alleviate this issue, we present a \textbf{F}orgery-aware \textbf{A}daptive \textbf{Vi}sion \textbf{T}ransformer (FA-ViT) under the adaptive learning paradigm, where the parameters in the pre-trained ViT are kept fixed while the designed adaptive modules are optimized to capture forgery features. Specifically, a global adaptive module is designed to model long-range interactions among input tokens, which takes advantage of self-attention mechanism to mine global forgery clues. To further explore essential local forgery clues, a local adaptive module is proposed to expose local inconsistencies by enhancing the local contextual association. In addition, we introduce a fine-grained adaptive learning module that emphasizes the common compact representation of genuine faces through relationship learning in fine-grained pairs, driving these proposed adaptive modules to be aware of fine-grained forgery-aware information. Extensive experiments demonstrate that our FA-ViT achieves state-of-the-arts results in the cross-dataset evaluation, and enhances the robustness against unseen perturbations. Particularly, FA-ViT achieves 93.83\% and 78.32\% AUC scores on Celeb-DF and DFDC datasets in the cross-dataset evaluation. The code and trained model have been released at: <a class="link-external link-https" href="https://github.com/LoveSiameseCat/FAViT" rel="external noopener nofollow">this https URL</a>.

Enhancing General Face Forgery Detection via Vision Transformer with Low-Rank Adaptation

Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer

Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method with Forgery Style Mixture

MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery Detection

Adaptive Transformers for Robust Few-shot Cross-domain Face Anti-spoofing

A Detail-Aware Transformer to Generalisable Face Forgery Detection

Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection

Unified Video and Image Representation for Boosted Video Face Forgery Detection

FakeFormer: Efficient Vulnerability-Driven Transformers for Generalisable Deepfake Detection

Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection

Latent Spatiotemporal Adaptation for Generalized Face Forgery Video Detection

FakeTransformer: Exposing Face Forgery From Spatial-Temporal Representation Modeled By Facial Pixel Variations

SA$$^3$$WT: Adaptive Wavelet-Based Transformer with Self-Paced Auto Augmentation for Face Forgery Detection

Video Forgery Detection Using Spatio-Temporal Dual Transformer.

Improving the generalization of face forgery detection via single domain augmentation

Face Forgery Detection Algorithm Based on Improved MobileViT Network

Face Forgery Detection with Long-Range Noise Features and Multilevel Frequency-Aware Clues

Liveness Detection in Computer Vision: Transformer-based Self-Supervised Learning for Face Anti-Spoofing

Learning Forgery Region-Aware and ID-Independent Features for Face Manipulation Detection

Temporal Consistency Based Deep Face Forgery Detection Network.

UIA-ViT: Unsupervised Inconsistency-Aware Method Based on Vision Transformer for Face Forgery Detection.