A Stable Vision Transformer for Out-of-Distribution Generalization.

Haoran Yu,Baodi Liu,Yingjie Wang,Kai Zhang,Dapeng Tao,Weifeng Liu
DOI: https://doi.org/10.1007/978-981-99-8543-2_27
2024-01-01
Abstract:Vision Transformer (ViT) has achieved amazing results in many visual applications where training and testing instances are drawn from the independent and identical distribution (I.I.D.). The performance will drop drastically when the distribution of testing instances is different from that of training ones in real open environments. To tackle this challenge, we propose a Stable Vision Transformer (SViT) for out-of-distribution (OOD) generalization. In particular, the SViT weights the samples to eliminate spurious correlations of token features in Vision Transformer and finally boosts the performance for OOD generalization. According to the structure and feature extraction characteristics of the ViT models, we design two forms of learning sample weights: SViT(C) and SViT(T). To demonstrate the effectiveness of two forms of SViT for OOD generalization, we conduct extensive experiments on the popular PACS and OfficeHome datasets and compare them with SOTA methods. The experimental results demonstrate the effectiveness of SViT(C) and SViT(T) for various OOD generalization tasks.
What problem does this paper attempt to address?