Abstract:By highlighting important features that contribute to model prediction, visual saliency is used as a natural form to interpret the working mechanism of deep neural networks. Numerous methods have been proposed to achieve better saliency results. However, we find that previous visual saliency methods are not reliable enough to provide meaningful interpretation through a simple sanity check: saliency methods are required to explain the output of non-maximum prediction classes, which are usually not ground-truth classes. For example, let the methods interpret an image of "dog" given a wrong class label "fish" as the query. This procedure can test whether these methods reliably interpret model's predictions based on existing features that appear in the data. Our experiments show that previous methods failed to pass the test by generating similar saliency maps or scattered patterns. This false saliency response can be dangerous in certain scenarios, such as medical diagnosis. We find that these failure cases are mainly due to the attribution vanishing and adversarial noise within these methods. In order to learn reliable visual saliency, we propose a simple method that requires the output of the model to be close to the original output while learning an explanatory saliency mask. To enhance the smoothness of the optimized saliency masks, we then propose a simple Hierarchical Attribution Fusion (HAF) technique. In order to fully evaluate the reliability of visual saliency methods, we propose a new task Disturbed Weakly Supervised Object Localization (D-WSOL) to measure whether these methods can correctly attribute the model's output to existing features. Experiments show that previous methods fail to meet this standard, and our approach helps to improve the reliability by suppressing false saliency responses. After observing a significant layout difference in saliency masks between real and adversarial samples. we propose to train a simple CNN on these learned hierarchical attribution masks to distinguish adversarial samples. Experiments show that our method can improve detection performance over other approaches significantly.

Saliency Prediction in the Deep Learning Era: Successes and Limitations

What Do Deep Saliency Models Learn about Visual Attention?

Revisiting Video Saliency Prediction in the Deep Learning Era

Saliency detection in deep learning era: trends of development

Weakly Supervised Visual Saliency Prediction

Revisiting Video Saliency: A Large-scale Benchmark and a New Model

Deep saliency models learn low-, mid-, and high-level features to predict scene attention

A Deep Model of Visual Attention for Saliency Detection on 3D Objects

Deep Learning for Video Saliency Detection.

SAL3D: a model for saliency prediction in 3D meshes

Training Better Deep Learning Models Using Human Saliency

Saliency Detection in Educational Videos: Analyzing the Performance of Current Models, Identifying Limitations and Advancement Directions

Data Augmentation via Latent Diffusion for Saliency Prediction

Bridging the Gap Between Saliency Prediction and Image Quality Assessment

Unified Image and Video Saliency Modeling

SID4VAM: A Benchmark Dataset with Synthetic Images for Visual Attention Modeling

Saliency Revisited: Analysis of Mouse Movements versus Fixations

Review of Visual Saliency Detection with Comprehensive Information

Learning Reliable Visual Saliency for Model Explanations

A Learning-Based Visual Saliency Prediction Model for Stereoscopic 3D Video (LBVS-3D)

Modern Learning Methodologies for Co-Saliency Detection