Abstract:Neural ranking models (NRMs) and dense retrieval (DR) models have given rise to substantial improvements in overall retrieval performance. In addition to their effectiveness, and motivated by the proven lack of robustness of deep learning-based approaches in other areas, there is growing interest in the robustness of deep learning-based approaches to the core retrieval problem. Adversarial attack methods that have so far been developed mainly focus on attacking NRMs, with very little attention being paid to the robustness of DR models. In this paper, we introduce the adversarial retrieval attack (AREA) task. The AREA task is meant to trick DR models into retrieving a target document that is outside the initial set of candidate documents retrieved by the DR model in response to a query. We consider the decision-based black-box adversarial setting, which is realistic in real-world search engines. To address the AREA task, we first employ existing adversarial attack methods designed for NRMs. We find that the promising results that have previously been reported on attacking NRMs, do not generalize to DR models: these methods underperform a simple term spamming method. We attribute the observed lack of generalizability to the interaction-focused architecture of NRMs, which emphasizes fine-grained relevance matching. DR models follow a different representation-focused architecture that prioritizes coarse-grained representations. We propose to formalize attacks on DR models as a contrastive learning problem in a multi-view representation space. The core idea is to encourage the consistency between each view representation of the target document and its corresponding viewer via view-wise supervision signals. Experimental results demonstrate that the proposed method can significantly outperform existing attack strategies in misleading the DR model with small indiscernible text perturbations.

Optimizing Dense Retrieval Model Training with Hard Negatives.

Learning To Retrieve: How to Train a Dense Retrieval Model Effectively and Efficiently

Hard Negatives or False Negatives: Correcting Pooling Bias in Training Neural Ranking Models

Enhancing Retrieval Performance: An Ensemble Approach For Hard Negative Mining

Enhancing the Ranking Context of Dense Retrieval Methods through Reciprocal Nearest Neighbors

Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval

Mitigating the Impact of False Negatives in Dense Retrieval with Contrastive Confidence Regularization

YES SIR!Optimizing Semantic Space of Negatives with Self-Involvement Ranker

Disentangled Modeling of Domain and Relevance for Adaptable Dense Retrieval

Constructing Tree-based Index for Efficient and Effective Dense Retrieval

Directly optimizing evaluation measures in learning to rank.

On the Theories Behind Hard Negative Sampling for Recommendation

Back to Basics: A Simple Recipe for Improving Out-of-Domain Retrieval in Dense Encoders

Black-box Adversarial Attacks against Dense Retrieval Models: A Multi-view Contrastive Learning Method

Directly Optimize Diversity Evaluation Measures: A New Approach to Search Result Diversification.

A Simple yet Effective Framework for Active Learning to Rank

Understanding the Ranking Loss for Recommendation with Sparse User Feedback

Combining Multiple Supervision for Robust Zero-Shot Dense Retrieval

Personalized Ranking with Importance Sampling.

TriSampler: A Better Negative Sampling Principle for Dense Retrieval

A Thorough Examination on Zero-shot Dense Retrieval