Abstract:Image Retrieval with Text Feedback (IRTF) is an emerging research topic where the query consists of an image and a text expressing a requested attribute modification. The goal is to retrieve the target images similar to the query text modified query image. The existing methods usually adopt feature fusion of the query image and text to match the target image. However, they ignore two crucial issues: overfitting and low diversity of training data, which make the feature fusion based IRTF task not generalizable. Conventional generation based data augmentation is an effective way to alleviate overfitting and improve diversity, but increases the volume of training data and generation model parameters, which is bound to bring huge computation costs. By rethinking the conventional data augmentation mechanism, we propose a plug-and-play Gradient Augmentation (GA) based regularization approach. Specifically, GA contains two items: 1) To alleviate model overfitting on the training set, we deduce an explicit adversarial gradient augmentation from the perspective of adversarial training, which challenges the “no free lunch” philosophy. 2) To improve the diversity of training set, we propose an implicit isotropic gradient augmentation from the perspective of gradient descent-based optimization, which achieves the goal of big gain but no pain. Besides, we introduce deep metric learning to train the model and provide theoretical insights of GA on generalisation. Finally, we propose a new evaluation protocol called Weighted Harmonic Mean (WHM) to assess the model generalisation. Experiments show that our GA outperforms the state-of-the-art methods by 6.2% and 4.7% on CSS and Fashion200k datasets, respectively, without bells and whistles.

Adversarial and Isotropic Gradient Augmentation for Image Retrieval with Text Feedback

TTIDA: Controllable Generative Data Augmentation via Text-to-Text and Text-to-Image Models

SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation

DIAGen: Diverse Image Augmentation with Generative Models

Gradient aggregation based fine-grained image retrieval : A unified viewpoint for CNN and Transformer

Feature Augmentation for Adversarial Robustness

Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators

DifAugGAN: A Practical Diffusion-style Data Augmentation for GAN-based Single Image Super-resolution

AGFSync: Leveraging AI-Generated Feedback for Preference Optimization in Text-to-Image Generation

A Good Data Augmentation Policy Is Not All You Need: A Multi-Task Learning Perspective

Learning to Augment: Hallucinating Data for Domain Generalized Segmentation

Data Augmentation for Image Classification using Generative AI

Semantic-aware Data Augmentation for Text-to-image Synthesis

Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image Classification

Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition

Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce Classification

Image Augmentations for GAN Training

FAGStyle: Feature Augmentation on Geodesic Surface for Zero-shot Text-guided Diffusion Image Style Transfer

A Simple Background Augmentation Method for Object Detection with Diffusion Model

Image retrieval outperforms diffusion models on data augmentation

DAug: Diffusion-based Channel Augmentation for Radiology Image Retrieval and Classification