Abstract:Attribute-specific fashion retrieval (ASFR) is a challenging information retrieval task, which has attracted increasing attention in recent years. Different from traditional fashion retrieval which mainly focuses on optimizing holistic similarity, the ASFR task concentrates on attribute-specific similarity, resulting in more fine-grained and interpretable retrieval results. As the attribute-specific similarity typically corresponds to the specific subtle regions of images, we propose a Region-to-Patch Framework (RPF) that consists of a region-aware branch and a patch-aware branch to extract fine-grained attribute-related visual features for precise retrieval in a coarse-to-fine manner. In particular, the region-aware branch is first to be utilized to locate the potential regions related to the semantic of the given attribute. Then, considering that the located region is coarse and still contains the background visual contents, the patch-aware branch is proposed to capture patch-wise attribute-related details from the previous amplified region. Such a hybrid architecture strikes a proper balance between region localization and feature extraction. Besides, different from previous works that solely focus on discriminating the attribute-relevant foreground visual features, we argue that the attribute-irrelevant background features are also crucial for distinguishing the detailed visual contexts in a contrastive manner. Therefore, a novel E-InfoNCE loss based on the foreground and background representations is further proposed to improve the discrimination of attribute-specific representation. Extensive experiments on three datasets demonstrate the effectiveness of our proposed framework, and also show a decent generalization of our RPF on out-of-domain fashion images. Our source code is available at <a class="link-external link-https" href="https://github.com/HuiGuanLab/RPF" rel="external noopener nofollow">this https URL</a>.

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training

FashionViL: Fashion-Focused Vision-and-Language Representation Learning

FashionSAP: Symbols and Attributes Prompt for Fine-grained Fashion Vision-Language Pre-training

FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning

FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks

A Fine-Grained Vision and Language Representation Framework with Graph-Based Fashion Semantic Knowledge

Describe Fashion Products via Local Sparse Self-Attention Mechanism and Attribute-based Re-sampling Strategy

A Deep-Learning-Based Fashion Attributes Detection Model

Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards

Learning Structured Relation Embeddings for Fine-Grained Fashion Attribute Recognition

Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion Analysis

Coarse-to-Fine Attribute Editing for Fashion Images.

From Region to Patch: Attribute-Aware Foreground-Background Contrastive Learning for Fine-Grained Fashion Retrieval

BigFashion: A Large-Scale Dataset for Fine-Grained Attributes Recognition

Fine-Grained Fashion Similarity Prediction By Attribute-Specific Embedding Learning

FashionERN: Enhance-and-Refine Network for Composed Fashion Image Retrieval

Real-Time Fashion-Guided Clothing Semantic Parsing: A Lightweight Multi-Scale Inception Neural Network and Benchmark.

Fashionformer: A Simple, Effective and Unified Baseline for Human Fashion Segmentation and Recognition.

FashionKLIP: Enhancing E-Commerce Image-Text Retrieval with Fashion Multi-Modal Conceptual Knowledge Graph

Masked Vision-Language Transformer in Fashion

FashionAI: A Hierarchical Dataset for Fashion Understanding