An End-to-End Framework for Clothing Collocation Based on Semantic Feature Fusion

Mingbo Zhao,Yu Liu,Xianrui Li,Zhao Zhang,Yue Zhang
DOI: https://doi.org/10.1109/mmul.2020.3024221
IF: 3.4911
2020-10-01
IEEE Multimedia
Abstract:In this article, we develop an end-to-end clothing collocation learning framework based on a bidirectional long short-term memories (Bi-LTSM) model, and propose new feature extraction and fusion modules. The feature extraction module uses Inception V3 to extract low-level feature information and the segmentation branches of Mask Region Convolutional Neural Network (RCNN) to extract high-level semantic information; whereas the feature fusion module creates a new reference vector for each image to fuse the two types of image feature information. As a result, the feature can involve both low-level image and high-level semantic feature information, so that the performance of Bi-LSTM can be enhanced. Extensive simulations are conducted based on Ployvore and DeepFashion2 datasets. Simulation results verify the effectiveness of the proposed method compared with other state-of-the-art clothing collocation methods.
computer science, information systems, theory & methods, software engineering, hardware & architecture
What problem does this paper attempt to address?