Object Localization Based on Natural Language Descriptions for Fine-Grained Image

Lijuan Duan,Mingliang Liang,Qing En,Yuanhua Qiao,Jun Miao,Longlong Ma
DOI: https://doi.org/10.1117/12.2579516
2020-01-01
Abstract:As a tool to express common semantics of objects, language can be used to describe the attributes and locations of objects within the scope of human vision. Searching for the location of an object in the field of vision through natural language is an important capability of the human. Proposing a mechanism to learn this ability of human is a major challenge for computer vision. Most existing object localization methods usually use strong supervised information of the training set to train the model. However, these models lack interpretability and require expensive labels which are difficult to obtain. Facing these challenges, we propose a new method for locating object by natural language descriptions for fine-grained image. Firstly, we propose a model that can learn the semantically relevant parts between fine-grained images and languages, and achieve ideal localization accuracy without using strong supervisory signal. In addition, we have improved the contrast loss function to make natural language descriptions better match target regions of fine-grained images.The multi-scale fusion techniques are utilized to improve the ability of capturing details on fine-grained images. Comprehensive experiments demonstrate that the proposed method achieves ideal localization results on the CUB200-2011 dataset. And the proposed model has strong zero-shot learning ability on untrained data.
What problem does this paper attempt to address?