A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images

Wang Zhang,Tingting Li,Yuntian Zhang,Gensheng Pei,Xiruo Jiang,Yazhou Yao
2024-04-30
Abstract:Matching visible and near-infrared (NIR) images remains a significant challenge in remote sensing image fusion. The nonlinear radiometric differences between heterogeneous remote sensing images make the image matching task even more difficult. Deep learning has gained substantial attention in computer vision tasks in recent years. However, many methods rely on supervised learning and necessitate large amounts of annotated data. Nevertheless, annotated data is frequently limited in the field of remote sensing image matching. To address this challenge, this paper proposes a novel keypoint descriptor approach that obtains robust feature descriptors via a self-supervised matching network. A light-weight transformer network, termed as LTFormer, is designed to generate deep-level feature descriptors. Furthermore, we implement an innovative triplet loss function, LT Loss, to enhance the matching performance further. Our approach outperforms conventional hand-crafted local feature descriptors and proves equally competitive compared to state-of-the-art deep learning-based methods, even amidst the shortage of annotated data.
Computer Vision and Pattern Recognition,Multimedia
What problem does this paper attempt to address?
The focus of this paper is on the problem of heterogeneous remote sensing image matching, particularly the matching of visible light and near-infrared (NIR) images. This task becomes particularly challenging due to the nonlinear radiometric differences between different sensors. Traditional methods such as region-based methods rely on surface information and are susceptible to geometric changes, brightness fluctuations, and other influences. On the other hand, feature-based methods achieve more effective matching by extracting local features. However, manually designed feature descriptors may perform poorly when dealing with heterogeneous images. Therefore, deep learning-based approaches are being explored as a potential solution.