Abstract:The existing RGB-D scene recognition approaches typically employ two separate and modality-specific networks to learn effective RGB and Depth representations respectively. This independent training scheme fails to capture the correlation of two modalities, and thus may be suboptimal for RGB-D scene recognition. To address this issue, this paper proposes a general and flexible framework to enhance RGB-D representation learning with a customized cross-modal pyramid translation branch, coined as TRecgNet. This framework unifies the tasks of cross-modal translation and modality-specific recognition with a shared feature encoder, and aims at leveraging the correspondence between two modalities to regularize the representation learning of each modality. Specifically, we present a cross-modal pyramid translation strategy to perform multi-scale image generation with a carefully designed layer-wise perceptual supervision. To improve the complementarity of cross-modal translation to modality specific scene recognition, we devise a feature selection module to adaptively enhance the discriminative information during the translation procedure. In addition, we train multiple auxiliary classifiers to further regularize the behavior of generated data to be consistent with its paired data on label prediction. Meanwhile, our translation branch enables us to generate cross-modal data for training data augmentation and further improve single modality scene recognition. Extensive experiments on benchmarks of SUN RGB-D and NYU Depth V2 demonstrate the superiority of the proposed method to the state-of-the-art RGB-D scene recognition methods. We also generalize the TRecgNet to the single modality scene recognition benchmark of MIT Indoor, and automatically synthesize a depth view to boost the final recognition accuracy.

Cross-Modal Matching and Adaptive Graph Attention Network for RGB-D Scene Recognition

ACM: Adaptive Cross-Modal Graph Convolutional Neural Networks for RGB-D Scene Recognition.

ACNET: Attention Based Network to Exploit Complementary Features for RGBD Semantic Segmentation.

MSN: Modality Separation Networks for RGB-D Scene Recognition

Cross-modal refined adjacent-guided network for RGB-D salient object detection

Cross-Modal Attentional Context Learning for RGB-D Object Detection

RGB-D Scene Recognition Via Spatial-Related Multi-Modal Feature Learning

Translate-to-Recognize Networks for RGB-D Scene Recognition

Cross-modal Attention Fusion Network for RGB-D Semantic Segmentation

Double Cross-Modality Progressively Guided Network for RGB-D Salient Object Detection

A Local-Global Self-attention Interaction Network for RGB-D Cross-Modal Person Re-identification.

Multi-scale Cross-Modal Transformer Network for RGB-D Object Detection

CMCLNet Cross-Modality Attention Fusion and Cross-Level Feature Interaction for RGBD salient object detection

Cross-Modal Pyramid Translation for RGB-D Scene Recognition

Modal-Adaptive Gated Recoding Network for RGB-D Salient Object Detection

Discriminative Cross-Modal Transfer Learning and Densely Cross-Level Feedback Fusion for RGB-D Salient Object Detection

M 2rnet: Multi-modal and Multi-Scale Refined Network for RGB-D Salient Object Detection

Mitigating Modality Discrepancies for RGB-T Semantic Segmentation

M2RNet: Multi-modal and Multi-scale Refined Network for RGB-D Salient Object Detection

Global Guided Cross-Modal Cross-Scale Network for RGB-D Salient Object Detection

Cross-modality Discrepant Interaction Network for RGB-D Salient Object Detection