Abstract:Scene depth super-resolution (DSR) poses an inherently ill-posed problem due to the extremely large space of one-to-many mapping functions from a given low-resolution (LR) depth map, which possesses limited depth information, to multiple plausible high-resolution (HR) depth maps. This characteristic renders the task highly challenging, as identifying an optimal solution becomes significantly intricate amidst this multitude of potential mappings. While simplistic constraints have been proposed to address the DSR task, the relationship between LR and HR depth maps and the color image has not been thoroughly investigated. In this paper, we introduce a novel mapping constraint network (MCNet) that incorporates additional constraints derived from both LR depth maps and color images. This integration aims to optimize the space of mapping functions and enhance the performance of DSR. Specifically, alongside the primary DSR network (DSRNet) dedicated to learning LR-to-HR mapping, we have developed an auxiliary degradation network (ADNet) that operates in reverse, generating the LR depth map from the reconstructed HR depth map to obtain depth features in LR space. To enhance the learning process of DSRNet in LR-to-HR mapping, we introduce two mapping constraints in LR space: 1) the cycle-consistent constraint, which offers additional supervision by establishing a closed loop between LR-to-HR and HR-to-LR mappings, and 2) the region-level contrastive constraint, aimed at reinforcing region-specific HR representations by explicitly modeling the consistency between LR and HR spaces. To leverage the color image effectively, we introduce a feature screening module (FSM) to adaptively fuse color features at different layers, which can simultaneously maintain strong structural context and suppress texture distraction through subspace generation and image projection. Comprehensive experimental results across synthetic and real-world benchmark datasets unequivocally demonstrate the superiority of our proposed method over state-of-the-art DSR methods. Our MCNet achieves an average MAD reduction of 3.7% and 7.5% over state-of-the-art DSR method for \(\times\) 8 and \(\times\) 16 cases on Milddleburry dataset, respectively, without incurring additional costs during inference.

ADRNet-S*: Asymmetric Depth Registration Network via Contrastive Knowledge Distillation for RGB-D Mirror Segmentation

ACNET: Attention Based Network to Exploit Complementary Features for RGBD Semantic Segmentation.

Semantic Progressive Guidance Network for RGB-D Mirror Segmentation

Depth Cue Enhancement and Guidance Network for RGB-D Salient Object Detection

Self-Knowledge Distillation-Based Staged Extraction and Multiview Collection Network for RGB-D Mirror Segmentation

Contrastive learning-based knowledge distillation for RGB-thermal urban scene semantic segmentation

ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-high Resolution Segmentation

ADRNet: Affine and Deformable Registration Networks for Multimodal Remote Sensing Images

TANet: Transformer-based Asymmetric Network for RGB-D Salient Object Detection

RGB×D: Learning Depth-Weighted RGB Patches for RGB-D Indoor Semantic Segmentation

Cross-modal refined adjacent-guided network for RGB-D salient object detection

Spatial-information Guided Adaptive Context-aware Network for Efficient RGB-D Semantic Segmentation

Mirrored X-Net: Joint Classification and Contrastive Learning for Weakly Supervised GA Segmentation in SD-OCT★

DAABNet: depth-wise asymmetric attention bottleneck for real-time semantic segmentation

Attention-Guided Multi-Modality Interaction Network for RGB-D Salient Object Detection

MoADNet: Mobile Asymmetric Dual-Stream Networks for Real-Time and Lightweight RGB-D Salient Object Detection

Digging into Depth and Color Spaces: A Mapping Constraint Network for Depth Super-Resolution

Mirror Segmentation via Semantic-aware Contextual Contrasted Feature Learning

Augmentation-Free Dense Contrastive Knowledge Distillation for Efficient Semantic Segmentation

Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes

CalibNet: Dual-branch Cross-modal Calibration for RGB-D Salient Instance Segmentation