Abstract:Exploring the semantic context in scene images is essential for indoor scene recognition. However, due to the diverse intra-class spatial layouts and the coexisting inter-class objects, modeling contextual relationships to adapt various image characteristics is a great challenge. Existing contextual modeling methods for scene recognition exhibit two limitations: 1) They typically model only one type of spatial relationship (order or metric) among objects within scenes, with limited exploration of diverse spatial layouts. 2) They often overlook the differences in coexisting objects across different scenes, suppressing scene recognition performance. To overcome these limitations, we propose SpaCoNet, which simultaneously models Spatial relation and Co-occurrence of objects guided by semantic segmentation. Firstly, the Semantic Spatial Relation Module (SSRM) is constructed to model scene spatial features. With the help of semantic segmentation, this module decouples spatial information from the scene image and thoroughly explores all spatial relationships among objects in an end-to-end manner, thereby obtaining semantic-based spatial features. Secondly, both spatial features from the SSRM and deep features from the Image Feature Extraction Module are allocated to each object, so as to distinguish the coexisting object across different scenes. Finally, utilizing the discriminative features above, we design a Global-Local Dependency Module to explore the long-range co-occurrence among objects, and further generate a semantic-guided feature representation for indoor scene recognition. Experimental results on three widely used scene datasets demonstrate the effectiveness and generality of the proposed method.

A Spindle Model For Contextual Object Detection

Detecting Human-Object Interactions with Object-Guided Cross-Modal Calibrated Semantics.

Context-Guided Super-Class Inference for Zero-Shot Detection

Object Tracking with Spatial Context Model

Exploit Spatiotemporal Contextual Information for 3D Single Object Tracking Via Memory Networks

Contextual Object Detection with Spatial Context Prototypes

Contextual Hypergraph Modeling for Salient Object Detection

Recognizing human-object interactions in still images by modeling the mutual context of objects and human poses

An Efficient Isolation Method for Contextual Object Detection

Human-Object Interaction Recognition by Modeling Context

Semantic-guided modeling of spatial relation and object co-occurrence for indoor scene recognition

Human Action Recognition with Contextual Constraints Using a RGB-D Sensor

A Semantic Context Model For Automatic Image Annotation

Exploring Dense Context for Salient Object Detection

ContextHOI: Spatial Context Learning for Human-Object Interaction Detection

Discovering spatial context prototypes for object detection

Context Dependent SVMs for Interconnected Image Network Annotation

Deep feature based contextual model for object detection

Context Modeling in 3D Human Pose Estimation: A Unified Perspective

Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object Detection

Learning discriminative context for salient object detection