Abstract:Combining Convolutional Neural Networks (CNNs) with Conditional Random Fields (CRFs) achieves great success among recent object segmentation methods. There are two advantages by such usage. First, CNNs can extract low-level features, which are very similar to the extracted features in primates’ primary visual cortex (V1). Second, CRFs can set up the relationship between input features and output labels in a direct way. In this paper, we extend the first advantage by using CNNs for low-level feature extraction and a Structured Random Forest (SRF)-based border ownership detector for high-level feature extraction, which are similar to the outputs of primates secondary visual cortex (V2). Compared to the CRF model, an improved Conditional Boltzmann Machine (CBM), which has a multi-channel visible layer, is proposed to model the relationship between predicted labels, local and global contexts of objects with multi-scale and multilevel features. Besides, our proposed CBM model is extended for object parsing by using multivisible branches instead of a single visible layer of CBM, which cannot only segment the whole body but also the parts of the body under. These visible branches use each branch for the segmentation of the whole body or one of the body parts. All branches share the same hidden layers of CBM and train the branches under an iterative way. By exploiting object parsing, the whole body segmentation performance of object is improved. To refine the segmentation output, two kinds of optimization algorithms are proposed. The superpixel-based algorithm can re-label the overlapped regions of multiple kinds of objects. The other curve correction algorithm corrects the edges of segmented object parts by using smooth edges under a curve similarity criterion. Experiments demonstrate that our models yield competitive results for object segmentation on the PASCAL VOC 2012 dataset and for object parsing on the PennFudan Pedestrian Parsing dataset, Pedestrian Parsing Surveillance Scenes dataset, Horse-Cow parsing dataset, and PASCAL Quadrupeds dataset.

A Top-Down Manner-Based DCNN Architecture for Semantic Image Segmentation.

Image Semantic Segmentation Based on Region and Deep Residual Network

Real-time Semantic Segmentation with Weighted Factorized-Depthwise Convolution

Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs

Semantic Segmentation of Aerial Imagery Via Split-Attention Networks with Disentangled Nonlocal and Edge Supervision

Real-Time High-Performance Semantic Image Segmentation of Urban Street Scenes

Better Image Segmentation by Exploiting Dense Semantic Predictions

Fully Convolutional Networks for Semantic Segmentation

DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs

Deep Convolutional Neural Networks with Spatial Regularization, Volume and Star-shape Priori for Image Segmentation

High-performance Semantic Segmentation Using Very Deep Fully Convolutional Networks

DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation

Densely Based Multi-Scale and Multi-Modal Fully Convolutional Networks for High-Resolution Remote-Sensing Image Semantic Segmentation

Discriminative Features Reconstruction Network For Semantic Segmentation

Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation

Learning Deep Conditional Neural Network for Image Segmentation

A Dual-Path and Lightweight Convolutional Neural Network for High-Resolution Aerial Image Segmentation

Dense Convolutional Networks for Semantic Segmentation.

An Open-Source Project for Real-Time Image Semantic Segmentation

STC: A Simple to Complex Framework for Weakly-Supervised Semantic Segmentation

Adaptive multi-scale dual attention network for semantic segmentation