Abstract:We present a novel and practical deep fully convolutional neural network architecture for semantic pixel-wise segmentation termed SegNet. This core trainable segmentation engine consists of an encoder network, a corresponding decoder network followed by a pixel-wise classification layer. The architecture of the encoder network is topologically identical to the 13 convolutional layers in the VGG16 network [1] . The role of the decoder network is to map the low resolution encoder feature maps to full input resolution feature maps for pixel-wise classification. The novelty of SegNet lies is in the manner in which the decoder upsamples its lower resolution input feature map(s). Specifically, the decoder uses pooling indices computed in the max-pooling step of the corresponding encoder to perform non-linear upsampling. This eliminates the need for learning to upsample. The upsampled maps are sparse and are then convolved with trainable filters to produce dense feature maps. We compare our proposed architecture with the widely adopted FCN [2] and also with the well known DeepLab-LargeFOV [3] , DeconvNet [4] architectures. This comparison reveals the memory versus accuracy trade-off involved in achieving good segmentation performance. SegNet was primarily motivated by scene understanding applications. Hence, it is designed to be efficient both in terms of memory and computational time during inference. It is also significantly smaller in the number of trainable parameters than other competing architectures and can be trained end-to-end using stochastic gradient descent. We also performed a controlled benchmark of SegNet and other architectures on both road scenes and SUN RGB-D indoor scene segmentation tasks. These quantitative assessments show that SegNet provides good performance with competitive inference time and most efficient inference memory-wise as compared to other architectures. We also provide a Caffe implementation of SegNet and a web demo at http://mi.eng.cam.ac.uk/projects/segnet.

SegBot: A Generic Neural Text Segmentation Model with Pointer Network.

SegFormer: A Topic Segmentation Model with Controllable Range of Attention.

A Sequence-to-Sequence Approach with Mixed Pointers to Topic Segmentation and Segment Labeling

GiTS: Gist-driven Text Segmentation

Text4Seg: Reimagining Image Segmentation as Text Generation

SegNet: A Deep Convolutional Encoder-Decoder Architecture for Robust Semantic Pixel-Wise Labelling

Tipster: A Topic-Guided Language Model for Topic-Aware Text Segmentation.

BAGNet: Branch Attention Guided Decoder for Semantic Segmentation

BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation

SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models

SegNet Network Architecture for Deep Learning Image Segmentation and Its Integrated Applications and Prospects

Bridging the Gap Between End-to-End and Two-Step Text Spotting

Rethinking Semantic Segmentation: A Prototype View

MetaSegNet: Metadata-collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing Images

Purify then Guide: A Bi-directional Bridge Network for Open-Vocabulary Semantic Segmentation

Accurate and Efficient Event-based Semantic Segmentation Using Adaptive Spiking Encoder-Decoder Network

Toward Fast and Accurate Neural Discourse Segmentation

CellSegNet: an Adaptive Multi-Resolution Hybrid Network for Cell Segmentation

BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation