Abstract:We present a novel and practical deep fully convolutional neural network architecture for semantic pixel-wise segmentation termed SegNet. This core trainable segmentation engine consists of an encoder network, a corresponding decoder network followed by a pixel-wise classification layer. The architecture of the encoder network is topologically identical to the 13 convolutional layers in the VGG16 network. The role of the decoder network is to map the low resolution encoder feature maps to full input resolution feature maps for pixel-wise classification. The novelty of SegNet lies is in the manner in which the decoder upsamples its lower resolution input feature map(s). Specifically, the decoder uses pooling indices computed in the max-pooling step of the corresponding encoder to perform non-linear upsampling. This eliminates the need for learning to upsample. The upsampled maps are sparse and are then convolved with trainable filters to produce dense feature maps. We compare our proposed architecture with the widely adopted FCN and also with the well known DeepLab-LargeFOV, DeconvNet architectures. This comparison reveals the memory versus accuracy trade-off involved in achieving good segmentation performance. SegNet was primarily motivated by scene understanding applications. Hence, it is designed to be efficient both in terms of memory and computational time during inference. It is also significantly smaller in the number of trainable parameters than other competing architectures. We also performed a controlled benchmark of SegNet and other architectures on both road scenes and SUN RGB-D indoor scene segmentation tasks. We show that SegNet provides good performance with competitive inference time and more efficient inference memory-wise as compared to other architectures. We also provide a Caffe implementation of SegNet and a web demo at http://mi.eng.cam.ac.uk/projects/segnet/.

TableSegNet: a fully convolutional network for table detection and segmentation in document images

CNN Based Page Object Detection in Document Images

HybridTabNet: Towards Better Table Detection in Scanned Document Images

TableDet: An end-to-end deep learning approach for table detection and table image classification in data sheet images

TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images

Robust Table Detection and Structure Recognition from Heterogeneous Document Images

U-SSD: Improved SSD Based on U-Net Architecture for End-to-End Table Detection in Document Images

Dual-branch dilated context convolutional for table detection transformer in the document images

A Deep Semantic Segmentation Model for Image-based Table Structure Recognition

CasTabDetectoRS: Cascade Network for Table Detection in Document Images with Recursive Feature Pyramid and Switchable Atrous Convolution

A Table Detection Method for PDF Documents Based on Convolutional Neural Networks

Table Detection for Visually Rich Document Images

Table Structure Extraction with Bi-directional Gated Recurrent Unit Networks

Image-based table recognition: data, model, and evaluation

DCSegNet: Deep Learning Framework Based on Divide-and-Conquer Method for Liver Segmentation

A Saliency-Based Convolutional Neural Network for Table and Chart Detection in Digitized Documents

SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation

UTTSR: A Novel Non-Structured Text Table Recognition Model Powered by Deep Learning Technology

End-to-End Semi-Supervised approach with Modulated Object Queries for Table Detection in Documents