Abstract:Sketch is a well-researched topic in the vision community by now. Sketch semantic segmentation in particular, serves as a fundamental step towards finer-level sketch interpretation. Recent works use various means of extracting discriminative features from sketches and have achieved considerable improvements on segmentation accuracy. Common approaches for this include attending to the sketch-image as a whole, its stroke-level representation or the sequence information embedded in it. However, they mostly focus on only a part of such multi-facet information. In this paper, we for the first time demonstrate that there is complementary information to be explored across all these three facets of sketch data, and that segmentation performance consequently benefits as a result of such exploration of sketch-specific information. Specifically, we propose the Sketch-Segformer, a transformer-based framework for sketch semantic segmentation that inherently treats sketches as stroke sequences other than pixel-maps. In particular, Sketch-Segformer introduces two types of self-attention modules having similar structures that work with different receptive fields (i.e., whole sketch or individual stroke). The order embedding is then further synergized with spatial embeddings learned from the entire sketch as well as localized stroke-level information. Extensive experiments show that our sketch-specific design is not only able to obtain state-of-the-art performance on traditional figurative sketches (such as SPG, SketchSeg-150K datasets), but also performs well on creative sketches that do not conform to conventional object semantics (CreativeSketch dataset) thanks for our usage of multi-facet sketch information. Ablation studies, visualizations, and invariance tests further justifies our design choice and the effectiveness of Sketch-Segformer. Codes are available at https://github.com/PRIS-CV/Sketch-SF.

SketchSegNet+: An End-to-End Learning of RNN for Multi-Class Sketch Semantic Segmentation

Sketchsegnet Plus : An End-To-End Learning Of Rnn For Multi-Class Sketch Semantic Segmentation

Sketchsegnet: A Rnn Model for Labeling Sketch Strokes.

Sketch-R2CNN: an Attentive Network for Vector Sketch Recognition

<i>Sketch-R2CNN</i>: An RNN-Rasterization-CNN Architecture for Vector Sketch Recognition

Stroke-based semantic segmentation for scene-level free-hand sketches

SFSegNet: Parse Freehand Sketches using Deep Fully Convolutional Networks

SketchGNN: Semantic Sketch Segmentation with Graph Neural Networks

SceneSketcher-v2: Fine-Grained Scene-Level Sketch-Based Image Retrieval Using Adaptive GCNs

Multi-column Point-CNN for Sketch Segmentation

SketchGCN: Semantic Sketch Segmentation with Graph Convolutional Networks

Sketch-Snet: Deeper Subdivision Of Temporal Cues For Sketch Recognition

ENDE-GNN: an Encoder-decoder GNN Framework for Sketch Semantic Segmentation

S3Net:Graph Representational Network for Sketch Recognition

Sketch-Segformer: Transformer-Based Segmentation for Figurative and Creative Sketches.

CreativeSeg: Semantic Segmentation of Creative Sketches

Segmentation of Online Sketching Using Geometric Feature

Scene Sketch Semantic Segmentation with Hierarchical Transformer.

AI-Sketcher : A Deep Generative Model for Producing High-Quality Sketches.

Stroke Classification for Sketch Segmentation by Fine-Tuning a Developmental VGGNet16.

On Learning Semantic Representations for Million-Scale Free-Hand Sketches