Abstract:Sketch is a well-researched topic in the vision community by now. Sketch semantic segmentation in particular, serves as a fundamental step towards finer-level sketch interpretation. Recent works use various means of extracting discriminative features from sketches and have achieved considerable improvements on segmentation accuracy. Common approaches for this include attending to the sketch-image as a whole, its stroke-level representation or the sequence information embedded in it. However, they mostly focus on only a part of such multi-facet information. In this paper, we for the first time demonstrate that there is complementary information to be explored across all these three facets of sketch data, and that segmentation performance consequently benefits as a result of such exploration of sketch-specific information. Specifically, we propose the Sketch-Segformer, a transformer-based framework for sketch semantic segmentation that inherently treats sketches as stroke sequences other than pixel-maps. In particular, Sketch-Segformer introduces two types of self-attention modules having similar structures that work with different receptive fields (i.e., whole sketch or individual stroke). The order embedding is then further synergized with spatial embeddings learned from the entire sketch as well as localized stroke-level information. Extensive experiments show that our sketch-specific design is not only able to obtain state-of-the-art performance on traditional figurative sketches (such as SPG, SketchSeg-150K datasets), but also performs well on creative sketches that do not conform to conventional object semantics (CreativeSketch dataset) thanks for our usage of multi-facet sketch information. Ablation studies, visualizations, and invariance tests further justifies our design choice and the effectiveness of Sketch-Segformer. Codes are available at https://github.com/PRIS-CV/Sketch-SF.

Scene Sketch Semantic Segmentation with Hierarchical Transformer.

Sketchformer++: A Hierarchical Transformer Architecture for Vector Sketch Representation

SceneSketcher-v2: Fine-Grained Scene-Level Sketch-Based Image Retrieval Using Adaptive GCNs

Sketch-Segformer: Transformer-Based Segmentation for Figurative and Creative Sketches.

Exploring Local Detail Perception for Scene Sketch Semantic Segmentation

Open Vocabulary Semantic Scene Sketch Understanding

Adapting Hierarchical Transformer for Scene-Level Sketch-Based Image Retrieval.

PSSD-Transformer: Powerful Sparse Spike-Driven Transformer for Image Semantic Segmentation

SFSegNet: Parse Freehand Sketches using Deep Fully Convolutional Networks

Sketch Face Recognition Based on Light Semantic Transformer Network

Stroke-based semantic segmentation for scene-level free-hand sketches

SketchSegNet+: An End-to-End Learning of RNN for Multi-Class Sketch Semantic Segmentation

On Learning Semantic Representations for Large-Scale Abstract Sketches

ContextSeg: Sketch Semantic Segmentation by Querying the Context with Attention

S 5 : Sketch-to-image Synthesis via Scene and Size Sensing

Multi-column Point-CNN for Sketch Segmentation

SSNet: A Novel Transformer and CNN Hybrid Network for Remote Sensing Semantic Segmentation

Sketchsegnet Plus : An End-To-End Learning Of Rnn For Multi-Class Sketch Semantic Segmentation

SketchGNN: Semantic Sketch Segmentation with Graph Neural Networks

On Learning Semantic Representations for Million-Scale Free-Hand Sketches

SketchGCN: Semantic Sketch Segmentation with Graph Convolutional Networks