Abstract:Although group convolution operators are increasingly used in deep convolutional neural networks to improve the computational efficiency and to reduce the number of parameters, most existing methods construct their group convolution architectures by a predefined partitioning of the filters of each convolutional layer into multiple regular filter groups with an equal spatial group size and data-independence, which prevents a full exploitation of their potential. To tackle this issue, we propose a novel method of designing self-grouping convolutional neural networks, called SG-CNN, in which the filters of each convolutional layer group themselves based on the similarity of their importance vectors. Concretely, for each filter, we first evaluate the importance value of their input channels to identify the importance vectors, and then group these vectors by clustering. Using the resulting data-dependent centroids, we prune the less important connections, which implicitly minimizes the accuracy loss of the pruning, thus yielding a set of diverse group convolution filters. Subsequently, we develop two fine-tuning schemes, i.e. (1) both local and global fine-tuning and (2) global only fine-tuning, which experimentally deliver comparable results, to recover the recognition capacity of the pruned network. Comprehensive experiments carried out on the CIFAR-10/100 and ImageNet datasets demonstrate that our self-grouping convolution method adapts to various state-of-the-art CNN architectures, such as ResNet and DenseNet, and delivers superior performance in terms of compression ratio, speedup and recognition accuracy. We demonstrate the ability of SG-CNN to generalise by transfer learning, including domain adaption and object detection, showing competitive results. Our source code is available at <a href="https://github.com/QingbeiGuo/SG-CNN.git">https://github.com/QingbeiGuo/SG-CNN.git</a>.

Group-Based Siamese Self-Supervised Learning

Siamese Image Modeling for Self-Supervised Vision Representation Learning

Siamese self-supervised learning for fine-grained visual classification

Learning Where to Learn in Cross-View Self-Supervised Learning

Mejigclu: more effective jigsaw clustering for unsupervised visual representation learning

Exploring Simple Siamese Representation Learning

Learning multi-view visual correspondences with self-supervision

Saliency Guided Contrastive Learning on Scene Images

Group Contrastive Self-Supervised Learning on Graphs

Self-grouping convolutional neural networks

Crafting Better Contrastive Views for Siamese Representation Learning

Multi-Scale Contrastive Siamese Networks for Self-Supervised Graph Representation Learning

Align Yourself: Self-supervised Pre-training for Fine-grained Recognition via Saliency Alignment.

A Visual Embedding for the Supervised Image Based on Self-Attention

Siamese Prototypical Contrastive Learning

Exploring the Equivalence of Siamese Self-Supervised Learning Via A Unified Gradient Framework

GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding

Self-supervised image co-saliency detection

Self-supervised Multi-view Stereo Via Effective Co-Segmentation and Data-Augmentation.

Self-supervised Semantic Segmentation Grounded in Visual Concepts

Group Identification Via Transitional Hypergraph Convolution with Cross-view Self-supervised Learning