Abstract:Image clustering is an important and open-challenging task in computer vision. Although many methods have been proposed to solve the image clustering task, they only explore images and uncover clusters according to the image features, thus being unable to distinguish visually similar but semantically different images. In this paper, we propose to investigate the task of image clustering with the help of a visual-language pre-training model. Different from the zero-shot setting, in which the class names are known, we only know the number of clusters in this setting. Therefore, how to map images to a proper semantic space and how to cluster images from both image and semantic spaces are two key problems. To solve the above problems, we propose a novel image clustering method guided by the visual-language pre-training model CLIP, named \textbf{Semantic-Enhanced Image Clustering (SIC)}. In this new method, we propose a method to map the given images to a proper semantic space first and efficient methods to generate pseudo-labels according to the relationships between images and semantics. Finally, we propose performing clustering with consistency learning in both image space and semantic space, in a self-supervised learning fashion. The theoretical result of convergence analysis shows that our proposed method can converge at a sublinear speed. Theoretical analysis of expectation risk also shows that we can reduce the expected risk by improving neighborhood consistency, increasing prediction confidence, or reducing neighborhood imbalance. Experimental results on five benchmark datasets clearly show the superiority of our new method.

Clustering Guided SVM for Semantic Image Retrieval

Semantic-oriented 3D model classification and retrieval using Gaussian processes

Semantic Image Retrieval Based on Multiple-Instance Learning

AN image retrieval method based on multiple hyperspheres OC-SVM hashing

3D Model Retrieval with Multi-Granular Semantics Based on Gaussian Process Classifier

Semantic-Enhanced Image Clustering

Image Annotations Based on Semi-supervised Clustering with Semantic Soft Constraints.

Semisupervised SVM Batch Mode Active Learning with Applications to Image Retrieval

Semantic-consistent cross-modal hashing for large-scale image retrieval

ImageSaker : A Semantic-based Image Retrieval System Refining with Concept Model

Support Vector Machine Learning for Image Retrieval

SGC-VQGAN: Towards Complex Scene Representation via Semantic Guided Clustering Codebook

Explanation guided cross-modal social image clustering

Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens

Probability fuzzy SVM for image retrieval

Supervised Clustering-Algorithm-based Visual Information Features Classification

Adaptive Gaussian Regularization Constrained Sparse Subspace Clustering for Image Segmentation.

Support Vector Machines For Region-Based Image Retrieval

Clustering Support Vector Machines for Unlabeled Data Classification

Group Sparse Representation for Image Categorization and Semantic Video Retrieval

Semantic-Spatial Matching for Image Classification