Abstract:Fine-grained image recognition is challenging because discriminative clues are usually fragmented, whether from a single image or multiple images. Despite their significant improvements, the majority of existing methods still focus on the most discriminative parts from a single image, ignoring informative details in other regions and lacking consideration of clues from other associated images. In this paper, we analyze the difficulties of fine-grained image recognition from a new perspective and propose a transformer architecture with the peak suppression module and knowledge guidance module, which respects the diversification of discriminative features in a single image and the aggregation of discriminative clues among multiple images. Specifically, the peak suppression module first utilizes a linear projection to convert the input image into sequential tokens. It then blocks the token based on the attention response generated by the transformer encoder. This module penalizes the attention to the most discriminative parts in the feature learning process, therefore, enhancing the information exploitation of the neglected regions. The knowledge guidance module compares the image-based representation generated from the peak suppression module with the learnable knowledge embedding set to obtain the knowledge response coefficients. Afterwards, it formalizes the knowledge learning as a classification problem using response coefficients as the classification scores. Knowledge embeddings and image-based representations are updated during training simultaneously so that the knowledge embedding includes a large number of discriminative clues for different images of the same category. Finally, we incorporate the acquired knowledge embeddings into the image-based representations as comprehensive representations, leading to significantly higher recognition performance. Extensive evaluations on the six popular datasets demonstrate the advantage of the proposed method in performance. The source code and models will be available online after the acceptance of the paper.

Multi-input trademark element recognition with transformer

Enhanced Multi-Scale Trademark Element Detection Using the Improved DETR

Asymmetric Vision Transformers for Multi-Label Classification

MAFormer: A transformer network with multi-scale attention fusion for visual recognition

Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image Recognition.

Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image Recognition

Trademark Text Recognition Combining SwinTransformer and Feature-Query Mechanisms

Transformer with peak suppression and knowledge guidance for fine-grained image recognition

Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification

Multi-level information fusion Transformer with background filter for fine-grained image recognition

MFF-Trans: Multi-level Feature Fusion Transformer for Fine-Grained Visual Classification

Vision transformer with multiple granularities for person re-identification

Visible-Infrared Person Re-Identification via Cross-Modality Interaction Transformer

TripleFormer: improving transformer-based image classification method using multiple self-attention inputs

TransFG: A Transformer Architecture for Fine-Grained Recognition

Transformer Based Multi-Grained Features for Unsupervised Person Re-Identification

SST: Spatial and Semantic Transformers for Multi-Label Image Recognition

A multimodal hyper-fusion transformer for remote sensing image classification

STMG: Swin transformer for multi-label image recognition with graph convolution network

Multi-level network based on transformer encoder for fine-grained image–text matching