Abstract:Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single $\textbf{Vi}$sion and $\textbf{S}$cene $\textbf{T}$ext $\textbf{A}$ggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least $\bf{8.4}\%$ at Recall@1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework.

A Reconstruction-based Visual-Acoustic-Semantic Embedding Method for Speech-Image Retrieval

AudioVSR: Enhancing Video Speech Recognition with Audio Data

Cross-modal Mask Fusion and Modality-Balanced Audio-Visual Speech Recognition

Multi-view visual semantic embedding for cross-modal image–text retrieval

Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement

Dynamic Visual Semantic Sub-Embeddings and Fast Re-Ranking

Cooperative Dual Attention for Audio-Visual Speech Enhancement with Facial Cues

Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval

Coarse-to-fine Alignment Makes Better Speech-image Retrieval

SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model

Incorporating Visual Information Reconstruction into Progressive Learning for Optimizing audio-visual Speech Enhancement

Catch Missing Details: Image Reconstruction with Frequency Augmented Variational Autoencoder

ViLEM: Visual-Language Error Modeling for Image-Text Retrieval

Semantic Retrieval Of Personal Photos Using A Deep Autoencoder Fusing Visual Features With Speech Annotations Represented As Word/Paragraph Vectors

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Unified Visual-Semantic Embeddings: Bridging Vision and Language with Structured Meaning Representations

Deep Visual Semantic Embedding with Text Data Augmentation and Word Embedding Initialization

Rethinking the visual cues in audio-visual speaker extraction

Multimodal Variational Auto-encoder based Audio-Visual Segmentation

For end-to-end audio-visual speech recognition