Abstract:This paper focuses on the task of Multi-Modal Summarization with Multi-Modal Output for China JD.COM e-commerce product description containing both source text and source images. In the context learning of multi-modal (text and image) input, there exists a semantic gap between text and image, especially in the cross-modal semantics of text and image. As a result, capturing shared cross-modal semantics earlier becomes crucial for multi-modal summarization. On the other hand, when generating the multi-modal summarization, based on the different contributions of input text and images, the relevance and irrelevance of multi-modal contexts to the target summary should be considered, so as to optimize the process of learning cross-modal context to guide the summary generation process and to emphasize the significant semantics within each modality. To address the aforementioned challenges, Multization has been proposed to enhance multi-modal semantic information by multi-contextually relevant and irrelevant attention alignment. Specifically, a Semantic Alignment Enhancement mechanism is employed to capture shared semantics between different modalities (text and image), so as to enhance the importance of crucial multi-modal information in the encoding stage. Additionally, the IR-Relevant Multi-Context Learning mechanism is utilized to observe the summary generation process from both relevant and irrelevant perspectives, so as to form a multi-modal context that incorporates both text and image semantic information. The experimental results in the China JD.COM e-commerce dataset demonstrate that the proposed Multization method effectively captures the shared semantics between the input source text and source images, and highlights essential semantics. It also successfully generates the multi-modal summary (including image and text) that comprehensively considers the semantics information of both text and image.

MSMO: Multimodal Summarization with Multimodal Output

Multimodal Summarization with Guidance of Multimodal Reference

An Unsupervised Video Summarization Method Based on Multimodal Representation.

VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles

MHMS: Multimodal Hierarchical Multimedia Summarization

MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos

Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment

Abstractive Sentence Summarization with Guidance of Selective Multimodal Reference.

TLDW: Extreme Multimodal Summarisation of News Videos

Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention Alignment

Heterogeneous graphormer for extractive multimodal summarization

A Modality-Enhanced Multi-Channel Attention Network for Multi-Modal Dialogue Summarization

CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization

Multi-Modal Summary Generation using Multi-Objective Optimization

Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization

CFSum: A Coarse-to-Fine Contribution Network for Multimodal Summarization

D$^2$TV: Dual Knowledge Distillation and Target-oriented Vision Modeling for Many-to-Many Multimodal Summarization

Weakening the Dominant Role of Text: CMOSI Dataset and Multimodal Semantic Enhancement Network

Self-Supervised Multimodal Opinion Summarization

VideoXum: Cross-modal Visual and Textural Summarization of Videos

Multi-modal Summarization for Video-containing Documents