Abstract:This paper focuses on the task of Multi-Modal Summarization with Multi-Modal Output for China JD.COM e-commerce product description containing both source text and source images. In the context learning of multi-modal (text and image) input, there exists a semantic gap between text and image, especially in the cross-modal semantics of text and image. As a result, capturing shared cross-modal semantics earlier becomes crucial for multi-modal summarization. On the other hand, when generating the multi-modal summarization, based on the different contributions of input text and images, the relevance and irrelevance of multi-modal contexts to the target summary should be considered, so as to optimize the process of learning cross-modal context to guide the summary generation process and to emphasize the significant semantics within each modality. To address the aforementioned challenges, Multization has been proposed to enhance multi-modal semantic information by multi-contextually relevant and irrelevant attention alignment. Specifically, a Semantic Alignment Enhancement mechanism is employed to capture shared semantics between different modalities (text and image), so as to enhance the importance of crucial multi-modal information in the encoding stage. Additionally, the IR-Relevant Multi-Context Learning mechanism is utilized to observe the summary generation process from both relevant and irrelevant perspectives, so as to form a multi-modal context that incorporates both text and image semantic information. The experimental results in the China JD.COM e-commerce dataset demonstrate that the proposed Multization method effectively captures the shared semantics between the input source text and source images, and highlights essential semantics. It also successfully generates the multi-modal summary (including image and text) that comprehensively considers the semantics information of both text and image.

CTNR: Compress-then-Reconstruct Approach for Multimodal Abstractive Summarization

An Unsupervised Video Summarization Method Based on Multimodal Representation.

SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization

Semantics-Consistent Cross-domain Summarization via Optimal Transport Alignment

Selective and Coverage Multi-head Attention for Abstractive Summarization

Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention Alignment

Multimodal Abstractive Summarization using bidirectional encoder representations from transformers with attention mechanism

VideoXum: Cross-modal Visual and Textural Summarization of Videos

MHMS: Multimodal Hierarchical Multimedia Summarization

TLDW: Extreme Multimodal Summarisation of News Videos

CISum: Learning Cross-modality Interaction to Enhance Multimodal Semantic Coverage for Multimodal Summarization

Multi-task Hierarchical Heterogeneous Fusion Framework for multimodal summarization

Inter- and Intra-Modal Contrastive Hybrid Learning Framework for Multimodal Abstractive Summarization

Learning Summary-Worthy Visual Representation for Abstractive Summarization in Video

Align and Attend: Multimodal Summarization with Dual Contrastive Losses

Unifying Cross-lingual Summarization and Machine Translation with Compression Rate

See, Hear, Read: Leveraging Multimodality with Guided Attention for Abstractive Text Summarization

CFSum: A Coarse-to-Fine Contribution Network for Multimodal Summarization

Multimodal Abstractive Summarization for How2 Videos

VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles

Topic-Guided Abstractive Text Summarization: a Joint Learning Approach