Data-driven Summarization of Scientific Articles

Nikola I. Nikolov,Michael Pfeiffer,Richard H.R. Hahnloser
DOI: https://doi.org/10.48550/arXiv.1804.08875
2018-04-24
Abstract:Data-driven approaches to sequence-to-sequence modelling have been successfully applied to short text summarization of news articles. Such models are typically trained on input-summary pairs consisting of only a single or a few sentences, partially due to limited availability of multi-sentence training data. Here, we propose to use scientific articles as a new milestone for text summarization: large-scale training data come almost for free with two types of high-quality summaries at different levels - the title and the abstract. We generate two novel multi-sentence summarization datasets from scientific articles and test the suitability of a wide range of existing extractive and abstractive neural network-based summarization approaches. Our analysis demonstrates that scientific papers are suitable for data-driven text summarization. Our results could serve as valuable benchmarks for scaling sequence-to-sequence models to very long sequences.
Computation and Language
What problem does this paper attempt to address?
The main problem that this paper attempts to solve is to explore the applicability of scientific articles as a new benchmark for data - driven text summarization. Specifically, the paper proposes two new multi - sentence summarization datasets, which are sourced from scientific articles, and tests the performance of a series of existing extractive and generative neural network summarization methods on these datasets. Through this method, the paper aims to evaluate whether scientific articles are suitable for data - driven text summarization and provide a valuable benchmark for extending sequence - to - sequence models to handle very long sequences. The paper mainly focuses on the following points: 1. **Scientific articles as a new source of summary data**: The paper proposes using scientific articles as a new milestone for data - driven text summarization, because scientific articles usually come with high - quality summaries and titles and can be used as training data. 2. **Constructing new datasets**: The paper constructs two new large - scale multi - sentence summarization datasets: - **title - gen**: It contains 5 million pairs of article titles and summaries in the biomedical field. - **abstract - gen**: It contains 900,000 pairs of article summaries and bodies. 3. **Evaluating existing methods**: The paper evaluates the performance of a series of existing extractive and generative neural network summarization methods on these new datasets, including extractive methods based on word embeddings and generative methods based on Recurrent Neural Networks (RNN) and Convolutional Neural Networks (CNN). 4. **Performance analysis**: The paper analyzes the outputs of these models through quantitative and qualitative methods, especially focusing on the performance of the models when dealing with long input / output sequence pairs. Overall, the goal of this paper is to promote research in the field of scientific article summarization and provide a basis for developing new models that can efficiently handle long input and output sequences.