Abstract:Summarization is an important natural language processing (NLP) task in identifying key information from text. For conversations, the summarization systems need to extract salient contents from spontaneous utterances by multiple speakers. In a special task-oriented scenario, namely medical conversations between patients and doctors, the symptoms, diagnoses, and treatments could be highly important because the nature of such conversation is to find a medical solution to the problem proposed by the patients. Especially consider that current online medical platforms provide millions of public available conversations between real patients and doctors, where the patients propose their medical problems and the registered doctors offer diagnosis and treatment, a conversation in most cases could be too long and the key information is hard to be located. Therefore, summarizations to the patients’ problems and the doctors’ treatments in the conversations can be highly useful, in terms of helping other patients with similar problems have a precise reference for potential medical solutions. In this paper, we focus on medical conversation summarization, using a dataset of medical conversations and corresponding summaries which were crawled from a well-known online healthcare service provider in China. We propose a hierarchical encoder-tagger model (HET) to generate summaries by identifying important utterances (with respect to problem proposing and solving) in the conversations. For the particular dataset used in this study, we show that high-quality summaries can be generated by extracting two types of utterances, namely, problem statements and treatment recommendations. Experimental results demonstrate that HET outperforms strong baselines and models from previous studies, and adding conversation-related features can further improve system performance.

Leveraging Multi-granularity Heterogeneous Graph for Chinese Electronic Medical Records Summarization

Summarizing Chinese Medical Answer with Graph Convolution Networks and Question-focused Dual Attention.

Hybrid Granularity-Based Medical Event Extraction in Chinese Electronic Medical Records

Enhanced Electronic Health Records Text Summarization Using Large Language Models

Multi-granularity heterogeneous graph attention networks for extractive document summarization

Leveraging Graph to Improve Abstractive Multi-Document Summarization.

Unsupervised Extractive Summarization with Heterogeneous Graph Embeddings for Chinese Document

A Multi-Granularity Heterogeneous Graph for Extractive Text Summarization

HetTreeSum: A Heterogeneous Tree Structure-based Extractive Summarization Model for Scientific Papers

Towards Efficient Medical Dialogue Summarization with Compacting-Abstractive Model.

Summarizing Medical Conversations via Identifying Important Utterances

Intelligent Multi-Document Summarization for Biomedical Literature by Word Embeddings and Graph-Based Ranking.

Multi-task Heterogeneous Graph Learning on Electronic Health Records

Medical Question Summarization with Entity-driven Contrastive Learning

Multi-View Metrics Enhanced Heterogeneous Graph Neural Network for Extractive Summarization

Optimizing Model Parameter for Entity Summarization Across Knowledge Graphs

A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model

A Joint Model for Chinese Medical Entity and Relation Extraction Based on Graph Convolutional Networks

Unsupervised multi-granular Chinese word segmentation and term discovery via graph partition

Question-Answering Based Summarization of Electronic Health Records using Retrieval Augmented Generation

Learning to Summarize Chinese Radiology Findings With a Pre-Trained Encoder