Abstract:Background: Research on disease-disease association (DDA), like comorbidity and complication, provides important insights into disease treatment and drug discovery, and a large body of the literature has been published in the field. However, using current search tools, it is not easy for researchers to retrieve information on the latest DDA findings. First, comorbidity and complication keywords pull up large numbers of PubMed studies. Second, disease is not highlighted in search results. Finally, DDA is not identified, as currently no disease-disease association extraction (DDAE) dataset or tools are available. Objective: As there are no available DDAE datasets or tools, this study aimed to develop (1) a DDAE dataset and (2) a neural network model for extracting DDA from the literature. Methods: In this study, we formulated DDAE as a supervised machine learning classification problem. To develop the system, we first built a DDAE dataset. We then employed two machine learning models, support vector machine and convolutional neural network, to extract DDA. Furthermore, we evaluated the effect of using the output layer as features of the support vector machine-based model. Finally, we implemented large margin context-aware convolutional neural network architecture to integrate context features and convolutional neural networks through the large margin function. Results: Our DDAE dataset consisted of 521 PubMed abstracts. Experiment results showed that the support vector machine-based approach achieved an F1 measure of 80.32%, which is higher than the convolutional neural network-based approach (73.32%). Using the output layer of convolutional neural network as a feature for the support vector machine does not further improve the performance of support vector machine. However, our large margin context-aware-convolutional neural network achieved the highest F1 measure of 84.18% and demonstrated that combining the hinge loss function of support vector machine with a convolutional neural network into a single neural network architecture outperforms other approaches. Conclusions: To facilitate the development of text-mining research for DDAE, we developed the first publicly available DDAE dataset consisting of disease mentions, Medical Subject Heading IDs, and relation annotations. We developed different conventional machine learning models and neural network architectures and evaluated their effects on our DDAE dataset. To further improve DDAE performance, we propose an large margin context-aware-convolutional neural network model for DDAE that outperforms other approaches.

Beyond associations: A benchmark Causal Relation Extraction Dataset (CRED) of disease-causing genes, its comparative evaluation, interpretation and application

BioRED: a rich biomedical relation extraction dataset

Causal relationships between diseases mined from the literature improve the use of polygenic risk scores

RENET2: high-performance full-text gene-disease relation extraction with iterative training data expansion

CRTRE: Causal Rule Generation with Target Trial Emulation Framework

TBGA: a large-scale Gene-Disease Association dataset for Biomedical Relation Extraction

Inferring disease correlation from healthcare data

Causal Gene Identification Using Non-linear Regression-based Independence Tests

MIMICause: Representation and automatic extraction of causal relation types from clinical notes

Overcoming Confounding Bias in Causal Discovery Using Minimum Redundancy and Maximum Relevancy Constraint

Development of a graphical model of causal gene regulatory networks using medical big data and Bayesian machine learning

Biomedical event causal relation extraction with deep knowledge fusion and Roberta-based data augmentation

Joint Learning-based Causal Relation Extraction from Biomedical Literature

REDDA: Integrating multiple biological relations to heterogeneous graph neural network for drug-disease association prediction

BioCreative V CDR task corpus: a resource for chemical disease relation extraction

Chemical-induced Disease Relation Extraction with Dependency Information and Prior Knowledge

Using a Large Margin Context-Aware Convolutional Neural Network to Automatically Extract Disease-Disease Association from Literature: Comparative Analytic Study

Decoding the gene-disease associations in type 2 diabetes: A curated dataset for text mining-based classification

Exploiting syntactic and semantics information for chemical–disease relation extraction

A Method for Improving the Reliability of Causal Inference from Large-Scale Data in Biomedicine.

Generating Gene Ontology-Disease Inferences to Explore Mechanisms of Human Disease at the Comparative Toxicogenomics Database