Abstract:This study proposes a text similarity model to help biocuration efforts of the Conserved Domain Database (CDD). CDD is a curated resource that catalogs annotated multiple sequence alignment models for ancient domains and full-length proteins. These models allow for fast searching and quick identification of conserved motifs in protein sequences via Reverse PSI-BLAST. In addition, CDD curators prepare summaries detailing the function of these conserved domains and specific protein families, based on published peer-reviewed articles. To facilitate information access for database users, it is desirable to specifically identify the referenced articles that support the assertions of curator-composed sentences. Moreover, CDD curators desire an alert system that scans the newly published literature and proposes related articles of relevance to the existing CDD records. Our approach to address these needs is a text similarity method that automatically maps a curator-written statement to candidate sentences extracted from the list of referenced articles, as well as the articles in the PubMed Central database. To evaluate this proposal, we paired CDD description sentences with the top 10 matching sentences from the literature, which were given to curators for review. Through this exercise, we discovered that we were able to map the articles in the reference list to the CDD description statements with an accuracy of 77%. In the dataset that was reviewed by curators, we were able to successfully provide references for 86% of the curator statements. In addition, we suggested new articles for curator review, which were accepted by curators to be added into the reference list at an acceptance rate of 50%. Through this process, we developed a substantial corpus of similar sentences from biomedical articles on protein sequence, structure and function research, which constitute the CDD text similarity corpus. This corpus contains 5159 sentence pairs judged for their similarity on a scale from 1 (low) to 5 (high) doubly annotated by four CDD curators. Curator-assigned similarity scores have a Pearson correlation coefficient of 0.70 and an inter-annotator agreement of 85%. To date, this is the largest biomedical text similarity resource that has been manually judged, evaluated and made publicly available to the community to foster research and development of text similarity algorithms.

Fusion Matrix–Based Text Similarity Measures for Clustering of Retrieval Results

Enhancing Medline Document Clustering by Incorporating Mesh Semantic Similarity

Text Similarity Measurement Method and Application of Online Medical Community Based on Density Peak Clustering

An Ensemble Semantic Textual Similarity Measure Based on Multiple Evidences for Biomedical Documents

Learning Based Combining Different Features for Medical Image Retrieval

Prospective Study for Semantic Inter-Media Fusion in Content-Based Medical Image Retrieval

Semantic Similarity Measures to Disambiguate Terms in Medical Text.

Double-target self-supervised clustering with multi-feature fusion for medical question texts

Medical Document Clustering Using Ontology-Based Term Similarity Measures

From "Identical" to "Similar": Fusing Retrieved Lists Based on Inter-Document Similarities

MeSH2Matrix: combining MeSH keywords and machine learning for biomedical relation classification based on PubMed

Bridging the gap: Incorporating a semantic similarity measure for effectively mapping PubMed queries to documents

Re-Structuring and Specific Similarity Computation of Electronic Medical Records

Efficient Semisupervised MEDLINE Document Clustering with MeSH-Semantic and Global-Content Constraints

Cited References and Medical Subject Headings (MeSH) as Two Different Knowledge Representations: Clustering and Mappings at the Paper Level

Discovering Thematically Coherent Biomedical Documents Using Contextualized Bidirectional Encoder Representations from Transformers-Based Clustering

Corpus domain effects on distributional semantic modeling of medical terms

Measures of Cluster Informativeness for Medical Evidence Aggregation and Dissemination

Fusion in medical imaging: theory, interests and industrial applications

PubMed Text Similarity Model and Its Application to Curation Efforts in the Conserved Domain Database.

Clustering cliques for graph-based summarization of the biomedical research literature