GP-HTNLoc: A Graph Prototype Head-Tail Network-based Model for Multi-label Subcellular Localization Prediction of ncRNAs

Shuangkai Han,Lin Liu
DOI: https://doi.org/10.1101/2024.03.04.583439
2024-03-07
Abstract:Numerous research findings demonstrated that understanding the subcellular localization of non-coding RNAs (ncRNAs) is pivotal in elucidating their roles and regulatory mechanisms in cells. Despite the existence of over ten computational models dedicated to predicting the subcellular localization of ncRNAs, a majority of these models are designed solely for single-label prediction. In reality, ncRNAs often exhibit localization across multiple subcellular compartments. Furthermore, the existing multi-label localization prediction models are insufficient in addressing the challenges posed by the scarcity of training samples and class imbalance in ncRNA dataset. This study addresses the limitations of existing models by introducing a novel multi-label localization prediction model for ncRNAs, termed GP-HTNLoc. To alleviate class imbalance, the model adopts a separate training approach for head and tail class labels. In GP-HTNLoc, a pioneering graph prototype module is introduced for capturing potential association of ncRNA samples with labels. This module efficiently learns the graph structure and aggregates sample features. Notably, only few samples are required to obtain label prototypes containing rich information. These prototypes are then utilized to train a transfer learner, facilitating the transfer of meta-knowledge from the head class to the tail class. Experimental results demonstrate that GP-HTNLoc surpasses current state-of-the-art models across all datasets. Ablation study underscore the vital role played by the graph prototype module in enhancing the performance of GP-HTNLoc. The user-friendly online GP-HTNLoc web server can be accessed at .
Bioinformatics
What problem does this paper attempt to address?