Species-specific model based on sequence and structural information for ubiquitination sites prediction

Weimin Li,Nan Chen,Jie Wang,Yin Luo,Huazhong Liu,Jihong Ding,Qun Jin
DOI: https://doi.org/10.1016/j.jmb.2024.168781
IF: 6.151
2024-09-09
Journal of Molecular Biology
Abstract:Ubiquitination is a common post-translational modification of proteins in eukaryotic cells, and it is also a significant method of regulating protein biological function. The use of computational methods for predicting ubiquitination sites can replace costly and time-consuming experimental methods. Existing computational methods often build classifiers based on protein sequence information, physical and chemical properties of amino acids, evolutionary information, and structural parameters. However, structural information about most proteins cannot be found in existing databases directly. The features of proteins differ among species, and some species have small amounts of ubiquitinated proteins. Therefore, it is necessary to develop species-specific models that can be applied to datasets with small sample sizes. To solve these problems, we propose a species-specific model (SSUbi) based on a capsule network, which integrates proteins' sequence and structural information. In this model, the feature extraction module is composed of two sub-modules that extract multi-dimensional features from sequence and structural information respectively. In the submodule, the convolution operation is used to extract encoding dimension features, and the channel attention mechanism is used to extract feature map dimension features. After the multi-dimensional features of the two kinds of information are integrated, the species-specific capsule network further converts the features into capsule vectors and classifies species-specific ubiquitination sites. The experimental results show that SSUbi can effectively improve the prediction performance of species with small sample sizes and outperform other models.
biochemistry & molecular biology
What problem does this paper attempt to address?