A Pipeline for Data-Driven Learning of Topological Features with Applications to Protein Stability Prediction

Amish Mishra,Francis Motta
2024-08-09
Abstract:In this paper, we propose a data-driven method to learn interpretable topological features of biomolecular data and demonstrate the efficacy of parsimonious models trained on topological features in predicting the stability of synthetic mini proteins. We compare models that leverage automatically-learned structural features against models trained on a large set of biophysical features determined by subject-matter experts (SME). Our models, based only on topological features of the protein structures, achieved 92%-99% of the performance of SME-based models in terms of the average precision score. By interrogating model performance and feature importance metrics, we extract numerous insights that uncover high correlations between topological features and SME features. We further showcase how combining topological features and SME features can lead to improved model performance over either feature set used in isolation, suggesting that, in some settings, topological features may provide new discriminating information not captured in existing SME features that are useful for protein stability prediction.
Machine Learning,Data Analysis, Statistics and Probability
What problem does this paper attempt to address?
This paper aims to address the problem of predicting protein stability and to explore interpretable topological features related to protein stability. Specifically, the authors propose a data-driven approach to learn interpretable topological features from biomolecular data and demonstrate the effectiveness of minimal models trained on these topological features in predicting the stability of synthetic mini-proteins. By comparing models that utilize automatically learned structural features with models based on a large number of biophysical features determined by subject matter experts (SME), the authors found that models using only the topological features of protein structures achieved 92% to 99% of the performance of SME models in terms of average precision scores. Furthermore, the study revealed a high correlation between topological features and SME features, and showed that combining topological features with SME features can further improve model performance. This suggests that in some cases, topological features may provide new information that existing SME features fail to capture, which is useful for predicting protein stability.