Semi-supervised Bayesian integration of multiple spatial proteomics datasets

Stephen D. Coleman,Lisa Breckels,Ross F. Waller,Kathryn S. Lilley,Chris Wallace,Oliver M. Crook,Paul D.W. Kirk
DOI: https://doi.org/10.1101/2024.02.08.579519
2024-04-02
Abstract:The subcellular localisation of proteins is a key determinant of their function. High-throughput analyses of these localisations can be performed using mass spectrometry-based spatial proteomics, which enables us to examine the localisation and relocalisation of proteins. Furthermore, complementary data sources can provide additional sources of functional or localisation information. Examples include protein annotations and other high-throughput ‘omic assays. Integrating these modalities can provide new insights as well as additional confidence in results, but existing approaches for integrative analyses of spatial proteomics datasets are limited in the types of data they can integrate and do not quantify uncertainty. Here we propose a semi-supervised Bayesian approach to integrate spatial proteomics datasets with other data sources, to improve the inference of protein sub-cellular localisation. We demonstrate our approach outperforms other transfer-learning methods and has greater flexibility in the data it can model. To demonstrate the flexibility of our approach, we apply our method to integrate spatial proteomics data generated for the parasite with time-course gene expression data generated over its cell cycle. Our findings suggest that proteins linked to invasion organelles are associated with expression programs that peak at the end of the first cell-cycle. Furthermore, this integrative analysis divides the dense granule proteins into heterogeneous populations suggestive of potentially different functions. Our method is disseminated via the mdir R package available on the lead author’s Github.
Systems Biology
What problem does this paper attempt to address?