Abstract:In mass spectrometry-based proteomics, the identification and quantification of peptides and proteins is usually done using database search algorithms or spectral library matching. The use of deep learning algorithms can help improve the identification rates of peptides and proteins through the generation of high-fidelity theoretical spectrum which can be used as the basis of a more complete spectral library than those presently available. Current methods focus on predicting only backbone ions, such as y- and b-ions. However, the inclusion of non-backbone ions is necessary to truly improve spectral library matching. Here we focus on providing a user-friendly machine learning workflow, which we call mplete pectrum ictor (CoSpred). Using CoSpred users can create their own machine learning compatible training dataset and then train a Machine Learning model to predict both backbone and non-backbone ions. For the model a transformer encoder architecture is used to predict the complete MS/MS spectrum from a given peptide sequence. This model does not require background knowledge of fragment ion annotations or fragmentation rules. The model outputs the set of pairs ( , ) where is the m/z (mass-to-charge ratio) of a peak in the spectrum and is the intensity of the peak. The model presented here for validation was trained on the dataset available in the MassIVE data repository and shows superior performance in terms of various metrics (e.g. precision/recall for mass, cosine similarity for peak intensity, etc) between the true and predicted spectra. Furthermore, CoSpred can be used to create custom models that allow for accurate spectrum prediction for different experimental conditions. In addition to the transformer model provided in the package, the code is built modularly to allow for alternate ML models to be easily “plugged in”. The CoSpred workflow (preprocessing->training->inference) provides a path for state-of-art ML capabilities to be more accessible to proteomics scientists.

Machine learning strategies to tackle data challenges in mass spectrometry-based proteomics

Recent Developments in Machine Learning for Mass Spectrometry

Machine learning meets mass spectrometry: a focused perspective

Application of Machine Learning in Spatial Proteomics

Making MS Omics Data ML-Ready: SpeCollate Protocols

Deep learning neural network tools for proteomics

Advancing the Prediction of MS/MS Spectra Using Machine Learning

A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models

Bioinformatics Methods for Mass Spectrometry-Based Proteomics Data Analysis

Prediction of peptide mass spectral libraries with machine learning.

Deep Learning in Proteomics

Bioinformatics analysis of mass spectrometry-based proteomics data sets

Computational and Statistical Analysis of Protein Mass Spectrometry Data

Mass spectrometry‐based high‐throughput proteomics and its role in biomedical studies and systems biology

Getting Started with Machine Learning for Experimental Biochemists and Other Molecular Scientists

Discovery of Proteomics based on Machine learning

Machine Learning in Complex Organic Mixtures: Applying Domain Knowledge Allows for Meaningful Performance with Small Datasets.

Deep Learning in Proteomics Informatics: Applications, Challenges, and Future Directions

Liquid chromatography mass spectrometry-based proteomics: Biological and technological aspects

CoSpred: Machine learning workflow to predict tandem mass spectrum in proteomics

Machine Learning in Complex Organic Mixtures: Applying Domain Knowledge Allows for Meaningful Performance with Small Data Sets