An Evaluation for Cross-Species Proteomics Research by Publicly Available Expressed Sequence Tag Database Search Using Tandem Mass Spectral Data

Mei Huang,Tong Chen,ZhuLong Chan
DOI: https://doi.org/10.1002/rcm.2631
2006-01-01
Abstract:With 1383 tandem mass spectra derived from 120 individual protein spots separated by the two-dimensional (2-D) gel electrophoresis of protein samples from three different species, comparative analyses were performed by searching the Expressed Sequence Tag (EST) database (DB) and the NCBI non-redundant (nr) DB of green plants, respectively, which uses the Mascot search engine to establish a statistical basis. It was confirmed that the former could identify more peptides manually validated by de novo sequencing (DNS) from fewer species in more closely phylogenetic relationships than the latter in a statistically significant manner. Our data demonstrated that correct peptide identifications were given low Mascot scores (e.g. 6-14) and incorrect peptide identifications were given high Mascot scores (e.g. 68-83). Our data also showed that the current evaluation approaches to protein assignments are unsatisfactory because a few 'false-positive' proteins are recognized and several 'false-negative' proteins are rescued by manual validation.
What problem does this paper attempt to address?