RobNorm: Model-Based Robust Normalization for High-Throughput Proteomics from Mass Spectrometry Platform

Meng Wang,Lihua Jiang,Ruiqi Jian,Joanne Chen,Michael P. Snyder,Hua Tang
DOI: https://doi.org/10.1101/770115
2019-01-01
Abstract:Motivation Data normalization is an important step in processing proteomics data generated in mass spectrometry (MS) experiments, which aims to reduce sample-level variation and facilitate comparisons of samples. Previously published methods for normalization primarily depend on the assumption that the distribution of protein expression is similar across all samples. However, this assumption fails when the protein expression data is generated from heterogenous samples, such as from various tissue types. This led us to develop a novel data-driven method for improved normalization to correct the systematic bias meanwhile maintaining underlying biological heterogeneity. Methods To robustly correct the systematic bias, we used the density-power-weight method to down-weigh outliers and extended the one-dimensional robust fitting method described in the previous work of ([Windham, 1995][1], [Fujisawa and Eguchi, 2008][2]) to our structured data. We then constructed a robustness criterion and developed a new normalization algorithm, called RobNorm. Results In simulation studies and analysis of real data from the genotype-tissue expression (GTEx) project, we compared and evaluated the performance of RobNorm against other normalization methods. We found that the RobNorm approach exhibits the greatest reduction in systematic bias while maintaining across-tissue variation, especially for datasets from highly heterogeneous samples. Availability Contact huatang{at}stanford.edu and mpsnyder{at}stanford.edu ### Competing Interest Statement M.P.S. is a cofounder and is on the scientific advisory board of Personalis, Filtircine, SensOmics, Qbio, January, Mirvie, Oralome and Proteus. He is also on the scientific advisory board (SAB) of Genapsys and Jupiter. The other authors declare no competing interests. [1]: #ref-30 [2]: #ref-11
What problem does this paper attempt to address?