Protein lipograms

Jason Laurie,Amit K. Chattopadhyay,Darren R. Flower
DOI: https://doi.org/10.1016/j.jtbi.2017.07.009
IF: 2.405
2017-10-01
Journal of Theoretical Biology
Abstract:Linguistic analysis of protein sequences is an underexploited technique. Here, we capitalize on the concept of the lipogram to characterize sequences at the proteome levels. A lipogram is a literary composition which omits one or more letters. A protein lipogram likewise omits one or more types of amino acid. In this article, we establish a usable terminology for the decomposition of a sequence collection in terms of the lipogram. Next, we characterize Uniref50 using a lipogram decomposition. At the global level, protein lipograms exhibit power-law properties. A clear correlation with metabolic cost is seen. Finally, we use the lipogram construction to assign proteomes to the four branches of the tree-of-life: archaea, bacteria, eukaryotes and viruses. We conclude from this pilot study that the lipogram demonstrates considerable potential as an additional tool for sequence analysis and proteome classification.
biology,mathematical & computational biology
What problem does this paper attempt to address?