High-throughput selection of human -emerged sORFs with high folding potential

Margaux Aubel,Filip Buchel,Brennen Heames,Alun Jones,Ondrej Honc,Erich Bornberg-Bauer,Klara Hlouchova
DOI: https://doi.org/10.1101/2024.01.22.576604
2024-01-24
Abstract:genes emerge from previously non-coding stretches of the genome. Their en-coded proteins are generally expected to be similar to random sequences and, accordingly, with no stable tertiary fold and high predicted disorder. However, structural properties of proteins and whether they differ during the stages of emergence and fixation have not been studied in depth and rely heavily on predictions. Here we generated a library of short human putative proteins of varying lengths and ages and sorted the candidates according to their structural compactness and disorder propensity. Using Förster resonance energy transfer (FRET) combined with Fluorescence-activated cell sorting (FACS) we were able to screen the library for most compact protein structures, as well as most elongated and flexible structures. Compact proteins are on average slightly shorter and contain lower predicted disorder than less compact ones. The predicted structures for most and least compact proteins correspond to expectations in that they contain more secondary structure content or higher disorder content, respectively. Our experiments indicate that older proteins have higher compactness and structural propensity compared to young ones. We discuss possible evolutionary scenarios and their implications underlying the age-dependencies of compactness and structural content of putative proteins.
Evolutionary Biology
What problem does this paper attempt to address?