Abstract:Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation models in genomics to learn generalizable features from unlabeled genome data that can then be fine-tuned for downstream tasks such as identifying regulatory elements. Due to the quadratic scaling of attention, previous Transformer-based genomic models have used 512 to 4k tokens as context (<0.001% of the human genome), significantly limiting the modeling of long-range interactions in DNA. In addition, these methods rely on tokenizers or fixed k-mers to aggregate meaningful DNA units, losing single nucleotide resolution where subtle genetic variations can completely alter protein function via single nucleotide polymorphisms (SNPs). Recently, Hyena, a large language model based on implicit convolutions was shown to match attention in quality while allowing longer context lengths and lower time complexity. Leveraging Hyena's new long-range capabilities, we present HyenaDNA, a genomic foundation model pretrained on the human reference genome with context lengths of up to 1 million tokens at the single nucleotide-level - an up to 500x increase over previous dense attention-based models. HyenaDNA scales sub-quadratically in sequence length (training up to 160x faster than Transformer), uses single nucleotide tokens, and has full global context at each layer. We explore what longer context enables - including the first use of in-context learning in genomics. On fine-tuned benchmarks from the Nucleotide Transformer, HyenaDNA reaches state-of-the-art (SotA) on 12 of 18 datasets using a model with orders of magnitude less parameters and pretraining data. On the GenomicBenchmarks, HyenaDNA surpasses SotA on 7 of 8 datasets on average by +10 accuracy points. Code at <a class="link-external link-https" href="https://github.com/HazyResearch/hyena-dna" rel="external noopener nofollow">this https URL</a>.

ProtHyena: A fast and efficient foundation protein language model at single amino acid Resolution

HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution

scHyena: Foundation Model for Full-Length Single-Cell RNA-Seq Analysis in Brain

How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena

DeProt: A protein language model with quantizied structure and disentangled attention

PeptideBERT: A Language Model based on Transformers for Peptide Property Prediction

ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing

Scavenging Hyena: Distilling Transformers into Long Convolution Models

HydraProt: A New Deep Learning Tool for Fast and Accurate Prediction of Water Molecule Positions for Protein Structures

ProteinAligner: A Multi-modal Pretraining Framework for Protein Foundation Models

Peptide Sequencing Via Protein Language Models

Unifying Sequences, Structures, and Descriptions for Any-to-Any Protein Generation with the Large Multimodal Model HelixProtX

ProtMamba: a homology-aware but alignment-free protein state space model

Efficient Inference, Training, and Fine-tuning of Protein Language Models

Retrieval-Enhanced Mutation Mastery: Augmenting Zero-Shot Prediction of Protein Language Model

Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling

xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein

In the twilight zone of protein sequence homology: do protein language models learn protein structure?

yHydra: Deep Learning enables an Ultra Fast Open Search by Jointly Embedding MS/MS Spectra and Peptides of Mass Spectrometry-based Proteomics

Modeling Protein Using Large-scale Pretrain Language Model