GRAMEP: an alignment-free method based on the Maximum Entropy Principle for identifying SNPs

Matheus Henrique Pimenta-Zanon,André Yoshiaki Kashiwabara,André Luís Laforga Vanzela,Fabricio Martins Lopes
2024-11-22
Abstract:Background: Advances in high throughput sequencing technologies provide a huge number of genomes to be analyzed. Thus, computational methods play a crucial role in analyzing and extracting knowledge from the data generated. Investigating genomic mutations is critical because of their impact on chromosomal evolution, genetic disorders, and diseases. It is common to adopt aligning sequences for analyzing genomic variations. However, this approach can be computationally expensive and restrictive in scenarios with large datasets. Results: We present a novel method for identifying single nucleotide polymorphisms (SNPs) in DNA sequences from assembled genomes. This study proposes GRAMEP, an alignment-free approach that adopts the principle of maximum entropy to discover the most informative k-mers specific to a genome or set of sequences under investigation. The informative k-mers enable the detection of variant-specific mutations in comparison to a reference genome or other set of sequences. In addition, our method offers the possibility of classifying novel sequences with no need for organism-specific information. GRAMEP demonstrated high accuracy in both in silico simulations and analyses of viral genomes, including Dengue, HIV, and SARS-CoV-2. Our approach maintained accurate SARS-CoV-2 variant identification while demonstrating a lower computational cost compared to methods with the same purpose. Conclusions: GRAMEP is an open and user-friendly software based on maximum entropy that provides an efficient alignment-free approach to identifying and classifying unique genomic subsequences and SNPs with high accuracy, offering advantages over comparative methods. The instructions for use, applicability, and usability of GRAMEP are open access at <a class="link-external link-https" href="https://github.com/omatheuspimenta/GRAMEP" rel="external noopener nofollow">this https URL</a>
Genomics,Information Theory,Applications
What problem does this paper attempt to address?
The problem that this paper attempts to solve is how to identify single - nucleotide polymorphisms (SNPs) efficiently and accurately in large - scale genomic data analysis. Specifically, the paper proposes an alignment - free method based on the maximum entropy principle - GRAMEP, which is used to identify SNPs in DNA sequences from assembled genomes. Traditional methods usually rely on sequence alignment, which is computationally costly and less efficient when dealing with large - scale data sets. GRAMEP aims to reduce computational costs by avoiding sequence alignment and improve the accuracy of SNPs identification, especially when analyzing genomes with high recombination and rearrangement rates such as RNA viruses (e.g., dengue virus, HIV, and SARS - CoV - 2). This method not only improves the accuracy of SNPs identification but also reduces the demand for computational resources, making it an effective tool for processing large - scale genomic data sets.