Abstract:Background Structural variants (SVs) are less common than single nucleotide polymorphisms and indels in the population, but collectively account for a significant fraction of genetic polymorphism and diseases. Base pair differences arising from SVs are on a much higher order (>100 fold) than point mutations; however, none of the current detection methods are comprehensive, and currently available methodologies are incapable of providing sufficient resolution and unambiguous information across complex regions in the human genome. To address these challenges, we applied a high-throughput, cost-effective genome mapping technology to comprehensively discover genome-wide SVs and characterize complex regions of the YH genome using long single molecules (>150 kb) in a global fashion. Results Utilizing nanochannel-based genome mapping technology, we obtained 708 insertions/deletions and 17 inversions larger than 1 kb. Excluding the 59 SVs (54 insertions/deletions, 5 inversions) that overlap with N-base gaps in the reference assembly hg19, 666 non-gap SVs remained, and 396 of them (60%) were verified by paired-end data from whole-genome sequencing-based re-sequencing or de novo assembly sequence from fosmid data. Of the remaining 270 SVs, 260 are insertions and 213 overlap known SVs in the Database of Genomic Variants. Overall, 609 out of 666 (90%) variants were supported by experimental orthogonal methods or historical evidence in public databases. At the same time, genome mapping also provides valuable information for complex regions with haplotypes in a straightforward fashion. In addition, with long single-molecule labeling patterns, exogenous viral sequences were mapped on a whole-genome scale, and sample heterogeneity was analyzed at a new level. Conclusion Our study highlights genome mapping technology as a comprehensive and cost-effective method for detecting structural variation and studying complex regions in the human genome, as well as deciphering viral integration into the host genome.

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

Mapping and characterization of structural variation in 17,795 deeply sequenced human genomes

Multi-platform Discovery of Haplotype-Resolved Structural Variation in Human Genomes

Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation

High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation

Mapping Copy Number Variation by Population-Scale Genome Sequencing.

Full characterization of unresolved structural variation through long-read sequencing and optical genome mapping

Impact and characterization of serial structural variations across humans and great apes

Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural variation on gene expression and DNA methylation

A high-quality reference panel reveals the complexity and distribution of structural genome changes in a human population

Paired-End Mapping Reveals Extensive Structural Variation in the Human Genome

SVLR: Genome Structure Variant Detection Using Long Read Sequencing Data

PGG.SV: a whole-genome-sequencing-based structural variant resource and data analysis platform

Large complex structural rearrangements in human genomes harbor cryptic structures

Assessing structural variation in a personal genome—towards a human reference diploid genome

Long-read sequencing of 945 Han individuals identifies novel structural variants associated with phenotypic diversity and disease susceptibility

A High-Quality Human Reference Panel Reveals the Complexity and Distribution of Genomic Structural Variants

Rapid Detection of Structural Variation in a Human Genome Using Nanochannel-Based Genome Mapping Technology.

Haplotype-resolved diverse human genomes and integrated analysis of structural variation

Comprehensive Evaluation of Structural Variant Genotyping Methods Based on Long-Read Sequencing Data.

SV-Pop: population-based structural variant analysis and visualization