Data Lakes, Clouds and Commons: A Review of Platforms for Analyzing and Sharing Genomic Data

Robert L. Grossman
DOI: https://doi.org/10.48550/arXiv.1809.01699
2018-12-25
Abstract:Data commons collate data with cloud computing infrastructure and commonly used software services, tools and applications to create biomedical resources for the large-scale management, analysis, harmonization, and sharing of biomedical data. Over the past few years, data commons have been used to analyze, harmonize and share large scale genomics datasets. Data ecosystems can be built by interoperating multiple data commons. It can be quite labor intensive to curate, import and analyze the data in a data commons. Data lakes provide an alternative to data commons and simply provide access to data, with the data curation and analysis deferred until later and delegated to those that access the data. We review software platforms for managing, analyzing and sharing genomic data, with an emphasis on data commons, but also covering data ecosystems and data lakes.
Genomics,Computers and Society
What problem does this paper attempt to address?