SciDataFlow: a tool for improving the flow of data through science

Vince Buffalo
DOI: https://doi.org/10.1093/bioinformatics/btad754
IF: 5.8
2024-01-01
Bioinformatics
Abstract:Abstract Motivation Managing data and code in open scientific research is complicated by two key problems: large datasets often cannot be stored alongside code in repository platforms like GitHub, and iterative analysis can lead to unnoticed changes to data, increasing the risk that analyses are based on older versions of data. Results SciDataFlow is a fast, concurrent command-line tool paired with a simple Data Manifest specification that streamlines tracking data changes, uploading data to remote repositories, and pulling in all data necessary to reproduce a computational analysis. Availability and implementation SciDataFlow is available at https://github.com/vsbuffalo/scidataflow.
biochemical research methods,biotechnology & applied microbiology,mathematical & computational biology
What problem does this paper attempt to address?