Abstract
Motivation Modern genomics laboratories generate massive volumes of sequencing data, often resulting in significant storage costs. Genomics storage consists of duplicate files, temporary processing files, and redundant intermediate data. Results We developed SeqManager, a web-based application that provides automated identification, classification, and management of sequencing data files with intelligent duplicate detection. It also detects intermediate sequencing files that can safely be removed. Evaluation across four genomics laboratory settings demonstrate that our tool is fast and has a very low memory footprint. Availability and implementation SeqManager is freely available under the MIT license at https://github.com/AIGeneRegulation/Sequencing-Data-Manager.
Cite
CITATION STYLE
Celerier, M., Oldfield, A. J., & Ritchie, W. (2025). SeqManager: a web-based tool for efficient sequencing data storage management and duplicate detection. Bioinformatics Advances, 5(1). https://doi.org/10.1093/bioadv/vbaf282
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.