Nanoq: ultra-fast quality control for nanopore reads

  • Steinig E
  • Coin L
N/ACitations
Citations of this article
62Readers
Mendeley users who have this article in their library.

Abstract

Nanopore sequencing is now routinely used in a variety of genomics applications, including whole genome assembly (Jain et al., 2018) and real-time infectious disease surveillance (Meredith et al., 2020). One of the first steps in many workflows is to assess the quality of reads, obtain summary statistics, and filter fragmented or low quality reads. With increasing throughput on scalable nanopore platforms like GridION or PromethION, fast quality control of sequence reads and the ability to generate summary statistics on-the-fly are required. Benchmarks indicate that nanoq is as fast as seqtk for small datasets (100,000 reads) and ~1.5x as fast for large datasets (3.5 million reads). Without quality scores, computing summary statistics is around ~2-3x faster than rust-bio-tools and seq-kit stats, 44x faster than seqtk, and up to ~450x faster than NanoStats (> 1.2 million reads per second). In read filtering applications, nanoq is considerably faster than other commonly used tools (NanoFilt, Filtlong). Memory consumption is consistent and tends to be lower than other applications (~5-10x). Nanoq offers nanopore-specific quality scores, read filtering options, and output compression. It can be applied to data from the public domain, as part of automated pipelines, in streaming applications, or to rapidly check progress of active sequencing runs.

Cite

CITATION STYLE

APA

Steinig, E., & Coin, L. (2022). Nanoq: ultra-fast quality control for nanopore reads. Journal of Open Source Software, 7(69), 2991. https://doi.org/10.21105/joss.02991

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free