Abstract
During the COVID-19 pandemic, wastewater-based epidemiology has progressively taken a central role as a pathogen surveillance tool. Tracking viral loads and variant outbreaks in sewage offers advantages over clinical surveillance methods by providing estimates not biased by testing practices and enabling early detection. However, wastewater-based epidemiology poses new computational research questions that need to be solved in order for this approach to be implemented broadly and successfully. Here, we address the variant deconvolution problem, where we aim to estimate the relative abundances of genomic variants from next-generation sequencing data of a mixed wastewater sample. We introduce LolliPop, a computational method to solve the variant deconvolution problem. LolliPop is tailored to wastewater time series sequencing data and applies temporal regularization in the form of a fused ridge penalty. We show that this regularization is equivalent to kernel smoothing and that it makes abundance estimates robust to very high levels of missing data, which is common for wastewater sequencing. We use the bootstrap to produce confidence intervals, and develop analytical standard errors that can produce similar confidence intervals at a fraction of the computational cost. We demonstrate the application of our method to data from the Swiss wastewater surveillance efforts as well as on simulated data.
Cite
CITATION STYLE
Dreifuss, D., Topolsky, I., Baykal, P. I., & Beerenwinkel, N. (2026). Tracking SARS-CoV-2 genomic variants in wastewater sequencing data with LolliPop. PLOS Computational Biology, 22(2), 1–19. https://doi.org/10.1371/journal.pcbi.1014003
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.