GPU-accelerated Kendall distance computation for large or sparse data

2Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Background: Current experimental practices typically produce large multidimensional datasets. Distance matrix calculation between elements (e.g., samples) for such data, although being often necessary in preprocessing for statistical inference or visualization, can be computationally demanding. Data sparsity, which is often observed in various experimental data modalities, such as single-cell sequencing in bioinformatics or collaborative filtering in recommendation systems, may pose additional algorithmic challenges. Results: We present GPU-Assisted Distance Estimation Software (GADES), a graphical processing unit (GPU)–enhanced package that allows for massively paralleled Kendall-τ distance matrices computation. The package’s architecture involves specific memory management, which lifts the limits for the data size imposed by GPU memory capacity. Additional algorithmic solutions provide a means to address the data sparsity problem and reinforce the acceleration effect for sparse datasets. Benchmarking against available central processing unit–based packages on simulated and real experimental single-cell RNA sequencing or single-cell ATAC sequencing datasets demonstrated significantly higher speed for GADES compared to other methods for both sparse and dense data processing, with additional performance boost for the sparse data. Conclusions: This work significantly contributes to the development of computational strategies for high-performance Kendall distance matrices computation and allows for the efficient processing of Big Data with the power of GPU. GADES is freely available at https://github.com/lab-medvedeva/GADES-main.

Cite

CITATION STYLE

APA

Akhtyamov, P., Nabi, A., Gafurov, V., Sizykh, A., Favorov, A., Medvedeva, Y., & Stupnikov, A. (2024). GPU-accelerated Kendall distance computation for large or sparse data. GigaScience, 13. https://doi.org/10.1093/gigascience/giae088

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free