Rye: Genetic ancestry inference at biobank scale

22Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Biobank projects are generating genomic data for many thousands of individuals. Computationalmethods are needed to handle these massive data sets, including genetic ancestry (GA) inference tools. Current methods for GA inference do not scale to biobank-size genomic datasets. We present Rye a new algorithm for GA inference at biobank scale. We compared the accuracy and runtime performance of Rye to the widely used RFMix, ADMIXTURE and iAdmix programs and applied it to a dataset of 488221 genome-wide variant samples from the UK Biobank. Rye infers GA based on principal component analysis of genomic variant samples from ancestral reference populations and query individuals. The algorithm s accuracy is powered by Metropolis-Hastings optimization and its speed is provided by non-negative least squares regression. Rye produces highly accurateGAestimates for threeway admixed populations African, European and Native American compared to RFMix and ADMIXTURE (R2 = 0.998 - 1.00), and shows 50× runtime improvement compared to ADMIXTURE on the UK Biobank dataset. Rye analysis of UK Biobank samples demonstrates how it can be used to infer GA at both continental and subcontinental levels. We discuss user consideration and options for the use of Rye; the program and its documentation are distributed on the GitHub repository: https://github.com/healthdisparities/rye.

Cite

CITATION STYLE

APA

Conley, A. B., Rishishwar, L., Ahmad, M., Sharma, S., Norris, E. T., Jordan, I. K., & Mariño-Ramírez, L. (2023). Rye: Genetic ancestry inference at biobank scale. Nucleic Acids Research, 51(8), e44. https://doi.org/10.1093/nar/gkad149

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free