A space and time efficient algorithm for constructing compressed suffix arrays

Tak Wah Lam; Kunihiko Sadakane; Wing Kin Sung; Siu Ming Yiu

Conference Proceedings

A space and time efficient algorithm for constructing compressed suffix arrays

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2002) 2387 401-410

DOI: 10.1007/3-540-45655-4_43

35Citations

26Readers

Get full text

Abstract

With the first Human DNA being decoded into a sequence of about 2.8 billion base pairs, many biological research has been centered on analyzing this sequence. Theoretically speaking, it is now feasible to accommodate an index for human DNA in main memory so that any pattern can be located efficiently. This is due to the recent breakthrough on compressed suffix arrays, which reduces the space requirement from O(n log n) bits to O(n) bits. However, constructing compressed suffix arrays is still not an easy task because we still have to compute suffix arrays first and need a working memory of O(n log n) bits (i.e., more than 13 Gigabytes for human DNA). This paper initiates the study of constructing compressed suffix arrays directly from text. The main contribution is a new construction algorithm that uses only O(n) bits of working memory, and more importantly, the time complexity remains the same as before, i.e., O(n log n).

Cite

CITATION STYLE

APA

Lam, T. W., Sadakane, K., Sung, W. K., & Yiu, S. M. (2002). A space and time efficient algorithm for constructing compressed suffix arrays. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 2387, pp. 401–410). Springer Verlag. https://doi.org/10.1007/3-540-45655-4_43

A space and time efficient algorithm for constructing compressed suffix arrays

Abstract

Cite

Register to see more suggestions