Efficient join query processing algorithm CHMJ based on hadoop

7Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

This paper proposes a join query processing algorithm CoLocationHashMapJoin (CHMJ). First the study designs a multi-copy consistency hash algorithm. The algorithm distributes the data of tables over the cluster according to the hash values of the join property, which improves the data locality while ensure data availability. Second, based on the multi-copy consistency hash algorithm, the study proposes a parallel join query processing algorithm called HashMapJoin. HashMapJoin improves the efficiency of join query significantly. CHMJ has been used in Tencent's data warehouse system, and plays an important role in Tencent's daily analysis tasks. The results show that CHMJ improves the efficiency of join query processing by five times comparing to Hive. ©2012 ISCAS.

Cite

CITATION STYLE

APA

Zhao, Y. R., Wang, W. P., Meng, D., Zhang, S. B., & Li, J. (2012). Efficient join query processing algorithm CHMJ based on hadoop. Ruan Jian Xue Bao/Journal of Software, 23(8), 2032–2041. https://doi.org/10.3724/SP.J.1001.2012.04124

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free