Parallel index and query for large scale data analysis

62Citations
Citations of this article
61Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Modern scientific datasets present numerous data management and analysis challenges. State-of-the-art index and query technologies are critical for facilitating interactive exploration of large datasets, but numerous challenges remain in terms of designing a system for processing general scientific datasets. The system needs to be able to run on distributed multi-core platforms, efficiently utilize underlying I/O infrastructure, and scale to massive datasets. We present FastQuery, a novel software framework that address these challenges. FastQuery utilizes a state-of-theart index and query technology (FastBit) and is designed to process massive datasets on modern supercomputing platforms. We apply FastQuery to processing of a massive 50TB dataset generated by a large scale accelerator modeling code. We demonstrate the scalability of the tool to 11,520 cores. Motivated by the scientific need to search for interesting particles in this dataset, we use our framework to reduce search time from hours to tens of seconds. Copyright 2011 ACM.

Cite

CITATION STYLE

APA

Chou, J., Wu, K., Rübel, O., Howison, M., Qiang, J., Prabhat, … Shoshani, A. (2011). Parallel index and query for large scale data analysis. In Proceedings of 2011 SC - International Conference for High Performance Computing, Networking, Storage and Analysis. Association for Computing Machinery. https://doi.org/10.1145/2063384.2063424

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free