Stencil computation optimization and auto-tuning on state-of-the-art multicore architectures

390Citations
Citations of this article
244Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Understanding the most efficient design and utilization of emerging multicore systems is one of the most challenging questions faced by the mainstream and scientific computing industries in several decades. Our work explores multicore stencil (nearest-neighbor) computations - a class of algorithms at the heart of many structured grid codes, including PDE solvers. We develop a number of effective optimization strategies, and build an auto-tuning environment that searches over our optimizations and their parameters to minimize runtime, while maximizing performance portability. To evaluate the effectiveness of these strategies we explore the broadest set of multicore architectures in the current HPC literature, including the Intel Clovertown, AMD Barcelona, Sun Victoria Falls, IBM QS22 PowerXCell 8i, and NVIDIA GTX280. Overall, our auto-tuning optimization methodology results in the fastest multicore stencil performance to date. Finally, we present several key insights into the architectural tradeoffs of emerging multicore designs and their implications on scientific algorithm development. © 2008 IEEE.

Cite

CITATION STYLE

APA

Datta, K., Murphy, M., Volkov, V., Williams, S., Carter, J., Oliker, L., … Yelick, K. (2008). Stencil computation optimization and auto-tuning on state-of-the-art multicore architectures. In 2008 SC - International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2008. IEEE Computer Society. https://doi.org/10.1109/SC.2008.5222004

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free