Abstract
Deploying large numbers of small, low-power cores has been gaining traction recently as a system design strategy in high performance computing (HPC). The ARM platform that dominates the embedded and mobile computing segments is now being considered as an alternative to high-end x86 processors that largely dominate HPC because peak performance per watt may be substantially improved using off-the-shelf commodity processors. In this work we methodically characterize the performance and energy of HPC computations drawn from a number of problem domains on current ARM and x86 processors. Unsurprisingly, we find that the performance, energy and energy-delay product of applications running on these platforms varies significantly across problem types and inputs. Using static program analysis we further show that this variation can be explained largely in terms of the capabilities of two processor subsystems: single instruction multiple data (SIMD)/floating point and the cache/memory hierarchy; and that static analysis of this kind is sufficient to predict which platform is best for a particular application/input pair. In the context of these findings, we evaluate how some of the key architectural changes being made for upcoming 64-bit ARM platforms may impact HPC application performance. © 2014 Springer International Publishing Switzerland.
Cite
CITATION STYLE
Laurenzano, M. A., Tiwari, A., Jundt, A., Peraza, J., Ward, W. A., Campbell, R., & Carrington, L. (2014). Characterizing the performance-energy tradeoff of small ARM Cores in HPC computation. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 8632 LNCS, pp. 124–137). Springer Verlag. https://doi.org/10.1007/978-3-319-09873-9_11
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.