Orion: A framework for GPU occupancy tuning

11Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

An important feature of modern GPU architectures is variable occupancy. Occupancy measures the ratio between the actual number of threads actively running on a GPU and the maximum number of threads that can be scheduled on a GPU. High-occupancy execution enables a large number of threads to run simultaneously and to hide memory latency, but may increase resource contention. Low-occupancy execution leads to less resource contention, but is less capable of hiding memory latency. Occupancy tuning is an important and challenging problem. A program running at two different occupancy levels can have three to four times difference in performance. We introduce Orion, the first GPU program occupancy tuning framework. The Orion framework automatically generates and chooses occupancy-adaptive code for any given GPU program. It is capable of finding the (near-)optimal occupancy level by combining static and dynamic tuning techniques. We demonstrate the efficiency of Orion with twelve representative benchmarks from the Rodinia benchmark suite and CUDA SDK evaluated on two different GPU architectures, obtaining up to 1.61 times speedup, 62.5% memory resource saving, and 6.7% energy saving compared to the baseline of optimized code compiled by nvcc.

Cite

CITATION STYLE

APA

Hayes, A. B., Li, L., Chavarría-Miranda, D., Song, S. L., & Zhang, E. Z. (2016). Orion: A framework for GPU occupancy tuning. In Proceedings of the 17th International Middleware Conference, Middleware 2016. Association for Computing Machinery, Inc. https://doi.org/10.1145/2988336.2988355

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free